# Query internet bandwidth usage?

**URL:** <https://discuss.elastic.co/t/query-internet-bandwidth-usage/32541>\
**Category:** Elasticsearch\
**Created:** [October 20, 2015, 1:49am UTC](https://discuss.elastic.co/t/query-internet-bandwidth-usage/32541 "2015-10-20T01:49:17Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![allenmchan](https://avatars.discourse-cdn.com/v4/letter/a/ec9cab/32.png) [@allenmchan](https://discuss.elastic.co/u/allenmchan)\
**Post date:** [October 20, 2015, 1:49am UTC](https://discuss.elastic.co/t/query-internet-bandwidth-usage/32541/1 "2015-10-20T01:49:17Z")

</div>

Hi all,

My company is about to move to multi-datacenter implementation of Elasticsearch. Each datacenter will have a local cluster that stores the data. We will have tribe nodes (that kibana connects to) in 1 or 2 of the datacenters to combine data from multiple datacenters.

My network team is asking what kind of bandwidth utilization will querying across datacenter take up? Each of the datacenters will have multiple TBs of data. The network team is worried about someone running \* query over a month of indices and ES data nodes trying to send that much data over the pipe.

I know there is a limit for recovery in "indices.recovery.max\_bytes\_per\_sec". Does ES cap how much bytes per sec is used for querys?

Thanks for reading

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [October 20, 2015, 4:47am UTC](https://discuss.elastic.co/t/query-internet-bandwidth-usage/32541/2 "2015-10-20T04:47:16Z")

</div>

There's no limits you can apply to the network traffic that ES sends, it'll just pull over everything that it needs, which can be quite a bit.

---

<div class="post-metadata">

**Author:** ![allenmchan](https://avatars.discourse-cdn.com/v4/letter/a/ec9cab/32.png) [@allenmchan](https://discuss.elastic.co/u/allenmchan)\
**Post date:** [October 20, 2015, 5:37am UTC](https://discuss.elastic.co/t/query-internet-bandwidth-usage/32541/3 "2015-10-20T05:37:09Z")

</div>

Interesting. Thanks for that. Any chance we can add this as a feature request? I assume most people who use elasticsearch in multidatacenters do not have a dedicated pipe between the datacenters just for elasticsearch

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [October 20, 2015, 5:38am UTC](https://discuss.elastic.co/t/query-internet-bandwidth-usage/32541/4 "2015-10-20T05:38:38Z")

</div>

I don't think it'd work. ES has no concept of what may be higher priority traffic (eg cluster state changes), so you'd end up with nodes dropping out of the cluster when someone ran a large query.

That said, feel free to raise the concern on GH, the core team may have other ideas 😄

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [October 20, 2015, 6:01pm UTC](https://discuss.elastic.co/t/query-internet-bandwidth-usage/32541/5 "2015-10-20T18:01:56Z")

</div>

In general, the amount of data passed between nodes is limited to the  
number of documents selected to be returned. Read the chapter called  
'Distributed Search Execution' in the official guide: [snip]

I would have posted a link, but the mailing list software does not allow  
it. Sigh, please bring back Google Groups.

From the deep pagination block:

"Remember that each shard must build a priority queue of length from +  
size, all of which need to be passed back to the coordinating node. And the  
coordinating node needs to sort through number\_of\_shards \* (from + size)  
documents in order to find the correct size documents."

Basically there will be (number\_of\_shards \* (from + size)) documents of  
data plus overhead following on the network. The main limitation for most  
cases is not the information contained in each query, but the number of  
queries.

AFAIK, there is not throttling of query bytes sent.

Cheers,

Ivan

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [October 20, 2015, 6:52pm UTC](https://discuss.elastic.co/t/query-internet-bandwidth-usage/32541/6 "2015-10-20T18:52:15Z")

</div>

Strange. Sounds like the spam filter found that you are a new user! So it forbids to post a first post with a link... o\_O

Obviously you're not a new user!

That should be fine now...

---

<div class="post-metadata">

**Author:** ![allenmchan](https://avatars.discourse-cdn.com/v4/letter/a/ec9cab/32.png) [@allenmchan](https://discuss.elastic.co/u/allenmchan)\
**Post date:** [October 20, 2015, 7:06pm UTC](https://discuss.elastic.co/t/query-internet-bandwidth-usage/32541/7 "2015-10-20T19:06:21Z")

</div>

Thanks Ivan. i just read through the doc you are referring to.

It seems like in the case of kibana, the documents returned are limited even if user queries \* from 30 days of indices. So the results are sent over the wire 1000 docs at a time (depending on the kibana config) as the user scroll through the results.  
Does that sound about right?

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [October 21, 2015, 6:12am UTC](https://discuss.elastic.co/t/query-internet-bandwidth-usage/32541/8 "2015-10-21T06:12:41Z")

</div>

Not sure what the defaults are in kibana, but that sounds correct.  
Remember, the number of shards also affects the internode communication  
since the coordinating node needs to sort the values.

Of course, your document size is also a factor. I have killed lightweight  
app servers before that were given 1gb result responses that elasticsearch  
handled with no issues.

Cheers,

Ivan

---

<div class="post-metadata">

**Author:** ![allenmchan](https://avatars.discourse-cdn.com/v4/letter/a/ec9cab/32.png) [@allenmchan](https://discuss.elastic.co/u/allenmchan)\
**Post date:** [October 22, 2015, 6:30am UTC](https://discuss.elastic.co/t/query-internet-bandwidth-usage/32541/9 "2015-10-22T06:30:41Z")

</div>

would a tiered tribe nodes architecture minimize the amount of traffic across the datacenter link?

For example each datacenter would have a dedicated tribe node and a "master" tribe node would talk to the datacenter tribe nodes.

A request from kibana would go to the master tribe node and it would propagate to the datacenter tribe nodes.

All the shards returned from the data nodes would be sorted by its datacenter tribe node and fully sorted document size would be sent over the datacenter link to the master tribe node.

Does this make sense or complicating it too much?

---

<div class="post-metadata">

**Author:** ![allenmchan](https://avatars.discourse-cdn.com/v4/letter/a/ec9cab/32.png) [@allenmchan](https://discuss.elastic.co/u/allenmchan)\
**Post date:** [October 22, 2015, 6:37am UTC](https://discuss.elastic.co/t/query-internet-bandwidth-usage/32541/10 "2015-10-22T06:37:49Z")

</div>

nevermind. looks like tiered tribe node architecture is not supported

> <https://github.com/elastic/elasticsearch/issues/12814>

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:43pm UTC](https://discuss.elastic.co/t/query-internet-bandwidth-usage/32541/11 "2017-07-05T23:43:22Z")

</div>


