# Memory utilization - predicting 'out of heap space' errors

**URL:** <https://discuss.elastic.co/t/memory-utilization-predicting-out-of-heap-space-errors/5546>\
**Category:** Elasticsearch\
**Created:** [October 9, 2011, 5:59pm UTC](https://discuss.elastic.co/t/memory-utilization-predicting-out-of-heap-space-errors/5546 "2011-10-09T17:59:14Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![Yooz](https://avatars.discourse-cdn.com/v4/letter/y/13edae/32.png) [@Yooz](https://discuss.elastic.co/u/Yooz)\
**Post date:** [October 9, 2011, 5:59pm UTC](https://discuss.elastic.co/t/memory-utilization-predicting-out-of-heap-space-errors/5546/1 "2011-10-09T17:59:14Z")

</div>

Hello,

We are running ES 0.17.7 on 4 16GB boxes. 14GB of memory is locked  
(dedicated) for ES. There are several indices, the largest ones range from  
50GB to 200GB of data @ 50 - 125 million documents. Currently there is no  
issue with memory, but the data size is continually growing. Because the  
memory is locked, it is hard to tell from the OS level how much memory ES  
actually needs. Is there a parameter exposed in the status API or a rule of  
thumb based on the data that shows if ES is running close to the limit?

I.e. because ES scales so well, we want to add more capacity ahead of time  
in order to avoid errors due to memory issues.

Also, Shay has mentioned in several posts that slicing the data up by date  
indices (daily, weekly), should minimize the amount of memory used by the  
Lucene indices. Is this optimization driven by the use-case? I.e. users  
would be most interested in the recent data and would query the past N days  
rather than search the entire index? What happens if they consistently  
query the entire date range, would this slicing scheme become inefficient  
compared to having one massive index?

Thanks!  
--young

---

<div class="post-metadata">

**Author:** ![AEvar\_Arnfjord\_Bjarm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aevar_arnfjord_bjarm/32/2746_2.png) [@AEvar\_Arnfjord\_Bjarm](https://discuss.elastic.co/u/AEvar_Arnfjord_Bjarm)\
**Post date:** [October 9, 2011, 9:32pm UTC](https://discuss.elastic.co/t/memory-utilization-predicting-out-of-heap-space-errors/5546/2 "2011-10-09T21:32:03Z")

</div>

Try Bigdesk, it gives you a graphical view of allocated / used memory.  
On Oct 9, 2011 7:59 PM, "Yooz" [youngmaeng@gmail.com](mailto:youngmaeng@gmail.com) wrote:

> Hello,
> 
> We are running ES 0.17.7 on 4 16GB boxes. 14GB of memory is locked  
> (dedicated) for ES. There are several indices, the largest ones range from  
> 50GB to 200GB of data @ 50 - 125 million documents. Currently there is no  
> issue with memory, but the data size is continually growing. Because the  
> memory is locked, it is hard to tell from the OS level how much memory ES  
> actually needs. Is there a parameter exposed in the status API or a rule of  
> thumb based on the data that shows if ES is running close to the limit?
> 
> I.e. because ES scales so well, we want to add more capacity ahead of time  
> in order to avoid errors due to memory issues.
> 
> Also, Shay has mentioned in several posts that slicing the data up by date  
> indices (daily, weekly), should minimize the amount of memory used by the  
> Lucene indices. Is this optimization driven by the use-case? I.e. users  
> would be most interested in the recent data and would query the past N days  
> rather than search the entire index? What happens if they consistently  
> query the entire date range, would this slicing scheme become inefficient  
> compared to having one massive index?
> 
> Thanks!  
> --young

---

<div class="post-metadata">

**Author:** ![Yooz](https://avatars.discourse-cdn.com/v4/letter/y/13edae/32.png) [@Yooz](https://discuss.elastic.co/u/Yooz)\
**Post date:** [October 10, 2011, 10:13pm UTC](https://discuss.elastic.co/t/memory-utilization-predicting-out-of-heap-space-errors/5546/3 "2011-10-10T22:13:59Z")

</div>

Thanks! Will try it out.

I've noticed search performance on some of the larger indices take 8-10  
seconds. For example on an index with 160 million documents @ 220gb size  
(not counting replication) w/15 shards, a text phrase search on a single  
field consistently took ~8500 milliseconds (changing the actual phrase to  
prevent caching results). There are no other clients on the cluster. The  
search query is a single phrase query with a range filter on a field called  
postdate (date of the document). No facets, no sorting. The query type is  
query\_and\_fetch.

Is this expected performance? Is there a way to decrease the latency down  
to sub-second response times for an index of this size?

Thanks again,  
--young

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [October 10, 2011, 10:34pm UTC](https://discuss.elastic.co/t/memory-utilization-predicting-out-of-heap-space-errors/5546/4 "2011-10-10T22:34:58Z")

</div>

> I've noticed search performance on some of the larger indices take  
> 8-10 seconds. For example on an index with 160 million documents @  
> 220gb size (not counting replication) w/15 shards, a text phrase  
> search on a single field consistently took ~8500 milliseconds  
> (changing the actual phrase to prevent caching results). There are no  
> other clients on the cluster. The search query is a single phrase  
> query with a range filter on a field called postdate (date of the  
> document). No facets, no sorting. The query type is query\_and\_fetch.

Try using a numeric\_range filter for postdate instead. numeric\_range is  
good for fields with many many unique terms (such as a datetime would  
have).

clint

---

<div class="post-metadata">

**Author:** ![Yooz](https://avatars.discourse-cdn.com/v4/letter/y/13edae/32.png) [@Yooz](https://discuss.elastic.co/u/Yooz)\
**Post date:** [October 11, 2011, 12:54am UTC](https://discuss.elastic.co/t/memory-utilization-predicting-out-of-heap-space-errors/5546/5 "2011-10-11T00:54:46Z")

</div>

Great suggestion Clint! This small change dropped down the average response  
times by more than half. It takes between 2500-4000 milliseconds per query  
(again with no facets, sorting) from ~8500 milliseconds.

When turning on 3 termsFacet fields (2 with cardinality in the 10's, 1 with  
cardinality in 100's, both types limited to 20 top results), and a  
dateHistogram facet with the interval set to "week", the performance becomes  
a bit more unpredictable. Response times range from 2000-10000  
milliseconds. I'm not sure what the variable here is (maybe the amount of  
results returned?) Monitoring on BigDesk seems to indicate that resources  
are sufficient (only small cpu spikes during the search query execution, no  
memory usage flux). Is it reasonable to expect sub-second query responses  
with large indices? 2-3 second return time is acceptable in our use case  
but it would be nice to know what the bottleneck is.

Cheers!  
--young

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [October 12, 2011, 7:14pm UTC](https://discuss.elastic.co/t/memory-utilization-predicting-out-of-heap-space-errors/5546/6 "2011-10-12T19:14:22Z")

</div>

Can you gist the full search request you make? With all the parameters. Most  
times, query\_then\_fetch is better then query\_and\_fetch.

On Tue, Oct 11, 2011 at 2:54 AM, Yooz [youngmaeng@gmail.com](mailto:youngmaeng@gmail.com) wrote:

> Great suggestion Clint! This small change dropped down the average  
> response times by more than half. It takes between 2500-4000 milliseconds  
> per query (again with no facets, sorting) from ~8500 milliseconds.
> 
> When turning on 3 termsFacet fields (2 with cardinality in the 10's, 1 with  
> cardinality in 100's, both types limited to 20 top results), and a  
> dateHistogram facet with the interval set to "week", the performance becomes  
> a bit more unpredictable. Response times range from 2000-10000  
> milliseconds. I'm not sure what the variable here is (maybe the amount of  
> results returned?) Monitoring on BigDesk seems to indicate that resources  
> are sufficient (only small cpu spikes during the search query execution, no  
> memory usage flux). Is it reasonable to expect sub-second query responses  
> with large indices? 2-3 second return time is acceptable in our use case  
> but it would be nice to know what the bottleneck is.
> 
> Cheers!  
> --young

---

<div class="post-metadata">

**Author:** ![Yooz](https://avatars.discourse-cdn.com/v4/letter/y/13edae/32.png) [@Yooz](https://discuss.elastic.co/u/Yooz)\
**Post date:** [October 13, 2011, 6:46am UTC](https://discuss.elastic.co/t/memory-utilization-predicting-out-of-heap-space-errors/5546/7 "2011-10-13T06:46:33Z")

</div>

Here is the gist generated directly from the java client:

> <https://gist.github.com/youngmaeng/1283492>

As the queries ramped up, the heap usage shot up (via bigdesk) from 3-4 GB  
per box to 12-14GB (14 is max). The filter cache settings are at default.

---

<div class="post-metadata">

**Author:** ![Yooz](https://avatars.discourse-cdn.com/v4/letter/y/13edae/32.png) [@Yooz](https://discuss.elastic.co/u/Yooz)\
**Post date:** [October 13, 2011, 10:00pm UTC](https://discuss.elastic.co/t/memory-utilization-predicting-out-of-heap-space-errors/5546/8 "2011-10-13T22:00:28Z")

</div>

I think I understand what the issue is with my setup. What I didn't mention  
was that I have 15 indices, each with a HUGE difference in the amount of  
documents they index. E.g. One index consists of 1000 documents, while  
another consists of 180 million documents. Because elasticsearch rebalances  
the cluster solely based on the shards per node, there is no consideration  
given to the amount of data that shard holds. Consequently, I've been  
getting the larger index shards allocated to the same node resulting in  
memory issues for that particular instance. In some cases this causes an  
out of memory error and brings that node to a wedged state (doesn't show up  
in es monitoring, logs continually spit out 'out of memory') and the data  
proceeds to be rebalanced to the other nodes. Unfortunately this results in  
a cascading failure because the same problem happens when the same load hits  
the next unlucky node.

Are there any plans to change the rebalancing scheme in the near future? If  
not, could I get a few entry points in the source on where I would patch the  
code to make index size aware rebalancing work?

Thanks!  
--young

---

<div class="post-metadata">

**Author:** ![Yooz](https://avatars.discourse-cdn.com/v4/letter/y/13edae/32.png) [@Yooz](https://discuss.elastic.co/u/Yooz)\
**Post date:** [October 13, 2011, 10:10pm UTC](https://discuss.elastic.co/t/memory-utilization-predicting-out-of-heap-space-errors/5546/9 "2011-10-13T22:10:20Z")

</div>

Also, would there be any issues scaling up to 10,000's of indices (most of  
them very small, on the order of 1,000 documents, but a few monolithic ones  
at 300 million documents)?

Thanks again!  
--young

---

<div class="post-metadata">

**Author:** ![ppearcy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ppearcy/32/980_2.png) [@ppearcy](https://discuss.elastic.co/u/ppearcy)\
**Post date:** [October 14, 2011, 1:52am UTC](https://discuss.elastic.co/t/memory-utilization-predicting-out-of-heap-space-errors/5546/10 "2011-10-14T01:52:39Z")

</div>

I'd recommend more RAM for the OS disk caches. I normally run a 50/50  
ES to OS RAM config. It's easy enough to tweak and test. We do this  
with 16GB servers and 48GB servers. Also, keep in mind initial queries  
aren't blazing fast, as disk caches and internal ES caches need to get  
built up.

Typically, that many indices is not recommended. Better to have a  
larger index with lots of shards and use the routing functionality.  
Only testing will tell how things perform with your setup and data,  
though 🙂

Best Regards,  
Paul

On Oct 13, 4:10 pm, Yooz [youngma...@gmail.com](mailto:youngma...@gmail.com) wrote:

> Also, would there be any issues scaling up to 10,000's of indices (most of  
> them very small, on the order of 1,000 documents, but a few monolithic ones  
> at 300 million documents)?
> 
> Thanks again!  
> --young

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [October 14, 2011, 1:27pm UTC](https://discuss.elastic.co/t/memory-utilization-predicting-out-of-heap-space-errors/5546/11 "2011-10-14T13:27:33Z")

</div>

Heya, yea, the balancing scheme elasticsearch uses is not perfect for this  
type of scenario. It was purposed to add one based on index size / document  
count, but its not there yet (though, a lot of work has going into building  
the basis for it in the future, which also allowed to do awareness  
allocation).

Its not a simple feature, though, I might implement a simpler feature for  
now (requires some thinking) of allowing to say that an index should be  
forced to rebalance on its own, without taking into account the rest of the  
indices.

On Fri, Oct 14, 2011 at 12:00 AM, Yooz [youngmaeng@gmail.com](mailto:youngmaeng@gmail.com) wrote:

> I think I understand what the issue is with my setup. What I didn't  
> mention was that I have 15 indices, each with a HUGE difference in the  
> amount of documents they index. E.g. One index consists of 1000 documents,  
> while another consists of 180 million documents. Because elasticsearch  
> rebalances the cluster solely based on the shards per node, there is no  
> consideration given to the amount of data that shard holds. Consequently,  
> I've been getting the larger index shards allocated to the same node  
> resulting in memory issues for that particular instance. In some cases this  
> causes an out of memory error and brings that node to a wedged state  
> (doesn't show up in es monitoring, logs continually spit out 'out of  
> memory') and the data proceeds to be rebalanced to the other nodes.  
> Unfortunately this results in a cascading failure because the same problem  
> happens when the same load hits the next unlucky node.
> 
> Are there any plans to change the rebalancing scheme in the near future?  
> If not, could I get a few entry points in the source on where I would patch  
> the code to make index size aware rebalancing work?
> 
> Thanks!  
> --young

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [October 14, 2011, 1:28pm UTC](https://discuss.elastic.co/t/memory-utilization-predicting-out-of-heap-space-errors/5546/12 "2011-10-14T13:28:31Z")

</div>

10k indices might be too many for a smaller set of nodes. Each index (well,  
actually shard, but even with 1 shard per index, thats 10k shards) is a  
Lucene indices, which requires resources. As was suggested, I recommend  
using larger amount of shards, and using routing instead.

On Fri, Oct 14, 2011 at 12:10 AM, Yooz [youngmaeng@gmail.com](mailto:youngmaeng@gmail.com) wrote:

> Also, would there be any issues scaling up to 10,000's of indices (most of  
> them very small, on the order of 1,000 documents, but a few monolithic ones  
> at 300 million documents)?
> 
> Thanks again!  
> --young

---

<div class="post-metadata">

**Author:** ![Yooz](https://avatars.discourse-cdn.com/v4/letter/y/13edae/32.png) [@Yooz](https://discuss.elastic.co/u/Yooz)\
**Post date:** [October 14, 2011, 6:17pm UTC](https://discuss.elastic.co/t/memory-utilization-predicting-out-of-heap-space-errors/5546/13 "2011-10-14T18:17:21Z")

</div>

@Paul: Thanks for the tip! I didn't think about the OS level caches and just  
dedicated all the memory to the ES process.

@Shay: Thanks for the clarification and suggestions!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:51am UTC](https://discuss.elastic.co/t/memory-utilization-predicting-out-of-heap-space-errors/5546/14 "2017-07-06T03:51:52Z")

</div>


