# Graceful way to un-overwhelm the ES

**URL:** <https://discuss.elastic.co/t/graceful-way-to-un-overwhelm-the-es/7643>\
**Category:** Elasticsearch\
**Created:** [May 10, 2012, 8:38pm UTC](https://discuss.elastic.co/t/graceful-way-to-un-overwhelm-the-es/7643 "2012-05-10T20:38:28Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![andym](https://avatars.discourse-cdn.com/v4/letter/a/7ab992/32.png) [@andym](https://discuss.elastic.co/u/andym)\
**Post date:** [May 10, 2012, 8:38pm UTC](https://discuss.elastic.co/t/graceful-way-to-un-overwhelm-the-es/7643/1 "2012-05-10T20:38:28Z")

</div>

Hi,

I am currently running indexing on c1.xlarge (with 4 ephemeral drives  
in RAID0 and gateway going to S3). Everything works great except every  
several hours ES gets overwhelmed and indexing slows significantly.  
From what I can see from bigdesk during “normal” ES operation “Heap  
Mem” window has jigsaw pattern but when it gets overwhelmed seems like  
no GC happens (no jigsaw pattern in bigdesk) and memory is maxed  
( configured at 5120M)

Doing ES service restart (through “bin/service/elasticsearch –  
restart”) solves the problem for a few hours but then the problem re-  
appears.

I wonder whether restarting ES when it is on such state going to lead  
to any data loss so I can put this into a cron job to assure indexing  
continues (or whether there are any better ways to address the  
problem)

Thanks,

-- Andy

P.S. Some background: I am running ES 19.2 with “refresh interval” set  
to zero and ES currently has about 400 million documents in 2 indexes  
with about 600G total index size (and I expect about 600M docs more  
with around 1T of data. The mapping has \_source set to compressed.)  
The data processing and insertion into ES is done by multiple threads  
on 20 or so m1.xlarge machines (when ES goes down or returns errors  
threads back-off with exponential timeout and restart when ES is back  
on-line). There are 8-12 threads per machine doing mostly data  
processing and if I were to trust that in bigdesk “HTTP channels”  
indicate number of active connections, then it means that 30-40  
threads are connected to ES at any given time. Indexing rate is about  
750 docs per second sometime maxing out at 10,000 docs per second. The  
average doc size is about 5000 bytes

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [May 10, 2012, 8:48pm UTC](https://discuss.elastic.co/t/graceful-way-to-un-overwhelm-the-es/7643/2 "2012-05-10T20:48:12Z")

</div>

We are talking about a single ES node, right? For the amount of data that  
you indexed, seems like you are hitting memory limits, 520mb for the amount  
of data you have is not enough, probably should go to 3.5 or 4 gb (out of  
the 7gb this instance type has) as ES\_HEAP\_SIZE.

On Thu, May 10, 2012 at 11:38 PM, andym [imwellnow@gmail.com](mailto:imwellnow@gmail.com) wrote:

> Hi,
> 
> I am currently running indexing on c1.xlarge (with 4 ephemeral drives  
> in RAID0 and gateway going to S3). Everything works great except every  
> several hours ES gets overwhelmed and indexing slows significantly.  
> From what I can see from bigdesk during “normal” ES operation “Heap  
> Mem” window has jigsaw pattern but when it gets overwhelmed seems like  
> no GC happens (no jigsaw pattern in bigdesk) and memory is maxed  
> ( configured at 5120M)
> 
> Doing ES service restart (through “bin/service/elasticsearch –  
> restart”) solves the problem for a few hours but then the problem re-  
> appears.
> 
> I wonder whether restarting ES when it is on such state going to lead  
> to any data loss so I can put this into a cron job to assure indexing  
> continues (or whether there are any better ways to address the  
> problem)
> 
> Thanks,
> 
> -- Andy
> 
> P.S. Some background: I am running ES 19.2 with “refresh interval” set  
> to zero and ES currently has about 400 million documents in 2 indexes  
> with about 600G total index size (and I expect about 600M docs more  
> with around 1T of data. The mapping has \_source set to compressed.)  
> The data processing and insertion into ES is done by multiple threads  
> on 20 or so m1.xlarge machines (when ES goes down or returns errors  
> threads back-off with exponential timeout and restart when ES is back  
> on-line). There are 8-12 threads per machine doing mostly data  
> processing and if I were to trust that in bigdesk “HTTP channels”  
> indicate number of active connections, then it means that 30-40  
> threads are connected to ES at any given time. Indexing rate is about  
> 750 docs per second sometime maxing out at 10,000 docs per second. The  
> average doc size is about 5000 bytes

---

<div class="post-metadata">

**Author:** ![andym](https://avatars.discourse-cdn.com/v4/letter/a/7ab992/32.png) [@andym](https://discuss.elastic.co/u/andym)\
**Post date:** [May 10, 2012, 8:59pm UTC](https://discuss.elastic.co/t/graceful-way-to-un-overwhelm-the-es/7643/3 "2012-05-10T20:59:07Z")

</div>

Yes, it's single node and ES\_HEAP\_SIZE is as 5120M (not 520M)

You're are right, I'll have to move to bigger machine or split it into  
2 machines as it started happening relatively recently (at ~300M docs)

My question is whether these restarts are safe at the moment and do  
not lead to data loss in ES (where ES would return "OK" to processing  
threads which would mark jobs as completed, but then ES would not  
persist them due to restart). ES is currently running with  
threadpool.index.type: cached  
threadpool.bulk.type: cached

I tried to make these "blocking" but then processing threads were idle  
most of the time just waiting for ES to return

On May 10, 4:48 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:

> We are talking about a single ES node, right? For the amount of data that  
> you indexed, seems like you are hitting memory limits, 520mb for the amount  
> of data you have is not enough, probably should go to 3.5 or 4 gb (out of  
> the 7gb this instance type has) as ES\_HEAP\_SIZE.
> 
> On Thu, May 10, 2012 at 11:38 PM, andym [imwell...@gmail.com](mailto:imwell...@gmail.com) wrote:
> 
> > Hi,
> 
> > I am currently running indexing on c1.xlarge (with 4 ephemeral drives  
> > in RAID0 and gateway going to S3). Everything works great except every  
> > several hours ES gets overwhelmed and indexing slows significantly.  
> > From what I can see from bigdesk during “normal” ES operation “Heap  
> > Mem” window has jigsaw pattern but when it gets overwhelmed seems like  
> > no GC happens (no jigsaw pattern in bigdesk) and memory is maxed  
> > ( configured at 5120M)
> 
> > Doing ES service restart (through “bin/service/elasticsearch –  
> > restart”) solves the problem for a few hours but then the problem re-  
> > appears.
> 
> > I wonder whether restarting ES when it is on such state going to lead  
> > to any data loss so I can put this into a cron job to assure indexing  
> > continues (or whether there are any better ways to address the  
> > problem)
> 
> > Thanks,
> 
> > -- Andy
> 
> > P.S. Some background: I am running ES 19.2 with “refresh interval” set  
> > to zero and ES currently has about 400 million documents in 2 indexes  
> > with about 600G total index size (and I expect about 600M docs more  
> > with around 1T of data. The mapping has \_source set to compressed.)  
> > The data processing and insertion into ES is done by multiple threads  
> > on 20 or so m1.xlarge machines (when ES goes down or returns errors  
> > threads back-off with exponential timeout and restart when ES is back  
> > on-line). There are 8-12 threads per machine doing mostly data  
> > processing and if I were to trust that in bigdesk “HTTP channels”  
> > indicate number of active connections, then it means that 30-40  
> > threads are connected to ES at any given time. Indexing rate is about  
> > 750 docs per second sometime maxing out at 10,000 docs per second. The  
> > average doc size is about 5000 bytes

---

<div class="post-metadata">

**Author:** ![Craig\_Brown](https://avatars.discourse-cdn.com/v4/letter/c/ce7236/32.png) [@Craig\_Brown](https://discuss.elastic.co/u/Craig_Brown)\
**Post date:** [May 10, 2012, 9:08pm UTC](https://discuss.elastic.co/t/graceful-way-to-un-overwhelm-the-es/7643/4 "2012-05-10T21:08:00Z")

</div>

We're on AWS and ran into some very similar problems on 2 nodes. I  
ended up using m2xl nodes with 12GB for ES across 2 nodes and indexing  
ran very well. We're up to 420m docs.

- Craig

On Thu, May 10, 2012 at 2:59 PM, andym [imwellnow@gmail.com](mailto:imwellnow@gmail.com) wrote:

> Yes, it's single node and ES\_HEAP\_SIZE is as 5120M (not 520M)
> 
> You're are right, I'll have to move to bigger machine or split it into  
> 2 machines as it started happening relatively recently (at ~300M docs)
> 
> My question is whether these restarts are safe at the moment and do  
> not lead to data loss in ES (where ES would return "OK" to processing  
> threads which would mark jobs as completed, but then ES would not  
> persist them due to restart). ES is currently running with  
> threadpool.index.type: cached  
> threadpool.bulk.type: cached
> 
> I tried to make these "blocking" but then processing threads were idle  
> most of the time just waiting for ES to return
> 
> On May 10, 4:48 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> 
> > We are talking about a single ES node, right? For the amount of data that  
> > you indexed, seems like you are hitting memory limits, 520mb for the amount  
> > of data you have is not enough, probably should go to 3.5 or 4 gb (out of  
> > the 7gb this instance type has) as ES\_HEAP\_SIZE.
> > 
> > On Thu, May 10, 2012 at 11:38 PM, andym [imwell...@gmail.com](mailto:imwell...@gmail.com) wrote:
> > 
> > > Hi,
> > 
> > > I am currently running indexing on c1.xlarge (with 4 ephemeral drives  
> > > in RAID0 and gateway going to S3). Everything works great except every  
> > > several hours ES gets overwhelmed and indexing slows significantly.  
> > > From what I can see from bigdesk during “normal” ES operation “Heap  
> > > Mem” window has jigsaw pattern but when it gets overwhelmed seems like  
> > > no GC happens (no jigsaw pattern in bigdesk) and memory is maxed  
> > > ( configured at 5120M)
> > 
> > > Doing ES service restart (through “bin/service/elasticsearch –  
> > > restart”) solves the problem for a few hours but then the problem re-  
> > > appears.
> > 
> > > I wonder whether restarting ES when it is on such state going to lead  
> > > to any data loss so I can put this into a cron job to assure indexing  
> > > continues (or whether there are any better ways to address the  
> > > problem)
> > 
> > > Thanks,
> > 
> > > -- Andy
> > 
> > > P.S. Some background: I am running ES 19.2 with “refresh interval” set  
> > > to zero and ES currently has about 400 million documents in 2 indexes  
> > > with about 600G total index size (and I expect about 600M docs more  
> > > with around 1T of data. The mapping has \_source set to compressed.)  
> > > The data processing and insertion into ES is done by multiple threads  
> > > on 20 or so m1.xlarge machines (when ES goes down or returns errors  
> > > threads back-off with exponential timeout and restart when ES is back  
> > > on-line). There are 8-12 threads per machine doing mostly data  
> > > processing and if I were to trust that in bigdesk “HTTP channels”  
> > > indicate number of active connections, then it means that 30-40  
> > > threads are connected to ES at any given time. Indexing rate is about  
> > > 750 docs per second sometime maxing out at 10,000 docs per second. The  
> > > average doc size is about 5000 bytes

--  
…  
CRAIG BROWN  
chief architect  
youwho, Inc.

[www.youwho.com](http://www.youwho.com)

T: 801.855. 0921  
M: 801.913. 0939

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [May 10, 2012, 9:11pm UTC](https://discuss.elastic.co/t/graceful-way-to-un-overwhelm-the-es/7643/5 "2012-05-10T21:11:36Z")

</div>

Since you are using the s3 gateway, you are safe up to the last checkpoint  
that happened (it hapens periodically but a checkpoint can take time).  
Sorry, I missed the size of heap allocated to ES, I typically recommend  
using ~50% of the machine memory to the ES\_HEAP\_SIZE.

One more thing, if you start another machine to form a cluster, you will  
now have 2 machines in the cluster. If you created the index / indices with  
default number of replicas, then it is set to 1, which means you will have  
2 copies of each shard. Once you start the second node, the replicas will  
be allocate on it, so you will end up with the same capacity problems.

If you don't care about replicas (less HA), then you can dynamically change  
the number of replicas to 0, otherwise, you will need to provision  
machines appropriately.

One last thing, I recommend using local (the default) gateway on AWS, not  
s3, because of the overhad it comes with the the time it can take to do a  
checkpoint. This means each node local drive (or EBS) is used for recovery.

On Thu, May 10, 2012 at 11:59 PM, andym [imwellnow@gmail.com](mailto:imwellnow@gmail.com) wrote:

> Yes, it's single node and ES\_HEAP\_SIZE is as 5120M (not 520M)
> 
> You're are right, I'll have to move to bigger machine or split it into  
> 2 machines as it started happening relatively recently (at ~300M docs)
> 
> My question is whether these restarts are safe at the moment and do  
> not lead to data loss in ES (where ES would return "OK" to processing  
> threads which would mark jobs as completed, but then ES would not  
> persist them due to restart). ES is currently running with  
> threadpool.index.type: cached  
> threadpool.bulk.type: cached
> 
> I tried to make these "blocking" but then processing threads were idle  
> most of the time just waiting for ES to return
> 
> On May 10, 4:48 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> 
> > We are talking about a single ES node, right? For the amount of data that  
> > you indexed, seems like you are hitting memory limits, 520mb for the  
> > amount  
> > of data you have is not enough, probably should go to 3.5 or 4 gb (out of  
> > the 7gb this instance type has) as ES\_HEAP\_SIZE.
> > 
> > On Thu, May 10, 2012 at 11:38 PM, andym [imwell...@gmail.com](mailto:imwell...@gmail.com) wrote:
> > 
> > > Hi,
> > 
> > > I am currently running indexing on c1.xlarge (with 4 ephemeral drives  
> > > in RAID0 and gateway going to S3). Everything works great except every  
> > > several hours ES gets overwhelmed and indexing slows significantly.  
> > > From what I can see from bigdesk during “normal” ES operation “Heap  
> > > Mem” window has jigsaw pattern but when it gets overwhelmed seems like  
> > > no GC happens (no jigsaw pattern in bigdesk) and memory is maxed  
> > > ( configured at 5120M)
> > 
> > > Doing ES service restart (through “bin/service/elasticsearch –  
> > > restart”) solves the problem for a few hours but then the problem re-  
> > > appears.
> > 
> > > I wonder whether restarting ES when it is on such state going to lead  
> > > to any data loss so I can put this into a cron job to assure indexing  
> > > continues (or whether there are any better ways to address the  
> > > problem)
> > 
> > > Thanks,
> > 
> > > -- Andy
> > 
> > > P.S. Some background: I am running ES 19.2 with “refresh interval” set  
> > > to zero and ES currently has about 400 million documents in 2 indexes  
> > > with about 600G total index size (and I expect about 600M docs more  
> > > with around 1T of data. The mapping has \_source set to compressed.)  
> > > The data processing and insertion into ES is done by multiple threads  
> > > on 20 or so m1.xlarge machines (when ES goes down or returns errors  
> > > threads back-off with exponential timeout and restart when ES is back  
> > > on-line). There are 8-12 threads per machine doing mostly data  
> > > processing and if I were to trust that in bigdesk “HTTP channels”  
> > > indicate number of active connections, then it means that 30-40  
> > > threads are connected to ES at any given time. Indexing rate is about  
> > > 750 docs per second sometime maxing out at 10,000 docs per second. The  
> > > average doc size is about 5000 bytes

---

<div class="post-metadata">

**Author:** ![Andrew\_at\_DataFeedFi](https://avatars.discourse-cdn.com/v4/letter/a/0ea827/32.png) [@Andrew\_at\_DataFeedFi](https://discuss.elastic.co/u/Andrew_at_DataFeedFi)\
**Post date:** [May 10, 2012, 10:34pm UTC](https://discuss.elastic.co/t/graceful-way-to-un-overwhelm-the-es/7643/6 "2012-05-10T22:34:59Z")

</div>

Andy,

Are just indexing only when the heap memory gets almost all occupied?  
Do you have any search queries running?

The reason I asked is because I have ran into this issue before where  
we have some searches with large facet field values, because facet does sort  
and loads all the field into cache (field cache) our heap was actually  
mostly filled  
with field caches. And since field caches are not considered garbage, GC  
never  
collected.

The server crawls until we restart it.

Use bigdesk to check your field cache size against your heap size. In case  
you  
are running into same issue.

Regards,

--Andrew

On Thursday, May 10, 2012 3:38:28 PM UTC-5, andym wrote:

> Hi,
> 
> I am currently running indexing on c1.xlarge (with 4 ephemeral drives  
> in RAID0 and gateway going to S3). Everything works great except every  
> several hours ES gets overwhelmed and indexing slows significantly.  
> From what I can see from bigdesk during “normal” ES operation “Heap  
> Mem” window has jigsaw pattern but when it gets overwhelmed seems like  
> no GC happens (no jigsaw pattern in bigdesk) and memory is maxed  
> ( configured at 5120M)
> 
> Doing ES service restart (through “bin/service/elasticsearch –  
> restart”) solves the problem for a few hours but then the problem re-  
> appears.
> 
> I wonder whether restarting ES when it is on such state going to lead  
> to any data loss so I can put this into a cron job to assure indexing  
> continues (or whether there are any better ways to address the  
> problem)
> 
> Thanks,
> 
> -- Andy
> 
> P.S. Some background: I am running ES 19.2 with “refresh interval” set  
> to zero and ES currently has about 400 million documents in 2 indexes  
> with about 600G total index size (and I expect about 600M docs more  
> with around 1T of data. The mapping has \_source set to compressed.)  
> The data processing and insertion into ES is done by multiple threads  
> on 20 or so m1.xlarge machines (when ES goes down or returns errors  
> threads back-off with exponential timeout and restart when ES is back  
> on-line). There are 8-12 threads per machine doing mostly data  
> processing and if I were to trust that in bigdesk “HTTP channels”  
> indicate number of active connections, then it means that 30-40  
> threads are connected to ES at any given time. Indexing rate is about  
> 750 docs per second sometime maxing out at 10,000 docs per second. The  
> average doc size is about 5000 bytes

---

<div class="post-metadata">

**Author:** ![andym](https://avatars.discourse-cdn.com/v4/letter/a/7ab992/32.png) [@andym](https://discuss.elastic.co/u/andym)\
**Post date:** [May 11, 2012, 2:09am UTC](https://discuss.elastic.co/t/graceful-way-to-un-overwhelm-the-es/7643/7 "2012-05-11T02:09:50Z")

</div>

Hi Andrew,

The server is used for indexing only at the moment without any other  
queries going into it and I think it’s hitting the memory limit. The  
only other process that makes queries into it is bigdesk.

I don’t think I paid enough attention memory to pattern in bigdesk  
since when indexing started as it was oscillating between 1-2G. The  
problem seemed to have started with much larger number of docs when  
oscillation would go from 1.7G to 3.9G (that was when I had ES  
originally configured at 4.1G, and since then I bumped it up to 5.2G  
and will move it to a bigger machine or split it as Shay and Craig had  
suggested) and sometimes it would just get stuck. It looked like some  
indexing was still happening at 10 or so docs per second and there  
were no errors returned from ES on new bulk inserts but I haven’t seen  
ES recover by doing GC collection even when completely shutting down  
processing threads (or maybe I just did not wait long enough).

The point you raise about GC not collecting I suspect might have  
something to do with how we set batch sizes – we split the text into  
smaller “documents” and send these “documents” to ES in batches but  
the batch size depends on the length of the original text. So in the  
case when many threads send very large texts at once and ES is close  
it the max memory limit, ES might get into state where the these  
batches are not garbage (from GC perspective) but ES has no other  
memory left to do any additional work.

-- Andy

On May 10, 6:34 pm, "Andrew[.:at:.][DataFeedFile.com](http://DataFeedFile.com)"  
[and...@datafeedfile.com](mailto:and...@datafeedfile.com) wrote:

> Andy,
> 
> Are just indexing only when the heap memory gets almost all occupied?  
> Do you have any search queries running?
> 
> The reason I asked is because I have ran into this issue before where  
> we have some searches with large facet field values, because facet does sort  
> and loads all the field into cache (field cache) our heap was actually  
> mostly filled  
> with field caches. And since field caches are not considered garbage, GC  
> never  
> collected.
> 
> The server crawls until we restart it.
> 
> Use bigdesk to check your field cache size against your heap size. In case  
> you  
> are running into same issue.
> 
> Regards,
> 
> --Andrew
> 
> On Thursday, May 10, 2012 3:38:28 PM UTC-5, andym wrote:
> 
> > Hi,
> 
> > I am currently running indexing on c1.xlarge (with 4 ephemeral drives  
> > in RAID0 and gateway going to S3). Everything works great except every  
> > several hours ES gets overwhelmed and indexing slows significantly.  
> > From what I can see from bigdesk during “normal” ES operation “Heap  
> > Mem” window has jigsaw pattern but when it gets overwhelmed seems like  
> > no GC happens (no jigsaw pattern in bigdesk) and memory is maxed  
> > ( configured at 5120M)
> 
> > Doing ES service restart (through “bin/service/elasticsearch –  
> > restart”) solves the problem for a few hours but then the problem re-  
> > appears.
> 
> > I wonder whether restarting ES when it is on such state going to lead  
> > to any data loss so I can put this into a cron job to assure indexing  
> > continues (or whether there are any better ways to address the  
> > problem)
> 
> > Thanks,
> 
> > -- Andy
> 
> > P.S. Some background: I am running ES 19.2 with “refresh interval” set  
> > to zero and ES currently has about 400 million documents in 2 indexes  
> > with about 600G total index size (and I expect about 600M docs more  
> > with around 1T of data. The mapping has \_source set to compressed.)  
> > The data processing and insertion into ES is done by multiple threads  
> > on 20 or so m1.xlarge machines (when ES goes down or returns errors  
> > threads back-off with exponential timeout and restart when ES is back  
> > on-line). There are 8-12 threads per machine doing mostly data  
> > processing and if I were to trust that in bigdesk “HTTP channels”  
> > indicate number of active connections, then it means that 30-40  
> > threads are connected to ES at any given time. Indexing rate is about  
> > 750 docs per second sometime maxing out at 10,000 docs per second. The  
> > average doc size is about 5000 bytes

---

<div class="post-metadata">

**Author:** ![drewr](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/drewr/32/7803_2.png) [@drewr](https://discuss.elastic.co/u/drewr)\
**Post date:** [May 11, 2012, 2:56pm UTC](https://discuss.elastic.co/t/graceful-way-to-un-overwhelm-the-es/7643/8 "2012-05-11T14:56:28Z")

</div>

andym wrote:

> The point you raise about GC not collecting I suspect might have  
> something to do with how we set batch sizes – we split the text  
> into smaller “documents” and send these “documents” to ES in  
> batches but the batch size depends on the length of the original  
> text. So in the case when many threads send very large texts at  
> once and ES is close it the max memory limit, ES might get into  
> state where the these batches are not garbage (from GC perspective)  
> but ES has no other memory left to do any additional work.

When the JVM runs out of memory, all bets are off. We've seen ES do  
interesting things after indexing it into the ground. Sometimes  
pings start timing out and it is banished from the cluster. Other  
times it OOMEs, leaves, rejoins, and keeps functioning. Most of the  
time, though, it does what you describe -- reaches a state where the  
node stays around and all you see are fruitless attempts at GC in the  
log.

You have to size your bulk requests by bytes instead of doc count so  
you can better predict what kind of batch you're throwing at the  
cluster. If you're really running at the margins, you should make  
sure your shard allocation is spread over the cluster really well, or  
index into separate indices whose shards are spread out. We run our  
ES JVM heaps 70% RAM and monitor for heap and RSS usage (the JVM  
leaks). We also have a sibling plugin to our elasticsearch-jetty  
project that adds a throttling filter for extra safety here. It  
rejects bulk requests when a node has met certain mem, cpu, disk, &  
request thresholds. We are close to open-sourcing it.

Best thing to do is monitor exactly what you're sending to ES  
overlayed with metrics from the data nodes and try to gather very  
specific information about your problem.

-Drew

---

<div class="post-metadata">

**Author:** ![Andrew\_at\_DataFeedFi](https://avatars.discourse-cdn.com/v4/letter/a/0ea827/32.png) [@Andrew\_at\_DataFeedFi](https://discuss.elastic.co/u/Andrew_at_DataFeedFi)\
**Post date:** [May 11, 2012, 4:30pm UTC](https://discuss.elastic.co/t/graceful-way-to-un-overwhelm-the-es/7643/9 "2012-05-11T16:30:29Z")

</div>

Hi Andy,

Yes, sounds like you are on the right track on investigating what is going  
on with  
the heap during indexing.

The batch size for bulk index is critical. I found out the hard way also  
that indexing  
10,000 docs at a time may not be a good idea. At 1000 at a time is much  
better.

BigDesk is great for initial quick look of the heap, the master version for  
0.19+ is much better  
than then 1.0.0. Even better is sematext's ES tool to give you even more  
info.

Good luck

--Andrew

On Thursday, May 10, 2012 9:09:50 PM UTC-5, andym wrote:

> Hi Andrew,
> 
> The server is used for indexing only at the moment without any other  
> queries going into it and I think it’s hitting the memory limit. The  
> only other process that makes queries into it is bigdesk.
> 
> I don’t think I paid enough attention memory to pattern in bigdesk  
> since when indexing started as it was oscillating between 1-2G. The  
> problem seemed to have started with much larger number of docs when  
> oscillation would go from 1.7G to 3.9G (that was when I had ES  
> originally configured at 4.1G, and since then I bumped it up to 5.2G  
> and will move it to a bigger machine or split it as Shay and Craig had  
> suggested) and sometimes it would just get stuck. It looked like some  
> indexing was still happening at 10 or so docs per second and there  
> were no errors returned from ES on new bulk inserts but I haven’t seen  
> ES recover by doing GC collection even when completely shutting down  
> processing threads (or maybe I just did not wait long enough).
> 
> The point you raise about GC not collecting I suspect might have  
> something to do with how we set batch sizes – we split the text into  
> smaller “documents” and send these “documents” to ES in batches but  
> the batch size depends on the length of the original text. So in the  
> case when many threads send very large texts at once and ES is close  
> it the max memory limit, ES might get into state where the these  
> batches are not garbage (from GC perspective) but ES has no other  
> memory left to do any additional work.
> 
> -- Andy
> 
> On May 10, 6:34 pm, "Andrew[.:at:.][DataFeedFile.com](http://DataFeedFile.com)"  
> [and...@datafeedfile.com](mailto:and...@datafeedfile.com) wrote:
> 
> > Andy,
> > 
> > Are just indexing only when the heap memory gets almost all occupied?  
> > Do you have any search queries running?
> > 
> > The reason I asked is because I have ran into this issue before where  
> > we have some searches with large facet field values, because facet does  
> > sort  
> > and loads all the field into cache (field cache) our heap was actually  
> > mostly filled  
> > with field caches. And since field caches are not considered garbage, GC  
> > never  
> > collected.
> > 
> > The server crawls until we restart it.
> > 
> > Use bigdesk to check your field cache size against your heap size. In  
> > case  
> > you  
> > are running into same issue.
> > 
> > Regards,
> > 
> > --Andrew
> > 
> > On Thursday, May 10, 2012 3:38:28 PM UTC-5, andym wrote:
> > 
> > > Hi,
> > 
> > > I am currently running indexing on c1.xlarge (with 4 ephemeral drives  
> > > in RAID0 and gateway going to S3). Everything works great except every  
> > > several hours ES gets overwhelmed and indexing slows significantly.  
> > > From what I can see from bigdesk during “normal” ES operation “Heap  
> > > Mem” window has jigsaw pattern but when it gets overwhelmed seems like  
> > > no GC happens (no jigsaw pattern in bigdesk) and memory is maxed  
> > > ( configured at 5120M)
> > 
> > > Doing ES service restart (through “bin/service/elasticsearch –  
> > > restart”) solves the problem for a few hours but then the problem re-  
> > > appears.
> > 
> > > I wonder whether restarting ES when it is on such state going to lead  
> > > to any data loss so I can put this into a cron job to assure indexing  
> > > continues (or whether there are any better ways to address the  
> > > problem)
> > 
> > > Thanks,
> > 
> > > -- Andy
> > 
> > > P.S. Some background: I am running ES 19.2 with “refresh interval” set  
> > > to zero and ES currently has about 400 million documents in 2 indexes  
> > > with about 600G total index size (and I expect about 600M docs more  
> > > with around 1T of data. The mapping has \_source set to compressed.)  
> > > The data processing and insertion into ES is done by multiple threads  
> > > on 20 or so m1.xlarge machines (when ES goes down or returns errors  
> > > threads back-off with exponential timeout and restart when ES is back  
> > > on-line). There are 8-12 threads per machine doing mostly data  
> > > processing and if I were to trust that in bigdesk “HTTP channels”  
> > > indicate number of active connections, then it means that 30-40  
> > > threads are connected to ES at any given time. Indexing rate is about  
> > > 750 docs per second sometime maxing out at 10,000 docs per second. The  
> > > average doc size is about 5000 bytes

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:29am UTC](https://discuss.elastic.co/t/graceful-way-to-un-overwhelm-the-es/7643/10 "2017-07-06T03:29:13Z")

</div>


