# Dealing with indexing spikes

**URL:** <https://discuss.elastic.co/t/dealing-with-indexing-spikes/45548>\
**Category:** Elasticsearch\
**Created:** [March 27, 2016, 9:11pm UTC](https://discuss.elastic.co/t/dealing-with-indexing-spikes/45548 "2016-03-27T21:11:55Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![jonatanzafar59](https://avatars.discourse-cdn.com/v4/letter/j/848f3c/32.png) [@jonatanzafar59](https://discuss.elastic.co/u/jonatanzafar59)\
**Post date:** [March 27, 2016, 9:11pm UTC](https://discuss.elastic.co/t/dealing-with-indexing-spikes/45548/1 "2016-03-27T21:11:55Z")

</div>

Hi,

We have an application that 3 times a day indexes 100K-200K documents per node in 2 minutes, which causes huge spikes in GC time and pressure on the cluster.

I tried 1MB bulks (will try to raise it), resizing the bulk threadpool, and changing refresh interval to 10s.

Does somebody have another suggestion?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [March 27, 2016, 10:38pm UTC](https://discuss.elastic.co/t/dealing-with-indexing-spikes/45548/2 "2016-03-27T22:38:54Z")

</div>

Ultimately you need to pay the "cost" of indexing, so other than spreading the load over a longer time, you could add more nodes during the indexing process.

---

<div class="post-metadata">

**Author:** ![lemon\_zmd](https://avatars.discourse-cdn.com/v4/letter/l/46a35a/32.png) [@lemon\_zmd](https://discuss.elastic.co/u/lemon_zmd)\
**Post date:** [March 29, 2016, 6:12am UTC](https://discuss.elastic.co/t/dealing-with-indexing-spikes/45548/3 "2016-03-29T06:12:33Z")

</div>

Better the ingestion can happen smoothly. Say , to make 2 mins to be 20 mins or even consume them from Kafka. But anyway , a profile could help to find the bottle neck. When heavy ingestion happen, cpu memory io can all consumed a lot. Simply you can do vmstat to see the possible bottle neck.  
Besides, I think G1 could also be tried. In my test with all default settings when ingest 2m docs to ES-2.2,  
G1 totally stop for 3s 318ms (362 times) while cms stop for 8s 24ms (1190 times). So if a lot of full gc happens even concurrent mode failure. G1 could be an option when jdk8 is being used.

---

<div class="post-metadata">

**Author:** ![jonatanzafar59](https://avatars.discourse-cdn.com/v4/letter/j/848f3c/32.png) [@jonatanzafar59](https://discuss.elastic.co/u/jonatanzafar59)\
**Post date:** [March 31, 2016, 10:09am UTC](https://discuss.elastic.co/t/dealing-with-indexing-spikes/45548/4 "2016-03-31T10:09:54Z")

</div>

Hi,

problem solved. For other users that encounter this problem:

- Obviously, the division of the bulks for a longer timeframe helped.  
Important note -  
We had 4 nodes for the cluster, with 4 CPU cores each. every node holds 10 shards (including replicas).  
We added 2 CPU cores for each node, which mitigated the GCs significantly. Important to say, that we didn't see the CPU working so hard, so this came as a surprise.  
My guess is that the indexing is per shard, so the CPU had to switch a lot, which delayed the concurrent searches. Alternatively, I guess we could have reindexed the cluster to a significantly smaller amount of shards, and it would have helped too, since the CPU wasn't even close to reaching its limit.

I hope someone can correct/approve my assumption.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:03pm UTC](https://discuss.elastic.co/t/dealing-with-indexing-spikes/45548/5 "2017-07-05T23:03:35Z")

</div>


