# ElasticSearch Node goes down

**URL:** <https://discuss.elastic.co/t/elasticsearch-node-goes-down/188152>\
**Category:** Elasticsearch\
**Created:** [June 29, 2019, 6:46pm UTC](https://discuss.elastic.co/t/elasticsearch-node-goes-down/188152 "2019-06-29T18:46:41Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Sourabh](https://avatars.discourse-cdn.com/v4/letter/s/9dc877/32.png) [@Sourabh](https://discuss.elastic.co/u/Sourabh)\
**Post date:** [June 29, 2019, 6:46pm UTC](https://discuss.elastic.co/t/elasticsearch-node-goes-down/188152/1 "2019-06-29T18:46:42Z")

</div>

I am having a cluster with 5 master nodes,12 coordinator nodes and 60 data nodes .Currently i am doing heavy indexing in this es cluster around 15 billion documents spread through the day.We have 3 index which are undergoing heavy indexing there are four rollover in a day for each indexes.Each indexes having 100 shards and the replica is set to 1.The nodes are up on a physical server having 200gb of RAM ,each nodes have around 32gb of heap and the translog durability is set to async.

The bulk indexing via bulk processor is happening very slowly with 32 clients and the batch size of the bulk processor is 7500 and bulk action is 20 and the bulk size is 25mb. All the interfaces have 10gb bandwidth.

- But the bulk indexing is happening very slowly.

- During indexing and searching these error are coming

- And in the logs i am getting failed to execute query phase(No search context found for id),gc errors , nodes being removed and added.

Please suggest .This is a very trivial issue we are facing.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 29, 2019, 6:54pm UTC](https://discuss.elastic.co/t/elasticsearch-node-goes-down/188152/2 "2019-06-29T18:54:07Z")

</div>

Why are you indexing into so many shards? Is that 100 primary or 100 primary and replica shards? What is the average shard size?

---

<div class="post-metadata">

**Author:** ![Sourabh](https://avatars.discourse-cdn.com/v4/letter/s/9dc877/32.png) [@Sourabh](https://discuss.elastic.co/u/Sourabh)\
**Post date:** [June 29, 2019, 6:55pm UTC](https://discuss.elastic.co/t/elasticsearch-node-goes-down/188152/3 "2019-06-29T18:55:29Z")

</div>

I have 100 primary shards and the replica is set to 1.The average shard size is around 30 to 40gb.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 29, 2019, 6:59pm UTC](https://discuss.elastic.co/t/elasticsearch-node-goes-down/188152/4 "2019-06-29T18:59:09Z")

</div>

If you have 100 primary shards per index and the average shard size is 30GB you are generating 72TB (30GB \* 100 primary shards \* 2 (1 replica) \* 3 indices \* 4 rollovers) of data on disk per day. That is 2400 shards generated per day. To me this sounds a bit strange. Are you sure those numbers are accurate? This does not sound slow to me...

What is your retention period?

---

<div class="post-metadata">

**Author:** ![Sourabh](https://avatars.discourse-cdn.com/v4/letter/s/9dc877/32.png) [@Sourabh](https://discuss.elastic.co/u/Sourabh)\
**Post date:** [June 29, 2019, 7:05pm UTC](https://discuss.elastic.co/t/elasticsearch-node-goes-down/188152/5 "2019-06-29T19:05:24Z")

</div>

Total number of shards is close to 8000 in the cluster at any given instant of time.Retention is d-2 days.Each document size in the index is approximately 1.8kb. We have opted for 100 shards so that the indexing is comparatively faster.The cluster is having index heavy load.Please guide where exactly the issue might be.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 29, 2019, 7:08pm UTC](https://discuss.elastic.co/t/elasticsearch-node-goes-down/188152/6 "2019-06-29T19:08:02Z")

</div>

Do you have any non-default settings? Can you confirm that the numbers given are accurate?

If the information is accurate I owuld recommend the following:

- Create the indices with 60 primary shards and 1 replica. Set rollover to cut over at an average shard size of over 50GB. This will reduce the number of shards you are indexing into as well as the number of shards in the cluster.
- As your nodes have prenty of RAM and you might be having heap pressure (check if this is the case) place 2 Elasticsearch nodes per host. This means each node will hold one shard per index on average.
- If you can, make sure each bulk request only indexes into one index. This will result in more documents being indexed per shard per request. As Elasticsearch syncs the transaction log per request, this should improve efficiency.
- If you have not already, install monitoring so you can see how heap usage looks like.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 27, 2019, 7:10pm UTC](https://discuss.elastic.co/t/elasticsearch-node-goes-down/188152/8 "2019-07-27T19:10:47Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
