# Cluster health and the amount of indexes and data

**URL:** <https://discuss.elastic.co/t/cluster-health-and-the-amount-of-indexes-and-data/40034>\
**Category:** Elasticsearch\
**Created:** [January 25, 2016, 2:52pm UTC](https://discuss.elastic.co/t/cluster-health-and-the-amount-of-indexes-and-data/40034 "2016-01-25T14:52:13Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Don\_Pich](https://avatars.discourse-cdn.com/v4/letter/d/b487fb/32.png) [@Don\_Pich](https://discuss.elastic.co/u/Don_Pich)\
**Post date:** [January 25, 2016, 2:52pm UTC](https://discuss.elastic.co/t/cluster-health-and-the-amount-of-indexes-and-data/40034/1 "2016-01-25T14:52:13Z")

</div>

So I have 4 nodes in my ELK stack. I have a logstash server that parses syslog messages into my cluster. I have 3 ES nodes. One is strictly a master node. The other two are master/data eligible.

So it's all working fine UNTIL I reach a 'magic point' in my data. At some point, the cluster goes red and I am not able to do anything to fix it. This seems to happen after the data nodes end up with a large amount of indexes. What I end up doing after snapshots is flushing all indexes that are 30 days and older, and the cluster goes green and the whole ELK stack starts functioning normally again.

I have the default shards (5) and 1 replica.

My question is this. I need to keep 90 days of active data plus 1 year of logs (Think PCI DSS). I have curator taking a snapshot of the data. My problem is that I can't get to 90 days with the raw amount of data.

Is this a sign of needing more shards? Do I need more data nodes?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [January 26, 2016, 1:50am UTC](https://discuss.elastic.co/t/cluster-health-and-the-amount-of-indexes-and-data/40034/2 "2016-01-26T01:50:10Z")

</div>

How much data per day? What are your node specs?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 26, 2016, 6:32am UTC](https://discuss.elastic.co/t/cluster-health-and-the-amount-of-indexes-and-data/40034/3 "2016-01-26T06:32:49Z")

</div>

How many indices and shards do you have in the cluster? What is your total amount of data?

---

<div class="post-metadata">

**Author:** ![Don\_Pich](https://avatars.discourse-cdn.com/v4/letter/d/b487fb/32.png) [@Don\_Pich](https://discuss.elastic.co/u/Don_Pich)\
**Post date:** [January 26, 2016, 3:35pm UTC](https://discuss.elastic.co/t/cluster-health-and-the-amount-of-indexes-and-data/40034/4 "2016-01-26T15:35:34Z")

</div>

There is approximately 1 Gig of log data on a daily basis with that looking like it's going to expand.

All servers are setup as follows in VMWare  
2 Sockets, 4 cores (8 CPUs) 16 Gigs of Ram/Heap 8 Gigs  
500 Gig of disk space

---

<div class="post-metadata">

**Author:** ![Don\_Pich](https://avatars.discourse-cdn.com/v4/letter/d/b487fb/32.png) [@Don\_Pich](https://discuss.elastic.co/u/Don_Pich)\
**Post date:** [January 26, 2016, 3:36pm UTC](https://discuss.elastic.co/t/cluster-health-and-the-amount-of-indexes-and-data/40034/5 "2016-01-26T15:36:37Z")

</div>

At this time, there are 5 indexes with 5 shards per index. We are capturing roughly 1 gig to 1.5 gig of daily information. Disk space on the partition that stores the indexes was roughly 85% when I experienced the issue.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [January 27, 2016, 11:11am UTC](https://discuss.elastic.co/t/cluster-health-and-the-amount-of-indexes-and-data/40034/6 "2016-01-27T11:11:24Z")

</div>

You should reduce your shard count to 1 primary with 1 replica, otherwise it's a waste.

---

<div class="post-metadata">

**Author:** ![Don\_Pich](https://avatars.discourse-cdn.com/v4/letter/d/b487fb/32.png) [@Don\_Pich](https://discuss.elastic.co/u/Don_Pich)\
**Post date:** [February 1, 2016, 2:48pm UTC](https://discuss.elastic.co/t/cluster-health-and-the-amount-of-indexes-and-data/40034/7 "2016-02-01T14:48:15Z")

</div>

So forgive the question, but how would that help the issue I'm having?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [February 1, 2016, 3:20pm UTC](https://discuss.elastic.co/t/cluster-health-and-the-amount-of-indexes-and-data/40034/8 "2016-02-01T15:20:28Z")

</div>

If I calculate correctly you have 5 indices \* 5 shards \* 2 copies \* 90 days, which gives a total of 4500 shards. That is 2250 shards per node with only 8GB of heap. I would guess this is most likely the cause of your problems.

Each shard is a Lucene index and carries with it overhead in terms of memory usage and file handles. By following Marks advice and reducing the number of shards per index, and possibly also the number of indices, e.g. by using weekly or monthly indices, you can reduce the overhead and get your nodes to handle more data. Having a large number of very small shards wastes system resources, so I would recommend aiming to have shard sizes in the range of a few hundred MB to a few GB in size. Read [this blog post](https://www.elastic.co/blog/found-crash-elasticsearch#too-many-shards-or-the-gazillion-shards-problem) for an example of how it is possible to crash an Elasticsearch cluster with too many shards even when no data has been indexed.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:19pm UTC](https://discuss.elastic.co/t/cluster-health-and-the-amount-of-indexes-and-data/40034/9 "2017-07-05T23:19:53Z")

</div>


