# Cluster deteriorates after a couple of days

**URL:** <https://discuss.elastic.co/t/cluster-deteriorates-after-a-couple-of-days/1285>\
**Category:** Elasticsearch\
**Created:** [May 26, 2015, 7:49am UTC](https://discuss.elastic.co/t/cluster-deteriorates-after-a-couple-of-days/1285 "2015-05-26T07:49:09Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![helium](https://avatars.discourse-cdn.com/v4/letter/h/2acd7d/32.png) [@helium](https://discuss.elastic.co/u/helium)\
**Post date:** [May 26, 2015, 7:49am UTC](https://discuss.elastic.co/t/cluster-deteriorates-after-a-couple-of-days/1285/1 "2015-05-26T07:49:09Z")

</div>

Hi,

My ES cluster works great, for a couple of days, and then becomes virtually unusable. It refuses to accept data from logstash and queries becomes extremely slow. Optimizing the indices doesn't seem to help but pruning old log entries makes it work again (keeping 3-4 days). The document count is usually around 50-60 Mil when it stops working. Any idea of what's going on?

I have a classic ELK setup with Elasticsearch-1.5.2, Logstash-1.5 and Kibana-4. 4 nodes: 3 ES data nodes, 1 ES no data/logstash/kibana.

ES config:`
node.name: ccdlog04
index.number_of_replicas: 2
discovery.zen.ping.unicast.hosts: ["ccdlog01","ccdlog02","ccdlog03","ccdlog04"]
node.master: true
node.data: false
http.cors.enabled: true`

Excerpts from LS config (in case it's the timestamps):  
`
grok {
match => [
"message", "%{YEAR:time}%{MONTHNUM:time}%{MONTHDAY:time} %{TIME:time}...GREEDYDATA:logMessage}"
...
mutate {
add_field => ["app_timestamp", "%{[time][0]}-%{[time][1]}-%{[time][2]}T%{[time][3]}Z" ]
}
date {
match => ["app_timestamp","ISO8601"]
}
`

LS log example of what happens after a couple of days:  
`
{:timestamp=>"2015-05-26T05:27:27.381000+0000", :message=>"retrying failed action with response code: 503", :level=>:warn}
{:timestamp=>"2015-05-26T05:27:27.381000+0000", :message=>"retrying failed action with response code: 503", :level=>:warn}
{:timestamp=>"2015-05-26T05:27:27.381000+0000", :message=>"retrying failed action with response code: 503", :level=>:warn}
{:timestamp=>"2015-05-26T05:27:27.381000+0000", :message=>"retrying failed action with response code: 503", :level=>:warn}
{:timestamp=>"2015-05-26T05:27:27.382000+0000", :message=>"retrying failed action with response code: 503", :level=>:warn}
{:timestamp=>"2015-05-26T05:27:27.382000+0000", :message=>"retrying failed action with response code: 503", :level=>:warn}
{:timestamp=>"2015-05-26T05:27:27.382000+0000", :message=>"retrying failed action with response code: 503", :level=>:warn}
{:timestamp=>"2015-05-26T05:27:27.382000+0000", :message=>"retrying failed action with response code: 503", :level=>:warn}
{:timestamp=>"2015-05-26T05:27:27.383000+0000", :message=>"retrying failed action with response code: 503", :level=>:warn}
{:timestamp=>"2015-05-26T05:27:27.383000+0000", :message=>"retrying failed action with response code: 503", :level=>:warn}
{:timestamp=>"2015-05-26T05:27:27.383000+0000", :message=>"retrying failed action with response code: 503", :level=>:warn}
{:timestamp=>"2015-05-26T05:27:27.383000+0000", :message=>"retrying failed action with response code: 503", :level=>:warn}
`

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [May 26, 2015, 8:20am UTC](https://discuss.elastic.co/t/cluster-deteriorates-after-a-couple-of-days/1285/2 "2015-05-26T08:20:10Z")

</div>

What's in the Elasticsearch logs? I suspect that you're running out of heap. How big is your heap?

---

<div class="post-metadata">

**Author:** ![simonrisberg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/simonrisberg/32/3513_2.png) [@simonrisberg](https://discuss.elastic.co/u/simonrisberg)\
**Post date:** [July 3, 2015, 7:26am UTC](https://discuss.elastic.co/t/cluster-deteriorates-after-a-couple-of-days/1285/3 "2015-07-03T07:26:03Z")

</div>

I am getting the same problem right now. It doesn't index any logs and I keep getting the same response code as he does.

---

<div class="post-metadata">

**Author:** ![helium](https://avatars.discourse-cdn.com/v4/letter/h/2acd7d/32.png) [@helium](https://discuss.elastic.co/u/helium)\
**Post date:** [July 3, 2015, 7:53am UTC](https://discuss.elastic.co/t/cluster-deteriorates-after-a-couple-of-days/1285/4 "2015-07-03T07:53:03Z")

</div>

I actually found the answer in [https://www.elastic.co/guide/en/elasticsearch/guide/current/\_limiting\_memory\_usage.html](https://www.elastic.co/guide/en/elasticsearch/guide/current/_limiting_memory_usage.html). ES will keep all the field data in memory until it chokes unless you control it. Configuring indices.fielddata.cache.size solved it for me.

---

<div class="post-metadata">

**Author:** ![simonrisberg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/simonrisberg/32/3513_2.png) [@simonrisberg](https://discuss.elastic.co/u/simonrisberg)\
**Post date:** [July 3, 2015, 8:55am UTC](https://discuss.elastic.co/t/cluster-deteriorates-after-a-couple-of-days/1285/5 "2015-07-03T08:55:09Z")

</div>

Where do I find this and how do I control it?

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [July 3, 2015, 10:51am UTC](https://discuss.elastic.co/t/cluster-deteriorates-after-a-couple-of-days/1285/6 "2015-07-03T10:51:20Z")

</div>

indices.fielddata.cache.size is a configuration parameter that you can set in elasticsearch.yml.

---

<div class="post-metadata">

**Author:** ![simonrisberg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/simonrisberg/32/3513_2.png) [@simonrisberg](https://discuss.elastic.co/u/simonrisberg)\
**Post date:** [July 3, 2015, 10:58am UTC](https://discuss.elastic.co/t/cluster-deteriorates-after-a-couple-of-days/1285/7 "2015-07-03T10:58:48Z")

</div>

I have done this. It still isn't really working. I put it to 40%. Yesterday I managed to index a few old logs and I'm guessing it became to much. The last log that was indexed into elasticsearch was indexed yesterday at 17:06. After that it just stopped indexing.

---

<div class="post-metadata">

**Author:** ![steverobbins](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/steverobbins/32/5961_2.png) [@steverobbins](https://discuss.elastic.co/u/steverobbins)\
**Post date:** [November 16, 2015, 9:13pm UTC](https://discuss.elastic.co/t/cluster-deteriorates-after-a-couple-of-days/1285/8 "2015-11-16T21:13:27Z")

</div>

Is there a way to offload some of the memory burden by storing on disk?

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [November 16, 2015, 9:38pm UTC](https://discuss.elastic.co/t/cluster-deteriorates-after-a-couple-of-days/1285/9 "2015-11-16T21:38:08Z")

</div>

@steverobbins, unless this is directly related to this old thread perhaps you can start a thread of your own? There's no silver bullet for reducing the memory pressure.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:38pm UTC](https://discuss.elastic.co/t/cluster-deteriorates-after-a-couple-of-days/1285/10 "2017-07-05T23:38:08Z")

</div>


