# ES becomes unresponsive!

**URL:** <https://discuss.elastic.co/t/es-becomes-unresponsive/56864>\
**Category:** Elasticsearch\
**Created:** [August 1, 2016, 7:34am UTC](https://discuss.elastic.co/t/es-becomes-unresponsive/56864 "2016-08-01T07:34:51Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![shahzaib](https://avatars.discourse-cdn.com/v4/letter/s/b5ac83/32.png) [@shahzaib](https://discuss.elastic.co/u/shahzaib)\
**Post date:** [August 1, 2016, 7:34am UTC](https://discuss.elastic.co/t/es-becomes-unresponsive/56864/1 "2016-08-01T07:34:51Z")

</div>

Hi,

We're using latest ES version 2.3.4 on a 2 nodes cluster. After each 7-8 hours elasticsearch gets halt. Telneting to ES port stucks in Trying .... & to make it stable again we've to kill java & start elastic-search again otherwise it doesn't restart using 'service elasticsearch restart' & becomes unresonsive. Most likely we've these in logs before halt take place.

[http://pastebin.com/Nx4C6ebJ](http://pastebin.com/Nx4C6ebJ)

We've one master & other data node. Both servers have 64GB memory 30GB of which is allocated to JAVA Heap. Please let me know if you need any more info.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [August 2, 2016, 10:07am UTC](https://discuss.elastic.co/t/es-becomes-unresponsive/56864/2 "2016-08-02T10:07:52Z")

</div>

There's not enough in your logs to help, we'd need to see more.

---

<div class="post-metadata">

**Author:** ![shahzaib](https://avatars.discourse-cdn.com/v4/letter/s/b5ac83/32.png) [@shahzaib](https://discuss.elastic.co/u/shahzaib)\
**Post date:** [August 2, 2016, 1:51pm UTC](https://discuss.elastic.co/t/es-becomes-unresponsive/56864/3 "2016-08-02T13:51:09Z")

</div>

Thanks for responding, well what we've found out is that during halt state of ES, no logs are appended which just seems like the whole ES service becomes unresponsive & unable to write any of the logs.

How are we supposed to troubleshoot if no logs are coming during that issue :(. No swap is used & there's nothing we can find suspicious as well.

Does it look like more of OS issue ?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [August 2, 2016, 9:08pm UTC](https://discuss.elastic.co/t/es-becomes-unresponsive/56864/4 "2016-08-02T21:08:44Z")

</div>

It doesn't look like anything as there is very little info to go on.

What does your config look like? What do the logs, whatever you have, look like?  
What OS?

---

<div class="post-metadata">

**Author:** ![shahzaib](https://avatars.discourse-cdn.com/v4/letter/s/b5ac83/32.png) [@shahzaib](https://discuss.elastic.co/u/shahzaib)\
**Post date:** [August 7, 2016, 7:21am UTC](https://discuss.elastic.co/t/es-becomes-unresponsive/56864/5 "2016-08-07T07:21:55Z")

</div>

Both nodes are now data+master nodes. Here are configs that we added :

bootstrap.mlockall: true  
indices.fielddata.cache.size: 60%  
indices.breaker.fielddata.limit: 70%  
index.max\_result\_window: 50000  
Java heap size is set to 31G out of 64G.

#Following queue sizes are configured over the cluster:

"threadpool.search.queue\_size" : 20000  
"threadpool.index.queue\_size" : 10000

#/etc/security/limits.conf:

- 

```
          soft nofile 700000

```

- 

```
          hard nofile 900000

```

elasticsearch soft memlock unlimited  
elasticsearch hard memlock unlimited

We've auto restart service check triggers after each 10mins to restart if ES goes down, according to server time ES again got restarted around 10:50 & you can see there are no logs just before the 10:50 while the last logged was around 8:00.

> **[Screenshot](https://prnt.sc/c2lowc)**
>
> Captured with Lightshot

I can attach the recent logfile if you want.

We're really being affected by this issue & in need of help ☹

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [August 7, 2016, 7:30am UTC](https://discuss.elastic.co/t/es-becomes-unresponsive/56864/6 "2016-08-07T07:30:00Z")

</div>

All of these;

> [@shahzaib](#):
>
> indices.fielddata.cache.size: 60%  
> indices.breaker.fielddata.limit: 70%  
> index.max\_result\_window: 50000  
> "threadpool.search.queue\_size" : 20000  
> "threadpool.index.queue\_size" : 10000

Are a Very Bad Idea and likely to be putting pressure on things.  
If you think you have needed to make these changes due to these problems, you probably just need more nodes or less data. But again, it's hard to say.

> [@shahzaib](#):
>
> We're really being affected by this issue & in need of help

This is a community based forum, people will offer assistance as best as possible.

---

<div class="post-metadata">

**Author:** ![shahzaib](https://avatars.discourse-cdn.com/v4/letter/s/b5ac83/32.png) [@shahzaib](https://discuss.elastic.co/u/shahzaib)\
**Post date:** [August 7, 2016, 2:52pm UTC](https://discuss.elastic.co/t/es-becomes-unresponsive/56864/7 "2016-08-07T14:52:56Z")

</div>

Thanks for response, well we've now removed these values from ES nodes & enabled gc logging, encountering lots of gc allocation failures here :

> **[Screenshot](https://prnt.sc/c2p7a9)**
>
> Captured with Lightshot

One more question, removing fielddata & breaker values from elasticsearch.yml will revert it to default values or i should make default values as well ?

---

<div class="post-metadata">

**Author:** ![shahzaib](https://avatars.discourse-cdn.com/v4/letter/s/b5ac83/32.png) [@shahzaib](https://discuss.elastic.co/u/shahzaib)\
**Post date:** [August 8, 2016, 8:55am UTC](https://discuss.elastic.co/t/es-becomes-unresponsive/56864/8 "2016-08-08T08:55:17Z")

</div>

We've found another thing, sometimes under /usr/share/elasticsearch directory large size of .hprof is created if ES becomes unresponsive. On googling it, this file is created if JVM gets crash. We've now updated ES to latest 2.3.5 version & current java version is :

openjdk version "1.8.0\_101"  
OpenJDK Runtime Environment (build 1.8.0\_101-b13)  
OpenJDK 64-Bit Server VM (build 25.101-b13, mixed mode)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:29pm UTC](https://discuss.elastic.co/t/es-becomes-unresponsive/56864/9 "2017-07-05T22:29:28Z")

</div>


