# EC2 instance hanging after a few hours

**URL:** <https://discuss.elastic.co/t/ec2-instance-hanging-after-a-few-hours/6167>\
**Category:** Elasticsearch\
**Created:** [December 15, 2011, 4:16pm UTC](https://discuss.elastic.co/t/ec2-instance-hanging-after-a-few-hours/6167 "2011-12-15T16:16:19Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Douglas\_Muth](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/douglas_muth/32/29006_2.png) [@Douglas\_Muth](https://discuss.elastic.co/u/Douglas_Muth)\
**Post date:** [December 15, 2011, 4:16pm UTC](https://discuss.elastic.co/t/ec2-instance-hanging-after-a-few-hours/6167/1 "2011-12-15T16:16:19Z")

</div>

Good Morning,

I'm a long-time listener, first-time caller. 🙂

I'm having some problems trying to get Elastic Search working at  
$WORK. Specifically, that the machine it's running on becomes  
unresponsive after a few hours of indexing. This is on a virgin  
system that doesn't run anything else. Here's my configuration:

- Amazon EC2 "Large" instance

- Ubuntu 10.10, EBS-backed (Kernel ID aki-427d952b, AMI ID  
ami-548c783d)

- Java 1.6.0\_20

- Max files set to 65k, and I verified this by watching the log  
messages when Elastic Search starts up

- Elastic Search is spawned by Daemontools, with the following command  
line: export ES\_JAVA\_OPTS="-server"; exec /usr/local/elasticsearch/bin/  
elasticsearch -f -Des.max-open-files=true

- Everything else is at its default setting

The issue is that Elastic search starts up fine, and runs fine, but if  
I start indexing documents, after some hours, the machine will hang.  
It still be responsive to ping, but any TCP connections such as SSH  
will time out. According to Cloudwatch, the CPU and Network usage  
drops to zero, so nothing on the machine is actually doing aynthing.  
Examining the machine after the fact does not yield anything  
interesting, as as /var/log/messages does not contain anything  
unusual. I'm guessing some weird kernel issue is coming up, but I  
cannot prove this.

I should also mention that we're subjecting Elastic Search to a VERY  
high write load. Specifically, I'm using node.js to read many small  
documents out of a database and add them to elastic search. We have  
about 50 million documents in total, stored into a single index, at  
the rate of 2,000 documents per second. In case it's worth  
mentioning, those documents have a field that is of the MySQL BINARY  
datatype. (since JSON usually isn't about binary data)

FWIW, the MySQL database is an RDS instance, and has given us zero  
problems.

I'm out of ideas, because I've never ran into a problem like this  
before. Does anyone have any suggestions for parameters I can check  
or tweak, or things I can log to get to the bottom of this? I've  
never really seen anything like this before, and any help

Thanks for your time,

-- Doug

---

<div class="post-metadata">

**Author:** ![Weiwei\_Wang](https://avatars.discourse-cdn.com/v4/letter/w/b19c9b/32.png) [@Weiwei\_Wang](https://discuss.elastic.co/u/Weiwei_Wang)\
**Post date:** [December 16, 2011, 2:32am UTC](https://discuss.elastic.co/t/ec2-instance-hanging-after-a-few-hours/6167/2 "2011-12-16T02:32:56Z")

</div>

same problem for me with large index rebuild(2620000+ documents), and  
also too much memory is used(5g+)

On Dec 16, 12:16 am, Douglas Muth [doug.m...@gmail.com](mailto:doug.m...@gmail.com) wrote:

> Good Morning,
> 
> I'm a long-time listener, first-time caller. 🙂
> 
> I'm having some problems trying to get Elastic Search working at  
> $WORK. Specifically, that the machine it's running on becomes  
> unresponsive after a few hours of indexing. This is on a virgin  
> system that doesn't run anything else. Here's my configuration:
> 
> - Amazon EC2 "Large" instance
> 
> - Ubuntu 10.10, EBS-backed (Kernel ID aki-427d952b, AMI ID  
> ami-548c783d)
> 
> - Java 1.6.0\_20
> 
> - Max files set to 65k, and I verified this by watching the log  
> messages when Elastic Search starts up
> 
> - Elastic Search is spawned by Daemontools, with the following command  
> line: export ES\_JAVA\_OPTS="-server"; exec /usr/local/elasticsearch/bin/  
> elasticsearch -f -Des.max-open-files=true
> 
> - Everything else is at its default setting
> 
> The issue is that Elastic search starts up fine, and runs fine, but if  
> I start indexing documents, after some hours, the machine will hang.  
> It still be responsive to ping, but any TCP connections such as SSH  
> will time out. According to Cloudwatch, the CPU and Network usage  
> drops to zero, so nothing on the machine is actually doing aynthing.  
> Examining the machine after the fact does not yield anything  
> interesting, as as /var/log/messages does not contain anything  
> unusual. I'm guessing some weird kernel issue is coming up, but I  
> cannot prove this.
> 
> I should also mention that we're subjecting Elastic Search to a VERY  
> high write load. Specifically, I'm using node.js to read many small  
> documents out of a database and add them to Elasticsearch. We have  
> about 50 million documents in total, stored into a single index, at  
> the rate of 2,000 documents per second. In case it's worth  
> mentioning, those documents have a field that is of the MySQL BINARY  
> datatype. (since JSON usually isn't about binary data)
> 
> FWIW, the MySQL database is an RDS instance, and has given us zero  
> problems.
> 
> I'm out of ideas, because I've never ran into a problem like this  
> before. Does anyone have any suggestions for parameters I can check  
> or tweak, or things I can log to get to the bottom of this? I've  
> never really seen anything like this before, and any help
> 
> Thanks for your time,
> 
> -- Doug

---

<div class="post-metadata">

**Author:** ![Douglas\_Muth](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/douglas_muth/32/29006_2.png) [@Douglas\_Muth](https://discuss.elastic.co/u/Douglas_Muth)\
**Post date:** [December 16, 2011, 2:38am UTC](https://discuss.elastic.co/t/ec2-instance-hanging-after-a-few-hours/6167/3 "2011-12-16T02:38:36Z")

</div>

On Thu, Dec 15, 2011 at 9:32 PM, Weiwei Wang [ww.wang.cs@gmail.com](mailto:ww.wang.cs@gmail.com) wrote:

> same problem for me with large index rebuild(2620000+ documents), and  
> also too much memory is used(5g+)

I can confirm in my case that memory is NOT an issue. The instance in  
question has ~7 GB of RAM,and according to my Munin graphs, total  
memory usage on that box doesn't go above 1.5 GB.

-- Doug

--

> **[Doug's Home On The Web](https://www.dmuth.org/)**
>
> Writings from a Philadelphia Software Engineer...

[http://twitter.com/dmuth](http://twitter.com/dmuth)

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [December 16, 2011, 5:41pm UTC](https://discuss.elastic.co/t/ec2-instance-hanging-after-a-few-hours/6167/4 "2011-12-16T17:41:40Z")

</div>

Few points:

1. I suggest using a newer Java version, 1.6\_20 is quite old.
2. The fact that the machine has 7gb does not mean elasticsearch will use  
it. I suggest to allocate to ES half the machine memory, in your case,  
~3.5gb. See more here:  
[Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/setup/installation.html).
3. If you can, use a larger instance (m1.xlarge), it "suffers" from  
noisy neighbors on aws less.

The fact that you can't connect to the machine might mean you are running  
out of sockets and they are being throttled by the OS. Can you monitor it?  
Are you using persistent connections to ES from node? netstat and lsof are  
your friends here to check it.

I also saw this behavior way back with the AWS problems with ubuntu 10.04,  
recently, I started to read that 10.10 has started to exhibit  
similar behavior (see the instagram blog:

> **[What Powers Instagram: Hundreds of Instances, Dozens of Technologies](https://instagram-engineering.tumblr.com/post/13649370142/what-powers-instagram-hundreds-of-instances)**
>
> One of the questions we always get asked at meet-ups and conversations with other engineers is, “what’s your stack?” We thought it would be fun to give a sense of all the systems that power Instagram,...

).

Things to monitor on ES is mainly the memory usage for this, node stats API  
is your friend here (or big desk plugin:  
[Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/modules/plugins.html)).

On Fri, Dec 16, 2011 at 4:38 AM, Douglas Muth [doug.muth@gmail.com](mailto:doug.muth@gmail.com) wrote:

> On Thu, Dec 15, 2011 at 9:32 PM, Weiwei Wang [ww.wang.cs@gmail.com](mailto:ww.wang.cs@gmail.com) wrote:
> 
> > same problem for me with large index rebuild(2620000+ documents), and  
> > also too much memory is used(5g+)
> 
> I can confirm in my case that memory is NOT an issue. The instance in  
> question has ~7 GB of RAM,and according to my Munin graphs, total  
> memory usage on that box doesn't go above 1.5 GB.
> 
> -- Doug
> 
> --  
> [http://www.dmuth.org/](http://www.dmuth.org/)  
> [http://twitter.com/dmuth](http://twitter.com/dmuth)

---

<div class="post-metadata">

**Author:** ![Douglas\_Muth](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/douglas_muth/32/29006_2.png) [@Douglas\_Muth](https://discuss.elastic.co/u/Douglas_Muth)\
**Post date:** [December 16, 2011, 6:08pm UTC](https://discuss.elastic.co/t/ec2-instance-hanging-after-a-few-hours/6167/5 "2011-12-16T18:08:07Z")

</div>

On Fri, Dec 16, 2011 at 12:41 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:

> Few points:
> 
> 1. I suggest using a newer Java version, 1.6\_20 is quite old.
> 2. The fact that the machine has 7gb does not mean elasticsearch will use  
> it. I suggest to allocate to ES half the machine memory, in your case,  
> ~3.5gb. See more  
> here: [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/setup/installation.html).
> 3. If you can, use a larger instance (m1.xlarge), it "suffers" from  
> noisy neighbors on aws less.

#1 and #2 are easy enough for me to do. #3 might be more of a  
challenge, since it will cost us more. 😛

> The fact that you can't connect to the machine might mean you are running  
> out of sockets and they are being throttled by the OS. Can you monitor it?  
> Are you using persistent connections to ES from node? netstat and lsof are  
> your friends here to check it.

The connections are not persistent, but I use the generic-pool module  
for node.js, which I use to limit to 10 slots, or 10 concurrent  
connections to Elastic Search. I did check things with netstat and  
lsof, and there are no issues there.

> I also saw this behavior way back with the AWS problems with ubuntu 10.04,  
> recently, I started to read that 10.10 has started to exhibit  
> similar behavior (see the instagram  
> blog: [What Powers Instagram: Hundreds of Instances, Dozens of Technologies - Instagram Engineering](http://instagram-engineering.tumblr.com/post/13649370142/what-powers-instagram-hundreds-of-instances-dozens-of)).

Interesting, as that's the first I heard of any concerns with 10.10.  
We're running our entire infrastructure on 10.10 and have an average  
of 1 freeze like this per machine per 6 months. (current issues  
excepted)

> Things to monitor on ES is mainly the memory usage for this, node stats API  
> is your friend here (or big desk  
> plugin: [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/modules/plugins.html)).

Big Desk looks very very cool. Many thanks for that, and the other pointers!

All the best,

-- Doug

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [December 16, 2011, 6:27pm UTC](https://discuss.elastic.co/t/ec2-instance-hanging-after-a-few-hours/6167/6 "2011-12-16T18:27:03Z")

</div>

Can you try and move to use persistent connections, see if it helps?

On Fri, Dec 16, 2011 at 8:08 PM, Douglas Muth [doug.muth@gmail.com](mailto:doug.muth@gmail.com) wrote:

> On Fri, Dec 16, 2011 at 12:41 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> 
> > Few points:
> > 
> > 1. I suggest using a newer Java version, 1.6\_20 is quite old.
> > 2. The fact that the machine has 7gb does not mean elasticsearch will use  
> > it. I suggest to allocate to ES half the machine memory, in your case,  
> > ~3.5gb. See more  
> > here:  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/setup/installation.html).
> > 3. If you can, use a larger instance (m1.xlarge), it "suffers" from  
> > noisy neighbors on aws less.
> 
> #1 and #2 are easy enough for me to do. #3 might be more of a  
> challenge, since it will cost us more. 😛
> 
> > The fact that you can't connect to the machine might mean you are running  
> > out of sockets and they are being throttled by the OS. Can you monitor  
> > it?  
> > Are you using persistent connections to ES from node? netstat and lsof  
> > are  
> > your friends here to check it.
> 
> The connections are not persistent, but I use the generic-pool module  
> for node.js, which I use to limit to 10 slots, or 10 concurrent  
> connections to Elastic Search. I did check things with netstat and  
> lsof, and there are no issues there.
> 
> > I also saw this behavior way back with the AWS problems with ubuntu  
> > 10.04,  
> > recently, I started to read that 10.10 has started to exhibit  
> > similar behavior (see the instagram  
> > blog:  
> > [What Powers Instagram: Hundreds of Instances, Dozens of Technologies - Instagram Engineering](http://instagram-engineering.tumblr.com/post/13649370142/what-powers-instagram-hundreds-of-instances-dozens-of)  
> > ).
> 
> Interesting, as that's the first I heard of any concerns with 10.10.  
> We're running our entire infrastructure on 10.10 and have an average  
> of 1 freeze like this per machine per 6 months. (current issues  
> excepted)
> 
> > Things to monitor on ES is mainly the memory usage for this, node stats  
> > API  
> > is your friend here (or big desk  
> > plugin:  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/modules/plugins.html)).
> 
> Big Desk looks very very cool. Many thanks for that, and the other  
> pointers!
> 
> All the best,
> 
> -- Doug

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:45am UTC](https://discuss.elastic.co/t/ec2-instance-hanging-after-a-few-hours/6167/7 "2017-07-06T03:45:08Z")

</div>


