# Problems with tcp connections

**URL:** <https://discuss.elastic.co/t/problems-with-tcp-connections/5885>\
**Category:** Elasticsearch\
**Created:** [November 16, 2011, 6:58am UTC](https://discuss.elastic.co/t/problems-with-tcp-connections/5885 "2011-11-16T06:58:15Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![darkyoung](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/darkyoung/32/2924_2.png) [@darkyoung](https://discuss.elastic.co/u/darkyoung)\
**Post date:** [November 16, 2011, 6:58am UTC](https://discuss.elastic.co/t/problems-with-tcp-connections/5885/1 "2011-11-16T06:58:15Z")

</div>

Hi

I have two servers running an elasticsearch cluster as our website's search  
engine. And use Elastica as our php client.

At beginning, the queries are sent directly to ES but the servers are very  
unstable and the tcp connections are about 500-600, so ES can't handle them  
quickly and always get timeout response (We set the timeout to 5s). So we  
added 5mins cache with memcached and the situation got better. The tcp  
connections are controlled around 10 (avg).

I found that if the connections over 100 then it will become very unstable.

Does this because the server can't handle too much request? Or I need to  
optimize my queries? (Most queries took about 50 ms)

Here is a gist of the node stats at some  
point. [https://gist.github.com/1369446](https://gist.github.com/1369446)

---

<div class="post-metadata">

**Author:** ![electic](https://avatars.discourse-cdn.com/v4/letter/e/ecccb3/32.png) [@electic](https://discuss.elastic.co/u/electic)\
**Post date:** [November 16, 2011, 7:30am UTC](https://discuss.elastic.co/t/problems-with-tcp-connections/5885/2 "2011-11-16T07:30:28Z")

</div>

Did you update your limits.conf? The number of acceptable connections  
might be maxed out and hence why you are getting the timeouts.

On Nov 15, 10:58 pm, Ocean Wu [darkyo...@gmail.com](mailto:darkyo...@gmail.com) wrote:

> Hi
> 
> I have two servers running an elasticsearch cluster as our website's search  
> engine. And use Elastica as our php client.
> 
> At beginning, the queries are sent directly to ES but the servers are very  
> unstable and the tcp connections are about 500-600, so ES can't handle them  
> quickly and always get timeout response (We set the timeout to 5s). So we  
> added 5mins cache with memcached and the situation got better. The tcp  
> connections are controlled around 10 (avg).
> 
> I found that if the connections over 100 then it will become very unstable.
> 
> Does this because the server can't handle too much request? Or I need to  
> optimize my queries? (Most queries took about 50 ms)
> 
> Here is a gist of the node stats at some  
> point.[https://gist.github.com/1369446](https://gist.github.com/1369446)

---

<div class="post-metadata">

**Author:** ![darkyoung](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/darkyoung/32/2924_2.png) [@darkyoung](https://discuss.elastic.co/u/darkyoung)\
**Post date:** [November 16, 2011, 8:08am UTC](https://discuss.elastic.co/t/problems-with-tcp-connections/5885/3 "2011-11-16T08:08:18Z")

</div>

Yes, the limits.conf set to 32000, and net.nf\_conntrack\_max set to 655360.

Thanks for reply.

---

<div class="post-metadata">

**Author:** ![electic](https://avatars.discourse-cdn.com/v4/letter/e/ecccb3/32.png) [@electic](https://discuss.elastic.co/u/electic)\
**Post date:** [November 16, 2011, 9:04am UTC](https://discuss.elastic.co/t/problems-with-tcp-connections/5885/4 "2011-11-16T09:04:04Z")

</div>

Then we are having the same issue:

[https://groups.google.com/group/elasticsearch/browse\_thread/thread/1861b5c253982c75](https://groups.google.com/group/elasticsearch/browse_thread/thread/1861b5c253982c75)

I notice when my total index size exceeds the RAM size (16GB of ram  
per machine) the queries start to take a bit longer. Once the  
connections pile up the entire cluster becomes massively unstable and  
crashes. I have a theory as the dataset goes up in size what was once  
a fast query suddenly is slow (my queries fetch data from a certain  
time and sort) and I think that might be killing the cluster.

-R

On Nov 16, 12:08 am, Ocean Wu [darkyo...@gmail.com](mailto:darkyo...@gmail.com) wrote:

> Yes, the limits.conf set to 32000, and net.nf\_conntrack\_max set to 655360.
> 
> Thanks for reply.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [November 16, 2011, 2:29pm UTC](https://discuss.elastic.co/t/problems-with-tcp-connections/5885/5 "2011-11-16T14:29:01Z")

</div>

Can you try and use 0.18.3, see if it helps? It might be related to the  
connection problem while searching fix.

On Wed, Nov 16, 2011 at 11:04 AM, electic [electic@gmail.com](mailto:electic@gmail.com) wrote:

> Then we are having the same issue:
> 
> [https://groups.google.com/group/elasticsearch/browse\_thread/thread/1861b5c253982c75](https://groups.google.com/group/elasticsearch/browse_thread/thread/1861b5c253982c75)
> 
> I notice when my total index size exceeds the RAM size (16GB of ram  
> per machine) the queries start to take a bit longer. Once the  
> connections pile up the entire cluster becomes massively unstable and  
> crashes. I have a theory as the dataset goes up in size what was once  
> a fast query suddenly is slow (my queries fetch data from a certain  
> time and sort) and I think that might be killing the cluster.
> 
> -R
> 
> On Nov 16, 12:08 am, Ocean Wu [darkyo...@gmail.com](mailto:darkyo...@gmail.com) wrote:
> 
> > Yes, the limits.conf set to 32000, and net.nf\_conntrack\_max set to
> 
> 1. 
> 
> > Thanks for reply.

---

<div class="post-metadata">

**Author:** ![electic](https://avatars.discourse-cdn.com/v4/letter/e/ecccb3/32.png) [@electic](https://discuss.elastic.co/u/electic)\
**Post date:** [November 16, 2011, 7:06pm UTC](https://discuss.elastic.co/t/problems-with-tcp-connections/5885/6 "2011-11-16T19:06:29Z")

</div>

Sweet. Okay, I am running a test now. Will report on any changes.

On Nov 16, 6:29 am, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:

> Can you try and use 0.18.3, see if it helps? It might be related to the  
> connection problem while searching fix.
> 
> On Wed, Nov 16, 2011 at 11:04 AM, electic [elec...@gmail.com](mailto:elec...@gmail.com) wrote:
> 
> > Then we are having the same issue:
> 
> > [https://groups.google.com/group/elasticsearch/browse\_thread/thread/18](https://groups.google.com/group/elasticsearch/browse_thread/thread/18)...
> 
> > I notice when my total index size exceeds the RAM size (16GB of ram  
> > per machine) the queries start to take a bit longer. Once the  
> > connections pile up the entire cluster becomes massively unstable and  
> > crashes. I have a theory as the dataset goes up in size what was once  
> > a fast query suddenly is slow (my queries fetch data from a certain  
> > time and sort) and I think that might be killing the cluster.
> 
> > -R
> 
> > On Nov 16, 12:08 am, Ocean Wu [darkyo...@gmail.com](mailto:darkyo...@gmail.com) wrote:
> > 
> > > Yes, the limits.conf set to 32000, and net.nf\_conntrack\_max set to
> > 
> > 1.
> 
> > > Thanks for reply.

---

<div class="post-metadata">

**Author:** ![darkyoung](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/darkyoung/32/2924_2.png) [@darkyoung](https://discuss.elastic.co/u/darkyoung)\
**Post date:** [November 17, 2011, 1:53am UTC](https://discuss.elastic.co/t/problems-with-tcp-connections/5885/7 "2011-11-17T01:53:12Z")

</div>

Seems better after I upgrade to 0.18.3.

在 2011年11月16日星期三UTC+8下午10时29分01秒，kimchy写道：

> Can you try and use 0.18.3, see if it helps? It might be related to the  
> connection problem while searching fix.
> 
> On Wed, Nov 16, 2011 at 11:04 AM, electic [ele...@gmail.com](mailto:ele...@gmail.com) wrote:
> 
> > Then we are having the same issue:
> > 
> > [https://groups.google.com/group/elasticsearch/browse\_thread/thread/1861b5c253982c75](https://groups.google.com/group/elasticsearch/browse_thread/thread/1861b5c253982c75)
> > 
> > I notice when my total index size exceeds the RAM size (16GB of ram  
> > per machine) the queries start to take a bit longer. Once the  
> > connections pile up the entire cluster becomes massively unstable and  
> > crashes. I have a theory as the dataset goes up in size what was once  
> > a fast query suddenly is slow (my queries fetch data from a certain  
> > time and sort) and I think that might be killing the cluster.
> > 
> > -R
> > 
> > On Nov 16, 12:08 am, Ocean Wu [dark...@gmail.com](mailto:dark...@gmail.com) wrote:
> > 
> > > Yes, the limits.conf set to 32000, and net.nf\_conntrack\_max set to
> > 
> > 1. 
> > 
> > > Thanks for reply.

---

<div class="post-metadata">

**Author:** ![electic](https://avatars.discourse-cdn.com/v4/letter/e/ecccb3/32.png) [@electic](https://discuss.elastic.co/u/electic)\
**Post date:** [November 17, 2011, 6:48pm UTC](https://discuss.elastic.co/t/problems-with-tcp-connections/5885/8 "2011-11-17T18:48:53Z")

</div>

So I still seem to be having the same issue. I have two machines with  
16GB RAM each on them. A 10,000 RPM drive. As the datasize increased  
to about 20GB total, 20 million documents, the queries seem to be  
taking longer and the connections start to backup until it no longer  
seems to be taking HTTP requests.

The logs show nothing. There are no huge CPU usage, or heap usage,  
just dead. Any ideas on what I can paste here in terms of logs to  
debug the issue?

On Nov 16, 5:53 pm, Ocean Wu [darkyo...@gmail.com](mailto:darkyo...@gmail.com) wrote:

> Seems better after I upgrade to 0.18.3.
> 
> 在 2011年11月16日星期三UTC+8下午10时29分01秒，kimchy写道：
> 
> > Can you try and use 0.18.3, see if it helps? It might be related to the  
> > connection problem while searching fix.
> 
> > On Wed, Nov 16, 2011 at 11:04 AM, electic [ele...@gmail.com](mailto:ele...@gmail.com) wrote:
> 
> > > Then we are having the same issue:
> 
> > > [https://groups.google.com/group/elasticsearch/browse\_thread/thread/18](https://groups.google.com/group/elasticsearch/browse_thread/thread/18)...
> 
> > > I notice when my total index size exceeds the RAM size (16GB of ram  
> > > per machine) the queries start to take a bit longer. Once the  
> > > connections pile up the entire cluster becomes massively unstable and  
> > > crashes. I have a theory as the dataset goes up in size what was once  
> > > a fast query suddenly is slow (my queries fetch data from a certain  
> > > time and sort) and I think that might be killing the cluster.
> 
> > > -R
> 
> > > On Nov 16, 12:08 am, Ocean Wu [dark...@gmail.com](mailto:dark...@gmail.com) wrote:
> > > 
> > > > Yes, the limits.conf set to 32000, and net.nf\_conntrack\_max set to
> > > 
> > > 1.
> 
> > > > Thanks for reply.

---

<div class="post-metadata">

**Author:** ![electic](https://avatars.discourse-cdn.com/v4/letter/e/ecccb3/32.png) [@electic](https://discuss.elastic.co/u/electic)\
**Post date:** [November 17, 2011, 7:21pm UTC](https://discuss.elastic.co/t/problems-with-tcp-connections/5885/9 "2011-11-17T19:21:14Z")

</div>

Here is my status after restarting the second node (the node that  
handles all the query requests):

[https://raw.github.com/gist/1374154/dc4df73f7fecb81491823ea7c51a6e00fa2c2ae3/gistfile1.txt](https://raw.github.com/gist/1374154/dc4df73f7fecb81491823ea7c51a6e00fa2c2ae3/gistfile1.txt)

On Nov 17, 10:48 am, electic [elec...@gmail.com](mailto:elec...@gmail.com) wrote:

> So I still seem to be having the same issue. I have two machines with  
> 16GB RAM each on them. A 10,000 RPM drive. As the datasize increased  
> to about 20GB total, 20 million documents, the queries seem to be  
> taking longer and the connections start to backup until it no longer  
> seems to be taking HTTP requests.
> 
> The logs show nothing. There are no huge CPU usage, or heap usage,  
> just dead. Any ideas on what I can paste here in terms of logs to  
> debug the issue?
> 
> On Nov 16, 5:53 pm, Ocean Wu [darkyo...@gmail.com](mailto:darkyo...@gmail.com) wrote:
> 
> > Seems better after I upgrade to 0.18.3.
> 
> > 在 2011年11月16日星期三UTC+8下午10时29分01秒，kimchy写道：
> 
> > > Can you try and use 0.18.3, see if it helps? It might be related to the  
> > > connection problem while searching fix.
> 
> > > On Wed, Nov 16, 2011 at 11:04 AM, electic [ele...@gmail.com](mailto:ele...@gmail.com) wrote:
> 
> > > > Then we are having the same issue:
> 
> > > > [https://groups.google.com/group/elasticsearch/browse\_thread/thread/18](https://groups.google.com/group/elasticsearch/browse_thread/thread/18)...
> 
> > > > I notice when my total index size exceeds the RAM size (16GB of ram  
> > > > per machine) the queries start to take a bit longer. Once the  
> > > > connections pile up the entire cluster becomes massively unstable and  
> > > > crashes. I have a theory as the dataset goes up in size what was once  
> > > > a fast query suddenly is slow (my queries fetch data from a certain  
> > > > time and sort) and I think that might be killing the cluster.
> 
> > > > -R
> 
> > > > On Nov 16, 12:08 am, Ocean Wu [dark...@gmail.com](mailto:dark...@gmail.com) wrote:
> > > > 
> > > > > Yes, the limits.conf set to 32000, and net.nf\_conntrack\_max set to
> > > > 
> > > > 1.
> 
> > > > > Thanks for reply.

---

<div class="post-metadata">

**Author:** ![electic](https://avatars.discourse-cdn.com/v4/letter/e/ecccb3/32.png) [@electic](https://discuss.elastic.co/u/electic)\
**Post date:** [November 17, 2011, 11:13pm UTC](https://discuss.elastic.co/t/problems-with-tcp-connections/5885/10 "2011-11-17T23:13:12Z")

</div>

I think this might have something to do with the merge policy. It is  
happening around 20GB. Any ideas?

On Nov 17, 11:21 am, electic [elec...@gmail.com](mailto:elec...@gmail.com) wrote:

> Here is my status after restarting the second node (the node that  
> handles all the query requests):
> 
> [https://raw.github.com/gist/1374154/dc4df73f7fecb81491823ea7c51a6e00f](https://raw.github.com/gist/1374154/dc4df73f7fecb81491823ea7c51a6e00f)...
> 
> On Nov 17, 10:48 am, electic [elec...@gmail.com](mailto:elec...@gmail.com) wrote:
> 
> > So I still seem to be having the same issue. I have two machines with  
> > 16GB RAM each on them. A 10,000 RPM drive. As the datasize increased  
> > to about 20GB total, 20 million documents, the queries seem to be  
> > taking longer and the connections start to backup until it no longer  
> > seems to be taking HTTP requests.
> 
> > The logs show nothing. There are no huge CPU usage, or heap usage,  
> > just dead. Any ideas on what I can paste here in terms of logs to  
> > debug the issue?
> 
> > On Nov 16, 5:53 pm, Ocean Wu [darkyo...@gmail.com](mailto:darkyo...@gmail.com) wrote:
> 
> > > Seems better after I upgrade to 0.18.3.
> 
> > > 在 2011年11月16日星期三UTC+8下午10时29分01秒，kimchy写道：
> 
> > > > Can you try and use 0.18.3, see if it helps? It might be related to the  
> > > > connection problem while searching fix.
> 
> > > > On Wed, Nov 16, 2011 at 11:04 AM, electic [ele...@gmail.com](mailto:ele...@gmail.com) wrote:
> 
> > > > > Then we are having the same issue:
> 
> > > > > [https://groups.google.com/group/elasticsearch/browse\_thread/thread/18](https://groups.google.com/group/elasticsearch/browse_thread/thread/18)...
> 
> > > > > I notice when my total index size exceeds the RAM size (16GB of ram  
> > > > > per machine) the queries start to take a bit longer. Once the  
> > > > > connections pile up the entire cluster becomes massively unstable and  
> > > > > crashes. I have a theory as the dataset goes up in size what was once  
> > > > > a fast query suddenly is slow (my queries fetch data from a certain  
> > > > > time and sort) and I think that might be killing the cluster.
> 
> > > > > -R
> 
> > > > > On Nov 16, 12:08 am, Ocean Wu [dark...@gmail.com](mailto:dark...@gmail.com) wrote:
> > > > > 
> > > > > > Yes, the limits.conf set to 32000, and net.nf\_conntrack\_max set to
> > > > > 
> > > > > 1.
> 
> > > > > > Thanks for reply.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [November 20, 2011, 8:16am UTC](https://discuss.elastic.co/t/problems-with-tcp-connections/5885/11 "2011-11-20T08:16:36Z")

</div>

Are you using connections with keep alive / persistent connection? If you  
open and close connections constantly, maybe the OS is throttling them?

On Thu, Nov 17, 2011 at 8:48 PM, electic [electic@gmail.com](mailto:electic@gmail.com) wrote:

> So I still seem to be having the same issue. I have two machines with  
> 16GB RAM each on them. A 10,000 RPM drive. As the datasize increased  
> to about 20GB total, 20 million documents, the queries seem to be  
> taking longer and the connections start to backup until it no longer  
> seems to be taking HTTP requests.
> 
> The logs show nothing. There are no huge CPU usage, or heap usage,  
> just dead. Any ideas on what I can paste here in terms of logs to  
> debug the issue?
> 
> On Nov 16, 5:53 pm, Ocean Wu [darkyo...@gmail.com](mailto:darkyo...@gmail.com) wrote:
> 
> > Seems better after I upgrade to 0.18.3.
> > 
> > 在 2011年11月16日星期三UTC+8下午10时29分01秒，kimchy写道：
> > 
> > > Can you try and use 0.18.3, see if it helps? It might be related to the  
> > > connection problem while searching fix.
> > 
> > > On Wed, Nov 16, 2011 at 11:04 AM, electic [ele...@gmail.com](mailto:ele...@gmail.com) wrote:
> > 
> > > > Then we are having the same issue:
> > 
> > > > [https://groups.google.com/group/elasticsearch/browse\_thread/thread/18](https://groups.google.com/group/elasticsearch/browse_thread/thread/18).  
> > > > ..
> > 
> > > > I notice when my total index size exceeds the RAM size (16GB of ram  
> > > > per machine) the queries start to take a bit longer. Once the  
> > > > connections pile up the entire cluster becomes massively unstable and  
> > > > crashes. I have a theory as the dataset goes up in size what was once  
> > > > a fast query suddenly is slow (my queries fetch data from a certain  
> > > > time and sort) and I think that might be killing the cluster.
> > 
> > > > -R
> > 
> > > > On Nov 16, 12:08 am, Ocean Wu [dark...@gmail.com](mailto:dark...@gmail.com) wrote:
> > > > 
> > > > > Yes, the limits.conf set to 32000, and net.nf\_conntrack\_max set to
> > > > 
> > > > 1.
> > 
> > > > > Thanks for reply.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:48am UTC](https://discuss.elastic.co/t/problems-with-tcp-connections/5885/12 "2017-07-06T03:48:09Z")

</div>


