# Performance problems

**URL:** <https://discuss.elastic.co/t/performance-problems/3205>\
**Category:** Elasticsearch\
**Created:** [August 10, 2010, 9:09pm UTC](https://discuss.elastic.co/t/performance-problems/3205 "2010-08-10T21:09:36Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![Michael\_Korbakov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/michael_korbakov/32/3199_2.png) [@Michael\_Korbakov](https://discuss.elastic.co/u/Michael_Korbakov)\
**Post date:** [August 10, 2010, 9:09pm UTC](https://discuss.elastic.co/t/performance-problems/3205/1 "2010-08-10T21:09:36Z")

</div>

Hi everyone.

I'm running elasticsearch-0.9 on the cluster of 5 EC2 instances  
(Cluster Compute Quadruple Extra Large Instance) with one of them  
configured as frontend (data: false). Index contains 10M of documents,  
4 shards, 2 replicas. Total stored size is 23GB and each node has 23  
GB of RAM, so it fit's to memory without any problems.  
Using JMeter as testing tool I'm getting as little as ~150 requests  
per second (that's for 4 worker nodes with 8 CPU cores each). I've  
checked with iostat that there's no disk activity on cluster nodes. I  
also see CPU utilization near 5-10% on worker nodes during performance  
testing. Looks somewhat strange to me.

I made performance measurements on the same cluster when index  
contained only 5M documents. I got nearly 1500 requests per second and  
CPU utilizations on worker nodes was close to 90%. After importing  
another million of documents performance started to degrade very  
rapidly.

Could anyone help me with this problem? I'm completely out of ideas  
now.

My queries are nothing complex: a single keyword search with a bunch  
of filter attached, also faceting by some fields. Typical query looks  
like this:

{  
"query":  
{  
"filtered":  
{  
"query":  
{  
"query\_string":  
{  
"fields":  
[  
"keywords.original\_keywords^2",  
"keywords.keywords"  
],  
"query": "bright"  
}  
},  
"filter":  
{  
"and":  
{  
"filters":  
[  
{  
"term":  
{  
"content.is\_offensive": false  
}  
},

```
                    {
                        "term":
                        {
                            "licenses.extended": true
                        }
                    },

                    {
                        "term":
                        {
                            "image.isolated": false
                        }
                    },

                    {
                        "term":
                        {
                            "content.orientation": "horizontal"
                        }
                    },

                    {
                        "term":
                        {
                            "categories.conceptual.depth2": "793"
                        }
                    }
                ]
            }
        }
    },
     "sort": "online.rating",
     "facets":
    {
        "representative_categories":
        {
            "terms":
            {
                "field": "categories.representative.depth2",
                 "size": 100
            }
        },
         "representative_categories":
        {
            "terms":
            {
                "field": "categories.conceptual.depth2",
                 "size": 100
            }
        },
         "licenses":
        {
            "terms":
            {
                "field": "licenses.size",
                 "size": 100
            }
        },
         "prices":
        {
            "histogram":
            {
                "field": "prices.min",
                 "interval": 1
            }
        }
    }
}

```

}

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 10, 2010, 9:41pm UTC](https://discuss.elastic.co/t/performance-problems/3205/2 "2010-08-10T21:41:53Z")

</div>

Hi,

```
Is there a chance that the response that you get is really large? It

```

seems like you are getting large result sets for the facets (not sure about  
the histogram facet of 1 for price, it depends on the range of it). Can you  
try and start with a simple query (no filters, no facets) and slowly add  
more to the search request? How much memory do you assign each node? It  
very strange that by moving from 5M docs to 6M docs, suddenly you get such  
different results, unless those 1M cause the facets to "explode" with the  
data they return?

-shay.banon

On Wed, Aug 11, 2010 at 12:09 AM, rmihael [rmihael@gmail.com](mailto:rmihael@gmail.com) wrote:

> Hi everyone.
> 
> I'm running elasticsearch-0.9 on the cluster of 5 EC2 instances  
> (Cluster Compute Quadruple Extra Large Instance) with one of them  
> configured as frontend (data: false). Index contains 10M of documents,  
> 4 shards, 2 replicas. Total stored size is 23GB and each node has 23  
> GB of RAM, so it fit's to memory without any problems.  
> Using JMeter as testing tool I'm getting as little as ~150 requests  
> per second (that's for 4 worker nodes with 8 CPU cores each). I've  
> checked with iostat that there's no disk activity on cluster nodes. I  
> also see CPU utilization near 5-10% on worker nodes during performance  
> testing. Looks somewhat strange to me.
> 
> I made performance measurements on the same cluster when index  
> contained only 5M documents. I got nearly 1500 requests per second and  
> CPU utilizations on worker nodes was close to 90%. After importing  
> another million of documents performance started to degrade very  
> rapidly.
> 
> Could anyone help me with this problem? I'm completely out of ideas  
> now.
> 
> My queries are nothing complex: a single keyword search with a bunch  
> of filter attached, also faceting by some fields. Typical query looks  
> like this:
> 
> {  
> "query":  
> {  
> "filtered":  
> {  
> "query":  
> {  
> "query\_string":  
> {  
> "fields":  
> [  
> "keywords.original\_keywords^2",  
> "keywords.keywords"  
> ],  
> "query": "bright"  
> }  
> },  
> "filter":  
> {  
> "and":  
> {  
> "filters":  
> [  
> {  
> "term":  
> {  
> "content.is\_offensive": false  
> }  
> },
> 
> ```
> {
> "term":
> {
> "licenses.extended": true
> }
> },
> 
> {
> "term":
> {
> "image.isolated": false
> }
> },
> 
> {
> "term":
> {
> "content.orientation": "horizontal"
> }
> },
> 
> {
> "term":
> {
> "categories.conceptual.depth2": "793"
> }
> }
> ]
> }
> }
> },
> "sort": "online.rating",
> "facets":
> {
> "representative_categories":
> {
> "terms":
> {
> "field": "categories.representative.depth2",
> "size": 100
> }
> },
> "representative_categories":
> {
> "terms":
> {
> "field": "categories.conceptual.depth2",
> "size": 100
> }
> },
> "licenses":
> {
> "terms":
> {
> "field": "licenses.size",
> "size": 100
> }
> },
> "prices":
> {
> "histogram":
> {
> "field": "prices.min",
> "interval": 1
> }
> }
> }
> 
> ```
> 
> }  
> }

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 10, 2010, 9:45pm UTC](https://discuss.elastic.co/t/performance-problems/3205/3 "2010-08-10T21:45:58Z")

</div>

One more thing, I would use the 5th node as data node as well, its a shame  
for it not to share the load.

On Wed, Aug 11, 2010 at 12:41 AM, Shay Banon  
[shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:

> Hi,
> 
> ```
> Is there a chance that the response that you get is really large? It
> 
> ```
> 
> seems like you are getting large result sets for the facets (not sure about  
> the histogram facet of 1 for price, it depends on the range of it). Can you  
> try and start with a simple query (no filters, no facets) and slowly add  
> more to the search request? How much memory do you assign each node? It  
> very strange that by moving from 5M docs to 6M docs, suddenly you get such  
> different results, unless those 1M cause the facets to "explode" with the  
> data they return?
> 
> -shay.banon
> 
> On Wed, Aug 11, 2010 at 12:09 AM, rmihael [rmihael@gmail.com](mailto:rmihael@gmail.com) wrote:
> 
> > Hi everyone.
> > 
> > I'm running elasticsearch-0.9 on the cluster of 5 EC2 instances  
> > (Cluster Compute Quadruple Extra Large Instance) with one of them  
> > configured as frontend (data: false). Index contains 10M of documents,  
> > 4 shards, 2 replicas. Total stored size is 23GB and each node has 23  
> > GB of RAM, so it fit's to memory without any problems.  
> > Using JMeter as testing tool I'm getting as little as ~150 requests  
> > per second (that's for 4 worker nodes with 8 CPU cores each). I've  
> > checked with iostat that there's no disk activity on cluster nodes. I  
> > also see CPU utilization near 5-10% on worker nodes during performance  
> > testing. Looks somewhat strange to me.
> > 
> > I made performance measurements on the same cluster when index  
> > contained only 5M documents. I got nearly 1500 requests per second and  
> > CPU utilizations on worker nodes was close to 90%. After importing  
> > another million of documents performance started to degrade very  
> > rapidly.
> > 
> > Could anyone help me with this problem? I'm completely out of ideas  
> > now.
> > 
> > My queries are nothing complex: a single keyword search with a bunch  
> > of filter attached, also faceting by some fields. Typical query looks  
> > like this:
> > 
> > {  
> > "query":  
> > {  
> > "filtered":  
> > {  
> > "query":  
> > {  
> > "query\_string":  
> > {  
> > "fields":  
> > [  
> > "keywords.original\_keywords^2",  
> > "keywords.keywords"  
> > ],  
> > "query": "bright"  
> > }  
> > },  
> > "filter":  
> > {  
> > "and":  
> > {  
> > "filters":  
> > [  
> > {  
> > "term":  
> > {  
> > "content.is\_offensive": false  
> > }  
> > },
> > 
> > ```
> > {
> > "term":
> > {
> > "licenses.extended": true
> > }
> > },
> > 
> > {
> > "term":
> > {
> > "image.isolated": false
> > }
> > },
> > 
> > {
> > "term":
> > {
> > "content.orientation": "horizontal"
> > }
> > },
> > 
> > {
> > "term":
> > {
> > "categories.conceptual.depth2": "793"
> > }
> > }
> > ]
> > }
> > }
> > },
> > "sort": "online.rating",
> > "facets":
> > {
> > "representative_categories":
> > {
> > "terms":
> > {
> > "field": "categories.representative.depth2",
> > "size": 100
> > }
> > },
> > "representative_categories":
> > {
> > "terms":
> > {
> > "field": "categories.conceptual.depth2",
> > "size": 100
> > }
> > },
> > "licenses":
> > {
> > "terms":
> > {
> > "field": "licenses.size",
> > "size": 100
> > }
> > },
> > "prices":
> > {
> > "histogram":
> > {
> > "field": "prices.min",
> > "interval": 1
> > }
> > }
> > }
> > 
> > ```
> > 
> > }  
> > }

---

<div class="post-metadata">

**Author:** ![Michael\_Korbakov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/michael_korbakov/32/3199_2.png) [@Michael\_Korbakov](https://discuss.elastic.co/u/Michael_Korbakov)\
**Post date:** [August 10, 2010, 10:11pm UTC](https://discuss.elastic.co/t/performance-problems/3205/4 "2010-08-10T22:11:10Z")

</div>

Typical 'total' is less then 100K items. Price ranging between 1 and  
10, so I don't think it can cause much problems. I'll try to remove  
faceting now and see how it goes.  
I'm bit confused with your question about memory. I didn't assigned  
any memory to nodes, it just runs as it is. This kind of EC2 instance  
have 23GB of memory if you mean it.  
Data I've been uploading to index are very uniform. In fact they are  
randomly generated and should cause any kind of statistical explosion.  
But I've certainly got this fast degradation between 5M and 6M. Very  
strange, looks like something in my setup is very broken.

BTW, I'm getting the following messages in log files:

[14:11:14,211][INFO][monitor.memory.alpha] [Sangre] [5]  
[Full] Ran after [2] consecutive clean swipes, memory\_to\_clean  
[155.7mb], lower\_memory\_threshold [809mb], upper\_memory\_threshold  
[960.6mb], used\_memory [964.7mb], total\_memory[1011.2mb],  
max\_memory[1011.2mb]  
[14:52:06,191][INFO][monitor.memory.alpha] [Sangre] [6]  
[Full] Ran after [2] consecutive clean swipes, memory\_to\_clean  
[179.2mb], lower\_memory\_threshold [809mb], upper\_memory\_threshold  
[960.6mb], used\_memory [988.2mb], total\_memory[1011.2mb],  
max\_memory[1011.2mb]

Is everything OK with it? "total\_memory[1011.2mb],  
max\_memory[1011.2mb]" part is confusing me, why it's so small?

On Aug 11, 12:41 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> Hi,
> 
> ```
> Is there a chance that the response that you get is really large? It
> 
> ```
> 
> seems like you are getting large result sets for the facets (not sure about  
> the histogram facet of 1 for price, it depends on the range of it). Can you  
> try and start with a simple query (no filters, no facets) and slowly add  
> more to the search request? How much memory do you assign each node? It  
> very strange that by moving from 5M docs to 6M docs, suddenly you get such  
> different results, unless those 1M cause the facets to "explode" with the  
> data they return?
> 
> -shay.banon
> 
> On Wed, Aug 11, 2010 at 12:09 AM, rmihael [rmih...@gmail.com](mailto:rmih...@gmail.com) wrote:
> 
> > Hi everyone.
> 
> > I'm running elasticsearch-0.9 on the cluster of 5 EC2 instances  
> > (Cluster Compute Quadruple Extra Large Instance) with one of them  
> > configured as frontend (data: false). Index contains 10M of documents,  
> > 4 shards, 2 replicas. Total stored size is 23GB and each node has 23  
> > GB of RAM, so it fit's to memory without any problems.  
> > Using JMeter as testing tool I'm getting as little as ~150 requests  
> > per second (that's for 4 worker nodes with 8 CPU cores each). I've  
> > checked with iostat that there's no disk activity on cluster nodes. I  
> > also see CPU utilization near 5-10% on worker nodes during performance  
> > testing. Looks somewhat strange to me.
> 
> > I made performance measurements on the same cluster when index  
> > contained only 5M documents. I got nearly 1500 requests per second and  
> > CPU utilizations on worker nodes was close to 90%. After importing  
> > another million of documents performance started to degrade very  
> > rapidly.
> 
> > Could anyone help me with this problem? I'm completely out of ideas  
> > now.
> 
> > My queries are nothing complex: a single keyword search with a bunch  
> > of filter attached, also faceting by some fields. Typical query looks  
> > like this:
> 
> > {  
> > "query":  
> > {  
> > "filtered":  
> > {  
> > "query":  
> > {  
> > "query\_string":  
> > {  
> > "fields":  
> > [  
> > "keywords.original\_keywords^2",  
> > "keywords.keywords"  
> > ],  
> > "query": "bright"  
> > }  
> > },  
> > "filter":  
> > {  
> > "and":  
> > {  
> > "filters":  
> > [  
> > {  
> > "term":  
> > {  
> > "content.is\_offensive": false  
> > }  
> > },
> 
> > ```
> > {
> > "term":
> > {
> > "licenses.extended": true
> > }
> > },
> > 
> > ```
> 
> > ```
> > {
> > "term":
> > {
> > "image.isolated": false
> > }
> > },
> > 
> > ```
> 
> > ```
> > {
> > "term":
> > {
> > "content.orientation": "horizontal"
> > }
> > },
> > 
> > ```
> 
> > ```
> > {
> > "term":
> > {
> > "categories.conceptual.depth2": "793"
> > }
> > }
> > ]
> > }
> > }
> > },
> > "sort": "online.rating",
> > "facets":
> > {
> > "representative_categories":
> > {
> > "terms":
> > {
> > "field": "categories.representative.depth2",
> > "size": 100
> > }
> > },
> > "representative_categories":
> > {
> > "terms":
> > {
> > "field": "categories.conceptual.depth2",
> > "size": 100
> > }
> > },
> > "licenses":
> > {
> > "terms":
> > {
> > "field": "licenses.size",
> > "size": 100
> > }
> > },
> > "prices":
> > {
> > "histogram":
> > {
> > "field": "prices.min",
> > "interval": 1
> > }
> > }
> > }
> > 
> > ```
> > 
> > }  
> > }

---

<div class="post-metadata">

**Author:** ![Michael\_Korbakov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/michael_korbakov/32/3199_2.png) [@Michael\_Korbakov](https://discuss.elastic.co/u/Michael_Korbakov)\
**Post date:** [August 10, 2010, 10:15pm UTC](https://discuss.elastic.co/t/performance-problems/3205/5 "2010-08-10T22:15:32Z")

</div>

EC2 setup is just for testing. We'll run the system on our own  
hardware and frontend node will be much weaker then worker nodes. I  
thought that it is the recommended use case:  
[http://www.elasticsearch.com/docs/elasticsearch/modules/node/data\_node/](http://www.elasticsearch.com/docs/elasticsearch/modules/node/data_node/)  
Do you think that we'll do better with homogeneous cluster? If so,  
then I'll try it too.

On Aug 11, 12:45 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> One more thing, I would use the 5th node as data node as well, its a shame  
> for it not to share the load.
> 
> On Wed, Aug 11, 2010 at 12:41 AM, Shay Banon  
> [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com)wrote:
> 
> > Hi,
> 
> > ```
> > Is there a chance that the response that you get is really large? It
> > 
> > ```
> > 
> > seems like you are getting large result sets for the facets (not sure about  
> > the histogram facet of 1 for price, it depends on the range of it). Can you  
> > try and start with a simple query (no filters, no facets) and slowly add  
> > more to the search request? How much memory do you assign each node? It  
> > very strange that by moving from 5M docs to 6M docs, suddenly you get such  
> > different results, unless those 1M cause the facets to "explode" with the  
> > data they return?
> 
> > -shay.banon
> 
> > On Wed, Aug 11, 2010 at 12:09 AM, rmihael [rmih...@gmail.com](mailto:rmih...@gmail.com) wrote:
> 
> > > Hi everyone.
> 
> > > I'm running elasticsearch-0.9 on the cluster of 5 EC2 instances  
> > > (Cluster Compute Quadruple Extra Large Instance) with one of them  
> > > configured as frontend (data: false). Index contains 10M of documents,  
> > > 4 shards, 2 replicas. Total stored size is 23GB and each node has 23  
> > > GB of RAM, so it fit's to memory without any problems.  
> > > Using JMeter as testing tool I'm getting as little as ~150 requests  
> > > per second (that's for 4 worker nodes with 8 CPU cores each). I've  
> > > checked with iostat that there's no disk activity on cluster nodes. I  
> > > also see CPU utilization near 5-10% on worker nodes during performance  
> > > testing. Looks somewhat strange to me.
> 
> > > I made performance measurements on the same cluster when index  
> > > contained only 5M documents. I got nearly 1500 requests per second and  
> > > CPU utilizations on worker nodes was close to 90%. After importing  
> > > another million of documents performance started to degrade very  
> > > rapidly.
> 
> > > Could anyone help me with this problem? I'm completely out of ideas  
> > > now.
> 
> > > My queries are nothing complex: a single keyword search with a bunch  
> > > of filter attached, also faceting by some fields. Typical query looks  
> > > like this:
> 
> > > {  
> > > "query":  
> > > {  
> > > "filtered":  
> > > {  
> > > "query":  
> > > {  
> > > "query\_string":  
> > > {  
> > > "fields":  
> > > [  
> > > "keywords.original\_keywords^2",  
> > > "keywords.keywords"  
> > > ],  
> > > "query": "bright"  
> > > }  
> > > },  
> > > "filter":  
> > > {  
> > > "and":  
> > > {  
> > > "filters":  
> > > [  
> > > {  
> > > "term":  
> > > {  
> > > "content.is\_offensive": false  
> > > }  
> > > },
> 
> > > ```
> > > {
> > > "term":
> > > {
> > > "licenses.extended": true
> > > }
> > > },
> > > 
> > > ```
> 
> > > ```
> > > {
> > > "term":
> > > {
> > > "image.isolated": false
> > > }
> > > },
> > > 
> > > ```
> 
> > > ```
> > > {
> > > "term":
> > > {
> > > "content.orientation": "horizontal"
> > > }
> > > },
> > > 
> > > ```
> 
> > > ```
> > > {
> > > "term":
> > > {
> > > "categories.conceptual.depth2": "793"
> > > }
> > > }
> > > ]
> > > }
> > > }
> > > },
> > > "sort": "online.rating",
> > > "facets":
> > > {
> > > "representative_categories":
> > > {
> > > "terms":
> > > {
> > > "field": "categories.representative.depth2",
> > > "size": 100
> > > }
> > > },
> > > "representative_categories":
> > > {
> > > "terms":
> > > {
> > > "field": "categories.conceptual.depth2",
> > > "size": 100
> > > }
> > > },
> > > "licenses":
> > > {
> > > "terms":
> > > {
> > > "field": "licenses.size",
> > > "size": 100
> > > }
> > > },
> > > "prices":
> > > {
> > > "histogram":
> > > {
> > > "field": "prices.min",
> > > "interval": 1
> > > }
> > > }
> > > }
> > > 
> > > ```
> > > 
> > > }  
> > > }

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 10, 2010, 10:22pm UTC](https://discuss.elastic.co/t/performance-problems/3205/6 "2010-08-10T22:22:11Z")

</div>

Good, thats what I was concerned about. When you run a Java virtual machine,  
you assign memory to it and it only consumes as much memory as you give it.  
By default, it is set to use 1g max memory. Certainly, with your machine,  
you can increase that quite significantly, I would say do 10G and see how it  
goes (you want to leave memory also for file system cache, and too large  
heaps can cause the JVM to hiccup). How to set the max memory is explained  
here: [http://www.elasticsearch.com/docs/elasticsearch/setup/installation/](http://www.elasticsearch.com/docs/elasticsearch/setup/installation/).  
For even better performance, set the minimum and the maximum to the same  
value.

One more thing, with your setup, if you use 2 replicas to try and increase  
the search performance, then 1 replicas should do. If you use 2 replicas to  
increase the availability aspect, then thats fine.

One cool thing to check how the JVM is behaving is to use something like  
visualvm to hook into it and check the memory consumption and GC activity.  
All that information is already exposed in the node stats API, and once I  
get around to build a nice management app for elasticsearch, it will be  
exposed there through the REST API.

-shay.banon

On Wed, Aug 11, 2010 at 1:11 AM, rmihael [rmihael@gmail.com](mailto:rmihael@gmail.com) wrote:

> Typical 'total' is less then 100K items. Price ranging between 1 and  
> 10, so I don't think it can cause much problems. I'll try to remove  
> faceting now and see how it goes.  
> I'm bit confused with your question about memory. I didn't assigned  
> any memory to nodes, it just runs as it is. This kind of EC2 instance  
> have 23GB of memory if you mean it.  
> Data I've been uploading to index are very uniform. In fact they are  
> randomly generated and should cause any kind of statistical explosion.  
> But I've certainly got this fast degradation between 5M and 6M. Very  
> strange, looks like something in my setup is very broken.
> 
> BTW, I'm getting the following messages in log files:
> 
> [14:11:14,211][INFO][monitor.memory.alpha] [Sangre] [5]  
> [Full] Ran after [2] consecutive clean swipes, memory\_to\_clean  
> [155.7mb], lower\_memory\_threshold [809mb], upper\_memory\_threshold  
> [960.6mb], used\_memory [964.7mb], total\_memory[1011.2mb],  
> max\_memory[1011.2mb]  
> [14:52:06,191][INFO][monitor.memory.alpha] [Sangre] [6]  
> [Full] Ran after [2] consecutive clean swipes, memory\_to\_clean  
> [179.2mb], lower\_memory\_threshold [809mb], upper\_memory\_threshold  
> [960.6mb], used\_memory [988.2mb], total\_memory[1011.2mb],  
> max\_memory[1011.2mb]
> 
> Is everything OK with it? "total\_memory[1011.2mb],  
> max\_memory[1011.2mb]" part is confusing me, why it's so small?
> 
> On Aug 11, 12:41 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > Hi,
> > 
> > ```
> > Is there a chance that the response that you get is really large? It
> > 
> > ```
> > 
> > seems like you are getting large result sets for the facets (not sure  
> > about  
> > the histogram facet of 1 for price, it depends on the range of it). Can  
> > you  
> > try and start with a simple query (no filters, no facets) and slowly add  
> > more to the search request? How much memory do you assign each node? It  
> > very strange that by moving from 5M docs to 6M docs, suddenly you get  
> > such  
> > different results, unless those 1M cause the facets to "explode" with the  
> > data they return?
> > 
> > -shay.banon
> > 
> > On Wed, Aug 11, 2010 at 12:09 AM, rmihael [rmih...@gmail.com](mailto:rmih...@gmail.com) wrote:
> > 
> > > Hi everyone.
> > 
> > > I'm running elasticsearch-0.9 on the cluster of 5 EC2 instances  
> > > (Cluster Compute Quadruple Extra Large Instance) with one of them  
> > > configured as frontend (data: false). Index contains 10M of documents,  
> > > 4 shards, 2 replicas. Total stored size is 23GB and each node has 23  
> > > GB of RAM, so it fit's to memory without any problems.  
> > > Using JMeter as testing tool I'm getting as little as ~150 requests  
> > > per second (that's for 4 worker nodes with 8 CPU cores each). I've  
> > > checked with iostat that there's no disk activity on cluster nodes. I  
> > > also see CPU utilization near 5-10% on worker nodes during performance  
> > > testing. Looks somewhat strange to me.
> > 
> > > I made performance measurements on the same cluster when index  
> > > contained only 5M documents. I got nearly 1500 requests per second and  
> > > CPU utilizations on worker nodes was close to 90%. After importing  
> > > another million of documents performance started to degrade very  
> > > rapidly.
> > 
> > > Could anyone help me with this problem? I'm completely out of ideas  
> > > now.
> > 
> > > My queries are nothing complex: a single keyword search with a bunch  
> > > of filter attached, also faceting by some fields. Typical query looks  
> > > like this:
> > 
> > > {  
> > > "query":  
> > > {  
> > > "filtered":  
> > > {  
> > > "query":  
> > > {  
> > > "query\_string":  
> > > {  
> > > "fields":  
> > > [  
> > > "keywords.original\_keywords^2",  
> > > "keywords.keywords"  
> > > ],  
> > > "query": "bright"  
> > > }  
> > > },  
> > > "filter":  
> > > {  
> > > "and":  
> > > {  
> > > "filters":  
> > > [  
> > > {  
> > > "term":  
> > > {  
> > > "content.is\_offensive": false  
> > > }  
> > > },
> > 
> > > ```
> > > {
> > > "term":
> > > {
> > > "licenses.extended": true
> > > }
> > > },
> > > 
> > > ```
> > 
> > > ```
> > > {
> > > "term":
> > > {
> > > "image.isolated": false
> > > }
> > > },
> > > 
> > > ```
> > 
> > > ```
> > > {
> > > "term":
> > > {
> > > "content.orientation": "horizontal"
> > > }
> > > },
> > > 
> > > ```
> > 
> > > ```
> > > {
> > > "term":
> > > {
> > > "categories.conceptual.depth2": "793"
> > > }
> > > }
> > > ]
> > > }
> > > }
> > > },
> > > "sort": "online.rating",
> > > "facets":
> > > {
> > > "representative_categories":
> > > {
> > > "terms":
> > > {
> > > "field": "categories.representative.depth2",
> > > "size": 100
> > > }
> > > },
> > > "representative_categories":
> > > {
> > > "terms":
> > > {
> > > "field": "categories.conceptual.depth2",
> > > "size": 100
> > > }
> > > },
> > > "licenses":
> > > {
> > > "terms":
> > > {
> > > "field": "licenses.size",
> > > "size": 100
> > > }
> > > },
> > > "prices":
> > > {
> > > "histogram":
> > > {
> > > "field": "prices.min",
> > > "interval": 1
> > > }
> > > }
> > > }
> > > 
> > > ```
> > > 
> > > }  
> > > }

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 10, 2010, 10:24pm UTC](https://discuss.elastic.co/t/performance-problems/3205/7 "2010-08-10T22:24:11Z")

</div>

It depends on how you index the data. If you do the load balancing yourself  
among the nodes (either using the native JVM client, or building on top of  
the exposed API like Elasticsearch.pm does), then most times, it makes sense  
to go with all data nodes.

-shay.banon

On Wed, Aug 11, 2010 at 1:15 AM, rmihael [rmihael@gmail.com](mailto:rmihael@gmail.com) wrote:

> EC2 setup is just for testing. We'll run the system on our own  
> hardware and frontend node will be much weaker then worker nodes. I  
> thought that it is the recommended use case:  
> [http://www.elasticsearch.com/docs/elasticsearch/modules/node/data\_node/](http://www.elasticsearch.com/docs/elasticsearch/modules/node/data_node/)  
> Do you think that we'll do better with homogeneous cluster? If so,  
> then I'll try it too.
> 
> On Aug 11, 12:45 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > One more thing, I would use the 5th node as data node as well, its a  
> > shame  
> > for it not to share the load.
> > 
> > On Wed, Aug 11, 2010 at 12:41 AM, Shay Banon  
> > [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com)wrote:
> > 
> > > Hi,
> > 
> > > ```
> > > Is there a chance that the response that you get is really large?
> > > 
> > > ```
> 
> It
> 
> > > seems like you are getting large result sets for the facets (not sure  
> > > about  
> > > the histogram facet of 1 for price, it depends on the range of it). Can  
> > > you  
> > > try and start with a simple query (no filters, no facets) and slowly  
> > > add  
> > > more to the search request? How much memory do you assign each node?  
> > > It  
> > > very strange that by moving from 5M docs to 6M docs, suddenly you get  
> > > such  
> > > different results, unless those 1M cause the facets to "explode" with  
> > > the  
> > > data they return?
> > 
> > > -shay.banon
> > 
> > > On Wed, Aug 11, 2010 at 12:09 AM, rmihael [rmih...@gmail.com](mailto:rmih...@gmail.com) wrote:
> > 
> > > > Hi everyone.
> > 
> > > > I'm running elasticsearch-0.9 on the cluster of 5 EC2 instances  
> > > > (Cluster Compute Quadruple Extra Large Instance) with one of them  
> > > > configured as frontend (data: false). Index contains 10M of documents,  
> > > > 4 shards, 2 replicas. Total stored size is 23GB and each node has 23  
> > > > GB of RAM, so it fit's to memory without any problems.  
> > > > Using JMeter as testing tool I'm getting as little as ~150 requests  
> > > > per second (that's for 4 worker nodes with 8 CPU cores each). I've  
> > > > checked with iostat that there's no disk activity on cluster nodes. I  
> > > > also see CPU utilization near 5-10% on worker nodes during performance  
> > > > testing. Looks somewhat strange to me.
> > 
> > > > I made performance measurements on the same cluster when index  
> > > > contained only 5M documents. I got nearly 1500 requests per second and  
> > > > CPU utilizations on worker nodes was close to 90%. After importing  
> > > > another million of documents performance started to degrade very  
> > > > rapidly.
> > 
> > > > Could anyone help me with this problem? I'm completely out of ideas  
> > > > now.
> > 
> > > > My queries are nothing complex: a single keyword search with a bunch  
> > > > of filter attached, also faceting by some fields. Typical query looks  
> > > > like this:
> > 
> > > > {  
> > > > "query":  
> > > > {  
> > > > "filtered":  
> > > > {  
> > > > "query":  
> > > > {  
> > > > "query\_string":  
> > > > {  
> > > > "fields":  
> > > > [  
> > > > "keywords.original\_keywords^2",  
> > > > "keywords.keywords"  
> > > > ],  
> > > > "query": "bright"  
> > > > }  
> > > > },  
> > > > "filter":  
> > > > {  
> > > > "and":  
> > > > {  
> > > > "filters":  
> > > > [  
> > > > {  
> > > > "term":  
> > > > {  
> > > > "content.is\_offensive": false  
> > > > }  
> > > > },
> > 
> > > > ```
> > > > {
> > > > "term":
> > > > {
> > > > "licenses.extended": true
> > > > }
> > > > },
> > > > 
> > > > ```
> > 
> > > > ```
> > > > {
> > > > "term":
> > > > {
> > > > "image.isolated": false
> > > > }
> > > > },
> > > > 
> > > > ```
> > 
> > > > ```
> > > > {
> > > > "term":
> > > > {
> > > > "content.orientation": "horizontal"
> > > > }
> > > > },
> > > > 
> > > > ```
> > 
> > > > ```
> > > > {
> > > > "term":
> > > > {
> > > > "categories.conceptual.depth2": "793"
> > > > }
> > > > }
> > > > ]
> > > > }
> > > > }
> > > > },
> > > > "sort": "online.rating",
> > > > "facets":
> > > > {
> > > > "representative_categories":
> > > > {
> > > > "terms":
> > > > {
> > > > "field": "categories.representative.depth2",
> > > > "size": 100
> > > > }
> > > > },
> > > > "representative_categories":
> > > > {
> > > > "terms":
> > > > {
> > > > "field": "categories.conceptual.depth2",
> > > > "size": 100
> > > > }
> > > > },
> > > > "licenses":
> > > > {
> > > > "terms":
> > > > {
> > > > "field": "licenses.size",
> > > > "size": 100
> > > > }
> > > > },
> > > > "prices":
> > > > {
> > > > "histogram":
> > > > {
> > > > "field": "prices.min",
> > > > "interval": 1
> > > > }
> > > > }
> > > > }
> > > > 
> > > > ```
> > > > 
> > > > }  
> > > > }

---

<div class="post-metadata">

**Author:** ![Michael\_Korbakov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/michael_korbakov/32/3199_2.png) [@Michael\_Korbakov](https://discuss.elastic.co/u/Michael_Korbakov)\
**Post date:** [August 10, 2010, 11:00pm UTC](https://discuss.elastic.co/t/performance-problems/3205/8 "2010-08-10T23:00:25Z")

</div>

Ok, I've set both ES\_MIN\_MEM and ES\_MAX\_MEM variables to 10g.  
Performance increased to ~300 requests per second and I don't see any  
garbage collection notifications in logs. CPU load of worker nodes  
still very low -- only 20% at most. It there any other parameters that  
can be tuned? May be some cache or buffer sizes? I need to get 1000  
requests per second before starting move to production deployment.  
I can add more servers to the pool but I have a feeling that four  
quite powerful machines should have enough capacity for it.

On Aug 11, 1:22 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> Good, thats what I was concerned about. When you run a Java virtual machine,  
> you assign memory to it and it only consumes as much memory as you give it.  
> By default, it is set to use 1g max memory. Certainly, with your machine,  
> you can increase that quite significantly, I would say do 10G and see how it  
> goes (you want to leave memory also for file system cache, and too large  
> heaps can cause the JVM to hiccup). How to set the max memory is explained  
> here:[http://www.elasticsearch.com/docs/elasticsearch/setup/installation/](http://www.elasticsearch.com/docs/elasticsearch/setup/installation/).  
> For even better performance, set the minimum and the maximum to the same  
> value.
> 
> One more thing, with your setup, if you use 2 replicas to try and increase  
> the search performance, then 1 replicas should do. If you use 2 replicas to  
> increase the availability aspect, then thats fine.
> 
> One cool thing to check how the JVM is behaving is to use something like  
> visualvm to hook into it and check the memory consumption and GC activity.  
> All that information is already exposed in the node stats API, and once I  
> get around to build a nice management app for elasticsearch, it will be  
> exposed there through the REST API.
> 
> -shay.banon
> 
> On Wed, Aug 11, 2010 at 1:11 AM, rmihael [rmih...@gmail.com](mailto:rmih...@gmail.com) wrote:
> 
> > Typical 'total' is less then 100K items. Price ranging between 1 and  
> > 10, so I don't think it can cause much problems. I'll try to remove  
> > faceting now and see how it goes.  
> > I'm bit confused with your question about memory. I didn't assigned  
> > any memory to nodes, it just runs as it is. This kind of EC2 instance  
> > have 23GB of memory if you mean it.  
> > Data I've been uploading to index are very uniform. In fact they are  
> > randomly generated and should cause any kind of statistical explosion.  
> > But I've certainly got this fast degradation between 5M and 6M. Very  
> > strange, looks like something in my setup is very broken.
> 
> > BTW, I'm getting the following messages in log files:
> 
> > [14:11:14,211][INFO][monitor.memory.alpha] [Sangre] [5]  
> > [Full] Ran after [2] consecutive clean swipes, memory\_to\_clean  
> > [155.7mb], lower\_memory\_threshold [809mb], upper\_memory\_threshold  
> > [960.6mb], used\_memory [964.7mb], total\_memory[1011.2mb],  
> > max\_memory[1011.2mb]  
> > [14:52:06,191][INFO][monitor.memory.alpha] [Sangre] [6]  
> > [Full] Ran after [2] consecutive clean swipes, memory\_to\_clean  
> > [179.2mb], lower\_memory\_threshold [809mb], upper\_memory\_threshold  
> > [960.6mb], used\_memory [988.2mb], total\_memory[1011.2mb],  
> > max\_memory[1011.2mb]
> 
> > Is everything OK with it? "total\_memory[1011.2mb],  
> > max\_memory[1011.2mb]" part is confusing me, why it's so small?
> 
> > On Aug 11, 12:41 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > 
> > > Hi,
> 
> > > ```
> > > Is there a chance that the response that you get is really large? It
> > > 
> > > ```
> > > 
> > > seems like you are getting large result sets for the facets (not sure  
> > > about  
> > > the histogram facet of 1 for price, it depends on the range of it). Can  
> > > you  
> > > try and start with a simple query (no filters, no facets) and slowly add  
> > > more to the search request? How much memory do you assign each node? It  
> > > very strange that by moving from 5M docs to 6M docs, suddenly you get  
> > > such  
> > > different results, unless those 1M cause the facets to "explode" with the  
> > > data they return?
> 
> > > -shay.banon
> 
> > > On Wed, Aug 11, 2010 at 12:09 AM, rmihael [rmih...@gmail.com](mailto:rmih...@gmail.com) wrote:
> > > 
> > > > Hi everyone.
> 
> > > > I'm running elasticsearch-0.9 on the cluster of 5 EC2 instances  
> > > > (Cluster Compute Quadruple Extra Large Instance) with one of them  
> > > > configured as frontend (data: false). Index contains 10M of documents,  
> > > > 4 shards, 2 replicas. Total stored size is 23GB and each node has 23  
> > > > GB of RAM, so it fit's to memory without any problems.  
> > > > Using JMeter as testing tool I'm getting as little as ~150 requests  
> > > > per second (that's for 4 worker nodes with 8 CPU cores each). I've  
> > > > checked with iostat that there's no disk activity on cluster nodes. I  
> > > > also see CPU utilization near 5-10% on worker nodes during performance  
> > > > testing. Looks somewhat strange to me.
> 
> > > > I made performance measurements on the same cluster when index  
> > > > contained only 5M documents. I got nearly 1500 requests per second and  
> > > > CPU utilizations on worker nodes was close to 90%. After importing  
> > > > another million of documents performance started to degrade very  
> > > > rapidly.
> 
> > > > Could anyone help me with this problem? I'm completely out of ideas  
> > > > now.
> 
> > > > My queries are nothing complex: a single keyword search with a bunch  
> > > > of filter attached, also faceting by some fields. Typical query looks  
> > > > like this:
> 
> > > > {  
> > > > "query":  
> > > > {  
> > > > "filtered":  
> > > > {  
> > > > "query":  
> > > > {  
> > > > "query\_string":  
> > > > {  
> > > > "fields":  
> > > > [  
> > > > "keywords.original\_keywords^2",  
> > > > "keywords.keywords"  
> > > > ],  
> > > > "query": "bright"  
> > > > }  
> > > > },  
> > > > "filter":  
> > > > {  
> > > > "and":  
> > > > {  
> > > > "filters":  
> > > > [  
> > > > {  
> > > > "term":  
> > > > {  
> > > > "content.is\_offensive": false  
> > > > }  
> > > > },
> 
> > > > ```
> > > > {
> > > > "term":
> > > > {
> > > > "licenses.extended": true
> > > > }
> > > > },
> > > > 
> > > > ```
> 
> > > > ```
> > > > {
> > > > "term":
> > > > {
> > > > "image.isolated": false
> > > > }
> > > > },
> > > > 
> > > > ```
> 
> > > > ```
> > > > {
> > > > "term":
> > > > {
> > > > "content.orientation": "horizontal"
> > > > }
> > > > },
> > > > 
> > > > ```
> 
> > > > ```
> > > > {
> > > > "term":
> > > > {
> > > > "categories.conceptual.depth2": "793"
> > > > }
> > > > }
> > > > ]
> > > > }
> > > > }
> > > > },
> > > > "sort": "online.rating",
> > > > "facets":
> > > > {
> > > > "representative_categories":
> > > > {
> > > > "terms":
> > > > {
> > > > "field": "categories.representative.depth2",
> > > > "size": 100
> > > > }
> > > > },
> > > > "representative_categories":
> > > > {
> > > > "terms":
> > > > {
> > > > "field": "categories.conceptual.depth2",
> > > > "size": 100
> > > > }
> > > > },
> > > > "licenses":
> > > > {
> > > > "terms":
> > > > {
> > > > "field": "licenses.size",
> > > > "size": 100
> > > > }
> > > > },
> > > > "prices":
> > > > {
> > > > "histogram":
> > > > {
> > > > "field": "prices.min",
> > > > "interval": 1
> > > > }
> > > > }
> > > > }
> > > > 
> > > > ```
> > > > 
> > > > }  
> > > > }

---

<div class="post-metadata">

**Author:** ![Andrew\_Harvey\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrew_harvey_2/32/3328_2.png) [@Andrew\_Harvey\_2](https://discuss.elastic.co/u/Andrew_Harvey_2)\
**Post date:** [August 10, 2010, 11:02pm UTC](https://discuss.elastic.co/t/performance-problems/3205/9 "2010-08-10T23:02:47Z")

</div>

I'm intrigued by the choice to only have 4 shards across your 5 machines.

Surely at least one primary shard per core would see a performance increase in distributing the search load?

Andrew

On 11/08/2010, at 9:00 AM, rmihael wrote:

> Ok, I've set both ES\_MIN\_MEM and ES\_MAX\_MEM variables to 10g.  
> Performance increased to ~300 requests per second and I don't see any  
> garbage collection notifications in logs. CPU load of worker nodes  
> still very low -- only 20% at most. It there any other parameters that  
> can be tuned? May be some cache or buffer sizes? I need to get 1000  
> requests per second before starting move to production deployment.  
> I can add more servers to the pool but I have a feeling that four  
> quite powerful machines should have enough capacity for it.
> 
> On Aug 11, 1:22 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > Good, thats what I was concerned about. When you run a Java virtual machine,  
> > you assign memory to it and it only consumes as much memory as you give it.  
> > By default, it is set to use 1g max memory. Certainly, with your machine,  
> > you can increase that quite significantly, I would say do 10G and see how it  
> > goes (you want to leave memory also for file system cache, and too large  
> > heaps can cause the JVM to hiccup). How to set the max memory is explained  
> > here:[http://www.elasticsearch.com/docs/elasticsearch/setup/installation/](http://www.elasticsearch.com/docs/elasticsearch/setup/installation/).  
> > For even better performance, set the minimum and the maximum to the same  
> > value.
> > 
> > One more thing, with your setup, if you use 2 replicas to try and increase  
> > the search performance, then 1 replicas should do. If you use 2 replicas to  
> > increase the availability aspect, then thats fine.
> > 
> > One cool thing to check how the JVM is behaving is to use something like  
> > visualvm to hook into it and check the memory consumption and GC activity.  
> > All that information is already exposed in the node stats API, and once I  
> > get around to build a nice management app for elasticsearch, it will be  
> > exposed there through the REST API.
> > 
> > -shay.banon
> > 
> > On Wed, Aug 11, 2010 at 1:11 AM, rmihael [rmih...@gmail.com](mailto:rmih...@gmail.com) wrote:
> > 
> > > Typical 'total' is less then 100K items. Price ranging between 1 and  
> > > 10, so I don't think it can cause much problems. I'll try to remove  
> > > faceting now and see how it goes.  
> > > I'm bit confused with your question about memory. I didn't assigned  
> > > any memory to nodes, it just runs as it is. This kind of EC2 instance  
> > > have 23GB of memory if you mean it.  
> > > Data I've been uploading to index are very uniform. In fact they are  
> > > randomly generated and should cause any kind of statistical explosion.  
> > > But I've certainly got this fast degradation between 5M and 6M. Very  
> > > strange, looks like something in my setup is very broken.
> > 
> > > BTW, I'm getting the following messages in log files:
> > 
> > > [14:11:14,211][INFO][monitor.memory.alpha] [Sangre] [5]  
> > > [Full] Ran after [2] consecutive clean swipes, memory\_to\_clean  
> > > [155.7mb], lower\_memory\_threshold [809mb], upper\_memory\_threshold  
> > > [960.6mb], used\_memory [964.7mb], total\_memory[1011.2mb],  
> > > max\_memory[1011.2mb]  
> > > [14:52:06,191][INFO][monitor.memory.alpha] [Sangre] [6]  
> > > [Full] Ran after [2] consecutive clean swipes, memory\_to\_clean  
> > > [179.2mb], lower\_memory\_threshold [809mb], upper\_memory\_threshold  
> > > [960.6mb], used\_memory [988.2mb], total\_memory[1011.2mb],  
> > > max\_memory[1011.2mb]
> > 
> > > Is everything OK with it? "total\_memory[1011.2mb],  
> > > max\_memory[1011.2mb]" part is confusing me, why it's so small?
> > 
> > > On Aug 11, 12:41 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > 
> > > > Hi,
> > 
> > > > Is there a chance that the response that you get is really large? It  
> > > > seems like you are getting large result sets for the facets (not sure  
> > > > about  
> > > > the histogram facet of 1 for price, it depends on the range of it). Can  
> > > > you  
> > > > try and start with a simple query (no filters, no facets) and slowly add  
> > > > more to the search request? How much memory do you assign each node? It  
> > > > very strange that by moving from 5M docs to 6M docs, suddenly you get  
> > > > such  
> > > > different results, unless those 1M cause the facets to "explode" with the  
> > > > data they return?
> > 
> > > > -shay.banon
> > 
> > > > On Wed, Aug 11, 2010 at 12:09 AM, rmihael [rmih...@gmail.com](mailto:rmih...@gmail.com) wrote:
> > > > 
> > > > > Hi everyone.
> > 
> > > > > I'm running elasticsearch-0.9 on the cluster of 5 EC2 instances  
> > > > > (Cluster Compute Quadruple Extra Large Instance) with one of them  
> > > > > configured as frontend (data: false). Index contains 10M of documents,  
> > > > > 4 shards, 2 replicas. Total stored size is 23GB and each node has 23  
> > > > > GB of RAM, so it fit's to memory without any problems.  
> > > > > Using JMeter as testing tool I'm getting as little as ~150 requests  
> > > > > per second (that's for 4 worker nodes with 8 CPU cores each). I've  
> > > > > checked with iostat that there's no disk activity on cluster nodes. I  
> > > > > also see CPU utilization near 5-10% on worker nodes during performance  
> > > > > testing. Looks somewhat strange to me.
> > 
> > > > > I made performance measurements on the same cluster when index  
> > > > > contained only 5M documents. I got nearly 1500 requests per second and  
> > > > > CPU utilizations on worker nodes was close to 90%. After importing  
> > > > > another million of documents performance started to degrade very  
> > > > > rapidly.
> > 
> > > > > Could anyone help me with this problem? I'm completely out of ideas  
> > > > > now.
> > 
> > > > > My queries are nothing complex: a single keyword search with a bunch  
> > > > > of filter attached, also faceting by some fields. Typical query looks  
> > > > > like this:
> > 
> > > > > {  
> > > > > "query":  
> > > > > {  
> > > > > "filtered":  
> > > > > {  
> > > > > "query":  
> > > > > {  
> > > > > "query\_string":  
> > > > > {  
> > > > > "fields":  
> > > > > [  
> > > > > "keywords.original\_keywords^2",  
> > > > > "keywords.keywords"  
> > > > > ],  
> > > > > "query": "bright"  
> > > > > }  
> > > > > },  
> > > > > "filter":  
> > > > > {  
> > > > > "and":  
> > > > > {  
> > > > > "filters":  
> > > > > [  
> > > > > {  
> > > > > "term":  
> > > > > {  
> > > > > "content.is\_offensive": false  
> > > > > }  
> > > > > },
> > 
> > > > > ```
> > > > > {
> > > > > "term":
> > > > > {
> > > > > "licenses.extended": true
> > > > > }
> > > > > },
> > > > > 
> > > > > ```
> > 
> > > > > ```
> > > > > {
> > > > > "term":
> > > > > {
> > > > > "image.isolated": false
> > > > > }
> > > > > },
> > > > > 
> > > > > ```
> > 
> > > > > ```
> > > > > {
> > > > > "term":
> > > > > {
> > > > > "content.orientation": "horizontal"
> > > > > }
> > > > > },
> > > > > 
> > > > > ```
> > 
> > > > > ```
> > > > > {
> > > > > "term":
> > > > > {
> > > > > "categories.conceptual.depth2": "793"
> > > > > }
> > > > > }
> > > > > ]
> > > > > }
> > > > > }
> > > > > },
> > > > > "sort": "online.rating",
> > > > > "facets":
> > > > > {
> > > > > "representative_categories":
> > > > > {
> > > > > "terms":
> > > > > {
> > > > > "field": "categories.representative.depth2",
> > > > > "size": 100
> > > > > }
> > > > > },
> > > > > "representative_categories":
> > > > > {
> > > > > "terms":
> > > > > {
> > > > > "field": "categories.conceptual.depth2",
> > > > > "size": 100
> > > > > }
> > > > > },
> > > > > "licenses":
> > > > > {
> > > > > "terms":
> > > > > {
> > > > > "field": "licenses.size",
> > > > > "size": 100
> > > > > }
> > > > > },
> > > > > "prices":
> > > > > {
> > > > > "histogram":
> > > > > {
> > > > > "field": "prices.min",
> > > > > "interval": 1
> > > > > }
> > > > > }
> > > > > }
> > > > > 
> > > > > ```
> > > > > 
> > > > > }  
> > > > > }

## Andrew Harvey / Developer lexer m/ t/ +61 2 9019 6379 w/ [http://lexer.com.au](http://lexer.com.au) Help put an end to whaling. Visit [http://www.givewhalesavoice.com.au/](http://www.givewhalesavoice.com.au/)

Please consider the environment before printing this email  
This email transmission is confidential and intended solely for the person or organisation to whom it is addressed. If you are not the intended recipient, you must not copy, distribute or disseminate the information, or take any action in relation to it and please delete this e-mail. Any views expressed in this message are those of the individual sender, except where the send specifically states them to be the views of any organisation or employer. If you have received this message in error, do not open any attachment but please notify the sender (above). This message has been checked for all known viruses powered by McAfee.

For further information visit [McAfee AI-Powered Antivirus, Scam, Identity, and Privacy Protection](http://www.mcafee.com/us/threat_center/default.asp)  
Please rely on your own virus check as no responsibility is taken by the sender for any damage rising out of any virus infection this communication may contain.

This message has been scanned for malware by Websense. [www.websense.com](http://www.websense.com)

---

<div class="post-metadata">

**Author:** ![Michael\_Korbakov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/michael_korbakov/32/3199_2.png) [@Michael\_Korbakov](https://discuss.elastic.co/u/Michael_Korbakov)\
**Post date:** [August 10, 2010, 11:07pm UTC](https://discuss.elastic.co/t/performance-problems/3205/10 "2010-08-10T23:07:48Z")

</div>

Among my 5 machines only 4 acts as data nodes. One node is performing  
solely as frontend, as described in  
[http://www.elasticsearch.com/docs/elasticsearch/modules/node/data\_node/](http://www.elasticsearch.com/docs/elasticsearch/modules/node/data_node/).  
So I have one primary shard per each data node.  
Do you suggest that each CPU core should have it's own shard? I have  
64 cores in total on data nodes, wouldn't it be too much to make 64  
shards?

On Aug 11, 2:02 am, Andrew Harvey [Andrew.Har...@lexer.com.au](mailto:Andrew.Har...@lexer.com.au) wrote:

> I'm intrigued by the choice to only have 4 shards across your 5 machines.
> 
> Surely at least one primary shard per core would see a performance increase in distributing the search load?
> 
> Andrew
> 
> On 11/08/2010, at 9:00 AM, rmihael wrote:
> 
> > Ok, I've set both ES\_MIN\_MEM and ES\_MAX\_MEM variables to 10g.  
> > Performance increased to ~300 requests per second and I don't see any  
> > garbage collection notifications in logs. CPU load of worker nodes  
> > still very low -- only 20% at most. It there any other parameters that  
> > can be tuned? May be some cache or buffer sizes? I need to get 1000  
> > requests per second before starting move to production deployment.  
> > I can add more servers to the pool but I have a feeling that four  
> > quite powerful machines should have enough capacity for it.
> 
> > On Aug 11, 1:22 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > 
> > > Good, thats what I was concerned about. When you run a Java virtual machine,  
> > > you assign memory to it and it only consumes as much memory as you give it.  
> > > By default, it is set to use 1g max memory. Certainly, with your machine,  
> > > you can increase that quite significantly, I would say do 10G and see how it  
> > > goes (you want to leave memory also for file system cache, and too large  
> > > heaps can cause the JVM to hiccup). How to set the max memory is explained  
> > > here:[http://www.elasticsearch.com/docs/elasticsearch/setup/installation/](http://www.elasticsearch.com/docs/elasticsearch/setup/installation/).  
> > > For even better performance, set the minimum and the maximum to the same  
> > > value.
> 
> > > One more thing, with your setup, if you use 2 replicas to try and increase  
> > > the search performance, then 1 replicas should do. If you use 2 replicas to  
> > > increase the availability aspect, then thats fine.
> 
> > > One cool thing to check how the JVM is behaving is to use something like  
> > > visualvm to hook into it and check the memory consumption and GC activity.  
> > > All that information is already exposed in the node stats API, and once I  
> > > get around to build a nice management app for elasticsearch, it will be  
> > > exposed there through the REST API.
> 
> > > -shay.banon
> 
> > > On Wed, Aug 11, 2010 at 1:11 AM, rmihael [rmih...@gmail.com](mailto:rmih...@gmail.com) wrote:
> > > 
> > > > Typical 'total' is less then 100K items. Price ranging between 1 and  
> > > > 10, so I don't think it can cause much problems. I'll try to remove  
> > > > faceting now and see how it goes.  
> > > > I'm bit confused with your question about memory. I didn't assigned  
> > > > any memory to nodes, it just runs as it is. This kind of EC2 instance  
> > > > have 23GB of memory if you mean it.  
> > > > Data I've been uploading to index are very uniform. In fact they are  
> > > > randomly generated and should cause any kind of statistical explosion.  
> > > > But I've certainly got this fast degradation between 5M and 6M. Very  
> > > > strange, looks like something in my setup is very broken.
> 
> > > > BTW, I'm getting the following messages in log files:
> 
> > > > [14:11:14,211][INFO][monitor.memory.alpha] [Sangre] [5]  
> > > > [Full] Ran after [2] consecutive clean swipes, memory\_to\_clean  
> > > > [155.7mb], lower\_memory\_threshold [809mb], upper\_memory\_threshold  
> > > > [960.6mb], used\_memory [964.7mb], total\_memory[1011.2mb],  
> > > > max\_memory[1011.2mb]  
> > > > [14:52:06,191][INFO][monitor.memory.alpha] [Sangre] [6]  
> > > > [Full] Ran after [2] consecutive clean swipes, memory\_to\_clean  
> > > > [179.2mb], lower\_memory\_threshold [809mb], upper\_memory\_threshold  
> > > > [960.6mb], used\_memory [988.2mb], total\_memory[1011.2mb],  
> > > > max\_memory[1011.2mb]
> 
> > > > Is everything OK with it? "total\_memory[1011.2mb],  
> > > > max\_memory[1011.2mb]" part is confusing me, why it's so small?
> 
> > > > On Aug 11, 12:41 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > > 
> > > > > Hi,
> 
> > > > > Is there a chance that the response that you get is really large? It  
> > > > > seems like you are getting large result sets for the facets (not sure  
> > > > > about  
> > > > > the histogram facet of 1 for price, it depends on the range of it). Can  
> > > > > you  
> > > > > try and start with a simple query (no filters, no facets) and slowly add  
> > > > > more to the search request? How much memory do you assign each node? It  
> > > > > very strange that by moving from 5M docs to 6M docs, suddenly you get  
> > > > > such  
> > > > > different results, unless those 1M cause the facets to "explode" with the  
> > > > > data they return?
> 
> > > > > -shay.banon
> 
> > > > > On Wed, Aug 11, 2010 at 12:09 AM, rmihael [rmih...@gmail.com](mailto:rmih...@gmail.com) wrote:
> > > > > 
> > > > > > Hi everyone.
> 
> > > > > > I'm running elasticsearch-0.9 on the cluster of 5 EC2 instances  
> > > > > > (Cluster Compute Quadruple Extra Large Instance) with one of them  
> > > > > > configured as frontend (data: false). Index contains 10M of documents,  
> > > > > > 4 shards, 2 replicas. Total stored size is 23GB and each node has 23  
> > > > > > GB of RAM, so it fit's to memory without any problems.  
> > > > > > Using JMeter as testing tool I'm getting as little as ~150 requests  
> > > > > > per second (that's for 4 worker nodes with 8 CPU cores each). I've  
> > > > > > checked with iostat that there's no disk activity on cluster nodes. I  
> > > > > > also see CPU utilization near 5-10% on worker nodes during performance  
> > > > > > testing. Looks somewhat strange to me.
> 
> > > > > > I made performance measurements on the same cluster when index  
> > > > > > contained only 5M documents. I got nearly 1500 requests per second and  
> > > > > > CPU utilizations on worker nodes was close to 90%. After importing  
> > > > > > another million of documents performance started to degrade very  
> > > > > > rapidly.
> 
> > > > > > Could anyone help me with this problem? I'm completely out of ideas  
> > > > > > now.
> 
> > > > > > My queries are nothing complex: a single keyword search with a bunch  
> > > > > > of filter attached, also faceting by some fields. Typical query looks  
> > > > > > like this:
> 
> > > > > > {  
> > > > > > "query":  
> > > > > > {  
> > > > > > "filtered":  
> > > > > > {  
> > > > > > "query":  
> > > > > > {  
> > > > > > "query\_string":  
> > > > > > {  
> > > > > > "fields":  
> > > > > > [  
> > > > > > "keywords.original\_keywords^2",  
> > > > > > "keywords.keywords"  
> > > > > > ],  
> > > > > > "query": "bright"  
> > > > > > }  
> > > > > > },  
> > > > > > "filter":  
> > > > > > {  
> > > > > > "and":  
> > > > > > {  
> > > > > > "filters":  
> > > > > > [  
> > > > > > {  
> > > > > > "term":  
> > > > > > {  
> > > > > > "content.is\_offensive": false  
> > > > > > }  
> > > > > > },
> 
> > > > > > ```
> > > > > > {
> > > > > > "term":
> > > > > > {
> > > > > > "licenses.extended": true
> > > > > > }
> > > > > > },
> > > > > > 
> > > > > > ```
> 
> > > > > > ```
> > > > > > {
> > > > > > "term":
> > > > > > {
> > > > > > "image.isolated": false
> > > > > > }
> > > > > > },
> > > > > > 
> > > > > > ```
> 
> > > > > > ```
> > > > > > {
> > > > > > "term":
> > > > > > {
> > > > > > "content.orientation": "horizontal"
> > > > > > }
> > > > > > },
> > > > > > 
> > > > > > ```
> 
> > > > > > ```
> > > > > > {
> > > > > > "term":
> > > > > > {
> > > > > > "categories.conceptual.depth2": "793"
> > > > > > }
> > > > > > }
> > > > > > ]
> > > > > > }
> > > > > > }
> > > > > > },
> > > > > > "sort": "online.rating",
> > > > > > "facets":
> > > > > > {
> > > > > > "representative_categories":
> > > > > > {
> > > > > > "terms":
> > > > > > {
> > > > > > "field": "categories.representative.depth2",
> > > > > > "size": 100
> > > > > > }
> > > > > > },
> > > > > > "representative_categories":
> > > > > > {
> > > > > > "terms":
> > > > > > {
> > > > > > "field": "categories.conceptual.depth2",
> > > > > > "size": 100
> > > > > > }
> > > > > > },
> > > > > > "licenses":
> > > > > > {
> > > > > > "terms":
> > > > > > {
> > > > > > "field": "licenses.size",
> > > > > > "size": 100
> > > > > > }
> > > > > > },
> > > > > > "prices":
> > > > > > {
> > > > > > "histogram":
> > > > > > {
> > > > > > "field": "prices.min",
> > > > > > "interval": 1
> > > > > > }
> > > > > > }
> > > > > > }
> > > > > > 
> > > > > > ```
> > > > > > 
> > > > > > }  
> > > > > > }
> 
> ## Andrew Harvey / Developer lexer m/ t/ +61 2 9019 6379 w/ [http://lexer.com.au](http://lexer.com.au) Help put an end to whaling. Visithttp://www.givewhalesavoice.com.au/
> 
> Please consider the environment before printing this email  
> This email transmission is confidential and intended solely for the person or organisation to whom it is addressed. If you are not the intended recipient, you must not copy, distribute or disseminate the information, or take any action in relation to it and please delete this e-mail. Any views expressed in this message are those of the individual sender, except where the send specifically states them to be the views of any organisation or employer. If you have received this message in error, do not open any attachment but please notify the sender (above). This message has been checked for all known viruses powered by McAfee.
> 
> For further information visithttp://www.mcafee.com/us/threat\_center/default.asp  
> Please rely on your own virus check as no responsibility is taken by the sender for any damage rising out of any virus infection this communication may contain.
> 
> This message has been scanned for malware by [Websense.www.websense.com](http://Websense.www.websense.com)

---

<div class="post-metadata">

**Author:** ![Andrew\_Harvey\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrew_harvey_2/32/3328_2.png) [@Andrew\_Harvey\_2](https://discuss.elastic.co/u/Andrew_Harvey_2)\
**Post date:** [August 10, 2010, 11:11pm UTC](https://discuss.elastic.co/t/performance-problems/3205/11 "2010-08-10T23:11:21Z")

</div>

I'm not recommending anything at his point, but in my experience, increasing the shard count helped with performance. This was across 3 large instances on EC2. So long as you've got enough memory configured (and now you do), it's probably worth trying. Of course, with the old gateway format it caused issues with hitting the limit of S3 buckets, but if you're looking for a local production environment, I'd expect you're going to be using something like NFS for that anyway.

Andrew

On 11/08/2010, at 9:07 AM, rmihael wrote:

> Among my 5 machines only 4 acts as data nodes. One node is performing  
> solely as frontend, as described in  
> [http://www.elasticsearch.com/docs/elasticsearch/modules/node/data\_node/](http://www.elasticsearch.com/docs/elasticsearch/modules/node/data_node/).  
> So I have one primary shard per each data node.  
> Do you suggest that each CPU core should have it's own shard? I have  
> 64 cores in total on data nodes, wouldn't it be too much to make 64  
> shards?
> 
> On Aug 11, 2:02 am, Andrew Harvey [Andrew.Har...@lexer.com.au](mailto:Andrew.Har...@lexer.com.au) wrote:
> 
> > I'm intrigued by the choice to only have 4 shards across your 5 machines.
> > 
> > Surely at least one primary shard per core would see a performance increase in distributing the search load?
> > 
> > Andrew
> > 
> > On 11/08/2010, at 9:00 AM, rmihael wrote:
> > 
> > > Ok, I've set both ES\_MIN\_MEM and ES\_MAX\_MEM variables to 10g.  
> > > Performance increased to ~300 requests per second and I don't see any  
> > > garbage collection notifications in logs. CPU load of worker nodes  
> > > still very low -- only 20% at most. It there any other parameters that  
> > > can be tuned? May be some cache or buffer sizes? I need to get 1000  
> > > requests per second before starting move to production deployment.  
> > > I can add more servers to the pool but I have a feeling that four  
> > > quite powerful machines should have enough capacity for it.
> > 
> > > On Aug 11, 1:22 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > 
> > > > Good, thats what I was concerned about. When you run a Java virtual machine,  
> > > > you assign memory to it and it only consumes as much memory as you give it.  
> > > > By default, it is set to use 1g max memory. Certainly, with your machine,  
> > > > you can increase that quite significantly, I would say do 10G and see how it  
> > > > goes (you want to leave memory also for file system cache, and too large  
> > > > heaps can cause the JVM to hiccup). How to set the max memory is explained  
> > > > here:[http://www.elasticsearch.com/docs/elasticsearch/setup/installation/](http://www.elasticsearch.com/docs/elasticsearch/setup/installation/).  
> > > > For even better performance, set the minimum and the maximum to the same  
> > > > value.
> > 
> > > > One more thing, with your setup, if you use 2 replicas to try and increase  
> > > > the search performance, then 1 replicas should do. If you use 2 replicas to  
> > > > increase the availability aspect, then thats fine.
> > 
> > > > One cool thing to check how the JVM is behaving is to use something like  
> > > > visualvm to hook into it and check the memory consumption and GC activity.  
> > > > All that information is already exposed in the node stats API, and once I  
> > > > get around to build a nice management app for elasticsearch, it will be  
> > > > exposed there through the REST API.
> > 
> > > > -shay.banon
> > 
> > > > On Wed, Aug 11, 2010 at 1:11 AM, rmihael [rmih...@gmail.com](mailto:rmih...@gmail.com) wrote:
> > > > 
> > > > > Typical 'total' is less then 100K items. Price ranging between 1 and  
> > > > > 10, so I don't think it can cause much problems. I'll try to remove  
> > > > > faceting now and see how it goes.  
> > > > > I'm bit confused with your question about memory. I didn't assigned  
> > > > > any memory to nodes, it just runs as it is. This kind of EC2 instance  
> > > > > have 23GB of memory if you mean it.  
> > > > > Data I've been uploading to index are very uniform. In fact they are  
> > > > > randomly generated and should cause any kind of statistical explosion.  
> > > > > But I've certainly got this fast degradation between 5M and 6M. Very  
> > > > > strange, looks like something in my setup is very broken.
> > 
> > > > > BTW, I'm getting the following messages in log files:
> > 
> > > > > [14:11:14,211][INFO][monitor.memory.alpha] [Sangre] [5]  
> > > > > [Full] Ran after [2] consecutive clean swipes, memory\_to\_clean  
> > > > > [155.7mb], lower\_memory\_threshold [809mb], upper\_memory\_threshold  
> > > > > [960.6mb], used\_memory [964.7mb], total\_memory[1011.2mb],  
> > > > > max\_memory[1011.2mb]  
> > > > > [14:52:06,191][INFO][monitor.memory.alpha] [Sangre] [6]  
> > > > > [Full] Ran after [2] consecutive clean swipes, memory\_to\_clean  
> > > > > [179.2mb], lower\_memory\_threshold [809mb], upper\_memory\_threshold  
> > > > > [960.6mb], used\_memory [988.2mb], total\_memory[1011.2mb],  
> > > > > max\_memory[1011.2mb]
> > 
> > > > > Is everything OK with it? "total\_memory[1011.2mb],  
> > > > > max\_memory[1011.2mb]" part is confusing me, why it's so small?
> > 
> > > > > On Aug 11, 12:41 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > > > 
> > > > > > Hi,
> > 
> > > > > > Is there a chance that the response that you get is really large? It  
> > > > > > seems like you are getting large result sets for the facets (not sure  
> > > > > > about  
> > > > > > the histogram facet of 1 for price, it depends on the range of it). Can  
> > > > > > you  
> > > > > > try and start with a simple query (no filters, no facets) and slowly add  
> > > > > > more to the search request? How much memory do you assign each node? It  
> > > > > > very strange that by moving from 5M docs to 6M docs, suddenly you get  
> > > > > > such  
> > > > > > different results, unless those 1M cause the facets to "explode" with the  
> > > > > > data they return?
> > 
> > > > > > -shay.banon
> > 
> > > > > > On Wed, Aug 11, 2010 at 12:09 AM, rmihael [rmih...@gmail.com](mailto:rmih...@gmail.com) wrote:
> > > > > > 
> > > > > > > Hi everyone.
> > 
> > > > > > > I'm running elasticsearch-0.9 on the cluster of 5 EC2 instances  
> > > > > > > (Cluster Compute Quadruple Extra Large Instance) with one of them  
> > > > > > > configured as frontend (data: false). Index contains 10M of documents,  
> > > > > > > 4 shards, 2 replicas. Total stored size is 23GB and each node has 23  
> > > > > > > GB of RAM, so it fit's to memory without any problems.  
> > > > > > > Using JMeter as testing tool I'm getting as little as ~150 requests  
> > > > > > > per second (that's for 4 worker nodes with 8 CPU cores each). I've  
> > > > > > > checked with iostat that there's no disk activity on cluster nodes. I  
> > > > > > > also see CPU utilization near 5-10% on worker nodes during performance  
> > > > > > > testing. Looks somewhat strange to me.
> > 
> > > > > > > I made performance measurements on the same cluster when index  
> > > > > > > contained only 5M documents. I got nearly 1500 requests per second and  
> > > > > > > CPU utilizations on worker nodes was close to 90%. After importing  
> > > > > > > another million of documents performance started to degrade very  
> > > > > > > rapidly.
> > 
> > > > > > > Could anyone help me with this problem? I'm completely out of ideas  
> > > > > > > now.
> > 
> > > > > > > My queries are nothing complex: a single keyword search with a bunch  
> > > > > > > of filter attached, also faceting by some fields. Typical query looks  
> > > > > > > like this:
> > 
> > > > > > > {  
> > > > > > > "query":  
> > > > > > > {  
> > > > > > > "filtered":  
> > > > > > > {  
> > > > > > > "query":  
> > > > > > > {  
> > > > > > > "query\_string":  
> > > > > > > {  
> > > > > > > "fields":  
> > > > > > > [  
> > > > > > > "keywords.original\_keywords^2",  
> > > > > > > "keywords.keywords"  
> > > > > > > ],  
> > > > > > > "query": "bright"  
> > > > > > > }  
> > > > > > > },  
> > > > > > > "filter":  
> > > > > > > {  
> > > > > > > "and":  
> > > > > > > {  
> > > > > > > "filters":  
> > > > > > > [  
> > > > > > > {  
> > > > > > > "term":  
> > > > > > > {  
> > > > > > > "content.is\_offensive": false  
> > > > > > > }  
> > > > > > > },
> > 
> > > > > > > ```
> > > > > > > {
> > > > > > > "term":
> > > > > > > {
> > > > > > > "licenses.extended": true
> > > > > > > }
> > > > > > > },
> > > > > > > 
> > > > > > > ```
> > 
> > > > > > > ```
> > > > > > > {
> > > > > > > "term":
> > > > > > > {
> > > > > > > "image.isolated": false
> > > > > > > }
> > > > > > > },
> > > > > > > 
> > > > > > > ```
> > 
> > > > > > > ```
> > > > > > > {
> > > > > > > "term":
> > > > > > > {
> > > > > > > "content.orientation": "horizontal"
> > > > > > > }
> > > > > > > },
> > > > > > > 
> > > > > > > ```
> > 
> > > > > > > ```
> > > > > > > {
> > > > > > > "term":
> > > > > > > {
> > > > > > > "categories.conceptual.depth2": "793"
> > > > > > > }
> > > > > > > }
> > > > > > > ]
> > > > > > > }
> > > > > > > }
> > > > > > > },
> > > > > > > "sort": "online.rating",
> > > > > > > "facets":
> > > > > > > {
> > > > > > > "representative_categories":
> > > > > > > {
> > > > > > > "terms":
> > > > > > > {
> > > > > > > "field": "categories.representative.depth2",
> > > > > > > "size": 100
> > > > > > > }
> > > > > > > },
> > > > > > > "representative_categories":
> > > > > > > {
> > > > > > > "terms":
> > > > > > > {
> > > > > > > "field": "categories.conceptual.depth2",
> > > > > > > "size": 100
> > > > > > > }
> > > > > > > },
> > > > > > > "licenses":
> > > > > > > {
> > > > > > > "terms":
> > > > > > > {
> > > > > > > "field": "licenses.size",
> > > > > > > "size": 100
> > > > > > > }
> > > > > > > },
> > > > > > > "prices":
> > > > > > > {
> > > > > > > "histogram":
> > > > > > > {
> > > > > > > "field": "prices.min",
> > > > > > > "interval": 1
> > > > > > > }
> > > > > > > }
> > > > > > > }
> > > > > > > 
> > > > > > > ```
> > > > > > > 
> > > > > > > }  
> > > > > > > }
> > 
> > ## Andrew Harvey / Developer lexer m/ t/ +61 2 9019 6379 w/ [http://lexer.com.au](http://lexer.com.au) Help put an end to whaling. Visithttp://www.givewhalesavoice.com.au/
> > 
> > Please consider the environment before printing this email  
> > This email transmission is confidential and intended solely for the person or organisation to whom it is addressed. If you are not the intended recipient, you must not copy, distribute or disseminate the information, or take any action in relation to it and please delete this e-mail. Any views expressed in this message are those of the individual sender, except where the send specifically states them to be the views of any organisation or employer. If you have received this message in error, do not open any attachment but please notify the sender (above). This message has been checked for all known viruses powered by McAfee.
> > 
> > For further information visithttp://www.mcafee.com/us/threat\_center/default.asp  
> > Please rely on your own virus check as no responsibility is taken by the sender for any damage rising out of any virus infection this communication may contain.
> > 
> > This message has been scanned for malware by [Websense.www.websense.com](http://Websense.www.websense.com)

## Andrew Harvey / Developer lexer m/ t/ +61 2 9019 6379 w/ [http://lexer.com.au](http://lexer.com.au) Help put an end to whaling. Visit [http://www.givewhalesavoice.com.au/](http://www.givewhalesavoice.com.au/)

Please consider the environment before printing this email  
This email transmission is confidential and intended solely for the person or organisation to whom it is addressed. If you are not the intended recipient, you must not copy, distribute or disseminate the information, or take any action in relation to it and please delete this e-mail. Any views expressed in this message are those of the individual sender, except where the send specifically states them to be the views of any organisation or employer. If you have received this message in error, do not open any attachment but please notify the sender (above). This message has been checked for all known viruses powered by McAfee.

For further information visit [McAfee AI-Powered Antivirus, Scam, Identity, and Privacy Protection](http://www.mcafee.com/us/threat_center/default.asp)  
Please rely on your own virus check as no responsibility is taken by the sender for any damage rising out of any virus infection this communication may contain.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 11, 2010, 7:49am UTC](https://discuss.elastic.co/t/performance-problems/3205/12 "2010-08-11T07:49:19Z")

</div>

Can you follow what I suggested before, an first start with a simple query,  
and then slowly add filters and facets, and see when you see  
performance degradation?

-shay.banon

On Wed, Aug 11, 2010 at 2:00 AM, rmihael [rmihael@gmail.com](mailto:rmihael@gmail.com) wrote:

> Ok, I've set both ES\_MIN\_MEM and ES\_MAX\_MEM variables to 10g.  
> Performance increased to ~300 requests per second and I don't see any  
> garbage collection notifications in logs. CPU load of worker nodes  
> still very low -- only 20% at most. It there any other parameters that  
> can be tuned? May be some cache or buffer sizes? I need to get 1000  
> requests per second before starting move to production deployment.  
> I can add more servers to the pool but I have a feeling that four  
> quite powerful machines should have enough capacity for it.
> 
> On Aug 11, 1:22 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > Good, thats what I was concerned about. When you run a Java virtual  
> > machine,  
> > you assign memory to it and it only consumes as much memory as you give  
> > it.  
> > By default, it is set to use 1g max memory. Certainly, with your machine,  
> > you can increase that quite significantly, I would say do 10G and see how  
> > it  
> > goes (you want to leave memory also for file system cache, and too large  
> > heaps can cause the JVM to hiccup). How to set the max memory is  
> > explained  
> > here:[http://www.elasticsearch.com/docs/elasticsearch/setup/installation/](http://www.elasticsearch.com/docs/elasticsearch/setup/installation/)  
> > .  
> > For even better performance, set the minimum and the maximum to the same  
> > value.
> > 
> > One more thing, with your setup, if you use 2 replicas to try and  
> > increase  
> > the search performance, then 1 replicas should do. If you use 2 replicas  
> > to  
> > increase the availability aspect, then thats fine.
> > 
> > One cool thing to check how the JVM is behaving is to use something like  
> > visualvm to hook into it and check the memory consumption and GC  
> > activity.  
> > All that information is already exposed in the node stats API, and once I  
> > get around to build a nice management app for elasticsearch, it will be  
> > exposed there through the REST API.
> > 
> > -shay.banon
> > 
> > On Wed, Aug 11, 2010 at 1:11 AM, rmihael [rmih...@gmail.com](mailto:rmih...@gmail.com) wrote:
> > 
> > > Typical 'total' is less then 100K items. Price ranging between 1 and  
> > > 10, so I don't think it can cause much problems. I'll try to remove  
> > > faceting now and see how it goes.  
> > > I'm bit confused with your question about memory. I didn't assigned  
> > > any memory to nodes, it just runs as it is. This kind of EC2 instance  
> > > have 23GB of memory if you mean it.  
> > > Data I've been uploading to index are very uniform. In fact they are  
> > > randomly generated and should cause any kind of statistical explosion.  
> > > But I've certainly got this fast degradation between 5M and 6M. Very  
> > > strange, looks like something in my setup is very broken.
> > 
> > > BTW, I'm getting the following messages in log files:
> > 
> > > [14:11:14,211][INFO][monitor.memory.alpha] [Sangre] [5]  
> > > [Full] Ran after [2] consecutive clean swipes, memory\_to\_clean  
> > > [155.7mb], lower\_memory\_threshold [809mb], upper\_memory\_threshold  
> > > [960.6mb], used\_memory [964.7mb], total\_memory[1011.2mb],  
> > > max\_memory[1011.2mb]  
> > > [14:52:06,191][INFO][monitor.memory.alpha] [Sangre] [6]  
> > > [Full] Ran after [2] consecutive clean swipes, memory\_to\_clean  
> > > [179.2mb], lower\_memory\_threshold [809mb], upper\_memory\_threshold  
> > > [960.6mb], used\_memory [988.2mb], total\_memory[1011.2mb],  
> > > max\_memory[1011.2mb]
> > 
> > > Is everything OK with it? "total\_memory[1011.2mb],  
> > > max\_memory[1011.2mb]" part is confusing me, why it's so small?
> > 
> > > On Aug 11, 12:41 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > 
> > > > Hi,
> > 
> > > > ```
> > > > Is there a chance that the response that you get is really large?
> > > > 
> > > > ```
> 
> It
> 
> > > > seems like you are getting large result sets for the facets (not sure  
> > > > about  
> > > > the histogram facet of 1 for price, it depends on the range of it).  
> > > > Can  
> > > > you  
> > > > try and start with a simple query (no filters, no facets) and slowly  
> > > > add  
> > > > more to the search request? How much memory do you assign each node?  
> > > > It  
> > > > very strange that by moving from 5M docs to 6M docs, suddenly you get  
> > > > such  
> > > > different results, unless those 1M cause the facets to "explode" with  
> > > > the  
> > > > data they return?
> > 
> > > > -shay.banon
> > 
> > > > On Wed, Aug 11, 2010 at 12:09 AM, rmihael [rmih...@gmail.com](mailto:rmih...@gmail.com) wrote:
> > > > 
> > > > > Hi everyone.
> > 
> > > > > I'm running elasticsearch-0.9 on the cluster of 5 EC2 instances  
> > > > > (Cluster Compute Quadruple Extra Large Instance) with one of them  
> > > > > configured as frontend (data: false). Index contains 10M of  
> > > > > documents,  
> > > > > 4 shards, 2 replicas. Total stored size is 23GB and each node has  
> > > > > 23  
> > > > > GB of RAM, so it fit's to memory without any problems.  
> > > > > Using JMeter as testing tool I'm getting as little as ~150 requests  
> > > > > per second (that's for 4 worker nodes with 8 CPU cores each). I've  
> > > > > checked with iostat that there's no disk activity on cluster nodes.  
> > > > > I  
> > > > > also see CPU utilization near 5-10% on worker nodes during  
> > > > > performance  
> > > > > testing. Looks somewhat strange to me.
> > 
> > > > > I made performance measurements on the same cluster when index  
> > > > > contained only 5M documents. I got nearly 1500 requests per second  
> > > > > and  
> > > > > CPU utilizations on worker nodes was close to 90%. After importing  
> > > > > another million of documents performance started to degrade very  
> > > > > rapidly.
> > 
> > > > > Could anyone help me with this problem? I'm completely out of ideas  
> > > > > now.
> > 
> > > > > My queries are nothing complex: a single keyword search with a  
> > > > > bunch  
> > > > > of filter attached, also faceting by some fields. Typical query  
> > > > > looks  
> > > > > like this:
> > 
> > > > > {  
> > > > > "query":  
> > > > > {  
> > > > > "filtered":  
> > > > > {  
> > > > > "query":  
> > > > > {  
> > > > > "query\_string":  
> > > > > {  
> > > > > "fields":  
> > > > > [  
> > > > > "keywords.original\_keywords^2",  
> > > > > "keywords.keywords"  
> > > > > ],  
> > > > > "query": "bright"  
> > > > > }  
> > > > > },  
> > > > > "filter":  
> > > > > {  
> > > > > "and":  
> > > > > {  
> > > > > "filters":  
> > > > > [  
> > > > > {  
> > > > > "term":  
> > > > > {  
> > > > > "content.is\_offensive": false  
> > > > > }  
> > > > > },
> > 
> > > > > ```
> > > > > {
> > > > > "term":
> > > > > {
> > > > > "licenses.extended": true
> > > > > }
> > > > > },
> > > > > 
> > > > > ```
> > 
> > > > > ```
> > > > > {
> > > > > "term":
> > > > > {
> > > > > "image.isolated": false
> > > > > }
> > > > > },
> > > > > 
> > > > > ```
> > 
> > > > > ```
> > > > > {
> > > > > "term":
> > > > > {
> > > > > "content.orientation": "horizontal"
> > > > > }
> > > > > },
> > > > > 
> > > > > ```
> > 
> > > > > ```
> > > > > {
> > > > > "term":
> > > > > {
> > > > > "categories.conceptual.depth2":
> > > > > 
> > > > > ```
> 
> "793"
> 
> > > > > ```
> > > > > }
> > > > > }
> > > > > ]
> > > > > }
> > > > > }
> > > > > },
> > > > > "sort": "online.rating",
> > > > > "facets":
> > > > > {
> > > > > "representative_categories":
> > > > > {
> > > > > "terms":
> > > > > {
> > > > > "field": "categories.representative.depth2",
> > > > > "size": 100
> > > > > }
> > > > > },
> > > > > "representative_categories":
> > > > > {
> > > > > "terms":
> > > > > {
> > > > > "field": "categories.conceptual.depth2",
> > > > > "size": 100
> > > > > }
> > > > > },
> > > > > "licenses":
> > > > > {
> > > > > "terms":
> > > > > {
> > > > > "field": "licenses.size",
> > > > > "size": 100
> > > > > }
> > > > > },
> > > > > "prices":
> > > > > {
> > > > > "histogram":
> > > > > {
> > > > > "field": "prices.min",
> > > > > "interval": 1
> > > > > }
> > > > > }
> > > > > }
> > > > > 
> > > > > ```
> > > > > 
> > > > > }  
> > > > > }

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:20am UTC](https://discuss.elastic.co/t/performance-problems/3205/13 "2017-07-06T04:20:57Z")

</div>


