# Large variance in query times

**URL:** <https://discuss.elastic.co/t/large-variance-in-query-times/12246>\
**Category:** Elasticsearch\
**Created:** [June 3, 2013, 10:47pm UTC](https://discuss.elastic.co/t/large-variance-in-query-times/12246 "2013-06-03T22:47:04Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![hazzadous](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hazzadous/32/1654_2.png) [@hazzadous](https://discuss.elastic.co/u/hazzadous)\
**Post date:** [June 3, 2013, 10:47pm UTC](https://discuss.elastic.co/t/large-variance-in-query-times/12246/1 "2013-06-03T22:47:04Z")

</div>

We're experiencing large variance in query times that I'm not sure how to  
diagnose. Our setup is as follows:

12 nodes (hexcore hyperthreaded, 64GB memory, 2x 3TB in RAID0 config)  
One index, 200 shards, 1 replica. ~20TB including replicas. ~160m docs.  
32GB JVM heap  
Indexing ~150 docs/s on average. Load ~1.5. Documents are a break

Aside from index and bulk threadpools (set to core counts, blocking)  
everything else is default.

Docs are ~as follows:

{  
"site": long,  
"countries": long,  
"text": {"standard": "string", "en": "string", "ru": "string", ... for  
all available analyzers, only indexed detected doc language},  
"publication\_date": date,  
... other longs and non analyzed terms  
}

Queries ~are:

{  
"query\_string": {"query": "...", "fields": ["text.standard", ...]},  
"facets": {  
"site": term facet,  
"countries": term facet,  
"publication\_date" histogram  
},  
"range\_filter": on publication date,  
"term\_filter": on sites and countries  
}

Currently queries take about 10-15 seconds, but ofter hit 75s (nginx  
timeout). We've had issues with failed merges resulting in shards that had  
huge segment counts. I use Lucene's CheckIndex to "-fix" these issues. I  
then \_optimised to 1 segment out of curiosity to see how much this affected  
performance. Search times were decreased to around the 1-5s mark. Great,  
but the process caused huge loads for about 6-8 hours. To try to keep the  
segment counts low I set optimize to run daily with a segment count of 3.  
Again there was a lot of instability.

I've keep details brief, assuming that the gist would probably highlight  
obvious wtf moments that will be highlighted. Really what I want to know,  
in no particular order, is:

(0. What sounds ridiculous in the above.)

1. Is it possible to get a breakdown of query execution (ie. took this  
long executing on shard x, it was merging at the time)
2. What's a good strategy for keeping segment count down:
  - Without killing the cluster. There are a lot of settings to  
throttle merges that sounds applicable, my concern is that just any merge  
is enough to cause massive query times.
  - Does this sound like something I should concentrate on?
  - perhaps more frequent merges will cause shorter freezes.
  - Is optimize something you should expect to have to run, or is there  
something wrong with the setup

3. Does the shard count sound "out there" for the doc count/size/etc
4. How do I optimize the heap/file system cache balance (when do I  
allocate more to the JVM, file system cache), does it sound like this would  
help?
5. How do other people go about profiling these types of issues

Details on request/interest.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![hazzadous](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hazzadous/32/1654_2.png) [@hazzadous](https://discuss.elastic.co/u/hazzadous)\
**Post date:** [June 4, 2013, 1:50pm UTC](https://discuss.elastic.co/t/large-variance-in-query-times/12246/2 "2013-06-04T13:50:20Z")

</div>

A few omission from the previous post:

- Running ubuntu 12.04 and OpenJDK (IcedTea 2.3.9).
- I'm using a refresh interval of 3600s although refresh is called  
explicitly when required, about every 10 mins.

On Monday, 3 June 2013 23:47:04 UTC+1, hazzadous wrote:

> We're experiencing large variance in query times that I'm not sure how to  
> diagnose. Our setup is as follows:
> 
> 12 nodes (hexcore hyperthreaded, 64GB memory, 2x 3TB in RAID0 config)  
> One index, 200 shards, 1 replica. ~20TB including replicas. ~160m docs.  
> 32GB JVM heap  
> Indexing ~150 docs/s on average. Load ~1.5. Documents are a break
> 
> Aside from index and bulk threadpools (set to core counts, blocking)  
> everything else is default.
> 
> Docs are ~as follows:
> 
> {  
> "site": long,  
> "countries": long,  
> "text": {"standard": "string", "en": "string", "ru": "string", ... for  
> all available analyzers, only indexed detected doc language},  
> "publication\_date": date,  
> ... other longs and non analyzed terms  
> }
> 
> Queries ~are:
> 
> {  
> "query\_string": {"query": "...", "fields": ["text.standard", ...]},  
> "facets": {  
> "site": term facet,  
> "countries": term facet,  
> "publication\_date" histogram  
> },  
> "range\_filter": on publication date,  
> "term\_filter": on sites and countries  
> }
> 
> Currently queries take about 10-15 seconds, but ofter hit 75s (nginx  
> timeout). We've had issues with failed merges resulting in shards that had  
> huge segment counts. I use Lucene's CheckIndex to "-fix" these issues. I  
> then \_optimised to 1 segment out of curiosity to see how much this affected  
> performance. Search times were decreased to around the 1-5s mark. Great,  
> but the process caused huge loads for about 6-8 hours. To try to keep the  
> segment counts low I set optimize to run daily with a segment count of 3.  
> Again there was a lot of instability.
> 
> I've keep details brief, assuming that the gist would probably highlight  
> obvious wtf moments that will be highlighted. Really what I want to know,  
> in no particular order, is:
> 
> (0. What sounds ridiculous in the above.)
> 
> 1. Is it possible to get a breakdown of query execution (ie. took this  
> long executing on shard x, it was merging at the time)
> 2. What's a good strategy for keeping segment count down:
> - Without killing the cluster. There are a lot of settings to  
> throttle merges that sounds applicable, my concern is that just any merge  
> is enough to cause massive query times.
> - Does this sound like something I should concentrate on?
> - perhaps more frequent merges will cause shorter freezes.
> - Is optimize something you should expect to have to run, or is there  
> something wrong with the setup
> 
> 3. Does the shard count sound "out there" for the doc count/size/etc
> 4. How do I optimize the heap/file system cache balance (when do I  
> allocate more to the JVM, file system cache), does it sound like this would  
> help?
> 5. How do other people go about profiling these types of issues
> 
> Details on request/interest.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![hazzadous](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hazzadous/32/1654_2.png) [@hazzadous](https://discuss.elastic.co/u/hazzadous)\
**Post date:** [June 4, 2013, 2:38pm UTC](https://discuss.elastic.co/t/large-variance-in-query-times/12246/3 "2013-06-04T14:38:58Z")

</div>

One more omission, running ES 0.20.4

On Tuesday, 4 June 2013 14:50:20 UTC+1, hazzadous wrote:

> A few omission from the previous post:
> 
> - Running ubuntu 12.04 and OpenJDK (IcedTea 2.3.9).
> - I'm using a refresh interval of 3600s although refresh is called  
> explicitly when required, about every 10 mins.
> 
> On Monday, 3 June 2013 23:47:04 UTC+1, hazzadous wrote:
> 
> > We're experiencing large variance in query times that I'm not sure how to  
> > diagnose. Our setup is as follows:
> > 
> > 12 nodes (hexcore hyperthreaded, 64GB memory, 2x 3TB in RAID0 config)  
> > One index, 200 shards, 1 replica. ~20TB including replicas. ~160m docs.  
> > 32GB JVM heap  
> > Indexing ~150 docs/s on average. Load ~1.5. Documents are a break
> > 
> > Aside from index and bulk threadpools (set to core counts, blocking)  
> > everything else is default.
> > 
> > Docs are ~as follows:
> > 
> > {  
> > "site": long,  
> > "countries": long,  
> > "text": {"standard": "string", "en": "string", "ru": "string", ... for  
> > all available analyzers, only indexed detected doc language},  
> > "publication\_date": date,  
> > ... other longs and non analyzed terms  
> > }
> > 
> > Queries ~are:
> > 
> > {  
> > "query\_string": {"query": "...", "fields": ["text.standard", ...]},  
> > "facets": {  
> > "site": term facet,  
> > "countries": term facet,  
> > "publication\_date" histogram  
> > },  
> > "range\_filter": on publication date,  
> > "term\_filter": on sites and countries  
> > }
> > 
> > Currently queries take about 10-15 seconds, but ofter hit 75s (nginx  
> > timeout). We've had issues with failed merges resulting in shards that had  
> > huge segment counts. I use Lucene's CheckIndex to "-fix" these issues. I  
> > then \_optimised to 1 segment out of curiosity to see how much this affected  
> > performance. Search times were decreased to around the 1-5s mark. Great,  
> > but the process caused huge loads for about 6-8 hours. To try to keep the  
> > segment counts low I set optimize to run daily with a segment count of 3.  
> > Again there was a lot of instability.
> > 
> > I've keep details brief, assuming that the gist would probably highlight  
> > obvious wtf moments that will be highlighted. Really what I want to know,  
> > in no particular order, is:
> > 
> > (0. What sounds ridiculous in the above.)
> > 
> > 1. Is it possible to get a breakdown of query execution (ie. took this  
> > long executing on shard x, it was merging at the time)
> > 2. What's a good strategy for keeping segment count down:
> > - Without killing the cluster. There are a lot of settings to  
> > throttle merges that sounds applicable, my concern is that just any merge  
> > is enough to cause massive query times.
> > - Does this sound like something I should concentrate on?
> > - perhaps more frequent merges will cause shorter freezes.
> > - Is optimize something you should expect to have to run, or is  
> > there something wrong with the setup
> > 
> > 3. Does the shard count sound "out there" for the doc count/size/etc
> > 4. How do I optimize the heap/file system cache balance (when do I  
> > allocate more to the JVM, file system cache), does it sound like this would  
> > help?
> > 5. How do other people go about profiling these types of issues
> > 
> > Details on request/interest.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![brian\_yoder](https://avatars.discourse-cdn.com/v4/letter/b/f1d935/32.png) [@brian\_yoder](https://discuss.elastic.co/u/brian_yoder)\
**Post date:** [June 4, 2013, 4:11pm UTC](https://discuss.elastic.co/t/large-variance-in-query-times/12246/4 "2013-06-04T16:11:35Z")

</div>

OpenJDK (IcedTea) != Java. Any time spent running ES on OpenJDK is time  
wasted.

Install and use real Oracle (aka Sun) Java. Works fine on Ubuntu too (which  
is the way I run ES on all systems: Mac, Ubuntu, and Solaris x86-64).

Brian

On Tuesday, June 4, 2013 9:50:20 AM UTC-4, hazzadous wrote:

> A few omission from the previous post:
> 
> - Running ubuntu 12.04 and OpenJDK (IcedTea 2.3.9).
> - I'm using a refresh interval of 3600s although refresh is called  
> explicitly when required, about every 10 mins.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![hazzadous](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hazzadous/32/1654_2.png) [@hazzadous](https://discuss.elastic.co/u/hazzadous)\
**Post date:** [June 4, 2013, 11:33pm UTC](https://discuss.elastic.co/t/large-variance-in-query-times/12246/5 "2013-06-04T23:33:56Z")

</div>

Thanks Brian will give it a try.

On Tuesday, 4 June 2013 17:11:35 UTC+1, InquiringMind wrote:

> OpenJDK (IcedTea) != Java. Any time spent running ES on OpenJDK is time  
> wasted.
> 
> Install and use real Oracle (aka Sun) Java. Works fine on Ubuntu too  
> (which is the way I run ES on all systems: Mac, Ubuntu, and Solaris x86-64).
> 
> Brian
> 
> On Tuesday, June 4, 2013 9:50:20 AM UTC-4, hazzadous wrote:
> 
> > A few omission from the previous post:
> > 
> > - Running ubuntu 12.04 and OpenJDK (IcedTea 2.3.9).
> > - I'm using a refresh interval of 3600s although refresh is called  
> > explicitly when required, about every 10 mins.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:32am UTC](https://discuss.elastic.co/t/large-variance-in-query-times/12246/6 "2017-07-06T02:32:58Z")

</div>


