# Acceptable Search Performance

**URL:** <https://discuss.elastic.co/t/acceptable-search-performance/12619>\
**Category:** Elasticsearch\
**Created:** [July 2, 2013, 2:57pm UTC](https://discuss.elastic.co/t/acceptable-search-performance/12619 "2013-07-02T14:57:24Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Devashish\_Tyagi](https://avatars.discourse-cdn.com/v4/letter/d/94ad74/32.png) [@Devashish\_Tyagi](https://discuss.elastic.co/u/Devashish_Tyagi)\
**Post date:** [July 2, 2013, 2:57pm UTC](https://discuss.elastic.co/t/acceptable-search-performance/12619/1 "2013-07-02T14:57:24Z")

</div>

I am using elasticsearch for indexing and searching around 6 million HTML  
documents. In order to do that I created an index with 10 shards. Earlier,  
I was running just one node on my physical machine (details below). I  
allocated 9 GB heap size to Elastic Search. I am performing just simple  
search queries (nothing fancy). You can see a typical search query here[https://gist.github.com/devashishtyagi/5909747](https://gist.github.com/devashishtyagi/5909747).  
The average response size from elasticsearch is ~80 KB. In order to test  
the performance of elasticsearch, I created an Apache Jmeter test. The test  
would read random words from a search term file I supplied and fetch back  
the response from the elasticsearch. The Jmeter tests were performed from a  
separate machine but located close by (so not much of network overhead).  
This is the result I got

1. _10 threads, 50 requests per thread_ - 2.3 QPS and average response  
time of \> 4 sec.
2. \*5 threads, 100 request per thread \*- 3.4 QPS and average response  
time of \> 1sec

Here are some of my index statistics

- Number of shards - 10
- Number of documents - 5174688
- Size of index - 56 GB
- Size of a typical shard - 5.5 GB
- Number of replicas - 0

My machine configuration

Amazon EC m3.xlarge

- RAM - 15 GB
- Compute Units - 13
- Hard Drive - 1 TB EBS Drive

I went through several search performance related mails on the group and it  
feels like that I am getting subpar performance. Or is it an acceptable  
search performance ?

During my tests I found out that elasticsearch was getting bottle necked on  
disk I/O. So I added 3 more EBS drives to the same machine and started up 3  
new elasticsearch nodes on same machine. So now I had 4 elasticsearch nodes  
running on the same server. Here the performance test results with this  
configuration

1. _10 threads, 50 requests per thread_ - 11.6 QPS and average response  
time of ~ 842 ms.
2. \*7 threads, 100 request per thread \*- 13.4 QPS and average response  
time of ~ 512 ms.
3. _8 threads, 100 requests per thread_ - 14.3 QPS and average response  
time of ~ 550 ms.

Although this seems like a huge improvement but with 4 drives too  
elasticsearch is getting bottle necked on Disk I/O. Is this is expected ?

P.S. I have come across various posts where it is mentioned that routing  
greatly improves performance but I have no idea how to use that in my use  
case.

Thanks in advance,  
Devashish Tyagi

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [July 2, 2013, 3:31pm UTC](https://discuss.elastic.co/t/acceptable-search-performance/12619/2 "2013-07-02T15:31:15Z")

</div>

Quick analysis

- you have highlighting on. This feature tends to work heavily on disk  
(doc fetching)
- you have size =100. This is an extreme setting. It creates additional  
burden for highlighting.
- to optimize highlighting, there are options that you do not use yet  
(hint: term vector, fast vector highlighter)  
[http://www.elasticsearch.org/guide/reference/api/search/highlighting/](http://www.elasticsearch.org/guide/reference/api/search/highlighting/)
- do not ramp up more than one node per machine, there is not much sense  
in it
- EBS drives are known to be slow (they go over 1Gbit network channels),  
ES is built to scale over machines, not only number of drives, so use  
more machines
- and, finally, use "query" instead of "filtered query" in the query  
unless you know what you want to test (you simply thrash your very large  
filter cache when load testing, which is bad for overall performance)

Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:28am UTC](https://discuss.elastic.co/t/acceptable-search-performance/12619/3 "2017-07-06T02:28:38Z")

</div>


