# Search two slow

**URL:** <https://discuss.elastic.co/t/search-two-slow/155106>\
**Category:** Elasticsearch\
**Created:** [November 2, 2018, 3:35am UTC](https://discuss.elastic.co/t/search-two-slow/155106 "2018-11-02T03:35:58Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![willam\_boss](https://avatars.discourse-cdn.com/v4/letter/w/3ab097/32.png) [@willam\_boss](https://discuss.elastic.co/u/willam_boss)\
**Post date:** [November 2, 2018, 3:35am UTC](https://discuss.elastic.co/t/search-two-slow/155106/1 "2018-11-02T03:35:59Z")

</div>

My index use ２T space, I have 10billion document， three nodes ,each node 4 cpus 13G mem , my match search is quite slow , it seem read io is use up , how to increase my search speed?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 2, 2018, 5:01am UTC](https://discuss.elastic.co/t/search-two-slow/155106/2 "2018-11-02T05:01:14Z")

</div>

How many shards you have?  
Could you start new nodes?

May I suggest you look at the following resources about sizing:

[https://www.elastic.co/elasticon/conf/2016/sf/quantitative-cluster-sizing](https://www.elastic.co/elasticon/conf/2016/sf/quantitative-cluster-sizing)

> **[How many shards should I have in my Elasticsearch cluster?
	  	 | Elastic](https://www.elastic.co/blog/how-many-shards-should-i-have-in-my-elasticsearch-cluster)**
>
> Elasticsearch is a very versatile platform, that supports a variety of use cases, and provides great flexibility around data organisation and replication strategies. This flexibility can however somet...

> **[NetSecureDay: Managing your Black Friday Logs](https://speakerdeck.com/elastic/netsecureday-managing-your-black-friday-logs)**
>
> Surveiller une application complexe n’est pas une tâche aisée, mais avec les bons outils, ce n’est pas si sorcier. Néanmoins, des périodes fortes telles que les opérations de type « Black Friday » (Vendredi noir) ou période de Noël peuvent pousser...

And [https://www.elastic.co/webinars/using-rally-to-get-your-elasticsearch-cluster-size-right](https://www.elastic.co/webinars/using-rally-to-get-your-elasticsearch-cluster-size-right)

---

<div class="post-metadata">

**Author:** ![willam\_boss](https://avatars.discourse-cdn.com/v4/letter/w/3ab097/32.png) [@willam\_boss](https://discuss.elastic.co/u/willam_boss)\
**Post date:** [November 2, 2018, 5:29am UTC](https://discuss.elastic.co/t/search-two-slow/155106/3 "2018-11-02T05:29:20Z")

</div>

50 shards ,each shard use 46G space, about 3300 segments, I have no server for new node,do you have any suggestion?I think 10 billion document is the reason . there is 'hit' in result of restful query , if I want to only little records does elasticsearch search the whole index for count this 'hits'? and it only return 10 records each time , i think this is a waste of time , can elasticsearch remove 'hits' for speed up query ?  
`{ "took" : 61051, "timed_out" : false, "_shards" : { "total" : 50, "successful" : 50, "skipped" : 0, "failed" : 0 }, "hits" : { "total" : 2716157, "max_score" : 1.0,`  
like this hits total

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 2, 2018, 7:19am UTC](https://discuss.elastic.co/t/search-two-slow/155106/4 "2018-11-02T07:19:43Z")

</div>

What is the use case? Are you continuously indexing into the index? Are you updating documents in the index? If so, what is the indexing/update rate?

What type of storage do you have?

What is the query that returns the response you provided?

---

<div class="post-metadata">

**Author:** ![willam\_boss](https://avatars.discourse-cdn.com/v4/letter/w/3ab097/32.png) [@willam\_boss](https://discuss.elastic.co/u/willam_boss)\
**Post date:** [November 2, 2018, 8:04am UTC](https://discuss.elastic.co/t/search-two-slow/155106/5 "2018-11-02T08:04:42Z")

</div>

use for query match a filed with my own analyzer , I load all data once with bulk api, now I am not putting data into elastisearch , all my host are from azure cloud, storage should be ssd , read io almost 128M /second,query string like  
`"query":{"match":{"name":{"query":"ijmlajip","operator":"and"}}`

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 2, 2018, 8:09am UTC](https://discuss.elastic.co/t/search-two-slow/155106/6 "2018-11-02T08:09:25Z")

</div>

If you are not indexing or updating data, I would recommend you [force merge the index down to a single segment](https://www.elastic.co/guide/en/elasticsearch/reference/6.4/indices-forcemerge.html#indices-forcemerge). If you still seem limited by disk I/O, I would recommend looking into getting faster storage for the nodes, but I am not very familiar with what is available on Azure.

---

<div class="post-metadata">

**Author:** ![willam\_boss](https://avatars.discourse-cdn.com/v4/letter/w/3ab097/32.png) [@willam\_boss](https://discuss.elastic.co/u/willam_boss)\
**Post date:** [November 2, 2018, 8:12am UTC](https://discuss.elastic.co/t/search-two-slow/155106/7 "2018-11-02T08:12:38Z")

</div>

you mean this ?  
`curl -X POST "es01:9200/myindex/_forcemerge?only_expunge_deletes=false&max_num_segments=100&flush=true"`  
now my index have 50 shards ,each shard has 20 segment , if I merge them to one segments I will have a big segment which take 46 GB space . is it too large ? are you sure this could improve search ?

set max\_num\_segments=1?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 2, 2018, 8:15am UTC](https://discuss.elastic.co/t/search-two-slow/155106/8 "2018-11-02T08:15:54Z")

</div>

It should. There are also additional recommendations provided [here](https://www.elastic.co/guide/en/elasticsearch/reference/6.4/tune-for-search-speed.html).

---

<div class="post-metadata">

**Author:** ![willam\_boss](https://avatars.discourse-cdn.com/v4/letter/w/3ab097/32.png) [@willam\_boss](https://discuss.elastic.co/u/willam_boss)\
**Post date:** [November 2, 2018, 8:20am UTC](https://discuss.elastic.co/t/search-two-slow/155106/9 "2018-11-02T08:20:36Z")

</div>

how much speed will increase ? as you see my query take almost 60 seconds, If I merge all segments into one , can I compelete one simple query like  
`"query":{"match":{"name":{"query":"ijmlajip","operator":"and"}}`  
in 10 seconds?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 2, 2018, 8:27am UTC](https://discuss.elastic.co/t/search-two-slow/155106/10 "2018-11-02T08:27:30Z")

</div>

I do not know. As you have indicated storage likely being the limitation, I would recommend upgrading that. This is probably what will make the biggest difference. From what I have heard, Azure Premium Storage is recommended for I/O intensive workloads.

---

<div class="post-metadata">

**Author:** ![willam\_boss](https://avatars.discourse-cdn.com/v4/letter/w/3ab097/32.png) [@willam\_boss](https://discuss.elastic.co/u/willam_boss)\
**Post date:** [November 2, 2018, 8:36am UTC](https://discuss.elastic.co/t/search-two-slow/155106/11 "2018-11-02T08:36:04Z")

</div>

each host has a limit of read io for 128m /second , if I have more hosts ,will search query be better? for index which take 2T space ,how many host should be ok ? I have 10 billion small documents ,do you really think 10 billion is not the bottleneck of my query ? if so how can i improve this ??

 ![2018-11-02%2014%3A22%3A08%E5%B1%8F%E5%B9%95%E6%88%AA%E5%9B%BE](https://us1.discourse-cdn.com/elastic/original/3X/7/2/72a25343a1db2fc6636e46f82e29df9d10bd1958.png)

this is a profile of my query

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 2, 2018, 8:40am UTC](https://discuss.elastic.co/t/search-two-slow/155106/12 "2018-11-02T08:40:25Z")

</div>

Elasticsearch typically performs a lot of small random reads during querying rather than large sequential ones, so fast storage that is able to handle this type of load is essential. Many throughput metrics for disks assume sequential reads, so might not be representative.

You need to test and benchmark. Have a look at the links David provided.

What type of Azure storage are you using at the moment?

---

<div class="post-metadata">

**Author:** ![willam\_boss](https://avatars.discourse-cdn.com/v4/letter/w/3ab097/32.png) [@willam\_boss](https://discuss.elastic.co/u/willam_boss)\
**Post date:** [November 2, 2018, 8:46am UTC](https://discuss.elastic.co/t/search-two-slow/155106/13 "2018-11-02T08:46:19Z")

</div>

![2018-11-02%2016%3A45%3A50%E5%B1%8F%E5%B9%95%E6%88%AA%E5%9B%BE](https://us1.discourse-cdn.com/elastic/original/3X/0/2/0201845bb6e7b40e9fd47545a71caae9b727f695.png)

> <https://unix.stackexchange.com/questions/65595/how-to-know-if-a-disk-is-an-ssd-or-an-hdd>

time for i in `seq 1 1000`; do

> ```
> dd bs=4k if=/dev/sdd count=1 skip=$(( $RANDOM * 128 )) >/dev/null 2>&1;
> 
> ```
> 
> done

real 0m1.931s  
user 0m0.622s  
sys 0m1.386s

the result show is ssd

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 2, 2018, 8:48am UTC](https://discuss.elastic.co/t/search-two-slow/155106/14 "2018-11-02T08:48:33Z")

</div>

As I am not very familiar with Azure, that does not really tell me anything.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 2, 2018, 9:10am UTC](https://discuss.elastic.co/t/search-two-slow/155106/15 "2018-11-02T09:10:21Z")

</div>

I had a look at the [Azure documentation about storage options](https://docs.microsoft.com/en-us/azure/virtual-machines/linux/about-disks-and-vhds). For Standard SSD disks (if that is what you are using) they state:

> Standard SSD disks combine elements of Premium SSD disks and Standard HDD disks to form a cost-effective solution best suited for applications like web servers that do not need high IOPS on disks.

Elasticsearch definitely require high IOPS, which is why Premium SSD disks are recommended.

---

<div class="post-metadata">

**Author:** ![willam\_boss](https://avatars.discourse-cdn.com/v4/letter/w/3ab097/32.png) [@willam\_boss](https://discuss.elastic.co/u/willam_boss)\
**Post date:** [November 2, 2018, 9:28am UTC](https://discuss.elastic.co/t/search-two-slow/155106/16 "2018-11-02T09:28:00Z")

</div>

merge segment quite slowly , can I reset some config to speed up this merge? just like `"index.refresh_interval": "-1", "index.translog.durability": "async", "index.translog.sync_interval": "60s"`  
?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 2, 2018, 9:36am UTC](https://discuss.elastic.co/t/search-two-slow/155106/17 "2018-11-02T09:36:14Z")

</div>

No, I do not think you can speed this up as it is quite I/O intensive.

---

<div class="post-metadata">

**Author:** ![willam\_boss](https://avatars.discourse-cdn.com/v4/letter/w/3ab097/32.png) [@willam\_boss](https://discuss.elastic.co/u/willam_boss)\
**Post date:** [November 5, 2018, 2:05am UTC](https://discuss.elastic.co/t/search-two-slow/155106/18 "2018-11-05T02:05:24Z")

</div>

thanks a lot ,your suggestions really help me ,my query response in ten seconds now

---

<div class="post-metadata">

**Author:** ![DidierB](https://avatars.discourse-cdn.com/v4/letter/d/ed8c4c/32.png) [@DidierB](https://discuss.elastic.co/u/DidierB)\
**Post date:** [November 8, 2018, 10:43am UTC](https://discuss.elastic.co/t/search-two-slow/155106/19 "2018-11-08T10:43:51Z")

</div>

Just passing by. Is it correct to say, to summarize, that the main improvement is to merge everything into one segment but it only works if you don't update the index?

In this case, if you have a realtime usecase and you make one index per day, it's actually a good idea to do a one-segment merge for all the past days indexes so the requests that spans across many indexes are faster?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 8, 2018, 12:17pm UTC](https://discuss.elastic.co/t/search-two-slow/155106/20 "2018-11-08T12:17:21Z")

</div>

> [@DidierB](#):
>
> In this case, if you have a realtime usecase and you make one index per day, it's actually a good idea to do a one-segment merge for all the past days indexes so the requests that spans across many indexes are faster?

Yes, it can make searches faster, but also result in lower heap usage due to reduced need for global ordinals as outlined in [this webinar](https://www.elastic.co/webinars/optimizing-storage-efficiency-in-elasticsearch).

[Next page](https://discuss.elastic.co/t/search-two-slow/155106.md?page=2)
