# Significance of two phases - "query then fetch" with default number of shards as 1

**URL:** <https://discuss.elastic.co/t/significance-of-two-phases-query-then-fetch-with-default-number-of-shards-as-1/233789>\
**Category:** Elasticsearch\
**Created:** [May 21, 2020, 6:22pm UTC](https://discuss.elastic.co/t/significance-of-two-phases-query-then-fetch-with-default-number-of-shards-as-1/233789 "2020-05-21T18:22:08Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![nages](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nages/32/46943_2.png) [@nages](https://discuss.elastic.co/u/nages)\
**Post date:** [May 21, 2020, 6:22pm UTC](https://discuss.elastic.co/t/significance-of-two-phases-query-then-fetch-with-default-number-of-shards-as-1/233789/1 "2020-05-21T18:22:08Z")

</div>

Is there any significance for search to be executed in two phases - "query then fetch" when the default number of shards as 1 ( starting from 7.x ) ? leaving the cases of considering replicas

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 21, 2020, 7:01pm UTC](https://discuss.elastic.co/t/significance-of-two-phases-query-then-fetch-with-default-number-of-shards-as-1/233789/2 "2020-05-21T19:01:06Z")

</div>

That is probably the only scenario where executing the query in two phases may not bring a lot of added benefits. I believe this is a quite rare scenario and likely one that generally performs quite well anyway, so I don't think it adds much overhead either. Optimizing this would thesefore likely bring very little benefit, but make the code more complex and difficult to maintain, which IMHO seems like a bad tradeoff.

---

<div class="post-metadata">

**Author:** ![nages](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nages/32/46943_2.png) [@nages](https://discuss.elastic.co/u/nages)\
**Post date:** [May 22, 2020, 2:03am UTC](https://discuss.elastic.co/t/significance-of-two-phases-query-then-fetch-with-default-number-of-shards-as-1/233789/3 "2020-05-22T02:03:24Z")

</div>

well. if we see 7.x, default number of shards is '1' , this decision was taken considering majority deployments were small / deployments with defaults. On the same note, does not it make sense to disable the strategy of doing aggregation in multiple places ( at least , once in data node and once in coordinating node) in the default scenarios

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 22, 2020, 2:59am UTC](https://discuss.elastic.co/t/significance-of-two-phases-query-then-fetch-with-default-number-of-shards-as-1/233789/4 "2020-05-22T02:59:52Z")

</div>

I do not think you would gain much as the same amount of work still need to be done, so it would add complexity for virtually no gain. If you have a small cluster it is also likely that the node serving the request also holds the shard which means there is not even a network hop to avoid in most cases, especially if you also use suitable preference setting.

I do not think this quite theoretical discussion around a very rare case which is usually very fast anyway is very useful so will leave it.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 19, 2020, 2:59am UTC](https://discuss.elastic.co/t/significance-of-two-phases-query-then-fetch-with-default-number-of-shards-as-1/233789/5 "2020-06-19T02:59:54Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
