# Slow query for large size values

**URL:** <https://discuss.elastic.co/t/slow-query-for-large-size-values/188719>\
**Category:** Elasticsearch\
**Created:** [July 3, 2019, 1:19pm UTC](https://discuss.elastic.co/t/slow-query-for-large-size-values/188719 "2019-07-03T13:19:35Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Victor\_Guimaraes](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/victor_guimaraes/32/46749_2.png) [@Victor\_Guimaraes](https://discuss.elastic.co/u/Victor_Guimaraes)\
**Post date:** [July 3, 2019, 1:19pm UTC](https://discuss.elastic.co/t/slow-query-for-large-size-values/188719/1 "2019-07-03T13:19:35Z")

</div>

Hello,

I have an ElasticSearch 6.4 index with 5 shards, 1 replica for each shard and 1.5 billions of documents.

I use Elastic from textual search, getting the result's ids andusing then to query on MongoDB.

We need get a roof of 350.000 documents for each query, and my problem start here. Setting query size between 20 and 50.000 documents, my search take less then 10 seconds to respond with more or less 10MB of non compressed JSON. But when I increase the size, the time increase exponential, and I can't get the results.

Any suggestions to resolve my problem? I can accept queries with 30 seconds, but I need a roof of 350.000 documents.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 3, 2019, 1:36pm UTC](https://discuss.elastic.co/t/slow-query-for-large-size-values/188719/2 "2019-07-03T13:36:38Z")

</div>

Extracting a lot of hits from elasticsearch consumes time.  
Some few things you can do:

- don't fetch the `_source` as you don't need it
- use the scroll API. After 10000 documents, the `_search` will refuse to work
- or use the `search_after`
- sort by `_doc` if possible (depends on your use case)

Some questions: why do you need to load then 350 000 hits from MongoDB? What is the use case for that? Asking that because we can may be propose another alternative.

---

<div class="post-metadata">

**Author:** ![Victor\_Guimaraes](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/victor_guimaraes/32/46749_2.png) [@Victor\_Guimaraes](https://discuss.elastic.co/u/Victor_Guimaraes)\
**Post date:** [July 3, 2019, 1:53pm UTC](https://discuss.elastic.co/t/slow-query-for-large-size-values/188719/3 "2019-07-03T13:53:40Z")

</div>

Thank you for your reply.

I will test use the scroll API, maybe can help me.

I sort my documents descending for created date, so I can't sort by \_doc.

I need this number of documents because my data and the aggregations are stored and perform on MongoDB. And with empiric tests we saw that this magic number, 350.000, give us a good approximation for the user filters once the textual search on MongoDB are very bad. In the past, we were using Solr for textual search, but we migrate recently to ElasticSearch and we are facing this problem to obtain the same number of documents we received from Solr.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 3, 2019, 2:41pm UTC](https://discuss.elastic.co/t/slow-query-for-large-size-values/188719/4 "2019-07-03T14:41:22Z")

</div>

Why not just running the aggregation on elasticsearch side on the whole dataset instead of a subset?

---

<div class="post-metadata">

**Author:** ![Victor\_Guimaraes](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/victor_guimaraes/32/46749_2.png) [@Victor\_Guimaraes](https://discuss.elastic.co/u/Victor_Guimaraes)\
**Post date:** [July 3, 2019, 3:11pm UTC](https://discuss.elastic.co/t/slow-query-for-large-size-values/188719/5 "2019-07-03T15:11:36Z")

</div>

We use aggregations heavily and the MongoDB performance was better than Elastic for them. Besides that all the legacy and new softwares are using MongoDB, and adapt all of then to Elastic was expensive and impossible for us.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 3, 2019, 3:27pm UTC](https://discuss.elastic.co/t/slow-query-for-large-size-values/188719/6 "2019-07-03T15:27:29Z")

</div>

Ok. Then you can't sadly do miracles with that mixed strategy.  
Anyway, I hope that the small advices I gave you can help to reduce a bit the overall time needed.

> [@Victor\_Guimaraes](#):
>
> We use aggregations heavily and the MongoDB performance was better than Elastic for them.

I'd try now to compare a search in ES + fetching 350.000 partial results + computing the aggs on MongoDB side with a single run on elasticsearch side with `size: 0`...  
I'm pretty sure about the winning architecture but well, I understand the "legacy" concern.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 31, 2019, 3:27pm UTC](https://discuss.elastic.co/t/slow-query-for-large-size-values/188719/7 "2019-07-31T15:27:35Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
