# Fetch million doc in seconds

**URL:** <https://discuss.elastic.co/t/fetch-million-doc-in-seconds/239611>\
**Category:** Elasticsearch\
**Created:** [July 2, 2020, 10:35am UTC](https://discuss.elastic.co/t/fetch-million-doc-in-seconds/239611 "2020-07-02T10:35:34Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![saurabhbhatt](https://avatars.discourse-cdn.com/v4/letter/s/97f17d/32.png) [@saurabhbhatt](https://discuss.elastic.co/u/saurabhbhatt)\
**Post date:** [July 2, 2020, 10:35am UTC](https://discuss.elastic.co/t/fetch-million-doc-in-seconds/239611/1 "2020-07-02T10:35:34Z")

</div>

Hi,

To give brief, I have an index "allreportingdataindex" (example) and it has over 7,50,000 records. Each document has about about 20 columns and about 15 columns with array values. The ES instance is installed on server and we access it over HTTP.

Now when I try to do "matchall" it takes about 17 minutes to get all data. If I only try to get 1 column of ID/number type, it takes 8 minutes. I need to fetch all the data within 5 seconds. Is this possible? What do I do?

Please help! 🙂

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 2, 2020, 10:47am UTC](https://discuss.elastic.co/t/fetch-million-doc-in-seconds/239611/2 "2020-07-02T10:47:07Z")

</div>

Not sure you can do it. May be it also depends on your hardware (SSD) and the network?

But you can use:

- the [`size` and `from` parameters](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-from-size.html) to display by default up to 10000 records to your users. If you want to change this limit, you can change `index.max_result_window` setting but be aware of the consequences (ie memory).
- the [search after](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-search-after.html) feature to do deep pagination.
- the [Scroll API](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-scroll.html) if you want to extract a resultset to be consumed by another tool later.

What did you do so far?

---

<div class="post-metadata">

**Author:** ![saurabhbhatt](https://avatars.discourse-cdn.com/v4/letter/s/97f17d/32.png) [@saurabhbhatt](https://discuss.elastic.co/u/saurabhbhatt)\
**Post date:** [July 2, 2020, 10:55am UTC](https://discuss.elastic.co/t/fetch-million-doc-in-seconds/239611/3 "2020-07-02T10:55:37Z")

</div>

So this is basically a Reporting project, so I will need all the data upfront and pass it on client side to datatable js. I am already using scroll API and I am getting all data but my main requirement is to get it faster.

Will using SSD make big difference to performance? Also this ES instance is hosted on VM.

Please suggest some way

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 2, 2020, 11:00am UTC](https://discuss.elastic.co/t/fetch-million-doc-in-seconds/239611/4 "2020-07-02T11:00:34Z")

</div>

If you can make sure the full index fits in the OS page cache you may also see an improvement in performance.

---

<div class="post-metadata">

**Author:** ![defalt](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/defalt/32/71379_2.png) [@defalt](https://discuss.elastic.co/u/defalt)\
**Post date:** [July 2, 2020, 1:04pm UTC](https://discuss.elastic.co/t/fetch-million-doc-in-seconds/239611/5 "2020-07-02T13:04:14Z")

</div>

Yes, an SSD makes a big difference.

---

<div class="post-metadata">

**Author:** ![saurabhbhatt](https://avatars.discourse-cdn.com/v4/letter/s/97f17d/32.png) [@saurabhbhatt](https://discuss.elastic.co/u/saurabhbhatt)\
**Post date:** [July 2, 2020, 2:56pm UTC](https://discuss.elastic.co/t/fetch-million-doc-in-seconds/239611/6 "2020-07-02T14:56:54Z")

</div>

But this data keeps updating. atleast twice a day. Also will caching this into OS page cache affect the server throughput/performance?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 2, 2020, 3:10pm UTC](https://discuss.elastic.co/t/fetch-million-doc-in-seconds/239611/7 "2020-07-02T15:10:22Z")

</div>

If you cannot cache it all you need very fast disks. I doubt you will get down to seconds though..

---

<div class="post-metadata">

**Author:** ![Vinayak\_Sapre](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vinayak_sapre/32/45939_2.png) [@Vinayak\_Sapre](https://discuss.elastic.co/u/Vinayak_Sapre)\
**Post date:** [July 3, 2020, 5:52am UTC](https://discuss.elastic.co/t/fetch-million-doc-in-seconds/239611/8 "2020-07-03T05:52:07Z")

</div>

Saurabh,  
What is the size of data (match\_all response or index) in bytes?

---

<div class="post-metadata">

**Author:** ![saurabhbhatt](https://avatars.discourse-cdn.com/v4/letter/s/97f17d/32.png) [@saurabhbhatt](https://discuss.elastic.co/u/saurabhbhatt)\
**Post date:** [July 3, 2020, 12:04pm UTC](https://discuss.elastic.co/t/fetch-million-doc-in-seconds/239611/9 "2020-07-03T12:04:35Z")

</div>

@Vinayak_Sapre Hi!

So the data is as below:  
813 MB = 813000000 Bytes (in decimal)  
813 MB = 852492288 Bytes (in binary)

---

<div class="post-metadata">

**Author:** ![Vinayak\_Sapre](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vinayak_sapre/32/45939_2.png) [@Vinayak\_Sapre](https://discuss.elastic.co/u/Vinayak_Sapre)\
**Post date:** [July 3, 2020, 5:15pm UTC](https://discuss.elastic.co/t/fetch-million-doc-in-seconds/239611/10 "2020-07-03T17:15:21Z")

</div>

Saurabh,

I did some experiments on my laptop. To put findings in context here is my setup

1. Single node cluster with 8GB heap
2. 20 indices / 36 shards / 15GB total data
3. Specific index used for experiment: 2 shards / 6.5M docs / 1.37GB index size / best\_compression codec / max\_result\_window = 1000,000
4. Nothing else was querying / ingesting during tests

match\_all query timing increase linearly with size parameter  
size = 50K took 2.4 seconds  
size = 100K took 4.6 seconds  
size = 200K took 9.2 seconds

msearch query with 2 match\_all queries with preference set to \_shards:0 and \_shards:1. Timing increased linearly with size. But timing remained comparable to single match\_all in previous test.  
size = 50K for each query, took 2.5 seconds (100K docs)  
size = 100K took 5.15 seconds (200K docs)  
size = 200K took 10.1 seconds (400K docs and response size 250MB)

- With index size \< 1GB, setting index.max\_result\_window to a high value as David suggested will reduce your round trips.
- msearch will allow you to run multiple queries concurrently.
- Multiple queries for msearch can be constructed by shards or scroll with slices or a natural partitioning key like timestamp
- If you have multiple nodes in the cluster, more shards will be a better choice.
- Instead of msearch, you can run same queries using multiple threads in your app. This will allow you to utilize multiple client nodes and client node will not have to aggregate all results.
- Since you are fetching all fields, retrieving \_source may be better than fetching 20 doc values.

---

<div class="post-metadata">

**Author:** ![Matthew\_Adams](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/matthew_adams/32/65554_2.png) [@Matthew\_Adams](https://discuss.elastic.co/u/Matthew_Adams)\
**Post date:** [July 6, 2020, 1:25pm UTC](https://discuss.elastic.co/t/fetch-million-doc-in-seconds/239611/11 "2020-07-06T13:25:00Z")

</div>

I really don't understand why you would try and do this and I don't see any way for this to be possible. You are not using any query, simply returning a whole copy of 800 MB of data and you want it to complete in 5s. Can you actually transfer 800 MB in 5s across the network?

If you are feeding a UI, I suggest you use a scroll and the javascript loads the data bit by bit as required e.g. as the user scrolls. I can't imaging the UI wants 800 MB of data to hold in memory either. You could definitely return a 'page' of results in 5 s.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 3, 2020, 1:30pm UTC](https://discuss.elastic.co/t/fetch-million-doc-in-seconds/239611/12 "2020-08-03T13:30:10Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
