# Retrieving millions of large documents

**URL:** <https://discuss.elastic.co/t/retrieving-millions-of-large-documents/341803>\
**Category:** Elasticsearch\
**Created:** [August 28, 2023, 12:19pm UTC](https://discuss.elastic.co/t/retrieving-millions-of-large-documents/341803 "2023-08-28T12:19:24Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Tomer\_Avira](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tomer_avira/32/124618_2.png) [@Tomer\_Avira](https://discuss.elastic.co/u/Tomer_Avira)\
**Post date:** [August 28, 2023, 12:19pm UTC](https://discuss.elastic.co/t/retrieving-millions-of-large-documents/341803/1 "2023-08-28T12:19:24Z")

</div>

Hello everyone, first time i am requesting your help.

I am working with elastic version 8.5.3 with java client 7.17.1, let me represent you with the problem I'm having.

I have daily indices with the largest of them holding unique data documents, the amount of documents I have are in the millions (about 15 million a day), with size of the documents reaching to tens of Gigabyte's.

For my use case the documents are timed and I need to aquire between times (sort of a recording and reviewing upon request sort of use case), I know this isn't the intended use of elastic but it's what we've got.

Untill now I have used the pagination process in order to retrieve all the records, and it worked fine untill the data set became too large, and because the web client is using the ping pong method (scroll and return) it can't really be multi threaded or splitting it to multiple identical services with load balancing.

I need to find a way to make the process quicker, there is also post processing after the retrieval of the documents.

For starters I have reduced the size of the documents and it helped but I fear that it wouldn't be enough.

Thanks in advance

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [August 28, 2023, 3:02pm UTC](https://discuss.elastic.co/t/retrieving-millions-of-large-documents/341803/2 "2023-08-28T15:02:10Z")

</div>

[Sliced search](https://www.elastic.co/guide/en/elasticsearch/reference/current/point-in-time-api.html#search-slicing) is there to let you parallelize the retrieval of a large data set, sounds like you want that.

---

<div class="post-metadata">

**Author:** ![Tomer\_Avira](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tomer_avira/32/124618_2.png) [@Tomer\_Avira](https://discuss.elastic.co/u/Tomer_Avira)\
**Post date:** [August 28, 2023, 3:17pm UTC](https://discuss.elastic.co/t/retrieving-millions-of-large-documents/341803/3 "2023-08-28T15:17:20Z")

</div>

Is the PIT id the same as the scroll\_id concept?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [August 28, 2023, 3:23pm UTC](https://discuss.elastic.co/t/retrieving-millions-of-large-documents/341803/4 "2023-08-28T15:23:23Z")

</div>

Similar, yes.

---

<div class="post-metadata">

**Author:** ![Keith\_Massey](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/keith_massey/32/83666_2.png) [@Keith\_Massey](https://discuss.elastic.co/u/Keith_Massey)\
**Post date:** [August 28, 2023, 3:43pm UTC](https://discuss.elastic.co/t/retrieving-millions-of-large-documents/341803/5 "2023-08-28T15:43:53Z")

</div>

If you have access to a hadoop or spark cluster, [es-hadoop](https://www.elastic.co/elasticsearch/hadoop) is another option..

---

<div class="post-metadata">

**Author:** ![Tomer\_Avira](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tomer_avira/32/124618_2.png) [@Tomer\_Avira](https://discuss.elastic.co/u/Tomer_Avira)\
**Post date:** [August 28, 2023, 3:52pm UTC](https://discuss.elastic.co/t/retrieving-millions-of-large-documents/341803/6 "2023-08-28T15:52:16Z")

</div>

I have a sort of follow up question for better understanding, while I'm not new to elastic I noticed something that puzzled me today that I have yet to notice.

As my explanation goes my client does a request that generates an original scroll\_id, I then return this scroll I'd to the client and then he sends a different request which performs the scroll (which returns the results and the next scroll\_id)

What seemed weird to me today is that I noticed that the scroll id remained the same, it worked as it should but I wondered if the scroll id should change or each query updates the "pointer" of the scroll id?

Just wondering.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [August 28, 2023, 6:06pm UTC](https://discuss.elastic.co/t/retrieving-millions-of-large-documents/341803/7 "2023-08-28T18:06:58Z")

</div>

> [@Tomer\_Avira](#):
>
> I wondered if the scroll id should change or each query updates the "pointer" of the scroll id?

No, it will (often) remain unchanged. But don't rely on that, treat it as different each time.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 25, 2023, 6:07pm UTC](https://discuss.elastic.co/t/retrieving-millions-of-large-documents/341803/8 "2023-09-25T18:07:14Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
