# How to pull large amount of data (all documents in the index) using elasticsearch python client?

**URL:** <https://discuss.elastic.co/t/how-to-pull-large-amount-of-data-all-documents-in-the-index-using-elasticsearch-python-client/302763>\
**Category:** Elasticsearch\
**Tags:** language-clients\
**Created:** [April 20, 2022, 2:57am UTC](https://discuss.elastic.co/t/how-to-pull-large-amount-of-data-all-documents-in-the-index-using-elasticsearch-python-client/302763 "2022-04-20T02:57:02Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![safarial.fatemeh](https://avatars.discourse-cdn.com/v4/letter/s/2bfe46/32.png) [@safarial.fatemeh](https://discuss.elastic.co/u/safarial.fatemeh)\
**Post date:** [April 20, 2022, 2:57am UTC](https://discuss.elastic.co/t/how-to-pull-large-amount-of-data-all-documents-in-the-index-using-elasticsearch-python-client/302763/1 "2022-04-20T02:57:02Z")

</div>

I'm using Elasticsearch.helpers.scan to pull down over 1M documents from Elasticsearch and I use match\_all query for that.  
the process is superslow (taking over 2hrs).  
is there a better way to pull down all the documents from Elasticsearch?

---

<div class="post-metadata">

**Author:** ![casterQ](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/casterq/32/93257_2.png) [@casterQ](https://discuss.elastic.co/u/casterQ)\
**Post date:** [April 20, 2022, 3:05am UTC](https://discuss.elastic.co/t/how-to-pull-large-amount-of-data-all-documents-in-the-index-using-elasticsearch-python-client/302763/2 "2022-04-20T03:05:51Z")

</div>

`PIT` or `Scroll`

---

<div class="post-metadata">

**Author:** ![safarial.fatemeh](https://avatars.discourse-cdn.com/v4/letter/s/2bfe46/32.png) [@safarial.fatemeh](https://discuss.elastic.co/u/safarial.fatemeh)\
**Post date:** [April 20, 2022, 4:01am UTC](https://discuss.elastic.co/t/how-to-pull-large-amount-of-data-all-documents-in-the-index-using-elasticsearch-python-client/302763/3 "2022-04-20T04:01:30Z")

</div>

thanks @casterQ  
can you elaborate? I'm using scan which I think is the wrapper utilizing Scroll. isn't it?  
and can you please provide some example about PIT and Scroll

---

<div class="post-metadata">

**Author:** ![casterQ](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/casterq/32/93257_2.png) [@casterQ](https://discuss.elastic.co/u/casterQ)\
**Post date:** [April 20, 2022, 6:22am UTC](https://discuss.elastic.co/t/how-to-pull-large-amount-of-data-all-documents-in-the-index-using-elasticsearch-python-client/302763/4 "2022-04-20T06:22:34Z")

</div>

1. snapshot(Can only be used on ES)
2. CCR(Platinum)
3. PIT or scroll(It is the way of pulling es, and it is recommended to use the later version of pit)

here is PIT doc:

> **[Point in time API | Elasticsearch Guide \[8.1\] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/point-in-time-api.html)**

and scroll can be accelerated using slice：

> **[Paginate search results | Elasticsearch Guide \[7.10\] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/7.10/paginate-search-results.html#slice-scroll)**

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 18, 2022, 6:22am UTC](https://discuss.elastic.co/t/how-to-pull-large-amount-of-data-all-documents-in-the-index-using-elasticsearch-python-client/302763/5 "2022-05-18T06:22:57Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
