# What is the fastest way to read all the records in an index?

**URL:** <https://discuss.elastic.co/t/what-is-the-fastest-way-to-read-all-the-records-in-an-index/635>\
**Category:** Elasticsearch\
**Created:** [May 13, 2015, 3:54pm UTC](https://discuss.elastic.co/t/what-is-the-fastest-way-to-read-all-the-records-in-an-index/635 "2015-05-13T15:54:36Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![ptrei](https://avatars.discourse-cdn.com/v4/letter/p/e480ec/32.png) [@ptrei](https://discuss.elastic.co/u/ptrei)\
**Post date:** [May 13, 2015, 3:54pm UTC](https://discuss.elastic.co/t/what-is-the-fastest-way-to-read-all-the-records-in-an-index/635/1 "2015-05-13T15:54:36Z")

</div>

New to the list, semi-new to ELK. Apologies if this is a FAQ, but I can't find an answer.

I'm processing proxy logs. Logstash does just fine in parsing them, and putting them into daily logstash-[date] indices.

However, I want to programmatically walk through all the records in each day's index, and collapse some of the data into historical summaries. The processing is beyond what I can do in an Elasticsearch query. I have a python program which performs this operation, but the performance isn't adequate.

The biggest bottleneck is reading the records from the logstash-[date] indices. I'm using the elasticsearch-py interface, which provides a wrapper for the 'scroll' API. Here's the relevant lines....

```
import elasticsearch
import elasticsearch.helpers as helpers

es = elasticsearch.Elasticsearch(retry_on_timeout=True)
# sets up global ES handle

#main processing loop 
def process_index(index_name)
   global es
   query_body = '{"size": 10000, "query": {"match_all":{}}}'
   scanResp = helpers.scan(client=es,query=query_body,scroll="5m",index=index_name,timeout="5m")
   resp={}
   for resp in scanResp:
      DO STUFF FOR ONE RECORD

```

The Bulk API appears to be useful only for writing, not reading. I'm using it to write out the collapsed data I generate.

But is there any way I can read all the records from an index with higher speed? At the moment, even if null out the processing, I'm reading about 2500 records/second. When my daily logs have 60 million entries, that's painfully slow.

thanks in advance,

Peter Trei

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [May 13, 2015, 10:16pm UTC](https://discuss.elastic.co/t/what-is-the-fastest-way-to-read-all-the-records-in-an-index/635/2 "2015-05-13T22:16:45Z")

</div>

Scroll is what you want, have you tried reducing the size to see if that helps performance?

---

<div class="post-metadata">

**Author:** ![AmitPandita](https://avatars.discourse-cdn.com/v4/letter/a/f6c823/32.png) [@AmitPandita](https://discuss.elastic.co/u/AmitPandita)\
**Post date:** [October 14, 2015, 10:39am UTC](https://discuss.elastic.co/t/what-is-the-fastest-way-to-read-all-the-records-in-an-index/635/3 "2015-10-14T10:39:34Z")

</div>

@ptrei Just curious to know the configuration of your server?

I am also running into the same issue and I am seeing even worse performance. It would be helpful in case you share your current configuration.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:44pm UTC](https://discuss.elastic.co/t/what-is-the-fastest-way-to-read-all-the-records-in-an-index/635/4 "2017-07-05T23:44:55Z")

</div>


