# Elasticsearch Bulk Write is slow using Scan and Scroll

**URL:** <https://discuss.elastic.co/t/elasticsearch-bulk-write-is-slow-using-scan-and-scroll/32139>\
**Category:** Elasticsearch\
**Created:** [October 14, 2015, 6:38am UTC](https://discuss.elastic.co/t/elasticsearch-bulk-write-is-slow-using-scan-and-scroll/32139 "2015-10-14T06:38:13Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![AmitPandita](https://avatars.discourse-cdn.com/v4/letter/a/f6c823/32.png) [@AmitPandita](https://discuss.elastic.co/u/AmitPandita)\
**Post date:** [October 14, 2015, 6:38am UTC](https://discuss.elastic.co/t/elasticsearch-bulk-write-is-slow-using-scan-and-scroll/32139/1 "2015-10-14T06:38:14Z")

</div>

Hi Group,

I am currently running into an issue on which i am really stuck.  
I am trying to work on a problem where I have to output the Elasticsearch documents and write them to csv. The docs range from 50,000 to 5 million.  
I am experience serious performance issues and I get a feeling that I am missing something here.

Right now I have a dataset to 400,000 documents on which I am trying to scan and scroll and which would ultimately be formatted and written to csv. But the time taken to just output is 20 mins!! That is insane.

Here is my script:

import elasticsearch  
import elasticsearch.exceptions  
import elasticsearch.helpers as helpers  
import time

es = elasticsearch.Elasticsearch(['[http://XX.XXX.XX.XXX:9200](http://XX.XXX.XX.XXX:9200)'],retry\_on\_timeout=True)

scanResp = helpers.scan(client=es,scroll="50m",index='MyDoc',doc\_type='MyDoc',timeout="50m",size=1000)

resp={}  
start\_time = time.time()  
for resp in scanResp:  
data = resp  
print data.values()[3]

print("--- %s seconds ---" % (time.time() - start\_time))

I am using a hosted AWS m3.medium server for Elasticsearch.

Can anyone please tell me what I might be doing wrong here?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [October 14, 2015, 9:54pm UTC](https://discuss.elastic.co/t/elasticsearch-bulk-write-is-slow-using-scan-and-scroll/32139/2 "2015-10-14T21:54:25Z")

</div>

So the `size` parameter is what it gets _from each shard_, so if you have (eg) 5 shards, that's 2 millions docs!  
I'd start by reducing that to something considerably smaller and see if it helps.

---

<div class="post-metadata">

**Author:** ![AmitPandita](https://avatars.discourse-cdn.com/v4/letter/a/f6c823/32.png) [@AmitPandita](https://discuss.elastic.co/u/AmitPandita)\
**Post date:** [October 15, 2015, 5:27am UTC](https://discuss.elastic.co/t/elasticsearch-bulk-write-is-slow-using-scan-and-scroll/32139/3 "2015-10-15T05:27:11Z")

</div>

@warkolm Yes i did that already, in fact i started the size from 10, then 50,100,150,200,300,500,100 ...... The best result was at 200 where i got the result in 18 seconds that too for just 4000 documents. That is a really bad figure. What else apart from the size do u think i might be missing?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [October 15, 2015, 5:50am UTC](https://discuss.elastic.co/t/elasticsearch-bulk-write-is-slow-using-scan-and-scroll/32139/4 "2015-10-15T05:50:42Z")

</div>

Are you monitoring statistics on the cluster?  
What do they tell you?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:44pm UTC](https://discuss.elastic.co/t/elasticsearch-bulk-write-is-slow-using-scan-and-scroll/32139/5 "2017-07-05T23:44:43Z")

</div>


