# How to improve performance of scan method?

**URL:** <https://discuss.elastic.co/t/how-to-improve-performance-of-scan-method/159425>\
**Category:** Elasticsearch\
**Created:** [December 4, 2018, 10:19pm UTC](https://discuss.elastic.co/t/how-to-improve-performance-of-scan-method/159425 "2018-12-04T22:19:23Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Samvid\_Kulkarni](https://avatars.discourse-cdn.com/v4/letter/s/439d5e/32.png) [@Samvid\_Kulkarni](https://discuss.elastic.co/u/Samvid_Kulkarni)\
**Post date:** [December 4, 2018, 10:19pm UTC](https://discuss.elastic.co/t/how-to-improve-performance-of-scan-method/159425/1 "2018-12-04T22:19:23Z")

</div>

I am trying to get data from elasticsearch using elasticsearch-dsl python library. I need to get all the data for last 15 min. Issue is retrieving data is extremely slow. It takes lot of time for 2.2 million hits. Here is my code

```
start_time = time.time()  
try:
	client = Elasticsearch(['IP_HERE'])
	s = Search(using=client, index="firewallv2-*", doc_type = 'doc').filter('range', **{'@timestamp': {'gte': 'now-15m' , 'lt': 'now'}})
	response = s.execute()
except Exception as e:
	print(e)
	print("error in getting data from FIREWALL")

try:
	for hit1 in s.scan():
		source_ip.append(hit1.to_dict().get('Source IP'))
		destination_ip.append(hit1.to_dict().get('Destination IP'))
		destination_port.append(hit1.to_dict().get('Destination Port'))
		source_port.append(hit1.to_dict().get('Source Port'))

except Exception as e:
	print("not able to parse json data")

elapsed_time = time.time() - start_time
print("Time to get data from server " + str(elapsed_time))

```

There is more to code but I am just posting the main slow component. Rest is pure python code. Below is my output

```
Time to get data from server 893.599892855
Time to store data into variables 27.647258997
Time to process for loop 9.32531404495

```

All time output is in seconds and you can see that it takes huge amount of time to retrieve 2.2 million hits.

I also tried using `bulk_size=10000` and even changing builk\_size to various values but not success.

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [December 4, 2018, 10:41pm UTC](https://discuss.elastic.co/t/how-to-improve-performance-of-scan-method/159425/2 "2018-12-04T22:41:36Z")

</div>

There are a few things you can do, like setting the size in the search higher. The default search size is more for regular search than scrolling. _Usually_ your better off trying to work these sorts of things into aggregations if you can. The documents are stored on disk in such a way that it is faster to run aggregations than it is to return the entire `_source`.

---

<div class="post-metadata">

**Author:** ![Samvid\_Kulkarni](https://avatars.discourse-cdn.com/v4/letter/s/439d5e/32.png) [@Samvid\_Kulkarni](https://discuss.elastic.co/u/Samvid_Kulkarni)\
**Post date:** [December 12, 2018, 9:19pm UTC](https://discuss.elastic.co/t/how-to-improve-performance-of-scan-method/159425/3 "2018-12-12T21:19:41Z")

</div>

Thank you very much for replying and sorry for late response. i have made the changes as you suggested and i was about to lower the time taken to 710 seconds but I am not able to lower it further.

We have only one node setup so could that be cause problem? or is there any other issue you can think of?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 9, 2019, 9:19pm UTC](https://discuss.elastic.co/t/how-to-improve-performance-of-scan-method/159425/4 "2019-01-09T21:19:50Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
