# Error while using .scan() function call

**URL:** <https://discuss.elastic.co/t/error-while-using-scan-function-call/343039>\
**Category:** Elasticsearch\
**Tags:** language-clients\
**Created:** [September 14, 2023, 9:54am UTC](https://discuss.elastic.co/t/error-while-using-scan-function-call/343039 "2023-09-14T09:54:44Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![hjazz6](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hjazz6/32/79007_2.png) [@hjazz6](https://discuss.elastic.co/u/hjazz6)\
**Post date:** [September 14, 2023, 9:54am UTC](https://discuss.elastic.co/t/error-while-using-scan-function-call/343039/1 "2023-09-14T09:54:44Z")

</div>

Hi,

I have the following Python code to query my ES cluster (v8.8.0).

```auto
from elasticsearch import Elasticsearch
from elasticsearch_dsl import Search

for i in range(0, 100):
    s = Search(using=es, index=index_name) \
          .filter('range', **{'@timestamp': {'gte': start_datetime, 'lt': end_datetime, 'format': 
    'strict_date_optional_time_nanos'}})

    if s.count() > 0:
       results = s.scan()
          for result in results:
              result_json = result.to_dict()
              res1, res2 = process_result(result_json)

```

The number of documents returned in each `scan` is between 1-2M.

During one of the queries, I got an error on ES:

`circuit_breaking_exception: [parent] Data too large, data for [<http_request>] would be [32283148058/30gb], which is larger than the limit of [31621696716/29.4gb], real usage: [32283146480/30gb], new bytes reserved: [1578/1.5kb]...`

It seems like ES ran out of memory when executing the query. Is there something I can specify in my query so as to avoid this issue?

I'm not sure if the `scan` I'm using is the same as what is mentioned [here](https://elasticsearch-py.readthedocs.io/en/7.x/helpers.html#scan) from `elasticsearch.helpers.scan`. If so, is there some setting I can specify? For example, should I set `scroll` or `request_timeout` to make sure the cache is cleared after each call to `scan`, maybe with an addition of `sleep` to make sure the timeout is met?

Also, if `clear_scroll=True` by default, then I would not need to explicitly clear the `scroll_id`, is that correct?

Thank you.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [September 14, 2023, 10:25am UTC](https://discuss.elastic.co/t/error-while-using-scan-function-call/343039/2 "2023-09-14T10:25:03Z")

</div>

I'm not a Python expert but I saw this [in the docs](https://elasticsearch-py.readthedocs.io/en/stable/helpers.html#scan):

> - **size** – size (per shard) of the batch send at each iteration.

> [@hjazz6](#):
>
> The number of documents returned in each `scan` is between 1-2M.

I'm curious about this. On how many shards is the scan operation running? To me it should not be more than 10000 per shard.  
Also, may be your documents (`_source` field) are super big? In which case, reducing the `size` could help?

---

<div class="post-metadata">

**Author:** ![hjazz6](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hjazz6/32/79007_2.png) [@hjazz6](https://discuss.elastic.co/u/hjazz6)\
**Post date:** [September 14, 2023, 11:24am UTC](https://discuss.elastic.co/t/error-while-using-scan-function-call/343039/3 "2023-09-14T11:24:21Z")

</div>

> [@dadoonet](#):
>
> On how many shards is the scan operation running? To me it should not be more than 10000 per shard.

On the query that failed, the error message was the "Scroll request has only succeeded on 11 shards out of 12".

> [@dadoonet](#):
>
> may be your documents (`_source` field) are super big? In which case, reducing the `size` could help?

Each document should be under 1KB. Is that considered big? I don't quite understand what this parameter mean.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [September 14, 2023, 12:05pm UTC](https://discuss.elastic.co/t/error-while-using-scan-function-call/343039/4 "2023-09-14T12:05:00Z")

</div>

`size` is the number of documents you are asking for each scroll request. Try to reduce it and see if it helps?

---

<div class="post-metadata">

**Author:** ![Quentin\_Pradet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/quentin_pradet/32/94192_2.png) [@Quentin\_Pradet](https://discuss.elastic.co/u/Quentin_Pradet)\
**Post date:** [September 14, 2023, 12:30pm UTC](https://discuss.elastic.co/t/error-while-using-scan-function-call/343039/5 "2023-09-14T12:30:36Z")

</div>

What @dadoonet said is correct. To reduce the size of the scan using the Elasticsearch DSL client, you need to use the params() function like this:

```python
from elasticsearch_dsl import Search

es = Elasticsearch(...)

s = Search(using=es, index="...")
for result in s.params(size=100).scan():
    ...

```

I also want to mention that you might want to look at point-in-time search ([Point in time API | Elasticsearch Guide [8.10] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/point-in-time-api.html)) which is more robust than scrolling.

---

<div class="post-metadata">

**Author:** ![hjazz6](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hjazz6/32/79007_2.png) [@hjazz6](https://discuss.elastic.co/u/hjazz6)\
**Post date:** [September 14, 2023, 6:18pm UTC](https://discuss.elastic.co/t/error-while-using-scan-function-call/343039/6 "2023-09-14T18:18:52Z")

</div>

Thanks, @dadoonet and @Quentin_Pradet I'll try to reduce `size` to see if it works, and also look at the Point in time API.

---

<div class="post-metadata">

**Author:** ![hjazz6](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hjazz6/32/79007_2.png) [@hjazz6](https://discuss.elastic.co/u/hjazz6)\
**Post date:** [September 15, 2023, 2:17am UTC](https://discuss.elastic.co/t/error-while-using-scan-function-call/343039/7 "2023-09-15T02:17:54Z")

</div>

I'm been thinking about the cause of the issue. I've ran the loop a couple of times (different post-processing of the results) prior to getting the error, and had no problem. The size of the data mentioned in the error (32GB) is also much larger than the size of the results from a single call of `scan`.

Someone else advised that it could be because the cache was not cleared, so ES eventually ran out of memory after so many calls. Although it seems like `scan` by default has `clear_scroll = True`.

If it is somehow due to the accumulative memory usage from multiple calls, how would reducing the `size` in `scan` help?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 13, 2023, 2:18am UTC](https://discuss.elastic.co/t/error-while-using-scan-function-call/343039/8 "2023-10-13T02:18:40Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
