# Data too large using BulkProcessor with size limit

**URL:** <https://discuss.elastic.co/t/data-too-large-using-bulkprocessor-with-size-limit/241779>\
**Category:** Elasticsearch\
**Created:** [July 19, 2020, 10:14am UTC](https://discuss.elastic.co/t/data-too-large-using-bulkprocessor-with-size-limit/241779 "2020-07-19T10:14:01Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Fabrizio\_Fortino](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fabrizio_fortino/32/1719_2.png) [@Fabrizio\_Fortino](https://discuss.elastic.co/u/Fabrizio_Fortino)\
**Post date:** [July 19, 2020, 10:14am UTC](https://discuss.elastic.co/t/data-too-large-using-bulkprocessor-with-size-limit/241779/1 "2020-07-19T10:14:01Z")

</div>

Hello,

We are using a BulkProcessor (ES 7.8.0) with the following properties:

- Actions: 250
- Size: 2097152 bytes (2MB)
- Flush Time: 3000 ms

During a load test we are getting the following exception:

```auto
        org.elasticsearch.ElasticsearchStatusException: Elasticsearch exception [type=circuit_breaking_exception, reason=[parent] Data too large, data for [<http_request>] would be [1026389342/978.8mb], which is larger than the limit of [1020054732/972.7mb], real usage: [1026388848/978.8mb], new bytes reserved: [494/494b], usages [request=0/0b, fielddata=0/0b, in_flight_requests=494/494b, accounting=96616/94.3kb]]
    	at org.elasticsearch.rest.BytesRestResponse.errorFromXContent(BytesRestResponse.java:177)
    	at org.elasticsearch.client.RestHighLevelClient.parseEntity(RestHighLevelClient.java:1897)
    	at org.elasticsearch.client.RestHighLevelClient.parseResponseException(RestHighLevelClient.java:1867)
    	at org.elasticsearch.client.RestHighLevelClient$1.onFailure(RestHighLevelClient.java:1783)
    	at org.elasticsearch.client.RestClient$FailureTrackingResponseListener.onDefinitiveFailure(RestClient.java:598)
    	at org.elasticsearch.client.RestClient$1.completed(RestClient.java:343)
    	at org.elasticsearch.client.RestClient$1.completed(RestClient.java:327)
    	at org.apache.http.concurrent.BasicFuture.completed(BasicFuture.java:122)
    	at org.apache.http.impl.nio.client.DefaultClientExchangeHandlerImpl.responseCompleted(DefaultClientExchangeHandlerImpl.java:181)
    	at org.apache.http.nio.protocol.HttpAsyncRequestExecutor.processResponse(HttpAsyncRequestExecutor.java:448)
    	at org.apache.http.nio.protocol.HttpAsyncRequestExecutor.inputReady(HttpAsyncRequestExecutor.java:338)
    	at org.apache.http.impl.nio.DefaultNHttpClientConnection.consumeInput(DefaultNHttpClientConnection.java:265)
    	at org.apache.http.impl.nio.client.InternalIODispatch.onInputReady(InternalIODispatch.java:81)
    	at org.apache.http.impl.nio.client.InternalIODispatch.onInputReady(InternalIODispatch.java:39)
    	at org.apache.http.impl.nio.reactor.AbstractIODispatch.inputReady(AbstractIODispatch.java:114)
    	at org.apache.http.impl.nio.reactor.BaseIOReactor.readable(BaseIOReactor.java:162)
    	at org.apache.http.impl.nio.reactor.AbstractIOReactor.processEvent(AbstractIOReactor.java:337)
    	at org.apache.http.impl.nio.reactor.AbstractIOReactor.processEvents(AbstractIOReactor.java:315)
    	at org.apache.http.impl.nio.reactor.AbstractIOReactor.execute(AbstractIOReactor.java:276)
    	at org.apache.http.impl.nio.reactor.BaseIOReactor.execute(BaseIOReactor.java:104)
    	at org.apache.http.impl.nio.reactor.AbstractMultiworkerIOReactor$Worker.run(AbstractMultiworkerIOReactor.java:591)
    	at java.base/java.lang.Thread.run(Thread.java:834)
    	Suppressed: org.elasticsearch.client.ResponseException: method [POST], host [http://localhost:32785], URI [/_bulk?timeout=1m], status line [HTTP/1.1 429 Too Many Requests]
    {"error":{"root_cause":[{"type":"circuit_breaking_exception","reason":"[parent] Data too large, data for [<http_request>] would be [1026389342/978.8mb], which is larger than the limit of [1020054732/972.7mb], real usage: [1026388848/978.8mb], new bytes reserved: [494/494b], usages [request=0/0b, fielddata=0/0b, in_flight_requests=494/494b, accounting=96616/94.3kb]","bytes_wanted":1026389342,"bytes_limit":1020054732,"durability":"PERMANENT"}],"type":"circuit_breaking_exception","reason":"[parent] Data too large, data for [<http_request>] would be [1026389342/978.8mb], which is larger than the limit of [1020054732/972.7mb], real usage: [1026388848/978.8mb], new bytes reserved: [494/494b], usages [request=0/0b, fielddata=0/0b, in_flight_requests=494/494b, accounting=96616/94.3kb]","bytes_wanted":1026389342,"bytes_limit":1020054732,"durability":"PERMANENT"},"status":429}
    		at org.elasticsearch.client.RestClient.convertResponse(RestClient.java:283)
    		at org.elasticsearch.client.RestClient.access$1700(RestClient.java:97)
    		at org.elasticsearch.client.RestClient$1.completed(RestClient.java:331)
    		... 16 common frames omitted

```

Shouldn't bulk processor prevent these kinds of errors?

Thanks,  
Fabrizio

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 19, 2020, 10:17am UTC](https://discuss.elastic.co/t/data-too-large-using-bulkprocessor-with-size-limit/241779/2 "2020-07-19T10:17:47Z")

</div>

> [@Fabrizio\_Fortino](#):
>
> `[HTTP/1.1 429 Too Many Requests]`

This is a sign you are overloading the cluster. Please see [this blog post](https://www.elastic.co/blog/why-am-i-seeing-bulk-rejections-in-my-elasticsearch-cluster) for further details.

---

<div class="post-metadata">

**Author:** ![Fabrizio\_Fortino](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fabrizio_fortino/32/1719_2.png) [@Fabrizio\_Fortino](https://discuss.elastic.co/u/Fabrizio_Fortino)\
**Post date:** [July 20, 2020, 11:58am UTC](https://discuss.elastic.co/t/data-too-large-using-bulkprocessor-with-size-limit/241779/3 "2020-07-20T11:58:07Z")

</div>

Thanks a lot for the useful resource.

I would expect that the failed requests would have been automatically retried since the BulkProcessor has a backoff exponential retry strategy. Instead it seems the retry mechanism does not kick in in this case. Is that the expected behaviour?

---

<div class="post-metadata">

**Author:** ![Steve\_Mushero](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/steve_mushero/32/22441_2.png) [@Steve\_Mushero](https://discuss.elastic.co/u/Steve_Mushero)\
**Post date:** [July 20, 2020, 2:13pm UTC](https://discuss.elastic.co/t/data-too-large-using-bulkprocessor-with-size-limit/241779/4 "2020-07-20T14:13:26Z")

</div>

@Christian_Dahlqvist - two questions from that blog:

1. Is a bulk request atomic, i.e. if a sub-request fails on a node, is the whole batch rejected, and if so, how does that happen as most nodes already indexed the docs?
2. Is bulk indexing multi-threaded, either on the coord node or the data node, i.e. what is the scaling ability vs. cores? My impression is it's better for the sender to send parallel bulk requests, implying it's single-threaded as the coord or data nodes, or both.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 20, 2020, 2:20pm UTC](https://discuss.elastic.co/t/data-too-large-using-bulkprocessor-with-size-limit/241779/5 "2020-07-20T14:20:08Z")

</div>

Parts of bulk requests can fail so it is not atomic. Indexing is as far as I know single-threaded per shard, but shards are processed in parallel.

---

<div class="post-metadata">

**Author:** ![Steve\_Mushero](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/steve_mushero/32/22441_2.png) [@Steve\_Mushero](https://discuss.elastic.co/u/Steve_Mushero)\
**Post date:** [July 20, 2020, 2:25pm UTC](https://discuss.elastic.co/t/data-too-large-using-bulkprocessor-with-size-limit/241779/6 "2020-07-20T14:25:36Z")

</div>

Does the error response indicate which docs failed, else I'd think the sender would not know how/what to re-submit? Sorry, my knowledge on this is poor.

Threaded per shard makes sense, as long as the node-level queue is shard-keyed, which I guess it would be as the coord node sent it there to a specific shard; so if my node has 4 shards for two indexes, it can use 4 threads to index in parallel. And parallel bulk batches won't scale more, but maybe not true on coord side.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 17, 2020, 2:25pm UTC](https://discuss.elastic.co/t/data-too-large-using-bulkprocessor-with-size-limit/241779/7 "2020-08-17T14:25:43Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
