# Circuit breaker exception

**URL:** <https://discuss.elastic.co/t/circuit-breaker-exception/305599>\
**Category:** Elasticsearch\
**Created:** [May 25, 2022, 11:23am UTC](https://discuss.elastic.co/t/circuit-breaker-exception/305599 "2022-05-25T11:23:20Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![charvi23](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/charvi23/32/103607_2.png) [@charvi23](https://discuss.elastic.co/u/charvi23)\
**Post date:** [May 25, 2022, 11:23am UTC](https://discuss.elastic.co/t/circuit-breaker-exception/305599/1 "2022-05-25T11:23:20Z")

</div>

I get the following exception while I am trying to insert data using ES-Hive Hadoop jar. I am currently inserting around 60 million data.

Following is the error I get

`Caused by: org.elasticsearch.hadoop.rest.EsHadoopInvalidRequest: org.elasticsearch.hadoop.rest.EsHadoopRemoteException: circuit_breaking_exception: [parent] Data too large, data for [<http_request>] would be [31820340712/29.6gb], which is larger than the limit of [31621696716/29.4gb], real usage: [31818275608/29.6gb], new bytes reserved: [2065104/1.9mb], usages [inflight_requests=105523462/100.6mb, request=0/0b, fielddata=0/0b, eql_sequence=0/0b, model_inference=0/0b]`

Elasticsearch.log shows only the following:

````auto
[2022-05-25T04:13:01,315][INFO][o.e.i.b.HierarchyCircuitBreakerService] [] GC did bring memory usage down, before [31636006904], after [30392212248], allocations [19], duration [137]
[2022-05-25T04:13:04,065][INFO][o.e.m.j.JvmGcMonitorService] [] [gc][159765] overhead, spent [300ms] collecting in the last [1s]
[2022-05-25T04:13:06,886][INFO][o.e.i.b.HierarchyCircuitBreakerService] [] attempting to trigger G1GC due to high heap usage [32059682216]
[2022-05-25T04:13:07,061][INFO][o.e.i.b.HierarchyCircuitBreakerService] [] GC did bring memory usage down, before [32059682216], after [30969353816], allocations [58], duration [175]
[2022-05-25T04:13:12,337][INFO][o.e.i.b.HierarchyCircuitBreakerService] [] attempting to trigger G1GC due to high heap usage [31690774104] ```

1. My shard allocation is 1.
2. Replica is 1.
3. JVM size allocated is max at i.e. 31 GB.

Stats below:

name id node.role heap.current heap.percent heap.max
xxxxx xx xxx 27.5gb 88 31gb

What can I do to fix it other than adding another node to the cluster.
````

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [May 25, 2022, 11:57pm UTC](https://discuss.elastic.co/t/circuit-breaker-exception/305599/2 "2022-05-25T23:57:49Z")

</div>

Reduce the size of your index request is the other alternative.

---

<div class="post-metadata">

**Author:** ![charvi23](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/charvi23/32/103607_2.png) [@charvi23](https://discuss.elastic.co/u/charvi23)\
**Post date:** [May 26, 2022, 4:28am UTC](https://discuss.elastic.co/t/circuit-breaker-exception/305599/3 "2022-05-26T04:28:17Z")

</div>

But I am using hive hadoop jar that uses bulk internally, can you please help in how can I reduce the size of index request in that case.

Also when I am moving data from hive the data size on disk is ~ 35 GB which when moved to Elasticsearch shows disk size of 500GB. Why is this happening. Is it something that is expected from this conversion?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [May 26, 2022, 4:44am UTC](https://discuss.elastic.co/t/circuit-breaker-exception/305599/4 "2022-05-26T04:44:53Z")

</div>

I'm not sure how do that sorry. Hopefully someone else can comment.

---

<div class="post-metadata">

**Author:** ![charvi23](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/charvi23/32/103607_2.png) [@charvi23](https://discuss.elastic.co/u/charvi23)\
**Post date:** [May 26, 2022, 4:50am UTC](https://discuss.elastic.co/t/circuit-breaker-exception/305599/5 "2022-05-26T04:50:43Z")

</div>

Thanks!! Any idea on this?  
Also when I am moving data from hive the data size on disk is ~ 35 GB which when moved to Elasticsearch shows disk size of 500GB. Why is this happening. Is it something that is expected from this conversion?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [May 26, 2022, 4:51am UTC](https://discuss.elastic.co/t/circuit-breaker-exception/305599/6 "2022-05-26T04:51:09Z")

</div>

You'd be best off making a new topic for that 🙂

---

<div class="post-metadata">

**Author:** ![charvi23](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/charvi23/32/103607_2.png) [@charvi23](https://discuss.elastic.co/u/charvi23)\
**Post date:** [May 26, 2022, 4:55am UTC](https://discuss.elastic.co/t/circuit-breaker-exception/305599/7 "2022-05-26T04:55:47Z")

</div>

Ok!! Thanks 🙂

---

<div class="post-metadata">

**Author:** ![charvi23](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/charvi23/32/103607_2.png) [@charvi23](https://discuss.elastic.co/u/charvi23)\
**Post date:** [June 6, 2022, 1:45pm UTC](https://discuss.elastic.co/t/circuit-breaker-exception/305599/8 "2022-06-06T13:45:35Z")

</div>

I added a new node to the ES cluster with 3 TB space, I am still stuck on the error:

`Caused by: org.elasticsearch.hadoop.rest.EsHadoopInvalidRequest: org.elasticsearch.hadoop.rest.EsHadoopRemoteException: circuit_breaking_exception: [parent] Data too large, data for [<http_request>] would be [31656845952/29.4gb], which is larger than the limit of [31621696716/29.4gb], real usage: [31654767576/29.4gb], new bytes reserved: [2078376/1.9mb], usages [eql_sequence=0/0b, fielddata=32168/31.4kb, request=0/0b, inflight_requests=333681778/318.2mb, model_inference=0/0b]`

But at the end of the error it also gives this error:

````nohighlight
, Vertex did not succeed due to OWN_TASK_FAILURE, failedTasks:1 killedTasks:1004, Vertex vertex_1654243481653_2108_13_00 [Map 1] killed/failed due to:OWN_TASK_FAILURE]DAG did not succeed due to VERTEX_FAILURE. failedVertices:1 killedVertices:0 (state=08S01,code=2)
Closing: 0: jdbc:hive2://datanode0..com:2181,datanode..com:2181,master..com:2181,master010..com:2181,master010..com:81/;serviceDiscoveryMode=zooKeeper;zooKeeperNamespace=hiveserver2 ```

Is this error due to space crunch of my cluster? Or is it due to ES circuit breaker? Do you have any idea on that, cause even after adding new node the error doesn't seem to go away.
````

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [June 6, 2022, 11:41pm UTC](https://discuss.elastic.co/t/circuit-breaker-exception/305599/9 "2022-06-06T23:41:26Z")

</div>

It's due to Elasticsearch still. Did you try reducing your request sizes?

---

<div class="post-metadata">

**Author:** ![charvi23](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/charvi23/32/103607_2.png) [@charvi23](https://discuss.elastic.co/u/charvi23)\
**Post date:** [June 7, 2022, 2:21am UTC](https://discuss.elastic.co/t/circuit-breaker-exception/305599/10 "2022-06-07T02:21:07Z")

</div>

How can I do that? As I mentioned I am using es-hadoop jar? Is there a way to do in that jar or set something explicitly?

---

<div class="post-metadata">

**Author:** ![charvi23](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/charvi23/32/103607_2.png) [@charvi23](https://discuss.elastic.co/u/charvi23)\
**Post date:** [June 7, 2022, 2:39am UTC](https://discuss.elastic.co/t/circuit-breaker-exception/305599/11 "2022-06-07T02:39:22Z")

</div>

Also how do you do it in Elasticsearch.. i don't know that either 😕

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [June 7, 2022, 2:41am UTC](https://discuss.elastic.co/t/circuit-breaker-exception/305599/12 "2022-06-07T02:41:00Z")

</div>

You don't do it in Elasticsearch, it's a client level approach. I don't know hadoop though sorry.

---

<div class="post-metadata">

**Author:** ![charvi23](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/charvi23/32/103607_2.png) [@charvi23](https://discuss.elastic.co/u/charvi23)\
**Post date:** [June 7, 2022, 3:02am UTC](https://discuss.elastic.co/t/circuit-breaker-exception/305599/13 "2022-06-07T03:02:31Z")

</div>

In case of ingestion in Elasticsearch can you help with an article may be, how can this be achieved?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [June 7, 2022, 3:06am UTC](https://discuss.elastic.co/t/circuit-breaker-exception/305599/14 "2022-06-07T03:06:32Z")

</div>

[Bulk API | Elasticsearch Guide [8.2] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/8.2/docs-bulk.html) might help, you need to tell your client to not include so many documents when it sends to Elasticsearch.

---

<div class="post-metadata">

**Author:** ![charvi23](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/charvi23/32/103607_2.png) [@charvi23](https://discuss.elastic.co/u/charvi23)\
**Post date:** [June 7, 2022, 4:57am UTC](https://discuss.elastic.co/t/circuit-breaker-exception/305599/15 "2022-06-07T04:57:09Z")

</div>

I found this documentation from ES-Hadoop jar, can this be helpful, if I try reducing the batch entry size/batch size bytes?

> **[Configuration | Elasticsearch for Apache Hadoop \[8.2\] | Elastic](https://www.elastic.co/guide/en/elasticsearch/hadoop/current/configuration.html#configuration-serialization)**
>
> Reference documentation of elasticsearch-hadoop

````auto
Size (in bytes) for batch writes using Elasticsearch bulk API. Note the bulk size is allocated per task instance. Always multiply by the number of tasks within a Hadoop job to get the total bulk size at runtime hitting Elasticsearch.
es.batch.size.entries (default 1000)
Size (in entries) for batch writes using Elasticsearch bulk API - (0 disables it). Companion to es.batch.size.bytes, once one matches, the batch update is executed. Similar to the size, this setting is per task instance; it gets multiplied at runtime by the total number of Hadoop tasks running.```
````

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2022, 4:57am UTC](https://discuss.elastic.co/t/circuit-breaker-exception/305599/16 "2022-07-05T04:57:44Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
