# Occasional Bulk Insert Failure (ES 2.4.4)

**URL:** <https://discuss.elastic.co/t/occasional-bulk-insert-failure-es-2-4-4/73287>\
**Category:** Elasticsearch\
**Created:** [January 30, 2017, 9:32pm UTC](https://discuss.elastic.co/t/occasional-bulk-insert-failure-es-2-4-4/73287 "2017-01-30T21:32:47Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![vyazici](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vyazici/32/39927_2.png) [@vyazici](https://discuss.elastic.co/u/vyazici)\
**Post date:** [January 30, 2017, 9:32pm UTC](https://discuss.elastic.co/t/occasional-bulk-insert-failure-es-2-4-4/73287/1 "2017-01-30T21:32:47Z")

</div>

Hi all!

In an application, I fetch rows from multiple tables over distinct JDBC connections in parallel, transform rows into document fields, and send it to ES in batches of size 1000 using the [Java Bulk API](https://www.elastic.co/guide/en/elasticsearch/client/java-api/current/java-docs-bulk.html). There I set the timeout of the bulk inserts to 15s and repeat at most 3 times on failure. The entire process takes \>6 hours with max. 6 fetches in parallel on an ES cluster of 3 beefy VMs. (1 data, 1 master node on each VM.) But occasionally some bulk inserts fail even after retries. How should I diagnose and tackle this problem? (For the records, index is created using indices.store.throttle.type=none, number\_of\_shards=6, number\_of\_replicas=0, index.refresh\_interval=-1, and translog.disable\_flush=true. Upon successful completion, we revert these to production settings.)

Best.

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [January 31, 2017, 8:56am UTC](https://discuss.elastic.co/t/occasional-bulk-insert-failure-es-2-4-4/73287/2 "2017-01-31T08:56:43Z")

</div>

Hey,

can you provide more information while the bulk inserts failed? Did they fail because of a server or client issue? Can you provide the responses?

--Alex

---

<div class="post-metadata">

**Author:** ![vyazici](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vyazici/32/39927_2.png) [@vyazici](https://discuss.elastic.co/u/vyazici)\
**Post date:** [January 31, 2017, 6:02pm UTC](https://discuss.elastic.co/t/occasional-bulk-insert-failure-es-2-4-4/73287/3 "2017-01-31T18:02:50Z")

</div>

Hey Alex!

Sorry for the misunderstanding. By "fail", I do mean that my wrapper Hystrix command timeouts after 15s. I even tried increasing timeout threshold to 30s. Even then, after a certain amount of inserts, occasionally some inserts just keep on waiting.

Best.

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [February 1, 2017, 8:24am UTC](https://discuss.elastic.co/t/occasional-bulk-insert-failure-es-2-4-4/73287/4 "2017-02-01T08:24:43Z")

</div>

Hey,

have you checked your Elasticsearch logs during that time? Is there a node doing garbage collection maybe? You will find that in the logs.

--Alex

---

<div class="post-metadata">

**Author:** ![vyazici](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vyazici/32/39927_2.png) [@vyazici](https://discuss.elastic.co/u/vyazici)\
**Post date:** [February 1, 2017, 12:46pm UTC](https://discuss.elastic.co/t/occasional-bulk-insert-failure-es-2-4-4/73287/5 "2017-02-01T12:46:49Z")

</div>

I think we found the culprit: translog flushes. Although I set translog.disable\_flush=true, apparently ES still prefers to do some:

 ![](https://us1.discourse-cdn.com/elastic/original/2X/8/88f0206e8a3a56ad951229cf1ef1a0dd7761f242.png)

See the idle state in the translog size? That's where the entire ES cluster gets busy with flushing the translog, which in the meantime causes \>2m delays in our bulk inserts. Is translog.disable\_flush=true not doing why I do expect it to do, or am I misinterpreting its function?

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [February 1, 2017, 1:16pm UTC](https://discuss.elastic.co/t/occasional-bulk-insert-failure-es-2-4-4/73287/6 "2017-02-01T13:16:22Z")

</div>

Hey,

is there any reason you decided to disable the flushing of the translog in the first place? This option was removed (and only useful in tests anyway).

--Alex

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 1, 2017, 1:16pm UTC](https://discuss.elastic.co/t/occasional-bulk-insert-failure-es-2-4-4/73287/7 "2017-03-01T13:16:39Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
