# Significant overhead of concurrency control for small docs

**URL:** <https://discuss.elastic.co/t/significant-overhead-of-concurrency-control-for-small-docs/60276>\
**Category:** Elasticsearch\
**Created:** [September 12, 2016, 11:03am UTC](https://discuss.elastic.co/t/significant-overhead-of-concurrency-control-for-small-docs/60276 "2016-09-12T11:03:33Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Zline](https://avatars.discourse-cdn.com/v4/letter/z/e5b9ba/32.png) [@Zline](https://discuss.elastic.co/u/Zline)\
**Post date:** [September 12, 2016, 11:03am UTC](https://discuss.elastic.co/t/significant-overhead-of-concurrency-control-for-small-docs/60276/1 "2016-09-12T11:03:33Z")

</div>

Hello,

I'm batch-indexing lots of small documents (200-250 bytes each, basically its DB rows) and I see a significant overhead of concurrency control (2-3 times greater than (quite simple) indexing itself):

![](https://us1.discourse-cdn.com/elastic/original/2X/0/0fca3ee05958ec5fa404920cb941e01cb9e39cf1.png)

Indexing performance compared to solr (all test conditions are equal, looks like solr does not perform any version control):  
 ![](https://us1.discourse-cdn.com/elastic/original/2X/8/89eef764b56d871007f27128931d83e388db449f.png)

My questions are:

1. Should't it be considered as a performance problem? Possibly there could be a solution which is based on performing loadCurrentVersionFromIndex in batches.
2. Given that I properly orchestrate clients and prevent concurrent updates, I suppose I could disable concurrency control for batch inserts by patching InternalEngine.java and providing some API. Am I right? What am I missing?

More info: indexing in batches with size 10 000, using TransportClient; bottleneck is CPU cycles (memory and disks are fine); GC activity is minimal;

Thanks!

---

<div class="post-metadata">

**Author:** ![javanna](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javanna/32/4698_2.png) [@javanna](https://discuss.elastic.co/u/javanna)\
**Post date:** [September 12, 2016, 12:38pm UTC](https://discuss.elastic.co/t/significant-overhead-of-concurrency-control-for-small-docs/60276/2 "2016-09-12T12:38:39Z")

</div>

This is something that we've been thinking about for a while. We started working on it also:

> <https://github.com/elastic/elasticsearch/issues/19813>

> <https://github.com/elastic/elasticsearch/pull/20211>

That's all to say that things will only get better with the next versions. The latter improvement will go out with 5.0.0.

---

<div class="post-metadata">

**Author:** ![Zline](https://avatars.discourse-cdn.com/v4/letter/z/e5b9ba/32.png) [@Zline](https://discuss.elastic.co/u/Zline)\
**Post date:** [September 12, 2016, 12:53pm UTC](https://discuss.elastic.co/t/significant-overhead-of-concurrency-control-for-small-docs/60276/3 "2016-09-12T12:53:45Z")

</div>

Thanks for your reply, Luca.  
Nice feature, I see even in [2.3.5](https://github.com/elastic/elasticsearch/blob/v2.3.5/core/src/main/java/org/elasticsearch/index/engine/InternalEngine.java#L360) one could leverage autogenerated IDs to bypass version control.  
Unfortunately, autogenerated IDs and append-only is not an option for me. In my case append-only stage is transient 😉

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:21pm UTC](https://discuss.elastic.co/t/significant-overhead-of-concurrency-control-for-small-docs/60276/4 "2017-07-05T22:21:04Z")

</div>


