# Bad bulk performance with self-generated id

**URL:** <https://discuss.elastic.co/t/bad-bulk-performance-with-self-generated-id/103344>\
**Category:** Elasticsearch\
**Created:** [October 10, 2017, 10:18am UTC](https://discuss.elastic.co/t/bad-bulk-performance-with-self-generated-id/103344 "2017-10-10T10:18:39Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![ginger](https://avatars.discourse-cdn.com/v4/letter/g/34f0e0/32.png) [@ginger](https://discuss.elastic.co/u/ginger)\
**Post date:** [October 10, 2017, 10:18am UTC](https://discuss.elastic.co/t/bad-bulk-performance-with-self-generated-id/103344/1 "2017-10-10T10:18:40Z")

</div>

Hi, all.

We recently use ES to store monitor data and depend on self-generated id to remove duplicate data. Our ES configure is as follows:

> 3 node(24core, 128GB, 3T SSD)  
> -Xms30g -Xmx30g  
> indices.memory.index\_buffer\_size: 15%  
> index.store.throttle.type: none

After running 24h, the bulk performance is about 5x slower than the auto-generated id.

We have adjusted our id to: **(timestamp/1000) + md5(monitor object fields) + (timestamp%1000)**, refering to the following blog:  
[Choosing a fast unique identifier (UUID) for Lucene](http://blog.mikemccandless.com/2014/05/choosing-fast-unique-identifier-uuid.html)  
I would add that the split of the timestamp is mainly for storage consideration. Without this, the disk usage rised by 25% percent.

According to jstack and jvmtop, the main cpu resource is mainly consumed in docid lookup. We are trying to generate larger initial segment and speed up merge process, so that less segements would need to lookup.

I hope I have explained clearly my question, and any help is appreciated.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 10, 2017, 10:46am UTC](https://discuss.elastic.co/t/bad-bulk-performance-with-self-generated-id/103344/2 "2017-10-10T10:46:47Z")

</div>

Have you plotted indexing throughput as a function of shard/index size? Is your data arriving in near real-time so that timestamps are largely sequential? How large portion of your data end up being updates?

---

<div class="post-metadata">

**Author:** ![ginger](https://avatars.discourse-cdn.com/v4/letter/g/34f0e0/32.png) [@ginger](https://discuss.elastic.co/u/ginger)\
**Post date:** [October 10, 2017, 11:29am UTC](https://discuss.elastic.co/t/bad-bulk-performance-with-self-generated-id/103344/3 "2017-10-10T11:29:21Z")

</div>

The data is mainly monitor data, which arrive continuously.  
Only very small portion will be updated.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 10, 2017, 11:32am UTC](https://discuss.elastic.co/t/bad-bulk-performance-with-self-generated-id/103344/4 "2017-10-10T11:32:08Z")

</div>

How large are the shards after the 24 hours?

---

<div class="post-metadata">

**Author:** ![ginger](https://avatars.discourse-cdn.com/v4/letter/g/34f0e0/32.png) [@ginger](https://discuss.elastic.co/u/ginger)\
**Post date:** [October 10, 2017, 11:33am UTC](https://discuss.elastic.co/t/bad-bulk-performance-with-self-generated-id/103344/5 "2017-10-10T11:33:57Z")

</div>

We hava 27 index. For some big index, each has multiple shards which is about 10GB, for small index, each will have 3 shards.

---

<div class="post-metadata">

**Author:** ![ginger](https://avatars.discourse-cdn.com/v4/letter/g/34f0e0/32.png) [@ginger](https://discuss.elastic.co/u/ginger)\
**Post date:** [October 11, 2017, 5:48am UTC](https://discuss.elastic.co/t/bad-bulk-performance-with-self-generated-id/103344/6 "2017-10-11T05:48:03Z")

</div>

@Christian_Dahlqvist Any better idea?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 11, 2017, 6:14am UTC](https://discuss.elastic.co/t/bad-bulk-performance-with-self-generated-id/103344/7 "2017-10-11T06:14:30Z")

</div>

What is the output of the [cluster stats API](https://www.elastic.co/guide/en/elasticsearch/reference/5.6/cluster-stats.html)? What is the size of your bulk requests?

Also, which Elasticsearch version are you using?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [October 11, 2017, 6:52am UTC](https://discuss.elastic.co/t/bad-bulk-performance-with-self-generated-id/103344/8 "2017-10-11T06:52:56Z")

</div>

If I understand the scenario correctly what you may be seeing is the added cost of doing a read on a growing pile of docs with every write. If you supply the ids there's no way for you to tell us "trust me, this doc doesn't exist". We'll always have to check that id is not present already.  
If you update rarely then maybe use auto generated ids and add a "my\_id" field and use update by query on that for those rare scenarios where you need to update. It'll be slower to update and you lose the only-one guarantee but the inserts should be faster.

---

<div class="post-metadata">

**Author:** ![ginger](https://avatars.discourse-cdn.com/v4/letter/g/34f0e0/32.png) [@ginger](https://discuss.elastic.co/u/ginger)\
**Post date:** [October 11, 2017, 8:07am UTC](https://discuss.elastic.co/t/bad-bulk-performance-with-self-generated-id/103344/9 "2017-10-11T08:07:16Z")

</div>

Each bulk request is 2k events. The cluster state is not available, what information do you want.

---

<div class="post-metadata">

**Author:** ![ginger](https://avatars.discourse-cdn.com/v4/letter/g/34f0e0/32.png) [@ginger](https://discuss.elastic.co/u/ginger)\
**Post date:** [October 11, 2017, 8:16am UTC](https://discuss.elastic.co/t/bad-bulk-performance-with-self-generated-id/103344/10 "2017-10-11T08:16:16Z")

</div>

We depend on id to remove duplicate data. so we always have to query by the id.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [October 11, 2017, 8:28am UTC](https://discuss.elastic.co/t/bad-bulk-performance-with-self-generated-id/103344/11 "2017-10-11T08:28:22Z")

</div>

> [@ginger](#):
>
> We depend on id to remove duplicate data. so we always have to query by the id.

You can have a field called "my\_id" (or whatever) and query on that.  
**Pros** :  
Elasticsearch couldn't care less about what you put in there so writes need not be slowed down with uniqueness checks for `my_id` values.  
**Cons** :  
Elasticsearch couldn't care less about what you put in there

It is possible to run a query for all docs with `my_id:foo` and either delete them or update them in order to perform an update. The downside of this is that unlike the elasticsearch-managed `id` field:

1. By default there is no fast-routing that knows which shard `foo` is on - all shards must be searched.
2. There are no guarantees that the index doesn't contain 2 docs with `my_id:foo` if your client app inserts the same data twice (forgetting to delete or update any prior doc).

(Perhaps worth pointing out this technique resolved a performance issue for a user indexing all-the-tweets-in-the-world using the original tweet id. That was an earlier version of elasticsearch and id lookups may have improved since but a design with no lookups will always be faster than one that requires them.)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 11, 2017, 8:37am UTC](https://discuss.elastic.co/t/bad-bulk-performance-with-self-generated-id/103344/12 "2017-10-11T08:37:59Z")

</div>

I am not asking for the cluster state, just cluster stats (statistics) to get on overview of the cluster.

Do you have monitoring installed?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [October 11, 2017, 8:46am UTC](https://discuss.elastic.co/t/bad-bulk-performance-with-self-generated-id/103344/13 "2017-10-11T08:46:47Z")

</div>

> The cluster state is not available

Why?

---

<div class="post-metadata">

**Author:** ![ginger](https://avatars.discourse-cdn.com/v4/letter/g/34f0e0/32.png) [@ginger](https://discuss.elastic.co/u/ginger)\
**Post date:** [October 11, 2017, 9:48am UTC](https://discuss.elastic.co/t/bad-bulk-performance-with-self-generated-id/103344/14 "2017-10-11T09:48:25Z")

</div>

Thanks for the reply. We have just stop the test today, I reproduce the environment latter and supply the cluster state.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [October 11, 2017, 10:00am UTC](https://discuss.elastic.co/t/bad-bulk-performance-with-self-generated-id/103344/15 "2017-10-11T10:00:17Z")

</div>

Not needed as @Christian_Dahlqvist said but I was just curious about why you said that.

---

<div class="post-metadata">

**Author:** ![ginger](https://avatars.discourse-cdn.com/v4/letter/g/34f0e0/32.png) [@ginger](https://discuss.elastic.co/u/ginger)\
**Post date:** [October 12, 2017, 2:40am UTC](https://discuss.elastic.co/t/bad-bulk-performance-with-self-generated-id/103344/16 "2017-10-12T02:40:06Z")

</div>

BTW, any advice about:

1. How to generate larger initial segment?
2. How to merge floor segments first?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [October 12, 2017, 4:35am UTC](https://discuss.elastic.co/t/bad-bulk-performance-with-self-generated-id/103344/17 "2017-10-12T04:35:32Z")

</div>

I’d not try to change elasticsearch behavior. Why do you think defaults are not good?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 9, 2017, 4:35am UTC](https://discuss.elastic.co/t/bad-bulk-performance-with-self-generated-id/103344/18 "2017-11-09T04:35:38Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
