# Is possible to extract \_id doc with Rally and custom track?

**URL:** <https://discuss.elastic.co/t/is-possible-to-extract-id-doc-with-rally-and-custom-track/319070>\
**Category:** Elasticsearch\
**Tags:** rally\
**Created:** [November 16, 2022, 11:39am UTC](https://discuss.elastic.co/t/is-possible-to-extract-id-doc-with-rally-and-custom-track/319070 "2022-11-16T11:39:27Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![rschirin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rschirin/32/45283_2.png) [@rschirin](https://discuss.elastic.co/u/rschirin)\
**Post date:** [November 16, 2022, 11:39am UTC](https://discuss.elastic.co/t/is-possible-to-extract-id-doc-with-rally-and-custom-track/319070/1 "2022-11-16T11:39:28Z")

</div>

Hey there,  
I would like to know if could be possible to extract the `_id` document when creating a `custom-track` with Rally.  
The goal is to test write performance when the `_id` is provided and not

thanks

---

<div class="post-metadata">

**Author:** ![RickBoyd](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rickboyd/32/80022_2.png) [@RickBoyd](https://discuss.elastic.co/u/RickBoyd)\
**Post date:** [November 16, 2022, 1:58pm UTC](https://discuss.elastic.co/t/is-possible-to-extract-id-doc-with-rally-and-custom-track/319070/2 "2022-11-16T13:58:43Z")

</div>

Unfortunately this is not implemented currently in Rally. We have an open enhancement request here: [Support includes-action-and-meta-data for track generation · Issue #1134 · elastic/rally · GitHub](https://github.com/elastic/rally/issues/1134)

Alternatively, you could change the extracted data as the user did [here](https://discuss.elastic.co/t/set-custom-document-ids-on-bulk-insert/258455/3)

---

<div class="post-metadata">

**Author:** ![RickBoyd](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rickboyd/32/80022_2.png) [@RickBoyd](https://discuss.elastic.co/u/RickBoyd)\
**Post date:** [November 16, 2022, 2:04pm UTC](https://discuss.elastic.co/t/is-possible-to-extract-id-doc-with-rally-and-custom-track/319070/3 "2022-11-16T14:04:39Z")

</div>

Also, note that if you don't need your own pre-existing IDs and just want to simulate provided vs auto-generated `_id` in general, the [documentation](https://esrally.readthedocs.io/en/stable/track.html#bulk) has this covered for the `bulk` operation parameters:

```auto
conflicts (optional): Type of index conflicts to simulate. If not specified, no conflicts will be simulated (also read below on how to use external index ids with no conflicts). Valid values are: ‘sequential’ (A document id is replaced with a document id with a sequentially increasing id), ‘random’ (A document id is replaced with a document id with a random other id).

conflict-probability (optional, defaults to 25 percent): A number between [0, 100] that defines how many of the documents will get replaced. Combining conflicts=sequential and conflict-probability=0 makes Rally generate index ids by itself, instead of relying on Elasticsearch’s automatic id generation.

```

So, by default we use auto generation, but if you want Rally to provide explicit `_id` you would configure the operation with `conflicts` `sequential` and `conflict-probability` `0`

---

<div class="post-metadata">

**Author:** ![rschirin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rschirin/32/45283_2.png) [@rschirin](https://discuss.elastic.co/u/rschirin)\
**Post date:** [November 16, 2022, 2:14pm UTC](https://discuss.elastic.co/t/is-possible-to-extract-id-doc-with-rally-and-custom-track/319070/4 "2022-11-16T14:14:01Z")

</div>

Thank you for the help 🙂  
With `conflicts` `sequential` , will Rally generate a first random `_id` (applied to the first doc) and then will increase sequentially it on the following documents?

---

<div class="post-metadata">

**Author:** ![RickBoyd](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rickboyd/32/80022_2.png) [@RickBoyd](https://discuss.elastic.co/u/RickBoyd)\
**Post date:** [November 16, 2022, 2:56pm UTC](https://discuss.elastic.co/t/is-possible-to-extract-id-doc-with-rally-and-custom-track/319070/5 "2022-11-16T14:56:37Z")

</div>

The first `_id` generated is `0`, so you will need a clean index in your track execution. Also note that if you have more than one `bulk` client that each client will get a batch of `_id`s to use in parallel, so the arrival of the `_id` in the system itself will not be sequential. If you need to run multiple `bulk` tasks in a row or execute against an existing index, `random` may be a better choice

---

<div class="post-metadata">

**Author:** ![rschirin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rschirin/32/45283_2.png) [@rschirin](https://discuss.elastic.co/u/rschirin)\
**Post date:** [November 17, 2022, 6:41pm UTC](https://discuss.elastic.co/t/is-possible-to-extract-id-doc-with-rally-and-custom-track/319070/6 "2022-11-17T18:41:53Z")

</div>

probably this isn't the correct thread anyway I have a doubt about `random/sequential _id`.  
I supposed that `bulk` operations with `sequential _id` should be at the end a little be slower than `random _id`. Should this consideration be correct?  
After few tests I saw that there is no relevant time differences. How should be possible?  
Probably should I fill the cluster with a lot of docs before to use Rally? In this way ES should check more docs to guarantee the \_id integrity?  
I am ingesting 1 million of docs.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 15, 2022, 6:42pm UTC](https://discuss.elastic.co/t/is-possible-to-extract-id-doc-with-rally-and-custom-track/319070/7 "2022-12-15T18:42:33Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
