# Accessing \_id in ingest pipeline

**URL:** <https://discuss.elastic.co/t/accessing-id-in-ingest-pipeline/176503>\
**Category:** Elasticsearch\
**Created:** [April 11, 2019, 9:23pm UTC](https://discuss.elastic.co/t/accessing-id-in-ingest-pipeline/176503 "2019-04-11T21:23:51Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Oleg-Arkhipov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/oleg-arkhipov/32/41041_2.png) [@Oleg-Arkhipov](https://discuss.elastic.co/u/Oleg-Arkhipov)\
**Post date:** [April 11, 2019, 9:23pm UTC](https://discuss.elastic.co/t/accessing-id-in-ingest-pipeline/176503/1 "2019-04-11T21:23:51Z")

</div>

[Search After documentation](https://www.elastic.co/guide/en/elasticsearch/reference/6.7/search-request-search-after.html) (as well as one or two other places in docs) state that aggregation and sorting on `_id` field is inefficient, and it is better to create a duplicate of `_id` in another ordinary field with `doc_values` enabled. Doc also suggests using ingest pipeline:

> Instead it is advised to duplicate (client side or with a [set ingest processor](https://www.elastic.co/guide/en/elasticsearch/reference/6.7/ingest-processors.html))

However, I haven't found a way myself or any online example showing how to do that. It looks like pipeline is executed before assigning autogenerated `_id`. The following processor:

```
{
  "set": {
    "field": "tie_breaker_id",
    "value": "{{_id}}"
  }
}

```

assigns empty string to `tie_breaker_id` field.

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [April 12, 2019, 6:19pm UTC](https://discuss.elastic.co/t/accessing-id-in-ingest-pipeline/176503/2 "2019-04-12T18:19:10Z")

</div>

Hey,

indeed. The id generation happens after the ingest pipeline is applied. I opened an issue at [https://github.com/elastic/elasticsearch/issues/41163](https://github.com/elastic/elasticsearch/issues/41163)

--Alex

---

<div class="post-metadata">

**Author:** ![Oleg-Arkhipov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/oleg-arkhipov/32/41041_2.png) [@Oleg-Arkhipov](https://discuss.elastic.co/u/Oleg-Arkhipov)\
**Post date:** [April 12, 2019, 6:31pm UTC](https://discuss.elastic.co/t/accessing-id-in-ingest-pipeline/176503/3 "2019-04-12T18:31:26Z")

</div>

Thanks for the reply. What would you suggest to implement desired `_id` field duplication? Only client-side processing (2 requests)?

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [April 12, 2019, 6:34pm UTC](https://discuss.elastic.co/t/accessing-id-in-ingest-pipeline/176503/4 "2019-04-12T18:34:10Z")

</div>

May I ask about more information of the use-case here? If it is logging, I am wondering if search after is needed, if it is something else I am curious to get to know more.

That said, client side id generation and configuring an additional field would work indeed.

---

<div class="post-metadata">

**Author:** ![Oleg-Arkhipov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/oleg-arkhipov/32/41041_2.png) [@Oleg-Arkhipov](https://discuss.elastic.co/u/Oleg-Arkhipov)\
**Post date:** [April 12, 2019, 8:07pm UTC](https://discuss.elastic.co/t/accessing-id-in-ingest-pipeline/176503/5 "2019-04-12T20:07:08Z")

</div>

It is not logging, it is storing some user-generated resources (like blog posts, for example) with different attributes for a powerful and fast search using Elasticsearch. I am using `search_after` for pagination (using `["_score", "_id"]` as sorting parameters), because as I understand it is the optimal (if not only) way for traditional real-time user pagination.

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [April 12, 2019, 8:12pm UTC](https://discuss.elastic.co/t/accessing-id-in-ingest-pipeline/176503/6 "2019-04-12T20:12:46Z")

</div>

if it is a blog post, maybe using the URL (or its slug) as the id might simplify things and also allow for stable lookups (and that field can also be part of the document itself), as that is another part of the data that should be unique?

---

<div class="post-metadata">

**Author:** ![Oleg-Arkhipov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/oleg-arkhipov/32/41041_2.png) [@Oleg-Arkhipov](https://discuss.elastic.co/u/Oleg-Arkhipov)\
**Post date:** [April 12, 2019, 8:23pm UTC](https://discuss.elastic.co/t/accessing-id-in-ingest-pipeline/176503/7 "2019-04-12T20:23:55Z")

</div>

Of course, I also have `id` field with autogenerated ID from my primary RDBMS. But I have a collection where each "blog post" may have multiple documents, which are duplicates except for one field with geopoint (this structure is used for aggregation with filters to display points on a map), so any otherwise unique attributes of "blog post" are not unique here. I guess it will be easier therefore to generate elastic id on the client. One more question: do you think it is better to populate both `_id` and `tie_breaker_id` with the same generated value, or to generate only `tie_breaker_id` and use it, leaving `_id` alone? I don't want to bring any problems by creating my own primary ids.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 10, 2019, 8:24pm UTC](https://discuss.elastic.co/t/accessing-id-in-ingest-pipeline/176503/8 "2019-05-10T20:24:07Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
