# Sort by \_id field

**URL:** <https://discuss.elastic.co/t/sort-by-id-field/169017>\
**Category:** Elasticsearch\
**Created:** [February 19, 2019, 12:10pm UTC](https://discuss.elastic.co/t/sort-by-id-field/169017 "2019-02-19T12:10:19Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![MillQK](https://avatars.discourse-cdn.com/v4/letter/m/aeb1de/32.png) [@MillQK](https://discuss.elastic.co/u/MillQK)\
**Post date:** [February 19, 2019, 12:10pm UTC](https://discuss.elastic.co/t/sort-by-id-field/169017/1 "2019-02-19T12:10:20Z")

</div>

Hello! I want to use `search after` to process all documents but document hasn't any unique field, only \_id. That's why I want use \_id for sorting. But there are some important note in [search after docs](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-search-after.html) about my case. I have several questions:

1. Is overhead really big?
2. Can I create script field with source `doc['_id'].value` or overhead will not disappear?
3. Can I set `doc_value = true` for \_id field (I using ids generated by myself, not auto-generated by elastic)?

Will appreciate for some advices.  
Thanks!

---

<div class="post-metadata">

**Author:** ![gbrown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gbrown/32/34482_2.png) [@gbrown](https://discuss.elastic.co/u/gbrown)\
**Post date:** [February 19, 2019, 7:27pm UTC](https://discuss.elastic.co/t/sort-by-id-field/169017/2 "2019-02-19T19:27:15Z")

</div>

Hi!

You mention wanting to use search-after to "process all documents" - before I answer your question, I want to ask if [the Scroll API](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-scroll.html) might be suited to your use case - this would alleviate the issues you're having with search-after. If you're processing all documents returned from a search all at once and just need the results returned in batches, consider using the scroll API instead.

Search-after is the correct choice, though, if your workload 1) has large delays between retrieving batches of results, or 2) has many clients which need to maintain independent contexts.

To answer your question:

1. The overhead is pretty significant - we generally don't make recommendations like that in our docs if we aren't pretty sure that it will cause problems.
2. I don't believe using a script field would be any better - the problem is how the `_id` field is stored on disk in comparison to fields with `doc_values` enabled.
3. No, unfortunately this is not currently possible, which is why recommend copying the `_id` into a regular document field.

---

<div class="post-metadata">

**Author:** ![MillQK](https://avatars.discourse-cdn.com/v4/letter/m/aeb1de/32.png) [@MillQK](https://discuss.elastic.co/u/MillQK)\
**Post date:** [February 21, 2019, 6:35am UTC](https://discuss.elastic.co/t/sort-by-id-field/169017/3 "2019-02-21T06:35:57Z")

</div>

Thanks for your great answer!

`Scroll API` documentation has note: _Scrolling is not intended for real time user requests_, but real time user requests is my case, that is why this api not good for me.

Well, then i will copy `_id` to doc field as recommended.

Can you tell me or may be share some article why `_id` has this significant restrictions?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 21, 2019, 6:50am UTC](https://discuss.elastic.co/t/sort-by-id-field/169017/4 "2019-03-21T06:50:06Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
