# Iterating through large datasets in various orders efficiently

**URL:** <https://discuss.elastic.co/t/iterating-through-large-datasets-in-various-orders-efficiently/362743>\
**Category:** Elasticsearch\
**Created:** [July 8, 2024, 10:41pm UTC](https://discuss.elastic.co/t/iterating-through-large-datasets-in-various-orders-efficiently/362743 "2024-07-08T22:41:35Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 9, 2024, 7:09am UTC](https://discuss.elastic.co/t/iterating-through-large-datasets-in-various-orders-efficiently/362743/2 "2024-07-09T07:09:15Z")

</div>

> [@Chervine\_Majeri](#):
>
> If I have a query with two filters, the engine will have two inverted indices to choose from, it will likely pick the one with a lower number of matching docs, and it will simply "filter out" in software the other one.

In Elasticsearch all fields are indexed individually and at query time the indices of multiple fields will be used, not just a single one. In the example you mentioned both indices will be queried and the sets of matches will be combined based on the logic of the query. The restriction that only one index can be used, which exists in some databases, does not apply to Elasticsearch.

> [@Chervine\_Majeri](#):
>
> I also understand that, even though "compound" inverted indices don't exist, we can kind of "make them up" using the "copy\_to" feature, to create a compound field, and then filter out on that field.

There is no need for this.

> [@Chervine\_Majeri](#):
>
> Now what I really don't understand however, is the sorting.  
> From what I understand, the main way data is sorted is in memory, after the entries have actually been fetched.

The shards return information about the matches they have identified to the node coordinating the query, which I believe is enough to score and sort. I believe only the documents that are to be returned are fetched from disk.

> [@Chervine\_Majeri](#):
>
> To solve this, there's the "sorted index" feature, which will (provided we don't care about number of matches) essentially only read as many entries as it returns (given a single shard). That is, it won't read all potential 10M entries to return a page of 10.

If I recall correctly (have not used this feature much), the sorted index feature is a very specialised feature that optimises data on disk if you have a specific sort order that you always use, and can cut latencies this way. The trade-off is that it adds overhead at indexing time and does not help for any other sort order. This is therefore something that you only typically use for very specific scenarios.

> [@Chervine\_Majeri](#):
>
> This also synergizes well with the "search\_after" feature, given the index is sorted, it's easier for elasticsearch to actually lookup the start of the data it needs to return.

This is in no way required for search after functionality.

> [@Chervine\_Majeri](#):
>
> But what happens if I want to query sorting under a different key?  
> Should I create a second sorted index, and insert all my entries twice? Is this a common use-case?

Sorting in different ways is standard and efficient in Elasticsearch and does not require any special handling. Just use a standard index.

---

_[View the full topic](https://discuss.elastic.co/t/iterating-through-large-datasets-in-various-orders-efficiently/362743)._
