# Sorting docs as input to facet phase

**URL:** <https://discuss.elastic.co/t/sorting-docs-as-input-to-facet-phase/13446>\
**Category:** Elasticsearch\
**Created:** [September 3, 2013, 1:31pm UTC](https://discuss.elastic.co/t/sorting-docs-as-input-to-facet-phase/13446 "2013-09-03T13:31:39Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Tikitu\_de\_Jager](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tikitu_de_jager/32/1258_2.png) [@Tikitu\_de\_Jager](https://discuss.elastic.co/u/Tikitu_de_Jager)\
**Post date:** [September 3, 2013, 1:31pm UTC](https://discuss.elastic.co/t/sorting-docs-as-input-to-facet-phase/13446/1 "2013-09-03T13:31:39Z")

</div>

Hi folks,

I'm building a custom facet that would benefit greatly if I could feed it  
its documents in a predefined order: the space requirements are _much_ smaller  
if I can guarantee that all documents that share the same value on a  
particular field pass through the facet collector in one bunch.

I.e., this ordering is cheap (grouping by "tweet"):

```
{"tweet": 3, "label": 5}
{"tweet": 3, "label": 7}
{"tweet": 4, "label": 3}
{"tweet": 4: "label": 5}

```

but this is expensive:

```
{"tweet": 3, "label": 5}
{"tweet": 4, "label": 3}
{"tweet": 3, "label": 7}
{"tweet": 4: "label": 5}

```

(I'm only interested in aggregate statistics across all "tweet" values, but  
I can't calculate the per-tweet value until I'm sure no more labels are  
coming -- the actual case is somewhat more complicated, and involves some  
timestamp calculations as well, but I think that's irrelevant.)

Is there any way to achieve this? I'm thinking maybe using nested documents  
("tweet" is actually a parent-doc ID, but I was hoping to use parent/child  
docs to avoid the reindexing requirement).

Regards,  
Tikitu

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Tikitu\_de\_Jager](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tikitu_de_jager/32/1258_2.png) [@Tikitu\_de\_Jager](https://discuss.elastic.co/u/Tikitu_de_Jager)\
**Post date:** [September 7, 2013, 10:13am UTC](https://discuss.elastic.co/t/sorting-docs-as-input-to-facet-phase/13446/2 "2013-09-07T10:13:04Z")

</div>

In case anyone finds this in the archives: as far as I can see this is  
basically not possible at all with parent/child docs (at least without  
reimplementing sorting yourself).

Nested docs, on the other hand, guarantee that the parent and children are  
adjacent in the same segment; it's quite easy to make a custom Collector  
which hands off both the parent and child docIds to (specialised versions  
of) doCollect():

```
/**
 * Modified from 

```

org.elasticsearch.index.search.nested.NestedChildrenCollector.java  
\*  
\* Collect root doc first, then all nested docs; send them to different  
methods.  
\*/  
@Override  
public void collect(int parentDoc) throws IOException {  
if (parentDoc == 0 || parentDocs == null) {  
return;  
}  
doCollectParent(parentDoc);  
int prevParentDoc = parentDocs.prevSetBit(parentDoc - 1);  
for (int i = (parentDoc - 1); i \> prevParentDoc; i--) {  
if (!currentReader.isDeleted(i) && childDocs.get(i)) {  
doCollectChild(i);  
}  
}  
}

On Tuesday, 3 September 2013 16:31:39 UTC+3, Tikitu de Jager wrote:

> Hi folks,
> 
> I'm building a custom facet that would benefit greatly if I could feed it  
> its documents in a predefined order: the space requirements are _much_ smaller  
> if I can guarantee that all documents that share the same value on a  
> particular field pass through the facet collector in one bunch.
> 
> I.e., this ordering is cheap (grouping by "tweet"):
> 
> ```
> {"tweet": 3, "label": 5}
> {"tweet": 3, "label": 7}
> {"tweet": 4, "label": 3}
> {"tweet": 4: "label": 5}
> 
> ```
> 
> but this is expensive:
> 
> ```
> {"tweet": 3, "label": 5}
> {"tweet": 4, "label": 3}
> {"tweet": 3, "label": 7}
> {"tweet": 4: "label": 5}
> 
> ```
> 
> (I'm only interested in aggregate statistics across all "tweet" values,  
> but I can't calculate the per-tweet value until I'm sure no more labels are  
> coming -- the actual case is somewhat more complicated, and involves some  
> timestamp calculations as well, but I think that's irrelevant.)
> 
> Is there any way to achieve this? I'm thinking maybe using nested  
> documents ("tweet" is actually a parent-doc ID, but I was hoping to use  
> parent/child docs to avoid the reindexing requirement).
> 
> Regards,  
> Tikitu

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:17am UTC](https://discuss.elastic.co/t/sorting-docs-as-input-to-facet-phase/13446/3 "2017-07-06T02:17:38Z")

</div>


