# Es-Hadoop Doc Value Access

**URL:** <https://discuss.elastic.co/t/es-hadoop-doc-value-access/21910>\
**Category:** Elasticsearch\
**Created:** [January 29, 2015, 6:52pm UTC](https://discuss.elastic.co/t/es-hadoop-doc-value-access/21910 "2015-01-29T18:52:10Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![chuck](https://avatars.discourse-cdn.com/v4/letter/c/8e7dd6/32.png) [@chuck](https://discuss.elastic.co/u/chuck)\
**Post date:** [January 29, 2015, 6:52pm UTC](https://discuss.elastic.co/t/es-hadoop-doc-value-access/21910/1 "2015-01-29T18:52:10Z")

</div>

I'm curious about reaching deeper into the lucene internals with es-hadoop,  
in a similar way that the aggregations module works. While aggregations  
are amazing, there are cases where they aren't an ideal solution, mainly  
due to the inability to shuffle/repartition the data as it moves through an  
analytic. I realize the current implementation can pull single fields by  
using an include/exclude on the query, but since this has to go to the  
source it does not strike me as a performant solution. With an es-spark  
interface that could pull doc values/doc ids in a similar way that  
aggregations do, it would be possible to create arbitrary analytics on any  
query context. Has any thought been given to this?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/9093e6de-370d-4f5d-a2ad-1f6f919f9d9f%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/9093e6de-370d-4f5d-a2ad-1f6f919f9d9f%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![chuck](https://avatars.discourse-cdn.com/v4/letter/c/8e7dd6/32.png) [@chuck](https://discuss.elastic.co/u/chuck)\
**Post date:** [January 30, 2015, 6:47pm UTC](https://discuss.elastic.co/t/es-hadoop-doc-value-access/21910/2 "2015-01-30T18:47:16Z")

</div>

So I guess I missed that fielddata fields could be specified in the search  
request body. That's pretty cool!

On Thursday, January 29, 2015 at 1:52:10 PM UTC-5, Elliott Bradshaw wrote:

> I'm curious about reaching deeper into the lucene internals with  
> es-hadoop, in a similar way that the aggregations module works. While  
> aggregations are amazing, there are cases where they aren't an ideal  
> solution, mainly due to the inability to shuffle/repartition the data as it  
> moves through an analytic. I realize the current implementation can pull  
> single fields by using an include/exclude on the query, but since this has  
> to go to the source it does not strike me as a performant solution. With  
> an es-spark interface that could pull doc values/doc ids in a similar way  
> that aggregations do, it would be possible to create arbitrary analytics on  
> any query context. Has any thought been given to this?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/3746683a-9823-4a04-aaf4-03d36c6d6b89%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/3746683a-9823-4a04-aaf4-03d36c6d6b89%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:35am UTC](https://discuss.elastic.co/t/es-hadoop-doc-value-access/21910/3 "2017-07-06T00:35:30Z")

</div>


