# Question regarding sharding when fetching term vector and TFIDF info

**URL:** https://discuss.elastic.co/t/question-regarding-sharding-when-fetching-term-vector-and-tfidf-info/13214
**Category:** Elasticsearch
**Created:** [August 14, 2013, 2:23pm UTC](https://discuss.elastic.co/t/question-regarding-sharding-when-fetching-term-vector-and-tfidf-info/13214 "2013-08-14T14:23:52Z")
**Posts on this page:** 5
**Page:** 1

<div class="post-metadata">

### Author: ![Clint\_Miller](https://avatars.discourse-cdn.com/v4/letter/c/ec9cab/32.png) [@Clint\_Miller](https://discuss.elastic.co/u/Clint_Miller)
#### Post date: [August 14, 2013, 2:23pm UTC](https://discuss.elastic.co/t/question-regarding-sharding-when-fetching-term-vector-and-tfidf-info/13214/1 "2013-08-14T14:23:52Z")

</div>

I've written a native script plugin to elasticsearch that extends  
AbstractSearchScript. My code grabs a handle to the  
org.apache.lucene.index.IndexReader and the docId. It then uses the  
IndexReader to fetch the term vector for the docId, the frequency of each  
term within the document, and the number of documents containing that term.  
All of that data is then stored in a json string which is sent back to the  
code performing the query.

We're currently on elasticsearch 0.20.5 (slightly old). I'm using the  
following IndexReader call to get the document count:

reader.docFreq()

I'm pretty sure my code will need to change when we upgrade to the latest  
elasticsearch as the Lucene IndexReader methods appear to have changed. No  
big deal.

My question: Is reader.docFreq() shard-aware? I'm assuming not. I'm  
assuming I'll only get the document count within the current shard.

If my assumption is correct, is there a way to get the document count for a  
term across all shards? Is there a way for my code to access the  
IndexReader's for the other shards? If there were, then I could just call  
reader.docFreq() for each shard and add the results.

(I'd happily upgrade to the latest elasticsearch if it makes it easier to  
solve this problem.)

Thanks for any help.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

### Author: ![Clint\_Miller](https://avatars.discourse-cdn.com/v4/letter/c/ec9cab/32.png) [@Clint\_Miller](https://discuss.elastic.co/u/Clint_Miller)
#### Post date: [August 14, 2013, 7:44pm UTC](https://discuss.elastic.co/t/question-regarding-sharding-when-fetching-term-vector-and-tfidf-info/13214/2 "2013-08-14T19:44:20Z")

</div>

I might have figured this out. Instead of getting the IndexReader by  
overriding the AbstractSearchScript method setNextReader(), I'm making the  
following call:

val searchContext = org.elasticsearch.search.internal.SearchContext.current  
searchContext.searcher.getIndexReader.docFreq(currTerm)

That seems to return a higher-level reader that rolls up the doc counts  
from a set of lower-level readers that might represent the shards?

I can also call searchContext.searcher.subReaders to get an array of child  
readers and call docFreq() on each. If I add all these up, I get the same  
value as searchContext.searcher.getIndexReader.docFreq().

Am I on the right track?

On Wednesday, August 14, 2013 9:23:52 AM UTC-5, Clint Miller wrote:

> I've written a native script plugin to elasticsearch that extends  
> AbstractSearchScript. My code grabs a handle to the  
> org.apache.lucene.index.IndexReader and the docId. It then uses the  
> IndexReader to fetch the term vector for the docId, the frequency of each  
> term within the document, and the number of documents containing that term.  
> All of that data is then stored in a json string which is sent back to the  
> code performing the query.
> 
> We're currently on elasticsearch 0.20.5 (slightly old). I'm using the  
> following IndexReader call to get the document count:
> 
> reader.docFreq()
> 
> I'm pretty sure my code will need to change when we upgrade to the latest  
> elasticsearch as the Lucene IndexReader methods appear to have changed. No  
> big deal.
> 
> My question: Is reader.docFreq() shard-aware? I'm assuming not. I'm  
> assuming I'll only get the document count within the current shard.
> 
> If my assumption is correct, is there a way to get the document count for  
> a term across all shards? Is there a way for my code to access the  
> IndexReader's for the other shards? If there were, then I could just call  
> reader.docFreq() for each shard and add the results.
> 
> (I'd happily upgrade to the latest elasticsearch if it makes it easier to  
> solve this problem.)
> 
> Thanks for any help.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

### Author: ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)
#### Post date: [August 14, 2013, 8:22pm UTC](https://discuss.elastic.co/t/question-regarding-sharding-when-fetching-term-vector-and-tfidf-info/13214/3 "2013-08-14T20:22:14Z")

</div>

IndexReader (or better AtomicReaderContext) is not exposed in the API for  
use from a (native) script. In 0.90+ the situation has not changed.

You are correct, you wouldn't be able to get the doc freq of the index, you  
would get the doc freq for the current shard only.

To compute the doc freq across all shards, you would have to write a plugin  
with a custom action that can distribute a request on index level to shard  
requests for the nodes that have the shards, and finally summarize the  
subresults. An example plugin is here

> **[jprante/elasticsearch-index-termlist](https://github.com/jprante/elasticsearch-index-termlist)**
>
> elasticsearch-index-termlist - Elasticsearch Index Termlist

Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

### Author: ![Clint\_Miller](https://avatars.discourse-cdn.com/v4/letter/c/ec9cab/32.png) [@Clint\_Miller](https://discuss.elastic.co/u/Clint_Miller)
#### Post date: [August 16, 2013, 1:23am UTC](https://discuss.elastic.co/t/question-regarding-sharding-when-fetching-term-vector-and-tfidf-info/13214/4 "2013-08-16T01:23:21Z")

</div>

Thank you very much. I was able to write a plugin based on your sample  
code, and it's working perfectly for me.

On Wednesday, August 14, 2013 3:22:14 PM UTC-5, Jörg Prante wrote:

> IndexReader (or better AtomicReaderContext) is not exposed in the API for  
> use from a (native) script. In 0.90+ the situation has not changed.
> 
> You are correct, you wouldn't be able to get the doc freq of the index,  
> you would get the doc freq for the current shard only.
> 
> To compute the doc freq across all shards, you would have to write a  
> plugin with a custom action that can distribute a request on index level to  
> shard requests for the nodes that have the shards, and finally summarize  
> the subresults. An example plugin is here  
> [GitHub - jprante/elasticsearch-index-termlist: Elasticsearch Index Termlist](https://github.com/jprante/elasticsearch-index-termlist)
> 
> Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 6, 2017, 2:20am UTC](https://discuss.elastic.co/t/question-regarding-sharding-when-fetching-term-vector-and-tfidf-info/13214/5 "2017-07-06T02:20:55Z")

</div>


