# Counting items in a list \[array\] returns (what we think) are incorrect counts via groovy

**URL:** <https://discuss.elastic.co/t/counting-items-in-a-list-array-returns-what-we-think-are-incorrect-counts-via-groovy/21547>\
**Category:** Elasticsearch\
**Created:** [January 9, 2015, 2:09am UTC](https://discuss.elastic.co/t/counting-items-in-a-list-array-returns-what-we-think-are-incorrect-counts-via-groovy/21547 "2015-01-09T02:09:00Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jeff\_Steinmetz](https://avatars.discourse-cdn.com/v4/letter/j/ecd19e/32.png) [@Jeff\_Steinmetz](https://discuss.elastic.co/u/Jeff_Steinmetz)\
**Post date:** [January 9, 2015, 2:09am UTC](https://discuss.elastic.co/t/counting-items-in-a-list-array-returns-what-we-think-are-incorrect-counts-via-groovy/21547/1 "2015-01-09T02:09:00Z")

</div>

Is there a better way to do this?

Please see this gist (or even better yet, run the script locally see the  
issue).

> <https://gist.github.com/jeffsteinmetz/2ea8329c667386c80fae>

You must have scripting enabled in your elasticsearch config for this to  
work.

This was originally based on some comments I found here:

> <https://stackoverflow.com/questions/17314123/search-by-size-of-object-type-field-elastic-search>

We would like to use a filtered query to only include documents that a  
small count of items in the list [aka array], filtering where  
values.size() \< 10

"script": "doc['titles'].values.size() \< 10"

Turns out the values.size() actually either counts tokenized (analyzed)  
words, or if the mapping turns off analysis, it still counts incorrectly if  
there are duplicates.  
If analyze is not turned off, it counts tokenized words, not the number of  
elements in the list.  
If analyze is turned off for a given field, it improves, but duplicates are  
missed.

For example, This comes back as size == 2  
"titles": ["one", "duplicate", "duplicate"]  
This comes back as size == 3, should be 4  
"titles": ["[http://bit.ly/abc](http://bit.ly/abc)", "[http://bit.ly/abc](http://bit.ly/abc)", "[http://bit.ly/def](http://bit.ly/def)",  
"[http://bit.ly/ghi](http://bit.ly/ghi)"]

Is this a bug, is there a better way, or is this just something that we  
don't understand about groovy and values.size()?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/f5e88338-8c4f-4cb8-b6c4-d7f47b365175%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/f5e88338-8c4f-4cb8-b6c4-d7f47b365175%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [January 9, 2015, 3:03am UTC](https://discuss.elastic.co/t/counting-items-in-a-list-array-returns-what-we-think-are-incorrect-counts-via-groovy/21547/2 "2015-01-09T03:03:57Z")

</div>

On Thu, Jan 8, 2015 at 9:09 PM, Jeff Steinmetz [jeffrey.steinmetz@gmail.com](mailto:jeffrey.steinmetz@gmail.com)  
wrote:

Is there a better way to do this?

> Please see this gist (or even better yet, run the script locally see the  
> issue).
> 
> [Determine list [array] size in elasticsearch issue · GitHub](https://gist.github.com/jeffsteinmetz/2ea8329c667386c80fae)
> 
> You must have scripting enabled in your elasticsearch config for this to  
> work.
> 
> This was originally based on some comments I found here:
> 
> [elasticsearch - Search by size of object type field elastic search - Stack Overflow](http://stackoverflow.com/questions/17314123/search-by-size-of-object-type-field-elastic-search)
> 
> We would like to use a filtered query to only include documents that a  
> small count of items in the list [aka array], filtering where  
> values.size() \< 10
> 
> "script": "doc['titles'].values.size() \< 10"
> 
> Turns out the values.size() actually either counts tokenized (analyzed)  
> words, or if the mapping turns off analysis, it still counts incorrectly if  
> there are duplicates.  
> If analyze is not turned off, it counts tokenized words, not the number of  
> elements in the list.  
> If analyze is turned off for a given field, it improves, but duplicates  
> are missed.
> 
> For example, This comes back as size == 2  
> "titles": ["one", "duplicate", "duplicate"]  
> This comes back as size == 3, should be 4  
> "titles": ["[Yahoo | Mail, Weather, Search, Politics, News, Finance, Sports & Videos](http://bit.ly/abc)", "[Yahoo | Mail, Weather, Search, Politics, News, Finance, Sports & Videos](http://bit.ly/abc)", "[Warning! | There might be a problem with the requested link](http://bit.ly/def)",  
> "[http://bit.ly/ghi](http://bit.ly/ghi)"]
> 
> Is this a bug, is there a better way, or is this just something that we  
> don't understand about groovy and values.size()?

I think that's just the way doc works. Try (but don't actually deploy)  
\_source['titles'].size() \< 10. That should do what you expect. Don't  
deploy that because its too slow. Try indexing the size and filtering on  
it. You can use a transform to add the size of the array as an integer  
field and just filter on it using a range filter. That'd probably be the  
fastest option.

Nik

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAPmjWd2d-KtOdV13trjnp3si\_7%2B%2BAnOd%2BTTeTN75jkBuMsywyQ%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAPmjWd2d-KtOdV13trjnp3si_7%2B%2BAnOd%2BTTeTN75jkBuMsywyQ%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Jeff\_Steinmetz](https://avatars.discourse-cdn.com/v4/letter/j/ecd19e/32.png) [@Jeff\_Steinmetz](https://discuss.elastic.co/u/Jeff_Steinmetz)\
**Post date:** [January 9, 2015, 5:03am UTC](https://discuss.elastic.co/t/counting-items-in-a-list-array-returns-what-we-think-are-incorrect-counts-via-groovy/21547/3 "2015-01-09T05:03:29Z")

</div>

Thank you, that worked.

I was curious about the speed, is running a script using \_source slower  
that doc ?

Totally understand a dynamic script is slower regardless of \_source vs  
doc.

Makes sense that having a count transformed up front during index to create  
a materialized value would certainly be much faster.

On Thursday, January 8, 2015 at 7:04:40 PM UTC-8, Nikolas Everett wrote:

> On Thu, Jan 8, 2015 at 9:09 PM, Jeff Steinmetz \<[jeffrey....@gmail.com](mailto:jeffrey....@gmail.com)  
> \<javascript:\>\> wrote:
> 
> Is there a better way to do this?
> 
> > Please see this gist (or even better yet, run the script locally see the  
> > issue).
> > 
> > [Determine list [array] size in elasticsearch issue · GitHub](https://gist.github.com/jeffsteinmetz/2ea8329c667386c80fae)
> > 
> > You must have scripting enabled in your elasticsearch config for this to  
> > work.
> > 
> > This was originally based on some comments I found here:
> > 
> > [elasticsearch - Search by size of object type field elastic search - Stack Overflow](http://stackoverflow.com/questions/17314123/search-by-size-of-object-type-field-elastic-search)
> > 
> > We would like to use a filtered query to only include documents that a  
> > small count of items in the list [aka array], filtering where  
> > values.size() \< 10
> > 
> > "script": "doc['titles'].values.size() \< 10"
> > 
> > Turns out the values.size() actually either counts tokenized (analyzed)  
> > words, or if the mapping turns off analysis, it still counts incorrectly if  
> > there are duplicates.  
> > If analyze is not turned off, it counts tokenized words, not the number  
> > of elements in the list.  
> > If analyze is turned off for a given field, it improves, but duplicates  
> > are missed.
> > 
> > For example, This comes back as size == 2  
> > "titles": ["one", "duplicate", "duplicate"]  
> > This comes back as size == 3, should be 4  
> > "titles": ["[Yahoo | Mail, Weather, Search, Politics, News, Finance, Sports & Videos](http://bit.ly/abc)", "[Yahoo | Mail, Weather, Search, Politics, News, Finance, Sports & Videos](http://bit.ly/abc)", "[Warning! | There might be a problem with the requested link](http://bit.ly/def)",  
> > "[http://bit.ly/ghi](http://bit.ly/ghi)"]
> > 
> > Is this a bug, is there a better way, or is this just something that we  
> > don't understand about groovy and values.size()?
> 
> I think that's just the way doc works. Try (but don't actually deploy)  
> \_source['titles'].size() \< 10. That should do what you expect. Don't  
> deploy that because its too slow. Try indexing the size and filtering on  
> it. You can use a transform to add the size of the array as an integer  
> field and just filter on it using a range filter. That'd probably be the  
> fastest option.
> 
> Nik

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/75736948-beac-43fc-84d4-25a94456d4ca%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/75736948-beac-43fc-84d4-25a94456d4ca%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [January 9, 2015, 5:15am UTC](https://discuss.elastic.co/t/counting-items-in-a-list-array-returns-what-we-think-are-incorrect-counts-via-groovy/21547/4 "2015-01-09T05:15:19Z")

</div>

Source is going to be pretty sloe, yeah. If its a one off then its probably  
fine but if you do it a lot probably best to index the count.  
On Jan 9, 2015 12:04 AM, "Jeff Steinmetz" [jeffrey.steinmetz@gmail.com](mailto:jeffrey.steinmetz@gmail.com)  
wrote:

> Thank you, that worked.
> 
> I was curious about the speed, is running a script using \_source slower  
> that doc ?
> 
> Totally understand a dynamic script is slower regardless of \_source vs  
> doc.
> 
> Makes sense that having a count transformed up front during index to  
> create a materialized value would certainly be much faster.
> 
> On Thursday, January 8, 2015 at 7:04:40 PM UTC-8, Nikolas Everett wrote:
> 
> > On Thu, Jan 8, 2015 at 9:09 PM, Jeff Steinmetz [jeffrey....@gmail.com](mailto:jeffrey....@gmail.com)  
> > wrote:
> > 
> > Is there a better way to do this?
> > 
> > > Please see this gist (or even better yet, run the script locally see the  
> > > issue).
> > > 
> > > [Determine list [array] size in elasticsearch issue · GitHub](https://gist.github.com/jeffsteinmetz/2ea8329c667386c80fae)
> > > 
> > > You must have scripting enabled in your elasticsearch config for this to  
> > > work.
> > > 
> > > This was originally based on some comments I found here:  
> > > [elasticsearch - Search by size of object type field elastic search - Stack Overflow](http://stackoverflow.com/questions/17314123/search-by-)  
> > > size-of-object-type-field-elastic-search
> > > 
> > > We would like to use a filtered query to only include documents that a  
> > > small count of items in the list [aka array], filtering where  
> > > values.size() \< 10
> > > 
> > > "script": "doc['titles'].values.size() \< 10"
> > > 
> > > Turns out the values.size() actually either counts tokenized (analyzed)  
> > > words, or if the mapping turns off analysis, it still counts incorrectly if  
> > > there are duplicates.  
> > > If analyze is not turned off, it counts tokenized words, not the number  
> > > of elements in the list.  
> > > If analyze is turned off for a given field, it improves, but duplicates  
> > > are missed.
> > > 
> > > For example, This comes back as size == 2  
> > > "titles": ["one", "duplicate", "duplicate"]  
> > > This comes back as size == 3, should be 4  
> > > "titles": ["[Yahoo | Mail, Weather, Search, Politics, News, Finance, Sports & Videos](http://bit.ly/abc)", "[Yahoo | Mail, Weather, Search, Politics, News, Finance, Sports & Videos](http://bit.ly/abc)", "[Warning! | There might be a problem with the requested link](http://bit.ly/def)",  
> > > "[http://bit.ly/ghi](http://bit.ly/ghi)"]
> > > 
> > > Is this a bug, is there a better way, or is this just something that we  
> > > don't understand about groovy and values.size()?
> > 
> > I think that's just the way doc works. Try (but don't actually deploy)  
> > \_source['titles'].size() \< 10. That should do what you expect. Don't  
> > deploy that because its too slow. Try indexing the size and filtering on  
> > it. You can use a transform to add the size of the array as an integer  
> > field and just filter on it using a range filter. That'd probably be the  
> > fastest option.
> > 
> > Nik
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/75736948-beac-43fc-84d4-25a94456d4ca%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/75736948-beac-43fc-84d4-25a94456d4ca%40googlegroups.com)  
> [https://groups.google.com/d/msgid/elasticsearch/75736948-beac-43fc-84d4-25a94456d4ca%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/75736948-beac-43fc-84d4-25a94456d4ca%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAPmjWd35LG%3Dki2jMigsfgwrojXVBTCkJH784wu7GbEcXvu3tRg%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAPmjWd35LG%3Dki2jMigsfgwrojXVBTCkJH784wu7GbEcXvu3tRg%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Jeff\_Steinmetz](https://avatars.discourse-cdn.com/v4/letter/j/ecd19e/32.png) [@Jeff\_Steinmetz](https://discuss.elastic.co/u/Jeff_Steinmetz)\
**Post date:** [January 9, 2015, 5:43am UTC](https://discuss.elastic.co/t/counting-items-in-a-list-array-returns-what-we-think-are-incorrect-counts-via-groovy/21547/5 "2015-01-09T05:43:35Z")

</div>

Transform worked well. Nice.

Curious how to get it to save to source? Tried this below, no go. (I can  
however do range queries agains title\_count, so transform was indexed and  
works well)

```
"transform" : {
  "script" : "ctx._source['\'title_count\''] = 

```

ctx.\_source[''titles''].size()",  
"lang": "groovy"  
},  
"properties": {  
"titles": { "type": "string", "index": "not\_analyzed" },  
"title\_count" : { "type": "integer", "store": "yes" }  
}  
}'

On Thursday, January 8, 2015 at 9:15:28 PM UTC-8, Nikolas Everett wrote:

> Source is going to be pretty sloe, yeah. If its a one off then its  
> probably fine but if you do it a lot probably best to index the count.  
> On Jan 9, 2015 12:04 AM, "Jeff Steinmetz" \<[jeffrey....@gmail.com](mailto:jeffrey....@gmail.com)  
> \<javascript:\>\> wrote:
> 
> > Thank you, that worked.
> > 
> > I was curious about the speed, is running a script using \_source slower  
> > that doc ?
> > 
> > Totally understand a dynamic script is slower regardless of \_source vs  
> > doc.
> > 
> > Makes sense that having a count transformed up front during index to  
> > create a materialized value would certainly be much faster.
> > 
> > On Thursday, January 8, 2015 at 7:04:40 PM UTC-8, Nikolas Everett wrote:
> > 
> > > On Thu, Jan 8, 2015 at 9:09 PM, Jeff Steinmetz [jeffrey....@gmail.com](mailto:jeffrey....@gmail.com)  
> > > wrote:
> > > 
> > > Is there a better way to do this?
> > > 
> > > > Please see this gist (or even better yet, run the script locally see  
> > > > the issue).
> > > > 
> > > > [Determine list [array] size in elasticsearch issue · GitHub](https://gist.github.com/jeffsteinmetz/2ea8329c667386c80fae)
> > > > 
> > > > You must have scripting enabled in your elasticsearch config for this  
> > > > to work.
> > > > 
> > > > This was originally based on some comments I found here:  
> > > > [elasticsearch - Search by size of object type field elastic search - Stack Overflow](http://stackoverflow.com/questions/17314123/search-by-)  
> > > > size-of-object-type-field-elastic-search
> > > > 
> > > > We would like to use a filtered query to only include documents that a  
> > > > small count of items in the list [aka array], filtering where  
> > > > values.size() \< 10
> > > > 
> > > > "script": "doc['titles'].values.size() \< 10"
> > > > 
> > > > Turns out the values.size() actually either counts tokenized (analyzed)  
> > > > words, or if the mapping turns off analysis, it still counts incorrectly if  
> > > > there are duplicates.  
> > > > If analyze is not turned off, it counts tokenized words, not the number  
> > > > of elements in the list.  
> > > > If analyze is turned off for a given field, it improves, but duplicates  
> > > > are missed.
> > > > 
> > > > For example, This comes back as size == 2  
> > > > "titles": ["one", "duplicate", "duplicate"]  
> > > > This comes back as size == 3, should be 4  
> > > > "titles": ["[Yahoo | Mail, Weather, Search, Politics, News, Finance, Sports & Videos](http://bit.ly/abc)", "[Yahoo | Mail, Weather, Search, Politics, News, Finance, Sports & Videos](http://bit.ly/abc)", "[Warning! | There might be a problem with the requested link](http://bit.ly/def)",  
> > > > "[http://bit.ly/ghi](http://bit.ly/ghi)"]
> > > > 
> > > > Is this a bug, is there a better way, or is this just something that we  
> > > > don't understand about groovy and values.size()?
> > > 
> > > I think that's just the way doc works. Try (but don't actually  
> > > deploy) \_source['titles'].size() \< 10. That should do what you expect.  
> > > Don't deploy that because its too slow. Try indexing the size and  
> > > filtering on it. You can use a transform to add the size of the array as  
> > > an integer field and just filter on it using a range filter. That'd  
> > > probably be the fastest option.
> > > 
> > > Nik
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/75736948-beac-43fc-84d4-25a94456d4ca%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/75736948-beac-43fc-84d4-25a94456d4ca%40googlegroups.com)  
> > [https://groups.google.com/d/msgid/elasticsearch/75736948-beac-43fc-84d4-25a94456d4ca%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/75736948-beac-43fc-84d4-25a94456d4ca%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > .  
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/00ff2bc1-94a9-4aa9-8c7e-ef5734affb4d%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/00ff2bc1-94a9-4aa9-8c7e-ef5734affb4d%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [January 9, 2015, 5:59am UTC](https://discuss.elastic.co/t/counting-items-in-a-list-array-returns-what-we-think-are-incorrect-counts-via-groovy/21547/6 "2015-01-09T05:59:48Z")

</div>

Transform never saves to source. You have to transform on the application  
side for that. It was designed for times when you wanted to index something  
like this that would just take up extra space in the source document. I  
imagine you could use a script field on the query if you need the result to  
contain the count. Or just count it on the result side.

Nik  
On Jan 9, 2015 12:43 AM, "Jeff Steinmetz" [jeffrey.steinmetz@gmail.com](mailto:jeffrey.steinmetz@gmail.com)  
wrote:

> Transform worked well. Nice.
> 
> Curious how to get it to save to source? Tried this below, no go. (I can  
> however do range queries agains title\_count, so transform was indexed and  
> works well)
> 
> ```
> "transform" : {
> "script" : "ctx._source['\'title_count\''] =
> 
> ```
> 
> ctx.\_source[''titles''].size()",  
> "lang": "groovy"  
> },  
> "properties": {  
> "titles": { "type": "string", "index": "not\_analyzed" },  
> "title\_count" : { "type": "integer", "store": "yes" }  
> }  
> }'
> 
> On Thursday, January 8, 2015 at 9:15:28 PM UTC-8, Nikolas Everett wrote:
> 
> > Source is going to be pretty sloe, yeah. If its a one off then its  
> > probably fine but if you do it a lot probably best to index the count.  
> > On Jan 9, 2015 12:04 AM, "Jeff Steinmetz" [jeffrey....@gmail.com](mailto:jeffrey....@gmail.com) wrote:
> > 
> > > Thank you, that worked.
> > > 
> > > I was curious about the speed, is running a script using \_source slower  
> > > that doc ?
> > > 
> > > Totally understand a dynamic script is slower regardless of \_source vs  
> > > doc.
> > > 
> > > Makes sense that having a count transformed up front during index to  
> > > create a materialized value would certainly be much faster.
> > > 
> > > On Thursday, January 8, 2015 at 7:04:40 PM UTC-8, Nikolas Everett wrote:
> > > 
> > > > On Thu, Jan 8, 2015 at 9:09 PM, Jeff Steinmetz [jeffrey....@gmail.com](mailto:jeffrey....@gmail.com)  
> > > > wrote:
> > > > 
> > > > Is there a better way to do this?
> > > > 
> > > > > Please see this gist (or even better yet, run the script locally see  
> > > > > the issue).
> > > > > 
> > > > > [Determine list [array] size in elasticsearch issue · GitHub](https://gist.github.com/jeffsteinmetz/2ea8329c667386c80fae)
> > > > > 
> > > > > You must have scripting enabled in your elasticsearch config for this  
> > > > > to work.
> > > > > 
> > > > > This was originally based on some comments I found here:  
> > > > > [elasticsearch - Search by size of object type field elastic search - Stack Overflow](http://stackoverflow.com/questions/17314123/search-by-size-)  
> > > > > of-object-type-field-elastic-search
> > > > > 
> > > > > We would like to use a filtered query to only include documents that a  
> > > > > small count of items in the list [aka array], filtering where  
> > > > > values.size() \< 10
> > > > > 
> > > > > "script": "doc['titles'].values.size() \< 10"
> > > > > 
> > > > > Turns out the values.size() actually either counts tokenized  
> > > > > (analyzed) words, or if the mapping turns off analysis, it still counts  
> > > > > incorrectly if there are duplicates.  
> > > > > If analyze is not turned off, it counts tokenized words, not the  
> > > > > number of elements in the list.  
> > > > > If analyze is turned off for a given field, it improves, but  
> > > > > duplicates are missed.
> > > > > 
> > > > > For example, This comes back as size == 2  
> > > > > "titles": ["one", "duplicate", "duplicate"]  
> > > > > This comes back as size == 3, should be 4  
> > > > > "titles": ["[Yahoo | Mail, Weather, Search, Politics, News, Finance, Sports & Videos](http://bit.ly/abc)", "[Yahoo | Mail, Weather, Search, Politics, News, Finance, Sports & Videos](http://bit.ly/abc)", "  
> > > > > [Warning! | There might be a problem with the requested link](http://bit.ly/def)", "[http://bit.ly/ghi](http://bit.ly/ghi)"]
> > > > > 
> > > > > Is this a bug, is there a better way, or is this just something that  
> > > > > we don't understand about groovy and values.size()?
> > > > 
> > > > I think that's just the way doc works. Try (but don't actually  
> > > > deploy) \_source['titles'].size() \< 10. That should do what you expect.  
> > > > Don't deploy that because its too slow. Try indexing the size and  
> > > > filtering on it. You can use a transform to add the size of the array as  
> > > > an integer field and just filter on it using a range filter. That'd  
> > > > probably be the fastest option.
> > > > 
> > > > Nik
> > > 
> > > --  
> > > You received this message because you are subscribed to the Google  
> > > Groups "elasticsearch" group.  
> > > To unsubscribe from this group and stop receiving emails from it, send  
> > > an email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com).  
> > > To view this discussion on the web visit [https://groups.google.com/d/](https://groups.google.com/d/)  
> > > msgid/elasticsearch/75736948-beac-43fc-84d4-25a94456d4ca%  
> > > [40googlegroups.com](http://40googlegroups.com)  
> > > [https://groups.google.com/d/msgid/elasticsearch/75736948-beac-43fc-84d4-25a94456d4ca%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/75736948-beac-43fc-84d4-25a94456d4ca%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > > .  
> > > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/00ff2bc1-94a9-4aa9-8c7e-ef5734affb4d%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/00ff2bc1-94a9-4aa9-8c7e-ef5734affb4d%40googlegroups.com)  
> > [https://groups.google.com/d/msgid/elasticsearch/00ff2bc1-94a9-4aa9-8c7e-ef5734affb4d%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/00ff2bc1-94a9-4aa9-8c7e-ef5734affb4d%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > .  
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAPmjWd1Z3H3xn255yTsvSoR-dhVRa7eGJCBcugt6oSb-MU9HHw%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAPmjWd1Z3H3xn255yTsvSoR-dhVRa7eGJCBcugt6oSb-MU9HHw%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Jeff\_Steinmetz](https://avatars.discourse-cdn.com/v4/letter/j/ecd19e/32.png) [@Jeff\_Steinmetz](https://discuss.elastic.co/u/Jeff_Steinmetz)\
**Post date:** [January 9, 2015, 7:19am UTC](https://discuss.elastic.co/t/counting-items-in-a-list-array-returns-what-we-think-are-incorrect-counts-via-groovy/21547/7 "2015-01-09T07:19:10Z")

</div>

Now that I am into the real wold scenario, it gets a bit tricker - I have  
nested objects (keys).  
I have to test the existence of the key in the Groovy script to avoid  
parsing errors on insert.

How do you access a nested object in groovy? and test for the existence of  
a nested object key?  
such as this example:

curl -XPOST 'http://'$NODE':9200/'$INDEX\_NAME'/post' -d '{  
"titles": ["title 1", "title 2", "title 3", "title 4"],  
"raw" : {  
"links" : ["[Yahoo | Mail, Weather, Search, Politics, News, Finance, Sports & Videos](http://bit.ly/abc)", "[Yahoo | Mail, Weather, Search, Politics, News, Finance, Sports & Videos](http://bit.ly/abc)",  
"[Warning! | There might be a problem with the requested link](http://bit.ly/def)", "[http://bit.ly/ghi](http://bit.ly/ghi)"]  
}  
}'

This doesn't seem to work (form what I can tell it never finds the key  
raw.links even when it does exist)

```
  "script" : "if (ctx._source.containsKey('raw.links') ) 

```

{ctx.\_source.links\_url\_count = ctx.\_source['raw.links''].size() } else {  
ctx.\_source.links\_url\_count = 0 }"

Simple keys work though like ctx.\_source.containsKey('title')

On Thursday, January 8, 2015 at 9:59:56 PM UTC-8, Nikolas Everett wrote:

> Transform never saves to source. You have to transform on the application  
> side for that. It was designed for times when you wanted to index something  
> like this that would just take up extra space in the source document. I  
> imagine you could use a script field on the query if you need the result to  
> contain the count. Or just count it on the result side.
> 
> Nik  
> On Jan 9, 2015 12:43 AM, "Jeff Steinmetz" \<[jeffrey....@gmail.com](mailto:jeffrey....@gmail.com)  
> \<javascript:\>\> wrote:
> 
> > Transform worked well. Nice.
> > 
> > Curious how to get it to save to source? Tried this below, no go. (I  
> > can however do range queries agains title\_count, so transform was indexed  
> > and works well)
> > 
> > ```
> > "transform" : {
> > "script" : "ctx._source['\'title_count\''] = 
> > 
> > ```
> > 
> > ctx.\_source[''titles''].size()",  
> > "lang": "groovy"  
> > },  
> > "properties": {  
> > "titles": { "type": "string", "index": "not\_analyzed" },  
> > "title\_count" : { "type": "integer", "store": "yes" }  
> > }  
> > }'
> > 
> > On Thursday, January 8, 2015 at 9:15:28 PM UTC-8, Nikolas Everett wrote:
> > 
> > > Source is going to be pretty sloe, yeah. If its a one off then its  
> > > probably fine but if you do it a lot probably best to index the count.  
> > > On Jan 9, 2015 12:04 AM, "Jeff Steinmetz" [jeffrey....@gmail.com](mailto:jeffrey....@gmail.com) wrote:
> > > 
> > > > Thank you, that worked.
> > > > 
> > > > I was curious about the speed, is running a script using \_source slower  
> > > > that doc ?
> > > > 
> > > > Totally understand a dynamic script is slower regardless of \_source vs  
> > > > doc.
> > > > 
> > > > Makes sense that having a count transformed up front during index to  
> > > > create a materialized value would certainly be much faster.
> > > > 
> > > > On Thursday, January 8, 2015 at 7:04:40 PM UTC-8, Nikolas Everett wrote:
> > > > 
> > > > > On Thu, Jan 8, 2015 at 9:09 PM, Jeff Steinmetz [jeffrey....@gmail.com](mailto:jeffrey....@gmail.com)  
> > > > > wrote:
> > > > > 
> > > > > Is there a better way to do this?
> > > > > 
> > > > > > Please see this gist (or even better yet, run the script locally see  
> > > > > > the issue).
> > > > > > 
> > > > > > [Determine list [array] size in elasticsearch issue · GitHub](https://gist.github.com/jeffsteinmetz/2ea8329c667386c80fae)
> > > > > > 
> > > > > > You must have scripting enabled in your elasticsearch config for this  
> > > > > > to work.
> > > > > > 
> > > > > > This was originally based on some comments I found here:  
> > > > > > [elasticsearch - Search by size of object type field elastic search - Stack Overflow](http://stackoverflow.com/questions/17314123/search-by-size-)  
> > > > > > of-object-type-field-elastic-search
> > > > > > 
> > > > > > We would like to use a filtered query to only include documents that  
> > > > > > a small count of items in the list [aka array], filtering where  
> > > > > > values.size() \< 10
> > > > > > 
> > > > > > "script": "doc['titles'].values.size() \< 10"
> > > > > > 
> > > > > > Turns out the values.size() actually either counts tokenized  
> > > > > > (analyzed) words, or if the mapping turns off analysis, it still counts  
> > > > > > incorrectly if there are duplicates.  
> > > > > > If analyze is not turned off, it counts tokenized words, not the  
> > > > > > number of elements in the list.  
> > > > > > If analyze is turned off for a given field, it improves, but  
> > > > > > duplicates are missed.
> > > > > > 
> > > > > > For example, This comes back as size == 2  
> > > > > > "titles": ["one", "duplicate", "duplicate"]  
> > > > > > This comes back as size == 3, should be 4  
> > > > > > "titles": ["[Yahoo | Mail, Weather, Search, Politics, News, Finance, Sports & Videos](http://bit.ly/abc)", "[Yahoo | Mail, Weather, Search, Politics, News, Finance, Sports & Videos](http://bit.ly/abc)", "  
> > > > > > [Warning! | There might be a problem with the requested link](http://bit.ly/def)", "[http://bit.ly/ghi](http://bit.ly/ghi)"]
> > > > > > 
> > > > > > Is this a bug, is there a better way, or is this just something that  
> > > > > > we don't understand about groovy and values.size()?
> > > > > 
> > > > > I think that's just the way doc works. Try (but don't actually  
> > > > > deploy) \_source['titles'].size() \< 10. That should do what you expect.  
> > > > > Don't deploy that because its too slow. Try indexing the size and  
> > > > > filtering on it. You can use a transform to add the size of the array as  
> > > > > an integer field and just filter on it using a range filter. That'd  
> > > > > probably be the fastest option.
> > > > > 
> > > > > Nik
> > > > 
> > > > --  
> > > > You received this message because you are subscribed to the Google  
> > > > Groups "elasticsearch" group.  
> > > > To unsubscribe from this group and stop receiving emails from it, send  
> > > > an email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com).  
> > > > To view this discussion on the web visit [https://groups.google.com/d/](https://groups.google.com/d/)  
> > > > msgid/elasticsearch/75736948-beac-43fc-84d4-25a94456d4ca%  
> > > > [40googlegroups.com](http://40googlegroups.com)  
> > > > [https://groups.google.com/d/msgid/elasticsearch/75736948-beac-43fc-84d4-25a94456d4ca%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/75736948-beac-43fc-84d4-25a94456d4ca%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > > > .  
> > > > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).
> > > 
> > > --  
> > > You received this message because you are subscribed to the Google Groups  
> > > "elasticsearch" group.  
> > > To unsubscribe from this group and stop receiving emails from it, send an  
> > > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > > To view this discussion on the web visit  
> > > [https://groups.google.com/d/msgid/elasticsearch/00ff2bc1-94a9-4aa9-8c7e-ef5734affb4d%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/00ff2bc1-94a9-4aa9-8c7e-ef5734affb4d%40googlegroups.com)  
> > > [https://groups.google.com/d/msgid/elasticsearch/00ff2bc1-94a9-4aa9-8c7e-ef5734affb4d%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/00ff2bc1-94a9-4aa9-8c7e-ef5734affb4d%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > > .  
> > > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/3717aecd-78c1-4e48-9771-acc49f8c730a%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/3717aecd-78c1-4e48-9771-acc49f8c730a%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [January 9, 2015, 10:37am UTC](https://discuss.elastic.co/t/counting-items-in-a-list-array-returns-what-we-think-are-incorrect-counts-via-groovy/21547/8 "2015-01-09T10:37:53Z")

</div>

"titles" : ["one","duplicate","duplicate"] is a short form and becomes  
"titles" : "one" and "titles":"duplicate" in the index.

With 'doc' scripts access the document in the indexed form, which is of  
course not 1:1 with the source document. Maybe it works to use 'source'  
to access the source field in the index to get the original form, but be  
warned, this is slow, because the whole source field must be loaded and  
decoded for each document.

Jörg

On Fri, Jan 9, 2015 at 3:09 AM, Jeff Steinmetz [jeffrey.steinmetz@gmail.com](mailto:jeffrey.steinmetz@gmail.com)  
wrote:

> Is there a better way to do this?
> 
> Please see this gist (or even better yet, run the script locally see the  
> issue).
> 
> [Determine list [array] size in elasticsearch issue · GitHub](https://gist.github.com/jeffsteinmetz/2ea8329c667386c80fae)
> 
> You must have scripting enabled in your elasticsearch config for this to  
> work.
> 
> This was originally based on some comments I found here:
> 
> [elasticsearch - Search by size of object type field elastic search - Stack Overflow](http://stackoverflow.com/questions/17314123/search-by-size-of-object-type-field-elastic-search)
> 
> We would like to use a filtered query to only include documents that a  
> small count of items in the list [aka array], filtering where  
> values.size() \< 10
> 
> "script": "doc['titles'].values.size() \< 10"
> 
> Turns out the values.size() actually either counts tokenized (analyzed)  
> words, or if the mapping turns off analysis, it still counts incorrectly if  
> there are duplicates.  
> If analyze is not turned off, it counts tokenized words, not the number of  
> elements in the list.  
> If analyze is turned off for a given field, it improves, but duplicates  
> are missed.
> 
> For example, This comes back as size == 2  
> "titles": ["one", "duplicate", "duplicate"]  
> This comes back as size == 3, should be 4  
> "titles": ["[Yahoo | Mail, Weather, Search, Politics, News, Finance, Sports & Videos](http://bit.ly/abc)", "[Yahoo | Mail, Weather, Search, Politics, News, Finance, Sports & Videos](http://bit.ly/abc)", "[Warning! | There might be a problem with the requested link](http://bit.ly/def)",  
> "[http://bit.ly/ghi](http://bit.ly/ghi)"]
> 
> Is this a bug, is there a better way, or is this just something that we  
> don't understand about groovy and values.size()?
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/f5e88338-8c4f-4cb8-b6c4-d7f47b365175%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/f5e88338-8c4f-4cb8-b6c4-d7f47b365175%40googlegroups.com)  
> [https://groups.google.com/d/msgid/elasticsearch/f5e88338-8c4f-4cb8-b6c4-d7f47b365175%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/f5e88338-8c4f-4cb8-b6c4-d7f47b365175%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoGtWd7c0zBDBvYtQ2j9F0%2ByKbJEPGSiFK5ni65kGHsAng%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoGtWd7c0zBDBvYtQ2j9F0%2ByKbJEPGSiFK5ni65kGHsAng%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:39am UTC](https://discuss.elastic.co/t/counting-items-in-a-list-array-returns-what-we-think-are-incorrect-counts-via-groovy/21547/9 "2017-07-06T00:39:53Z")

</div>


