# Faceting on a field with very many unique values, on a very large index

**URL:** https://discuss.elastic.co/t/faceting-on-a-field-with-very-many-unique-values-on-a-very-large-index/9144
**Category:** Elasticsearch
**Created:** [September 26, 2012, 12:02am UTC](https://discuss.elastic.co/t/faceting-on-a-field-with-very-many-unique-values-on-a-very-large-index/9144 "2012-09-26T00:02:18Z")
**Posts on this page:** 6
**Page:** 1

<div class="post-metadata">

### Author: ![Mark\_MacGillivray](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_macgillivray/32/2702_2.png) [@Mark\_MacGillivray](https://discuss.elastic.co/u/Mark_MacGillivray)
#### Post date: [September 26, 2012, 12:02am UTC](https://discuss.elastic.co/t/faceting-on-a-field-with-very-many-unique-values-on-a-very-large-index/9144/1 "2012-09-26T00:02:18Z")

</div>

Hi there, I am using elasticsearch (which is totally brilliant by the way)  
for an index with approximately 21 million records in it, and within those  
records I have one particular field that has between 1 and perhaps 10  
values, and those values are often unique to just that record. The values  
are text strings - names of people. I am using a dynamic mapping.

I would like to be able to facet on this field, but whatever I do, I just  
crash my index. So I am looking for further suggestions.

I have stored this field unanalysed, and I have tried the field cache field  
type set to soft and not set at all, and tried field cache max size to  
various values ranging from 1 to 10,000,000.

I have run this on a single machine with 60gb memory reserved to  
elasticsearch. It eventually fails with an Out of Memory error and tries to  
dump the heap.

I have also tried running it on a cluster of 8 machines with 6gb for  
elasticsearch on each, trying with between 1 and 16 shards, and between 1  
and 8 replicas. Also on a cluster of 4 machines with 12gb each. However it  
again fails with OOM, a bit sooner than the one big machine.

Are other people running facets on fields with this many potentially unique  
values - on the order of 70,000,000? Am I just pushing elasticsearch too  
far, or is it worth trying with more machines / one even bigger machine /  
many even bigger machines?

Any feedback from people doing this sort of scale of faceting would be  
appreciated, or any other settings suggestions you can provide would be  
great, so that I can get an idea if it is worth trying any further or just  
give up faceting on this field.

Thanks!

--

---

<div class="post-metadata">

### Author: ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)
#### Post date: [September 26, 2012, 12:06pm UTC](https://discuss.elastic.co/t/faceting-on-a-field-with-very-many-unique-values-on-a-very-large-index/9144/2 "2012-09-26T12:06:05Z")

</div>

Hi Mark,

As you have seen it, faceting in 0.19.x with huge dataset is a problem.  
Although Shay wrote before that there will be some memory usage optimization  
when using facets in 0.20, I'm not sure that you will be able to facet on  
70,000,000 unique values (I did not test 0.20 SNAPSHOT yet myself).

But, I'm wondering about your use case. What are you trying to achieve here? If  
all values are unique, why do you need to group and count them as you will have  
probably 1 count per value?

In the past, the only way I found to avoid OOM was to restrict with  
filters/queries the number of documents to facet on.

So, don't you have in your documents somewhere a field that could help you to  
reduce the number of documents? (a date field for example?)

Not sure my answer helps... ☹  
David.

Le 26 septembre 2012 à 02:02, Mark MacGillivray [mark@cottagelabs.com](mailto:mark@cottagelabs.com) a écrit :

> Hi there, I am using elasticsearch (which is totally brilliant by the way) for  
> an index with approximately 21 million records in it, and within those records  
> I have one particular field that has between 1 and perhaps 10 values, and  
> those values are often unique to just that record. The values are text strings
> 
> - names of people. I am using a dynamic mapping.
> 
> I would like to be able to facet on this field, but whatever I do, I just  
> crash my index. So I am looking for further suggestions.
> 
> I have stored this field unanalysed, and I have tried the field cache field  
> type set to soft and not set at all, and tried field cache max size to various  
> values ranging from 1 to 10,000,000.
> 
> I have run this on a single machine with 60gb memory reserved to  
> elasticsearch. It eventually fails with an Out of Memory error and tries to  
> dump the heap.
> 
> I have also tried running it on a cluster of 8 machines with 6gb for  
> elasticsearch on each, trying with between 1 and 16 shards, and between 1 and  
> 8 replicas. Also on a cluster of 4 machines with 12gb each. However it again  
> fails with OOM, a bit sooner than the one big machine.
> 
> Are other people running facets on fields with this many potentially unique  
> values - on the order of 70,000,000? Am I just pushing elasticsearch too far,  
> or is it worth trying with more machines / one even bigger machine / many even  
> bigger machines?
> 
> Any feedback from people doing this sort of scale of faceting would be  
> appreciated, or any other settings suggestions you can provide would be great,  
> so that I can get an idea if it is worth trying any further or just give up  
> faceting on this field.
> 
> Thanks!
> 
> --

--  
David Pilato  
[http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
[http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

--

---

<div class="post-metadata">

### Author: ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)
#### Post date: [September 26, 2012, 1:29pm UTC](https://discuss.elastic.co/t/faceting-on-a-field-with-very-many-unique-values-on-a-very-large-index/9144/3 "2012-09-26T13:29:16Z")

</div>

Hi Mark,

If your field is stored, you can try using script\_field instead of standard  
field terms facet. Assuming that your field is called "authors", you can  
try something like this:

"facets": {  
"authors\_facet": {  
"terms" : {  
"script\_field" : "\_fields.authors.values"  
}  
}  
}

Start with really small result set. Running this request on all 70 mln  
records will take really long time. However, if your result sets  
are relatively small, you might get acceptable performance out of it.

Igor

On Wednesday, September 26, 2012 8:06:10 AM UTC-4, David Pilato wrote:

> Hi Mark,
> 
> As you have seen it, faceting in 0.19.x with huge dataset is a problem.  
> Although Shay wrote before that there will be some memory usage  
> optimization when using facets in 0.20, I'm not sure that you will be able  
> to facet on 70,000,000 unique values (I did not test 0.20 SNAPSHOT yet  
> myself).
> 
> But, I'm wondering about your use case. What are you trying to achieve  
> here? If all values are unique, why do you need to group and count them as  
> you will have probably 1 count per value?
> 
> In the past, the only way I found to avoid OOM was to restrict with  
> filters/queries the number of documents to facet on.
> 
> So, don't you have in your documents somewhere a field that could help  
> you to reduce the number of documents? (a date field for example?)
> 
> Not sure my answer helps... ☹  
> David.
> 
> Le 26 septembre 2012 à 02:02, Mark MacGillivray [mark@cottagelabs.com](mailto:mark@cottagelabs.com) a  
> écrit :
> 
> Hi there, I am using elasticsearch (which is totally brilliant by the way)  
> for an index with approximately 21 million records in it, and within those  
> records I have one particular field that has between 1 and perhaps 10  
> values, and those values are often unique to just that record. The values  
> are text strings - names of people. I am using a dynamic mapping.
> 
> I would like to be able to facet on this field, but whatever I do, I just  
> crash my index. So I am looking for further suggestions.
> 
> I have stored this field unanalysed, and I have tried the field cache  
> field type set to soft and not set at all, and tried field cache max size  
> to various values ranging from 1 to 10,000,000.
> 
> I have run this on a single machine with 60gb memory reserved to  
> elasticsearch. It eventually fails with an Out of Memory error and tries to  
> dump the heap.
> 
> I have also tried running it on a cluster of 8 machines with 6gb for  
> elasticsearch on each, trying with between 1 and 16 shards, and between 1  
> and 8 replicas. Also on a cluster of 4 machines with 12gb each. However it  
> again fails with OOM, a bit sooner than the one big machine.
> 
> Are other people running facets on fields with this many potentially  
> unique values - on the order of 70,000,000? Am I just pushing elasticsearch  
> too far, or is it worth trying with more machines / one even bigger machine  
> / many even bigger machines?
> 
> Any feedback from people doing this sort of scale of faceting would be  
> appreciated, or any other settings suggestions you can provide would be  
> great, so that I can get an idea if it is worth trying any further or just  
> give up faceting on this field.
> 
> Thanks!
> 
> --
> 
> --  
> David Pilato  
> [http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
> [http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
> Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

--

---

<div class="post-metadata">

### Author: ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)
#### Post date: [September 27, 2012, 1:44am UTC](https://discuss.elastic.co/t/faceting-on-a-field-with-very-many-unique-values-on-a-very-large-index/9144/4 "2012-09-27T01:44:42Z")

</div>

Hi,

I was going to suggest to try limiting faceting to only top N results.  
There is an issue open for this, but no implementation. If that is an  
option (and this means that you don't need exact counts) then you could try  
doing this in the client, too - just get the results including authors,  
chop up the authors field and count the names.

Maybe limiting to top N is something that could be accomplished through the  
script\_field Igor mentioned?

## Otis

Search Analytics - [Cloud Monitoring Tools & Services | Sematext](http://sematext.com/search-analytics/index.html)  
Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)

On Wednesday, September 26, 2012 9:29:16 AM UTC-4, Igor Motov wrote:

> Hi Mark,
> 
> If your field is stored, you can try using script\_field instead of  
> standard field terms facet. Assuming that your field is called "authors",  
> you can try something like this:
> 
> "facets": {  
> "authors\_facet": {  
> "terms" : {  
> "script\_field" : "\_fields.authors.values"  
> }  
> }  
> }
> 
> Start with really small result set. Running this request on all 70 mln  
> records will take really long time. However, if your result sets  
> are relatively small, you might get acceptable performance out of it.
> 
> Igor
> 
> On Wednesday, September 26, 2012 8:06:10 AM UTC-4, David Pilato wrote:
> 
> > Hi Mark,
> > 
> > As you have seen it, faceting in 0.19.x with huge dataset is a problem.  
> > Although Shay wrote before that there will be some memory usage  
> > optimization when using facets in 0.20, I'm not sure that you will be able  
> > to facet on 70,000,000 unique values (I did not test 0.20 SNAPSHOT yet  
> > myself).
> > 
> > But, I'm wondering about your use case. What are you trying to achieve  
> > here? If all values are unique, why do you need to group and count them as  
> > you will have probably 1 count per value?
> > 
> > In the past, the only way I found to avoid OOM was to restrict with  
> > filters/queries the number of documents to facet on.
> > 
> > So, don't you have in your documents somewhere a field that could help  
> > you to reduce the number of documents? (a date field for example?)
> > 
> > Not sure my answer helps... ☹  
> > David.
> > 
> > Le 26 septembre 2012 à 02:02, Mark MacGillivray \<[ma...@cottagelabs.com](mailto:ma...@cottagelabs.com)\<javascript:\>\>  
> > a écrit :
> > 
> > Hi there, I am using elasticsearch (which is totally brilliant by the  
> > way) for an index with approximately 21 million records in it, and within  
> > those records I have one particular field that has between 1 and perhaps 10  
> > values, and those values are often unique to just that record. The values  
> > are text strings - names of people. I am using a dynamic mapping.
> > 
> > I would like to be able to facet on this field, but whatever I do, I  
> > just crash my index. So I am looking for further suggestions.
> > 
> > I have stored this field unanalysed, and I have tried the field cache  
> > field type set to soft and not set at all, and tried field cache max size  
> > to various values ranging from 1 to 10,000,000.
> > 
> > I have run this on a single machine with 60gb memory reserved to  
> > elasticsearch. It eventually fails with an Out of Memory error and tries to  
> > dump the heap.
> > 
> > I have also tried running it on a cluster of 8 machines with 6gb for  
> > elasticsearch on each, trying with between 1 and 16 shards, and between 1  
> > and 8 replicas. Also on a cluster of 4 machines with 12gb each. However it  
> > again fails with OOM, a bit sooner than the one big machine.
> > 
> > Are other people running facets on fields with this many potentially  
> > unique values - on the order of 70,000,000? Am I just pushing elasticsearch  
> > too far, or is it worth trying with more machines / one even bigger machine  
> > / many even bigger machines?
> > 
> > Any feedback from people doing this sort of scale of faceting would be  
> > appreciated, or any other settings suggestions you can provide would be  
> > great, so that I can get an idea if it is worth trying any further or just  
> > give up faceting on this field.
> > 
> > Thanks!
> > 
> > --
> > 
> > --  
> > David Pilato  
> > [http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
> > [http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
> > Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

--

---

<div class="post-metadata">

### Author: ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)
#### Post date: [September 27, 2012, 7:06pm UTC](https://discuss.elastic.co/t/faceting-on-a-field-with-very-many-unique-values-on-a-very-large-index/9144/5 "2012-09-27T19:06:46Z")

</div>

Are the names of people constantly changing or is it static? I facet only  
on numeric keys (ints and longs) and convert the values to the proper text  
values using a mapping on the client side.

--  
Ivan

On Wed, Sep 26, 2012 at 6:44 PM, Otis Gospodnetic \<  
[otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com)\> wrote:

> Hi,
> 
> I was going to suggest to try limiting faceting to only top N results.  
> There is an issue open for this, but no implementation. If that is an  
> option (and this means that you don't need exact counts) then you could try  
> doing this in the client, too - just get the results including authors,  
> chop up the authors field and count the names.
> 
> Maybe limiting to top N is something that could be accomplished through  
> the script\_field Igor mentioned?
> 
> ## Otis
> 
> Search Analytics - [Cloud Monitoring Tools & Services | Sematext](http://sematext.com/search-analytics/index.html)  
> Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> 
> On Wednesday, September 26, 2012 9:29:16 AM UTC-4, Igor Motov wrote:
> 
> > Hi Mark,
> > 
> > If your field is stored, you can try using script\_field instead of  
> > standard field terms facet. Assuming that your field is called "authors",  
> > you can try something like this:
> > 
> > "facets": {  
> > "authors\_facet": {  
> > "terms" : {  
> > "script\_field" : "\_fields.authors.values"  
> > }  
> > }  
> > }
> > 
> > Start with really small result set. Running this request on all 70 mln  
> > records will take really long time. However, if your result sets  
> > are relatively small, you might get acceptable performance out of it.
> > 
> > Igor
> > 
> > On Wednesday, September 26, 2012 8:06:10 AM UTC-4, David Pilato wrote:
> > 
> > > Hi Mark,
> > > 
> > > As you have seen it, faceting in 0.19.x with huge dataset is a problem.  
> > > Although Shay wrote before that there will be some memory usage  
> > > optimization when using facets in 0.20, I'm not sure that you will be able  
> > > to facet on 70,000,000 unique values (I did not test 0.20 SNAPSHOT yet  
> > > myself).
> > > 
> > > But, I'm wondering about your use case. What are you trying to achieve  
> > > here? If all values are unique, why do you need to group and count them as  
> > > you will have probably 1 count per value?
> > > 
> > > In the past, the only way I found to avoid OOM was to restrict with  
> > > filters/queries the number of documents to facet on.
> > > 
> > > So, don't you have in your documents somewhere a field that could help  
> > > you to reduce the number of documents? (a date field for example?)
> > > 
> > > Not sure my answer helps... ☹  
> > > David.
> > > 
> > > Le 26 septembre 2012 à 02:02, Mark MacGillivray [ma...@cottagelabs.com](mailto:ma...@cottagelabs.com)  
> > > a écrit :
> > > 
> > > Hi there, I am using elasticsearch (which is totally brilliant by the  
> > > way) for an index with approximately 21 million records in it, and within  
> > > those records I have one particular field that has between 1 and perhaps 10  
> > > values, and those values are often unique to just that record. The values  
> > > are text strings - names of people. I am using a dynamic mapping.
> > > 
> > > I would like to be able to facet on this field, but whatever I do, I  
> > > just crash my index. So I am looking for further suggestions.
> > > 
> > > I have stored this field unanalysed, and I have tried the field cache  
> > > field type set to soft and not set at all, and tried field cache max size  
> > > to various values ranging from 1 to 10,000,000.
> > > 
> > > I have run this on a single machine with 60gb memory reserved to  
> > > elasticsearch. It eventually fails with an Out of Memory error and tries to  
> > > dump the heap.
> > > 
> > > I have also tried running it on a cluster of 8 machines with 6gb for  
> > > elasticsearch on each, trying with between 1 and 16 shards, and between 1  
> > > and 8 replicas. Also on a cluster of 4 machines with 12gb each. However it  
> > > again fails with OOM, a bit sooner than the one big machine.
> > > 
> > > Are other people running facets on fields with this many potentially  
> > > unique values - on the order of 70,000,000? Am I just pushing elasticsearch  
> > > too far, or is it worth trying with more machines / one even bigger machine  
> > > / many even bigger machines?
> > > 
> > > Any feedback from people doing this sort of scale of faceting would be  
> > > appreciated, or any other settings suggestions you can provide would be  
> > > great, so that I can get an idea if it is worth trying any further or just  
> > > give up faceting on this field.
> > > 
> > > Thanks!
> > > 
> > > --
> > > 
> > > --  
> > > David Pilato  
> > > [http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
> > > [http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
> > > Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs
> > 
> > --

--

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 6, 2017, 3:11am UTC](https://discuss.elastic.co/t/faceting-on-a-field-with-very-many-unique-values-on-a-very-large-index/9144/6 "2017-07-06T03:11:07Z")

</div>


