# Significant term aggregation with Snowball analyzer

**URL:** <https://discuss.elastic.co/t/significant-term-aggregation-with-snowball-analyzer/160881>\
**Category:** Elasticsearch\
**Created:** [December 14, 2018, 10:29am UTC](https://discuss.elastic.co/t/significant-term-aggregation-with-snowball-analyzer/160881 "2018-12-14T10:29:28Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![pramod.kumar2](https://avatars.discourse-cdn.com/v4/letter/p/977dab/32.png) [@pramod.kumar2](https://discuss.elastic.co/u/pramod.kumar2)\
**Post date:** [December 14, 2018, 10:29am UTC](https://discuss.elastic.co/t/significant-term-aggregation-with-snowball-analyzer/160881/1 "2018-12-14T10:29:28Z")

</div>

Hi,

I am using elasticsearch snowball analyzer for product field in an index. I need to get significant terms from elasticsearch aggregation(significant terms aggregation), but the results are not correct. The problem is that the resultant terms I am getting are with not exact as in my resultant part.  
For example - productDescription in hits are like -  
**"SWEET BISCUITS - MILK BIKIS MILK CREAM"**  
**"BRITTANNIA PRODUCTS: MILK BIKIES CREAM 1 00GM X 100NOS"**

and significant term I am getting is -

{  
**"key": "biki",**  
"doc\_count": 4,  
"score": 553.7252991452992,  
"bg\_count": 260  
}

**Please suggest how can I get correct results like ("BIKIES", "BIKIS")**

**Below is the query sample -**  
{  
"\_source": {  
"include": [  
"productDescription"  
]  
},  
"query": {  
"bool": {  
"filter": [{  
"query\_string": {  
"default\_field": "productDescription.SnowField",  
"default\_operator": "AND",  
"query": "(milk cream)"  
}  
},  
{  
"term": {  
"isUnique": true  
}  
},  
{  
"range": {  
"date": {  
"gte": "2018-01-01",  
"lte": "2018-12-31"  
}  
}  
}  
],  
"must": ,  
"must\_not":   
}  
},  
"sort": [{  
"date": {  
"order": "desc"  
}  
}],  
"size": 500,  
"aggs": {  
"my\_sample": {  
"sampler": {  
"shard\_size": 20  
},  
"aggregations": {  
"keywords": {  
"significant\_text": {  
"field": "productDescription.SnowField",  
"size": 10,  
"filter\_duplicate\_text": true  
}  
}  
}  
}  
}  
}

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [December 14, 2018, 10:46am UTC](https://discuss.elastic.co/t/significant-term-aggregation-with-snowball-analyzer/160881/2 "2018-12-14T10:46:45Z")

</div>

> [@pramod.kumar2](#):
>
> Please suggest how can I get correct results like ("BIKIES", "BIKIS")

Use a field with an analyzer that doesn't stem or lowercase e.g. the "whitespace" analyzer

---

<div class="post-metadata">

**Author:** ![pramod.kumar2](https://avatars.discourse-cdn.com/v4/letter/p/977dab/32.png) [@pramod.kumar2](https://discuss.elastic.co/u/pramod.kumar2)\
**Post date:** [December 14, 2018, 11:35am UTC](https://discuss.elastic.co/t/significant-term-aggregation-with-snowball-analyzer/160881/3 "2018-12-14T11:35:43Z")

</div>

Hi Mark,

Thanks for the reply. The problem is not lowercase results, the problem is - I got "biki" from significant term aggregation while I need it as it is like "bikies" and "bikis" as you see it is in hits returned from . This was sample result, as I checked with different searches, I got many words which were mis-spelled(removed s/es i.e. without plural parts). But required is to get meaningful words(suggestions).

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [December 14, 2018, 11:41am UTC](https://discuss.elastic.co/t/significant-term-aggregation-with-snowball-analyzer/160881/4 "2018-12-14T11:41:06Z")

</div>

> [@pramod.kumar2](#):
>
> , I got many words which were mis-spelled(removed s/es i.e. without plural parts)

That's what "stemming" does.

> **[Stemming](https://en.wikipedia.org/wiki/Stemming)**
>
> In linguistic morphology and information retrieval, stemming is the process of reducing inflected (or sometimes derived) words to their word stem, base or root form—generally a written word form. The stem need not be identical to the morphological root of the word; it is usually sufficient that related words map to the same stem, even if this stem is not in itself a valid root. Algorithms for stemming have been studied in computer science since the 1960s. Many search engines treat words with the ...

> **[Snowball Token Filter | Elasticsearch Guide \[6.5\] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/6.5/analysis-snowball-tokenfilter.html)**

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [December 14, 2018, 11:47am UTC](https://discuss.elastic.co/t/significant-term-aggregation-with-snowball-analyzer/160881/5 "2018-12-14T11:47:36Z")

</div>

Check out [this blog](https://www.elastic.co/blog/significant-terms-aggregation) which includes an example of taking potentially stemmed significant terms and using them in a `terms` query with a highlighter to show KWIC (Keywords In Context) examples of the discovered terms in text.  
Note it talks about `significant_terms` rather than the new `significant_text` aggregation but the same principles still hold.

---

<div class="post-metadata">

**Author:** ![pramod.kumar2](https://avatars.discourse-cdn.com/v4/letter/p/977dab/32.png) [@pramod.kumar2](https://discuss.elastic.co/u/pramod.kumar2)\
**Post date:** [December 14, 2018, 11:52am UTC](https://discuss.elastic.co/t/significant-term-aggregation-with-snowball-analyzer/160881/6 "2018-12-14T11:52:32Z")

</div>

Hi Mark,

Yes I knew it. That is due to snowball analyzer as I mentioned above. I am using snowball in query as I need to include sound like words in results. And I also tried with removing SnowBall analyzer from aggregation and tried keyword analyzer as well in aggregation field but did not got exact results.

But is there any way I can get exact results like if any way if I need to reindex data with any other analyzer to get significant results or something else by which I can get aggregation results as they exists in productDescription field?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [December 14, 2018, 12:04pm UTC](https://discuss.elastic.co/t/significant-term-aggregation-with-snowball-analyzer/160881/7 "2018-12-14T12:04:32Z")

</div>

It depends.  
If your docs were orders where you wanted to know "which products are typically also bought with pasta?" then you might use a `keyword` field and `significant_terms` because you'd be examining significant patterns in repeated orders for exactly the same product.  
If your docs were products you'd (hopefully) only ever have exactly one unique product description so the `keyword` field would be of no use with any significance analysis (everything occurs once). If you were looking at some of the ingredients in the text of these descriptions (eg. common ingredients mentioned in high-fat products) then you might use an analyzed text field and significant\_text. Maybe indexing with shingles would help too. Remember the indexed field you search on can be different (eg stemmed) from the indexed field you use for significant\_text analysis (e.g. whitespace)

---

<div class="post-metadata">

**Author:** ![Nitesh\_Kumar\_dcpl](https://avatars.discourse-cdn.com/v4/letter/n/b77776/32.png) [@Nitesh\_Kumar\_dcpl](https://discuss.elastic.co/u/Nitesh_Kumar_dcpl)\
**Post date:** [December 17, 2018, 10:39am UTC](https://discuss.elastic.co/t/significant-term-aggregation-with-snowball-analyzer/160881/8 "2018-12-17T10:39:29Z")

</div>

Hi Mark,

If i am using the significant term on multiple indexes, so how can we specify the missing terms

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [December 17, 2018, 10:45am UTC](https://discuss.elastic.co/t/significant-term-aggregation-with-snowball-analyzer/160881/9 "2018-12-17T10:45:37Z")

</div>

> [@Nitesh\_Kumar\_dcpl](#):
>
> how can we specify the missing terms

Significant terms is a tool for _discovering_ terms - I don't follow why you're asking a question about _specifying_ them?

---

<div class="post-metadata">

**Author:** ![Nitesh\_Kumar\_dcpl](https://avatars.discourse-cdn.com/v4/letter/n/b77776/32.png) [@Nitesh\_Kumar\_dcpl](https://discuss.elastic.co/u/Nitesh_Kumar_dcpl)\
**Post date:** [December 17, 2018, 11:30am UTC](https://discuss.elastic.co/t/significant-term-aggregation-with-snowball-analyzer/160881/10 "2018-12-17T11:30:37Z")

</div>

Hi Mark,

Let me explain you few things.. I have 2 indexes.. I created same alias name on these so that I can search on these at once. In one index, i have field name productDescription and in second index it is productDesc. So the issue in getting significant terms is that when I pass productDescription field name in aggregation, it says that - "Aggregation [keywords] cannot process field [productDescription.StandardField] since it is not present". So is there any way by which I can pass two fields in significant term aggregation or otherwise can ignore it anyhow(Like we pass "missing" property in terms aggregation, but that is not supported in significant terms aggregation.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [December 17, 2018, 11:35am UTC](https://discuss.elastic.co/t/significant-term-aggregation-with-snowball-analyzer/160881/11 "2018-12-17T11:35:52Z")

</div>

Ah. So missing "fields".  
If the overall goal is to blend the term stats from 2 fields in 2 indices the answer is "no".  
Generally, significant terms will work best on a single index and single shard since all of the stats are available in one place. If you're trying to use it to spot low-frequency terms (e.g. something that only occurs twice) in a distributed system that makes life hard because every single-occurrence term on a local shard (of which there are typically many) suddenly becomes a candidate for global consideration.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 14, 2019, 11:35am UTC](https://discuss.elastic.co/t/significant-term-aggregation-with-snowball-analyzer/160881/12 "2019-01-14T11:35:53Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
