# Analyzer, Fuzzy Query?

**URL:** <https://discuss.elastic.co/t/analyzer-fuzzy-query/8499>\
**Category:** Elasticsearch\
**Created:** [July 24, 2012, 7:42am UTC](https://discuss.elastic.co/t/analyzer-fuzzy-query/8499 "2012-07-24T07:42:52Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![maik2102](https://avatars.discourse-cdn.com/v4/letter/m/d2c977/32.png) [@maik2102](https://discuss.elastic.co/u/maik2102)\
**Post date:** [July 24, 2012, 7:42am UTC](https://discuss.elastic.co/t/analyzer-fuzzy-query/8499/1 "2012-07-24T07:42:52Z")

</div>

Hi toghether,

I'm wondering if its possible for elasticsearch to solve the following  
problems:

1. Productname contains "Media Player", Customer searches for "mediaplayer"  
(0 hits) or "media player" (lots of hits)
2. Productname contains "dl380", Customer searches for "dl 380" (0 hits) or  
"dl380" (lots of hits)

As today the name is analyzed with the standard analyzer and its queried by  
a ftl query.

How to analyze the string or query the index to get similar results for  
both searches, with blank and without it?  
I know synonyms, but I hope there is a better, more general solution.

Thanks in advance  
Greetings  
maik

---

<div class="post-metadata">

**Author:** ![simonw\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/simonw_2/32/1130_2.png) [@simonw\_2](https://discuss.elastic.co/u/simonw_2)\
**Post date:** [July 24, 2012, 7:04pm UTC](https://discuss.elastic.co/t/analyzer-fuzzy-query/8499/2 "2012-07-24T19:04:31Z")

</div>

hey,

in your case I'd likely use a shingle filter that builds token n-grams for  
you. ie. the document "my super media player" would be filtered into "my",  
"mysuper", "super", "supermedia", "media", "mediaplayer"

here is a  
link: [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/shingle-tokenfilter.html)

set max\_shingle\_size = 2 & output\_unigrams = true (that is actually the  
default)

this would match for "mediaplayer" as well as "media player" & your dl380  
problem woudl be solved as well.  
This might create a ton more tokens but should work just fine!

simon

On Tuesday, July 24, 2012 9:42:52 AM UTC+2, maik wrote:

> Hi toghether,
> 
> I'm wondering if its possible for elasticsearch to solve the following  
> problems:
> 
> 1. Productname contains "Media Player", Customer searches for  
> "mediaplayer" (0 hits) or "media player" (lots of hits)
> 2. Productname contains "dl380", Customer searches for "dl 380" (0 hits)  
> or "dl380" (lots of hits)
> 
> As today the name is analyzed with the standard analyzer and its queried  
> by a ftl query.
> 
> How to analyze the string or query the index to get similar results for  
> both searches, with blank and without it?  
> I know synonyms, but I hope there is a better, more general solution.
> 
> Thanks in advance  
> Greetings  
> maik

---

<div class="post-metadata">

**Author:** ![maik2102](https://avatars.discourse-cdn.com/v4/letter/m/d2c977/32.png) [@maik2102](https://discuss.elastic.co/u/maik2102)\
**Post date:** [July 25, 2012, 5:53am UTC](https://discuss.elastic.co/t/analyzer-fuzzy-query/8499/3 "2012-07-25T05:53:09Z")

</div>

Hi simon,

thank you for your help.

I read the documentation of the shingle filter but didn't find that it will  
create "media player" as well as "mediaplayer".  
In the example it only catches 2 words WITH the blank between them.

Did I get it right?

maik

On Tuesday, July 24, 2012 9:04:31 PM UTC+2, simonw wrote:

> hey,
> 
> in your case I'd likely use a shingle filter that builds token n-grams for  
> you. ie. the document "my super media player" would be filtered into "my",  
> "mysuper", "super", "supermedia", "media", "mediaplayer"
> 
> here is a link:  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/shingle-tokenfilter.html)
> 
> set max\_shingle\_size = 2 & output\_unigrams = true (that is actually the  
> default)
> 
> this would match for "mediaplayer" as well as "media player" & your dl380  
> problem woudl be solved as well.  
> This might create a ton more tokens but should work just fine!
> 
> simon
> 
> On Tuesday, July 24, 2012 9:42:52 AM UTC+2, maik wrote:
> 
> > Hi toghether,
> > 
> > I'm wondering if its possible for elasticsearch to solve the following  
> > problems:
> > 
> > 1. Productname contains "Media Player", Customer searches for  
> > "mediaplayer" (0 hits) or "media player" (lots of hits)
> > 2. Productname contains "dl380", Customer searches for "dl 380" (0 hits)  
> > or "dl380" (lots of hits)
> > 
> > As today the name is analyzed with the standard analyzer and its queried  
> > by a ftl query.
> > 
> > How to analyze the string or query the index to get similar results for  
> > both searches, with blank and without it?  
> > I know synonyms, but I hope there is a better, more general solution.
> > 
> > Thanks in advance  
> > Greetings  
> > maik

---

<div class="post-metadata">

**Author:** ![simonw\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/simonw_2/32/1130_2.png) [@simonw\_2](https://discuss.elastic.co/u/simonw_2)\
**Post date:** [July 25, 2012, 6:45am UTC](https://discuss.elastic.co/t/analyzer-fuzzy-query/8499/4 "2012-07-25T06:45:24Z")

</div>

hey you are right,

Elasticsearch doesn't expose all the functionality the shingle filter has  
like specifying the token separator etc. I will open an issue and add the  
functionality you need so you can specify a token separator instead of a  
blank.

simon

On Wednesday, July 25, 2012 7:53:09 AM UTC+2, maik wrote:

> Hi simon,
> 
> thank you for your help.
> 
> I read the documentation of the shingle filter but didn't find that it  
> will create "media player" as well as "mediaplayer".  
> In the example it only catches 2 words WITH the blank between them.
> 
> Did I get it right?
> 
> maik
> 
> On Tuesday, July 24, 2012 9:04:31 PM UTC+2, simonw wrote:
> 
> > hey,
> > 
> > in your case I'd likely use a shingle filter that builds token n-grams  
> > for you. ie. the document "my super media player" would be filtered into  
> > "my", "mysuper", "super", "supermedia", "media", "mediaplayer"
> > 
> > here is a link:  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/shingle-tokenfilter.html)
> > 
> > set max\_shingle\_size = 2 & output\_unigrams = true (that is actually the  
> > default)
> > 
> > this would match for "mediaplayer" as well as "media player" & your dl380  
> > problem woudl be solved as well.  
> > This might create a ton more tokens but should work just fine!
> > 
> > simon
> > 
> > On Tuesday, July 24, 2012 9:42:52 AM UTC+2, maik wrote:
> > 
> > > Hi toghether,
> > > 
> > > I'm wondering if its possible for elasticsearch to solve the following  
> > > problems:
> > > 
> > > 1. Productname contains "Media Player", Customer searches for  
> > > "mediaplayer" (0 hits) or "media player" (lots of hits)
> > > 2. Productname contains "dl380", Customer searches for "dl 380" (0 hits)  
> > > or "dl380" (lots of hits)
> > > 
> > > As today the name is analyzed with the standard analyzer and its queried  
> > > by a ftl query.
> > > 
> > > How to analyze the string or query the index to get similar results for  
> > > both searches, with blank and without it?  
> > > I know synonyms, but I hope there is a better, more general solution.
> > > 
> > > Thanks in advance  
> > > Greetings  
> > > maik

---

<div class="post-metadata">

**Author:** ![maik2102](https://avatars.discourse-cdn.com/v4/letter/m/d2c977/32.png) [@maik2102](https://discuss.elastic.co/u/maik2102)\
**Post date:** [July 25, 2012, 6:49am UTC](https://discuss.elastic.co/t/analyzer-fuzzy-query/8499/5 "2012-07-25T06:49:15Z")

</div>

Hi Simon,

sound good! Thank you.

Nevertheless, i tried the shingle filter, and it creates "better" results.  
I turned on explanations and while searching for "dl 380" the explanation  
says, he found the "dl380" in one of my fields.  
So I think the filter still replaces the whitespace?

For now, the results are ok for me.  
But if its not that big work, more functionality isn't that bad I think 🙂

Greetings  
maik

On Wednesday, July 25, 2012 8:45:24 AM UTC+2, simonw wrote:

> hey you are right,
> 
> Elasticsearch doesn't expose all the functionality the shingle filter has  
> like specifying the token separator etc. I will open an issue and add the  
> functionality you need so you can specify a token separator instead of a  
> blank.
> 
> simon
> 
> On Wednesday, July 25, 2012 7:53:09 AM UTC+2, maik wrote:
> 
> > Hi simon,
> > 
> > thank you for your help.
> > 
> > I read the documentation of the shingle filter but didn't find that it  
> > will create "media player" as well as "mediaplayer".  
> > In the example it only catches 2 words WITH the blank between them.
> > 
> > Did I get it right?
> > 
> > maik
> > 
> > On Tuesday, July 24, 2012 9:04:31 PM UTC+2, simonw wrote:
> > 
> > > hey,
> > > 
> > > in your case I'd likely use a shingle filter that builds token n-grams  
> > > for you. ie. the document "my super media player" would be filtered into  
> > > "my", "mysuper", "super", "supermedia", "media", "mediaplayer"
> > > 
> > > here is a link:  
> > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/shingle-tokenfilter.html)
> > > 
> > > set max\_shingle\_size = 2 & output\_unigrams = true (that is actually the  
> > > default)
> > > 
> > > this would match for "mediaplayer" as well as "media player" & your  
> > > dl380 problem woudl be solved as well.  
> > > This might create a ton more tokens but should work just fine!
> > > 
> > > simon
> > > 
> > > On Tuesday, July 24, 2012 9:42:52 AM UTC+2, maik wrote:
> > > 
> > > > Hi toghether,
> > > > 
> > > > I'm wondering if its possible for elasticsearch to solve the following  
> > > > problems:
> > > > 
> > > > 1. Productname contains "Media Player", Customer searches for  
> > > > "mediaplayer" (0 hits) or "media player" (lots of hits)
> > > > 2. Productname contains "dl380", Customer searches for "dl 380"  
> > > > (0 hits) or "dl380" (lots of hits)
> > > > 
> > > > As today the name is analyzed with the standard analyzer and its  
> > > > queried by a ftl query.
> > > > 
> > > > How to analyze the string or query the index to get similar results for  
> > > > both searches, with blank and without it?  
> > > > I know synonyms, but I hope there is a better, more general solution.
> > > > 
> > > > Thanks in advance  
> > > > Greetings  
> > > > maik

---

<div class="post-metadata">

**Author:** ![simonw\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/simonw_2/32/1130_2.png) [@simonw\_2](https://discuss.elastic.co/u/simonw_2)\
**Post date:** [July 25, 2012, 6:53am UTC](https://discuss.elastic.co/t/analyzer-fuzzy-query/8499/6 "2012-07-25T06:53:01Z")

</div>

On Wednesday, July 25, 2012 8:49:15 AM UTC+2, maik wrote:

> Hi Simon,
> 
> sound good! Thank you.
> 
> Nevertheless, i tried the shingle filter, and it creates "better" results.  
> I turned on explanations and while searching for "dl 380" the explanation  
> says, he found the "dl380" in one of my fields.  
> So I think the filter still replaces the whitespace?

hmm I am not sure If I understand this 🙂 can you post the explain output?

I opened [ShingleTokenFilterFactory doesn't expose all relevant settings · Issue #2116 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/2116) for this

simon

> For now, the results are ok for me.  
> But if its not that big work, more functionality isn't that bad I think 🙂
> 
> Greetings  
> maik
> 
> On Wednesday, July 25, 2012 8:45:24 AM UTC+2, simonw wrote:
> 
> > hey you are right,
> > 
> > Elasticsearch doesn't expose all the functionality the shingle filter has  
> > like specifying the token separator etc. I will open an issue and add the  
> > functionality you need so you can specify a token separator instead of a  
> > blank.
> > 
> > simon
> > 
> > On Wednesday, July 25, 2012 7:53:09 AM UTC+2, maik wrote:
> > 
> > > Hi simon,
> > > 
> > > thank you for your help.
> > > 
> > > I read the documentation of the shingle filter but didn't find that it  
> > > will create "media player" as well as "mediaplayer".  
> > > In the example it only catches 2 words WITH the blank between them.
> > > 
> > > Did I get it right?
> > > 
> > > maik
> > > 
> > > On Tuesday, July 24, 2012 9:04:31 PM UTC+2, simonw wrote:
> > > 
> > > > hey,
> > > > 
> > > > in your case I'd likely use a shingle filter that builds token n-grams  
> > > > for you. ie. the document "my super media player" would be filtered into  
> > > > "my", "mysuper", "super", "supermedia", "media", "mediaplayer"
> > > > 
> > > > here is a link:  
> > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/shingle-tokenfilter.html)
> > > > 
> > > > set max\_shingle\_size = 2 & output\_unigrams = true (that is actually the  
> > > > default)
> > > > 
> > > > this would match for "mediaplayer" as well as "media player" & your  
> > > > dl380 problem woudl be solved as well.  
> > > > This might create a ton more tokens but should work just fine!
> > > > 
> > > > simon
> > > > 
> > > > On Tuesday, July 24, 2012 9:42:52 AM UTC+2, maik wrote:
> > > > 
> > > > > Hi toghether,
> > > > > 
> > > > > I'm wondering if its possible for elasticsearch to solve the following  
> > > > > problems:
> > > > > 
> > > > > 1. Productname contains "Media Player", Customer searches for  
> > > > > "mediaplayer" (0 hits) or "media player" (lots of hits)
> > > > > 2. Productname contains "dl380", Customer searches for "dl 380"  
> > > > > (0 hits) or "dl380" (lots of hits)
> > > > > 
> > > > > As today the name is analyzed with the standard analyzer and its  
> > > > > queried by a ftl query.
> > > > > 
> > > > > How to analyze the string or query the index to get similar results  
> > > > > for both searches, with blank and without it?  
> > > > > I know synonyms, but I hope there is a better, more general solution.
> > > > > 
> > > > > Thanks in advance  
> > > > > Greetings  
> > > > > maik

---

<div class="post-metadata">

**Author:** ![maik2102](https://avatars.discourse-cdn.com/v4/letter/m/d2c977/32.png) [@maik2102](https://discuss.elastic.co/u/maik2102)\
**Post date:** [July 30, 2012, 12:40pm UTC](https://discuss.elastic.co/t/analyzer-fuzzy-query/8499/7 "2012-07-30T12:40:49Z")

</div>

Hi Simon,

sorry für my delayed answer, not that much time at the moment 🙂

We have products which have "dl 380" in their names. Before putting the  
shingle filter into the analysis process, with "dl380" you didn't find them.  
With the single filter you can find them.

I don't have an explanation output at my hands, but it said it found  
"dl380" in the name.  
So I think it already works in the way I need it.

If you're interested in further information, please let me know.

Greetings  
maik

On Wednesday, July 25, 2012 8:53:01 AM UTC+2, simonw wrote:

> On Wednesday, July 25, 2012 8:49:15 AM UTC+2, maik wrote:
> 
> > Hi Simon,
> > 
> > sound good! Thank you.
> > 
> > Nevertheless, i tried the shingle filter, and it creates "better" results.  
> > I turned on explanations and while searching for "dl 380" the explanation  
> > says, he found the "dl380" in one of my fields.  
> > So I think the filter still replaces the whitespace?
> 
> hmm I am not sure If I understand this 🙂 can you post the explain output?
> 
> I opened [ShingleTokenFilterFactory doesn't expose all relevant settings · Issue #2116 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/2116) for  
> this
> 
> simon
> 
> > For now, the results are ok for me.  
> > But if its not that big work, more functionality isn't that bad I think  
> > 🙂
> > 
> > Greetings  
> > maik
> > 
> > On Wednesday, July 25, 2012 8:45:24 AM UTC+2, simonw wrote:
> > 
> > > hey you are right,
> > > 
> > > Elasticsearch doesn't expose all the functionality the shingle filter  
> > > has like specifying the token separator etc. I will open an issue and add  
> > > the functionality you need so you can specify a token separator instead of  
> > > a blank.
> > > 
> > > simon
> > > 
> > > On Wednesday, July 25, 2012 7:53:09 AM UTC+2, maik wrote:
> > > 
> > > > Hi simon,
> > > > 
> > > > thank you for your help.
> > > > 
> > > > I read the documentation of the shingle filter but didn't find that it  
> > > > will create "media player" as well as "mediaplayer".  
> > > > In the example it only catches 2 words WITH the blank between them.
> > > > 
> > > > Did I get it right?
> > > > 
> > > > maik
> > > > 
> > > > On Tuesday, July 24, 2012 9:04:31 PM UTC+2, simonw wrote:
> > > > 
> > > > > hey,
> > > > > 
> > > > > in your case I'd likely use a shingle filter that builds token n-grams  
> > > > > for you. ie. the document "my super media player" would be filtered into  
> > > > > "my", "mysuper", "super", "supermedia", "media", "mediaplayer"
> > > > > 
> > > > > here is a link:  
> > > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/shingle-tokenfilter.html)
> > > > > 
> > > > > set max\_shingle\_size = 2 & output\_unigrams = true (that is actually  
> > > > > the default)
> > > > > 
> > > > > this would match for "mediaplayer" as well as "media player" & your  
> > > > > dl380 problem woudl be solved as well.  
> > > > > This might create a ton more tokens but should work just fine!
> > > > > 
> > > > > simon
> > > > > 
> > > > > On Tuesday, July 24, 2012 9:42:52 AM UTC+2, maik wrote:
> > > > > 
> > > > > > Hi toghether,
> > > > > > 
> > > > > > I'm wondering if its possible for elasticsearch to solve the  
> > > > > > following problems:
> > > > > > 
> > > > > > 1. Productname contains "Media Player", Customer searches for  
> > > > > > "mediaplayer" (0 hits) or "media player" (lots of hits)
> > > > > > 2. Productname contains "dl380", Customer searches for "dl 380"  
> > > > > > (0 hits) or "dl380" (lots of hits)
> > > > > > 
> > > > > > As today the name is analyzed with the standard analyzer and its  
> > > > > > queried by a ftl query.
> > > > > > 
> > > > > > How to analyze the string or query the index to get similar results  
> > > > > > for both searches, with blank and without it?  
> > > > > > I know synonyms, but I hope there is a better, more general solution.
> > > > > > 
> > > > > > Thanks in advance  
> > > > > > Greetings  
> > > > > > maik

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:18am UTC](https://discuss.elastic.co/t/analyzer-fuzzy-query/8499/8 "2017-07-06T03:18:29Z")

</div>


