# Synonyms with Keyword Tokenizer

**URL:** <https://discuss.elastic.co/t/synonyms-with-keyword-tokenizer/13757>\
**Category:** Elasticsearch\
**Created:** [September 25, 2013, 8:26pm UTC](https://discuss.elastic.co/t/synonyms-with-keyword-tokenizer/13757 "2013-09-25T20:26:53Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Anthony\_Campagna](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anthony_campagna/32/2079_2.png) [@Anthony\_Campagna](https://discuss.elastic.co/u/Anthony_Campagna)\
**Post date:** [September 25, 2013, 8:26pm UTC](https://discuss.elastic.co/t/synonyms-with-keyword-tokenizer/13757/1 "2013-09-25T20:26:53Z")

</div>

_Goal:_ To seamlessly autocomplete addresses while utilizing synonyms.

In my mapping/index settings I have a filter for synonym and a filter for  
edgeNgrams. I then use an index analyzer to utilize the synonym filter and  
edgeNgram. Obviously this doesn't work because we tokenize on the whole  
string and the string doesn't match a synonym. This is very problematic  
because, for example, we want to have a synonym that says "street" = "st".

So the question is, how do I accomplish this? Is there a way to do a  
standard tokenization, apply synonyms, then concat the tokens into a single  
token before applying the edgeNGram filter? Maybe there is something else?

I have tried to use a standard tokenizer and utilize a match\_phrase\_prefix  
but that gives me issues. Two examples:

- If I type in "500 m" or "500 ma" it will not return the result i'm  
looking for. This is because "madison" is far down the expressions list. I  
have to go up to around 750 max expressions in order to get this to work  
properly
- If I type in "500 madison a" it will return no results. This is because  
it can't get to "ave" within it's max expressions. I have to go up to  
around 7500 max expressions in order for this to work properly.

And that's just not a reasonable solution for autocomplete.

_Synonym Filter:_  
"synonym": {  
"type": "synonym",  
"synonyms\_path": "analysis/address\_syms.txt"  
}

_EdgeNGram Filter:_  
"substring": {  
"type": "edgeNGram",  
"min\_gram": 1,  
"max\_gram": 50,  
"side": "front"  
},

_Analyzer:_  
"str\_index\_analyzer": {  
"tokenizer": "keyword",  
"char\_filter": [  
"my\_filter"  
],  
"filter": [  
"lowercase",  
"synonym",  
"substring"  
]  
}

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Anthony\_Campagna](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anthony_campagna/32/2079_2.png) [@Anthony\_Campagna](https://discuss.elastic.co/u/Anthony_Campagna)\
**Post date:** [October 4, 2013, 6:59pm UTC](https://discuss.elastic.co/t/synonyms-with-keyword-tokenizer/13757/2 "2013-10-04T18:59:28Z")

</div>

Nobody has a solution for using the synonym filter with the keyword  
tokenizer? I find this hard to believe. It seems like something that would  
be very useful for many users of elasticsearch.

On Wednesday, September 25, 2013 4:26:53 PM UTC-4, Anthony Campagna wrote:

> _Goal:_ To seamlessly autocomplete addresses while utilizing synonyms.
> 
> In my mapping/index settings I have a filter for synonym and a filter for  
> edgeNgrams. I then use an index analyzer to utilize the synonym filter and  
> edgeNgram. Obviously this doesn't work because we tokenize on the whole  
> string and the string doesn't match a synonym. This is very problematic  
> because, for example, we want to have a synonym that says "street" = "st".
> 
> So the question is, how do I accomplish this? Is there a way to do a  
> standard tokenization, apply synonyms, then concat the tokens into a single  
> token before applying the edgeNGram filter? Maybe there is something else?
> 
> I have tried to use a standard tokenizer and utilize a match\_phrase\_prefix  
> but that gives me issues. Two examples:
> 
> - If I type in "500 m" or "500 ma" it will not return the result i'm  
> looking for. This is because "madison" is far down the expressions list. I  
> have to go up to around 750 max expressions in order to get this to work  
> properly
> - If I type in "500 madison a" it will return no results. This is because  
> it can't get to "ave" within it's max expressions. I have to go up to  
> around 7500 max expressions in order for this to work properly.
> 
> And that's just not a reasonable solution for autocomplete.
> 
> _Synonym Filter:_  
> "synonym": {  
> "type": "synonym",  
> "synonyms\_path": "analysis/address\_syms.txt"  
> }
> 
> _EdgeNGram Filter:_  
> "substring": {  
> "type": "edgeNGram",  
> "min\_gram": 1,  
> "max\_gram": 50,  
> "side": "front"  
> },
> 
> _Analyzer:_  
> "str\_index\_analyzer": {  
> "tokenizer": "keyword",  
> "char\_filter": [  
> "my\_filter"  
> ],  
> "filter": [  
> "lowercase",  
> "synonym",  
> "substring"  
> ]  
> }

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:13am UTC](https://discuss.elastic.co/t/synonyms-with-keyword-tokenizer/13757/3 "2017-07-06T02:13:44Z")

</div>


