# Synonym multi words search

**URL:** <https://discuss.elastic.co/t/synonym-multi-words-search/10964>\
**Category:** Elasticsearch\
**Created:** [March 1, 2013, 12:02pm UTC](https://discuss.elastic.co/t/synonym-multi-words-search/10964 "2013-03-01T12:02:40Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Rajesh\_Tarle](https://avatars.discourse-cdn.com/v4/letter/r/3da27b/32.png) [@Rajesh\_Tarle](https://discuss.elastic.co/u/Rajesh_Tarle)\
**Post date:** [March 1, 2013, 12:02pm UTC](https://discuss.elastic.co/t/synonym-multi-words-search/10964/1 "2013-03-01T12:02:40Z")

</div>

hi,

i am using elastic search 0.20.5 version.

I have mapped in synonym file as follow

abc,oracle america  
xyz,abc  
xyz,eloqua automation

and my analyzer mapping as following

{ \n"  
+ " "index" : {\n"  
+ ""analysis" : {\n"  
+ ""analyzer" : {\n"  
+ " "mainindexanalyzer" : {\n"  
+ ""type":"custom",\n" //whitespace standard  
+ ""tokenizer" : "standard",\n" //,"myshingle"  
,"mystopword" "length", "length",["lowercase"  
,"asciifolding","myworddelimiter","my\_snowball"]  
+ ""filter" :  
["lowercase","asciifolding","length","mystopword","my\_snowball","mystemmer","myshingle","mysynonym"],\n"  
+ ""char\_filter" :["html\_strip"]\n"  
+ " },\n"  
+ ""mainsearchanalyzer" : {\n"  
+ ""type":"custom",\n"  
+ " "tokenizer" :  
"whitespace",\n"//["lowercase","asciifolding","mystopword","mysynonym","myworddelimiter","my\_snowball"]  
+ ""filter" :  
["lowercase","asciifolding","mystopword","mystemmer","myshingle","mysynonym","myworddelimiter"],\n"  
+ ""char\_filter" :["html\_strip"]\n"  
+ "}\n"  
+ "},\n"  
+ ""filter" : {\n"  
+ ""mystemmer":{\n"  
+ " "type" : "stemmer",\n"  
+ ""name" :"english"\n"  
+ "},\n"  
+ ""my\_snowball" : {\n"  
+ ""type" : "snowball",\n"  
+ ""language" : "English"\n"  
+ "},\n"  
+ ""mystopword": {\n"  
+ " "type" : "stop",\n"  
+ ""stopwords\_path" :"" + stopwordfilepath + "" ,\n"  
//+ ""stopwords\_path"  
:"F:/resources/stopwordeng.txt" ,\n"  
+ ""ignore\_case":true\n"  
+ "},\n"  
+ ""myshingle":{\n"  
+ " "type" : "shingle",\n"  
+ ""max\_shingle\_size" :100,\n"  
+ ""min\_shingle\_size":2,\n"  
+ ""output\_unigrams":true \n"  
+ "},\n"  
+ ""mysynonym": {\n"  
+ " "type" : "synonym",\n"  
+ ""synonyms\_path" :"" + synonymfilepaths\_to\_p + ""  
,\n"  
+ ""ignore\_case":true,\n"  
+ ""expand":true\n"  
+ "},\n"  
+ ""myworddelimiter":{\n"  
+ " "type" : "word\_delimiter",\n"  
+ ""generate\_word\_parts" :true ,\n"  
+ ""generate\_number\_parts" :true ,\n"  
+ ""catenate\_words" :true ,\n"  
+ ""catenate\_numbers" :false ,\n"  
+ ""catenate\_all" :true ,\n"  
+ ""split\_on\_case\_change" :true ,\n"  
+ ""preserve\_original" :true ,\n"  
+ ""split\_on\_numerics":true ,\n"  
+ ""stem\_english\_possessive":true,\n"  
+ ""protected\_words\_path " : "" +  
protectedwordfilepath + "",\n"  
+ ""type\_table\_path " : "" + typetablefilepath +  
""\n"  
+ "}\n"  
+ "}\n"  
+ "}\n"  
+ "}\n"  
+ "}";

and my index are following.  
{id=4, Description=abc}  
{id=5, Description=xyz}  
{id=1, Description=Eloqua Automation}  
{id=3, Description=Oracle america}

when I search using "_oracle america_" i get '_abc_' and '_oracle america_'  
it's ok. but when searching using '_oracle_' i got only '_oracle america_'  
it's also ok.  
but I search using _'america_' I got "_oracle america_" and "\*abc" \*both.  
why both value are come when using america it my doubt.what it wrong.

Please help.

Thanks  
Rajesh

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![ppearcy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ppearcy/32/980_2.png) [@ppearcy](https://discuss.elastic.co/u/ppearcy)\
**Post date:** [March 1, 2013, 11:34pm UTC](https://discuss.elastic.co/t/synonym-multi-words-search/10964/2 "2013-03-01T23:34:31Z")

</div>

I find the following approach effective if you are doing multi-word  
synonyms (synonym phrases):

- Only apply the synonym expansion at index time
- Don't have the synonym filter applied search
- Use directional synonyms where appropriate. You want to make sure that  
you're not injecting terms that are too general.

For example, you probably want:  
oracle america =\> abc

Otherwise, a more general term "america" will get injected when you see  
something specific.

If you provide a reproducible curl based gist with your current and  
expected behavior, I could provide more details.

Best Regards,  
Paul

On Friday, March 1, 2013 5:02:40 AM UTC-7, raj wrote:

> hi,
> 
> i am using Elasticsearch 0.20.5 version.
> 
> I have mapped in synonym file as follow
> 
> abc,oracle america  
> xyz,abc  
> xyz,eloqua automation
> 
> and my analyzer mapping as following
> 
> { \n"  
> + " "index" : {\n"  
> + ""analysis" : {\n"  
> + ""analyzer" : {\n"  
> + " "mainindexanalyzer" : {\n"  
> + ""type":"custom",\n" //whitespace standard  
> + ""tokenizer" : "standard",\n" //,"myshingle"  
> ,"mystopword" "length", "length",["lowercase"  
> ,"asciifolding","myworddelimiter","my\_snowball"]  
> + ""filter" :  
> ["lowercase","asciifolding","length","mystopword","my\_snowball","mystemmer","myshingle","mysynonym"],\n"  
> + ""char\_filter" :["html\_strip"]\n"  
> + " },\n"  
> + ""mainsearchanalyzer" : {\n"  
> + ""type":"custom",\n"  
> + " "tokenizer" :  
> "whitespace",\n"//["lowercase","asciifolding","mystopword","mysynonym","myworddelimiter","my\_snowball"]  
> + ""filter" :  
> ["lowercase","asciifolding","mystopword","mystemmer","myshingle","mysynonym","myworddelimiter"],\n"  
> + ""char\_filter" :["html\_strip"]\n"  
> + "}\n"  
> + "},\n"  
> + ""filter" : {\n"  
> + ""mystemmer":{\n"  
> + " "type" : "stemmer",\n"  
> + ""name" :"english"\n"  
> + "},\n"  
> + ""my\_snowball" : {\n"  
> + ""type" : "snowball",\n"  
> + ""language" : "English"\n"  
> + "},\n"  
> + ""mystopword": {\n"  
> + " "type" : "stop",\n"  
> + ""stopwords\_path" :"" + stopwordfilepath + ""  
> ,\n"  
> //+ ""stopwords\_path"  
> :"F:/resources/stopwordeng.txt" ,\n"  
> + ""ignore\_case":true\n"  
> + "},\n"  
> + ""myshingle":{\n"  
> + " "type" : "shingle",\n"  
> + ""max\_shingle\_size" :100,\n"  
> + ""min\_shingle\_size":2,\n"  
> + ""output\_unigrams":true \n"  
> + "},\n"  
> + ""mysynonym": {\n"  
> + " "type" : "synonym",\n"  
> + ""synonyms\_path" :"" + synonymfilepaths\_to\_p +  
> "" ,\n"  
> + ""ignore\_case":true,\n"  
> + ""expand":true\n"  
> + "},\n"  
> + ""myworddelimiter":{\n"  
> + " "type" : "word\_delimiter",\n"  
> + ""generate\_word\_parts" :true ,\n"  
> + ""generate\_number\_parts" :true ,\n"  
> + ""catenate\_words" :true ,\n"  
> + ""catenate\_numbers" :false ,\n"  
> + ""catenate\_all" :true ,\n"  
> + ""split\_on\_case\_change" :true ,\n"  
> + ""preserve\_original" :true ,\n"  
> + ""split\_on\_numerics":true ,\n"  
> + ""stem\_english\_possessive":true,\n"  
> + ""protected\_words\_path " : "" +  
> protectedwordfilepath + "",\n"  
> + ""type\_table\_path " : "" + typetablefilepath +  
> ""\n"  
> + "}\n"  
> + "}\n"  
> + "}\n"  
> + "}\n"  
> + "}";
> 
> and my index are following.  
> {id=4, Description=abc}  
> {id=5, Description=xyz}  
> {id=1, Description=Eloqua Automation}  
> {id=3, Description=Oracle america}
> 
> when I search using "_oracle america_" i get '_abc_' and '_oracle america_  
> '  
> it's ok. but when searching using '_oracle_' i got only '_oracle america_  
> '  
> it's also ok.  
> but I search using _'america_' I got "_oracle america_" and "\*abc" \*both.  
> why both value are come when using america it my doubt.what it wrong.
> 
> Please help.
> 
> Thanks  
> Rajesh

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Rajesh\_Tarle](https://avatars.discourse-cdn.com/v4/letter/r/3da27b/32.png) [@Rajesh\_Tarle](https://discuss.elastic.co/u/Rajesh_Tarle)\
**Post date:** [March 4, 2013, 8:43am UTC](https://discuss.elastic.co/t/synonym-multi-words-search/10964/3 "2013-03-04T08:43:19Z")

</div>

On Friday, March 1, 2013 5:32:40 PM UTC+5:30, raj wrote:

> hi,
> 
> i am using Elasticsearch 0.20.5 version.
> 
> I have mapped in synonym file as follow
> 
> abc,oracle america  
> xyz,abc  
> xyz,eloqua automation
> 
> and my analyzer mapping as following
> 
> { \n"  
> + " "index" : {\n"  
> + ""analysis" : {\n"  
> + ""analyzer" : {\n"  
> + " "mainindexanalyzer" : {\n"  
> + ""type":"custom",\n" //whitespace standard  
> + ""tokenizer" : "standard",\n" //,"myshingle"  
> ,"mystopword" "length", "length",["lowercase"  
> ,"asciifolding","myworddelimiter","my\_snowball"]  
> + ""filter" :  
> ["lowercase","asciifolding","length","mystopword","my\_snowball","mystemmer","myshingle","mysynonym"],\n"  
> + ""char\_filter" :["html\_strip"]\n"  
> + " },\n"  
> + ""mainsearchanalyzer" : {\n"  
> + ""type":"custom",\n"  
> + " "tokenizer" :  
> "whitespace",\n"//["lowercase","asciifolding","mystopword","mysynonym","myworddelimiter","my\_snowball"]  
> + ""filter" :  
> ["lowercase","asciifolding","mystopword","mystemmer","myshingle","mysynonym","myworddelimiter"],\n"  
> + ""char\_filter" :["html\_strip"]\n"  
> + "}\n"  
> + "},\n"  
> + ""filter" : {\n"  
> + ""mystemmer":{\n"  
> + " "type" : "stemmer",\n"  
> + ""name" :"english"\n"  
> + "},\n"  
> + ""my\_snowball" : {\n"  
> + ""type" : "snowball",\n"  
> + ""language" : "English"\n"  
> + "},\n"  
> + ""mystopword": {\n"  
> + " "type" : "stop",\n"  
> + ""stopwords\_path" :"" + stopwordfilepath + ""  
> ,\n"  
> //+ ""stopwords\_path"  
> :"F:/resources/stopwordeng.txt" ,\n"  
> + ""ignore\_case":true\n"  
> + "},\n"  
> + ""myshingle":{\n"  
> + " "type" : "shingle",\n"  
> + ""max\_shingle\_size" :100,\n"  
> + ""min\_shingle\_size":2,\n"  
> + ""output\_unigrams":true \n"  
> + "},\n"  
> + ""mysynonym": {\n"  
> + " "type" : "synonym",\n"  
> + ""synonyms\_path" :"" + synonymfilepaths\_to\_p +  
> "" ,\n"  
> + ""ignore\_case":true,\n"  
> + ""expand":true\n"  
> + "},\n"  
> + ""myworddelimiter":{\n"  
> + " "type" : "word\_delimiter",\n"  
> + ""generate\_word\_parts" :true ,\n"  
> + ""generate\_number\_parts" :true ,\n"  
> + ""catenate\_words" :true ,\n"  
> + ""catenate\_numbers" :false ,\n"  
> + ""catenate\_all" :true ,\n"  
> + ""split\_on\_case\_change" :true ,\n"  
> + ""preserve\_original" :true ,\n"  
> + ""split\_on\_numerics":true ,\n"  
> + ""stem\_english\_possessive":true,\n"  
> + ""protected\_words\_path " : "" +  
> protectedwordfilepath + "",\n"  
> + ""type\_table\_path " : "" + typetablefilepath +  
> ""\n"  
> + "}\n"  
> + "}\n"  
> + "}\n"  
> + "}\n"  
> + "}";
> 
> and my index are following.  
> {id=4, Description=abc}  
> {id=5, Description=xyz}  
> {id=1, Description=Eloqua Automation}  
> {id=3, Description=Oracle america}
> 
> when I search using "_oracle america_" i get '_abc_' and '_oracle america_  
> '  
> it's ok. but when searching using '_oracle_' i got only '_oracle america_  
> '  
> it's also ok.  
> but I search using _'america_' I got "_oracle america_" and "\*abc" \*both.  
> why both value are come when using america it my doubt.what it wrong.
> 
> Please help.
> 
> Thanks  
> Rajesh

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Rajesh\_Tarle](https://avatars.discourse-cdn.com/v4/letter/r/3da27b/32.png) [@Rajesh\_Tarle](https://discuss.elastic.co/u/Rajesh_Tarle)\
**Post date:** [March 4, 2013, 8:51am UTC](https://discuss.elastic.co/t/synonym-multi-words-search/10964/4 "2013-03-04T08:51:53Z")

</div>

hi Paul,

Thankyou for reply,  
I am new for Elasticsearch.  
please may you explain "synonym expansion at index time"  
and if I using "directional synonyms" then I not get appropriated result.  
For example  
when I using this  
oracle america =\> abc  
and search using "oracle america" I got only "oracle america"  
i can't got abc.

"abc" is available in index document.  
please help.

Thankyou  
Rajesh

On Saturday, March 2, 2013 5:04:31 AM UTC+5:30, ppearcy wrote:

> I find the following approach effective if you are doing multi-word  
> synonyms (synonym phrases):
> 
> - Only apply the synonym expansion at index time
> - Don't have the synonym filter applied search
> - Use directional synonyms where appropriate. You want to make sure that  
> you're not injecting terms that are too general.
> 
> For example, you probably want:  
> oracle america =\> abc
> 
> Otherwise, a more general term "america" will get injected when you see  
> something specific.
> 
> If you provide a reproducible curl based gist with your current and  
> expected behavior, I could provide more details.
> 
> Best Regards,  
> Paul
> 
> On Friday, March 1, 2013 5:02:40 AM UTC-7, raj wrote:
> 
> > hi,
> > 
> > i am using Elasticsearch 0.20.5 version.
> > 
> > I have mapped in synonym file as follow
> > 
> > abc,oracle america  
> > xyz,abc  
> > xyz,eloqua automation
> > 
> > and my analyzer mapping as following
> > 
> > { \n"  
> > + " "index" : {\n"  
> > + ""analysis" : {\n"  
> > + ""analyzer" : {\n"  
> > + " "mainindexanalyzer" : {\n"  
> > + ""type":"custom",\n" //whitespace standard  
> > + ""tokenizer" : "standard",\n" //,"myshingle"  
> > ,"mystopword" "length", "length",["lowercase"  
> > ,"asciifolding","myworddelimiter","my\_snowball"]  
> > + ""filter" :  
> > ["lowercase","asciifolding","length","mystopword","my\_snowball","mystemmer","myshingle","mysynonym"],\n"  
> > + ""char\_filter" :["html\_strip"]\n"  
> > + " },\n"  
> > + ""mainsearchanalyzer" : {\n"  
> > + ""type":"custom",\n"  
> > + " "tokenizer" :  
> > "whitespace",\n"//["lowercase","asciifolding","mystopword","mysynonym","myworddelimiter","my\_snowball"]  
> > + ""filter" :  
> > ["lowercase","asciifolding","mystopword","mystemmer","myshingle","mysynonym","myworddelimiter"],\n"  
> > + ""char\_filter" :["html\_strip"]\n"  
> > + "}\n"  
> > + "},\n"  
> > + ""filter" : {\n"  
> > + ""mystemmer":{\n"  
> > + " "type" : "stemmer",\n"  
> > + ""name" :"english"\n"  
> > + "},\n"  
> > + ""my\_snowball" : {\n"  
> > + ""type" : "snowball",\n"  
> > + ""language" : "English"\n"  
> > + "},\n"  
> > + ""mystopword": {\n"  
> > + " "type" : "stop",\n"  
> > + ""stopwords\_path" :"" + stopwordfilepath + ""  
> > ,\n"  
> > //+ ""stopwords\_path"  
> > :"F:/resources/stopwordeng.txt" ,\n"  
> > + ""ignore\_case":true\n"  
> > + "},\n"  
> > + ""myshingle":{\n"  
> > + " "type" : "shingle",\n"  
> > + ""max\_shingle\_size" :100,\n"  
> > + ""min\_shingle\_size":2,\n"  
> > + ""output\_unigrams":true \n"  
> > + "},\n"  
> > + ""mysynonym": {\n"  
> > + " "type" : "synonym",\n"  
> > + ""synonyms\_path" :"" + synonymfilepaths\_to\_p +  
> > "" ,\n"  
> > + ""ignore\_case":true,\n"  
> > + ""expand":true\n"  
> > + "},\n"  
> > + ""myworddelimiter":{\n"  
> > + " "type" : "word\_delimiter",\n"  
> > + ""generate\_word\_parts" :true ,\n"  
> > + ""generate\_number\_parts" :true ,\n"  
> > + ""catenate\_words" :true ,\n"  
> > + ""catenate\_numbers" :false ,\n"  
> > + ""catenate\_all" :true ,\n"  
> > + ""split\_on\_case\_change" :true ,\n"  
> > + ""preserve\_original" :true ,\n"  
> > + ""split\_on\_numerics":true ,\n"  
> > + ""stem\_english\_possessive":true,\n"  
> > + ""protected\_words\_path " : "" +  
> > protectedwordfilepath + "",\n"  
> > + ""type\_table\_path " : "" + typetablefilepath +  
> > ""\n"  
> > + "}\n"  
> > + "}\n"  
> > + "}\n"  
> > + "}\n"  
> > + "}";
> > 
> > and my index are following.  
> > {id=4, Description=abc}  
> > {id=5, Description=xyz}  
> > {id=1, Description=Eloqua Automation}  
> > {id=3, Description=Oracle america}
> > 
> > when I search using "_oracle america_" i get '_abc_' and '\*oracle america  
> > \*'  
> > it's ok. but when searching using '_oracle_' i got only '\*oracle america  
> > \*'  
> > it's also ok.  
> > but I search using _'america_' I got "_oracle america_" and "\*abc" \*both.  
> > why both value are come when using america it my doubt.what it wrong.
> > 
> > Please help.
> > 
> > Thanks  
> > Rajesh

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![ppearcy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ppearcy/32/980_2.png) [@ppearcy](https://discuss.elastic.co/u/ppearcy)\
**Post date:** [March 4, 2013, 9:00pm UTC](https://discuss.elastic.co/t/synonym-multi-words-search/10964/5 "2013-03-04T21:00:30Z")

</div>

For a specific field's mappings, you have the ability to specify a  
index\_analyzer and a search\_analyzer. For, index time synonym expansion,  
you include the synonym token filter in the index analyzer and not the  
search analyzer.

Keep in mind the search engine is term based. So, if your synonym list is:  
abc, oracle america

Here is what happens to various data in the document:  
abc -\> abc oracle america  
oracle america -\> oracle america abc  
oracle -\> oracle  
america -\> america

Play around with index time synonym expansion and let us know if you have  
more specific questions and recreations of examples.

Thanks,  
Paul

On Monday, March 4, 2013 1:51:53 AM UTC-7, raj wrote:

> hi Paul,
> 
> Thankyou for reply,  
> I am new for Elasticsearch.  
> please may you explain "synonym expansion at index time"  
> and if I using "directional synonyms" then I not get appropriated result.  
> For example  
> when I using this  
> oracle america =\> abc  
> and search using "oracle america" I got only "oracle america"  
> i can't got abc.
> 
> "abc" is available in index document.  
> please help.
> 
> Thankyou  
> Rajesh
> 
> On Saturday, March 2, 2013 5:04:31 AM UTC+5:30, ppearcy wrote:
> 
> > I find the following approach effective if you are doing multi-word  
> > synonyms (synonym phrases):
> > 
> > - Only apply the synonym expansion at index time
> > - Don't have the synonym filter applied search
> > - Use directional synonyms where appropriate. You want to make sure that  
> > you're not injecting terms that are too general.
> > 
> > For example, you probably want:  
> > oracle america =\> abc
> > 
> > Otherwise, a more general term "america" will get injected when you see  
> > something specific.
> > 
> > If you provide a reproducible curl based gist with your current and  
> > expected behavior, I could provide more details.
> > 
> > Best Regards,  
> > Paul
> > 
> > On Friday, March 1, 2013 5:02:40 AM UTC-7, raj wrote:
> > 
> > > hi,
> > > 
> > > i am using Elasticsearch 0.20.5 version.
> > > 
> > > I have mapped in synonym file as follow
> > > 
> > > abc,oracle america  
> > > xyz,abc  
> > > xyz,eloqua automation
> > > 
> > > and my analyzer mapping as following
> > > 
> > > { \n"  
> > > + " "index" : {\n"  
> > > + ""analysis" : {\n"  
> > > + ""analyzer" : {\n"  
> > > + " "mainindexanalyzer" : {\n"  
> > > + ""type":"custom",\n" //whitespace standard  
> > > + ""tokenizer" : "standard",\n" //,"myshingle"  
> > > ,"mystopword" "length", "length",["lowercase"  
> > > ,"asciifolding","myworddelimiter","my\_snowball"]  
> > > + ""filter" :  
> > > ["lowercase","asciifolding","length","mystopword","my\_snowball","mystemmer","myshingle","mysynonym"],\n"  
> > > + ""char\_filter" :["html\_strip"]\n"  
> > > + " },\n"  
> > > + ""mainsearchanalyzer" : {\n"  
> > > + ""type":"custom",\n"  
> > > + " "tokenizer" :  
> > > "whitespace",\n"//["lowercase","asciifolding","mystopword","mysynonym","myworddelimiter","my\_snowball"]  
> > > + ""filter" :  
> > > ["lowercase","asciifolding","mystopword","mystemmer","myshingle","mysynonym","myworddelimiter"],\n"  
> > > + ""char\_filter" :["html\_strip"]\n"  
> > > + "}\n"  
> > > + "},\n"  
> > > + ""filter" : {\n"  
> > > + ""mystemmer":{\n"  
> > > + " "type" : "stemmer",\n"  
> > > + ""name" :"english"\n"  
> > > + "},\n"  
> > > + ""my\_snowball" : {\n"  
> > > + ""type" : "snowball",\n"  
> > > + ""language" : "English"\n"  
> > > + "},\n"  
> > > + ""mystopword": {\n"  
> > > + " "type" : "stop",\n"  
> > > + ""stopwords\_path" :"" + stopwordfilepath + ""  
> > > ,\n"  
> > > //+ ""stopwords\_path"  
> > > :"F:/resources/stopwordeng.txt" ,\n"  
> > > + ""ignore\_case":true\n"  
> > > + "},\n"  
> > > + ""myshingle":{\n"  
> > > + " "type" : "shingle",\n"  
> > > + ""max\_shingle\_size" :100,\n"  
> > > + ""min\_shingle\_size":2,\n"  
> > > + ""output\_unigrams":true \n"  
> > > + "},\n"  
> > > + ""mysynonym": {\n"  
> > > + " "type" : "synonym",\n"  
> > > + ""synonyms\_path" :"" + synonymfilepaths\_to\_p +  
> > > "" ,\n"  
> > > + ""ignore\_case":true,\n"  
> > > + ""expand":true\n"  
> > > + "},\n"  
> > > + ""myworddelimiter":{\n"  
> > > + " "type" : "word\_delimiter",\n"  
> > > + ""generate\_word\_parts" :true ,\n"  
> > > + ""generate\_number\_parts" :true ,\n"  
> > > + ""catenate\_words" :true ,\n"  
> > > + ""catenate\_numbers" :false ,\n"  
> > > + ""catenate\_all" :true ,\n"  
> > > + ""split\_on\_case\_change" :true ,\n"  
> > > + ""preserve\_original" :true ,\n"  
> > > + ""split\_on\_numerics":true ,\n"  
> > > + ""stem\_english\_possessive":true,\n"  
> > > + ""protected\_words\_path " : "" +  
> > > protectedwordfilepath + "",\n"  
> > > + ""type\_table\_path " : "" + typetablefilepath +  
> > > ""\n"  
> > > + "}\n"  
> > > + "}\n"  
> > > + "}\n"  
> > > + "}\n"  
> > > + "}";
> > > 
> > > and my index are following.  
> > > {id=4, Description=abc}  
> > > {id=5, Description=xyz}  
> > > {id=1, Description=Eloqua Automation}  
> > > {id=3, Description=Oracle america}
> > > 
> > > when I search using "_oracle america_" i get '_abc_' and '_oracle  
> > > america_'  
> > > it's ok. but when searching using '_oracle_' i got only '_oracle  
> > > america_'  
> > > it's also ok.  
> > > but I search using _'america_' I got "_oracle america_" and "\*abc" \*  
> > > both.  
> > > why both value are come when using america it my doubt.what it wrong.
> > > 
> > > Please help.
> > > 
> > > Thanks  
> > > Rajesh

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Rajesh\_Tarle](https://avatars.discourse-cdn.com/v4/letter/r/3da27b/32.png) [@Rajesh\_Tarle](https://discuss.elastic.co/u/Rajesh_Tarle)\
**Post date:** [March 5, 2013, 5:07am UTC](https://discuss.elastic.co/t/synonym-multi-words-search/10964/6 "2013-03-05T05:07:28Z")

</div>

hi paul,

i want to use synonym "_abc_" to "_oracle america_" and "_oracle america_"  
to "_abc_" only.  
i don't want to match "_oracle_" or "_america_" to "_abc_" and reverse. _how  
to define it in synonym file? please explain._  
I am using synonym in only in index time not in search time according to  
you.

Thanks  
Rajesh

On Tuesday, March 5, 2013 2:30:30 AM UTC+5:30, ppearcy wrote:

> For a specific field's mappings, you have the ability to specify a  
> index\_analyzer and a search\_analyzer. For, index time synonym expansion,  
> you include the synonym token filter in the index analyzer and not the  
> search analyzer.
> 
> Keep in mind the search engine is term based. So, if your synonym list is:  
> abc, oracle america
> 
> Here is what happens to various data in the document:  
> abc -\> abc oracle america  
> oracle america -\> oracle america abc  
> oracle -\> oracle  
> america -\> america
> 
> Play around with index time synonym expansion and let us know if you have  
> more specific questions and recreations of examples.
> 
> Thanks,  
> Paul
> 
> On Monday, March 4, 2013 1:51:53 AM UTC-7, raj wrote:
> 
> > hi Paul,
> > 
> > Thankyou for reply,  
> > I am new for Elasticsearch.  
> > please may you explain "synonym expansion at index time"  
> > and if I using "directional synonyms" then I not get appropriated result.  
> > For example  
> > when I using this  
> > oracle america =\> abc  
> > and search using "oracle america" I got only "oracle america"  
> > i can't got abc.
> > 
> > "abc" is available in index document.  
> > please help.
> > 
> > Thankyou  
> > Rajesh
> > 
> > On Saturday, March 2, 2013 5:04:31 AM UTC+5:30, ppearcy wrote:
> > 
> > > I find the following approach effective if you are doing multi-word  
> > > synonyms (synonym phrases):
> > > 
> > > - Only apply the synonym expansion at index time
> > > - Don't have the synonym filter applied search
> > > - Use directional synonyms where appropriate. You want to make sure that  
> > > you're not injecting terms that are too general.
> > > 
> > > For example, you probably want:  
> > > oracle america =\> abc
> > > 
> > > Otherwise, a more general term "america" will get injected when you see  
> > > something specific.
> > > 
> > > If you provide a reproducible curl based gist with your current and  
> > > expected behavior, I could provide more details.
> > > 
> > > Best Regards,  
> > > Paul
> > > 
> > > On Friday, March 1, 2013 5:02:40 AM UTC-7, raj wrote:
> > > 
> > > > hi,
> > > > 
> > > > i am using Elasticsearch 0.20.5 version.
> > > > 
> > > > I have mapped in synonym file as follow
> > > > 
> > > > abc,oracle america  
> > > > xyz,abc  
> > > > xyz,eloqua automation
> > > > 
> > > > and my analyzer mapping as following
> > > > 
> > > > { \n"  
> > > > + " "index" : {\n"  
> > > > + ""analysis" : {\n"  
> > > > + ""analyzer" : {\n"  
> > > > + " "mainindexanalyzer" : {\n"  
> > > > + ""type":"custom",\n" //whitespace standard  
> > > > + ""tokenizer" : "standard",\n"  
> > > > //,"myshingle" ,"mystopword" "length", "length",["lowercase"  
> > > > ,"asciifolding","myworddelimiter","my\_snowball"]  
> > > > + ""filter" :  
> > > > ["lowercase","asciifolding","length","mystopword","my\_snowball","mystemmer","myshingle","mysynonym"],\n"  
> > > > + ""char\_filter" :["html\_strip"]\n"  
> > > > + " },\n"  
> > > > + ""mainsearchanalyzer" : {\n"  
> > > > + ""type":"custom",\n"  
> > > > + " "tokenizer" :  
> > > > "whitespace",\n"//["lowercase","asciifolding","mystopword","mysynonym","myworddelimiter","my\_snowball"]  
> > > > + ""filter" :  
> > > > ["lowercase","asciifolding","mystopword","mystemmer","myshingle","mysynonym","myworddelimiter"],\n"  
> > > > + ""char\_filter" :["html\_strip"]\n"  
> > > > + "}\n"  
> > > > + "},\n"  
> > > > + ""filter" : {\n"  
> > > > + ""mystemmer":{\n"  
> > > > + " "type" : "stemmer",\n"  
> > > > + ""name" :"english"\n"  
> > > > + "},\n"  
> > > > + ""my\_snowball" : {\n"  
> > > > + ""type" : "snowball",\n"  
> > > > + ""language" : "English"\n"  
> > > > + "},\n"  
> > > > + ""mystopword": {\n"  
> > > > + " "type" : "stop",\n"  
> > > > + ""stopwords\_path" :"" + stopwordfilepath + ""  
> > > > ,\n"  
> > > > //+ ""stopwords\_path"  
> > > > :"F:/resources/stopwordeng.txt" ,\n"  
> > > > + ""ignore\_case":true\n"  
> > > > + "},\n"  
> > > > + ""myshingle":{\n"  
> > > > + " "type" : "shingle",\n"  
> > > > + ""max\_shingle\_size" :100,\n"  
> > > > + ""min\_shingle\_size":2,\n"  
> > > > + ""output\_unigrams":true \n"  
> > > > + "},\n"  
> > > > + ""mysynonym": {\n"  
> > > > + " "type" : "synonym",\n"  
> > > > + ""synonyms\_path" :"" + synonymfilepaths\_to\_p +  
> > > > "" ,\n"  
> > > > + ""ignore\_case":true,\n"  
> > > > + ""expand":true\n"  
> > > > + "},\n"  
> > > > + ""myworddelimiter":{\n"  
> > > > + " "type" : "word\_delimiter",\n"  
> > > > + ""generate\_word\_parts" :true ,\n"  
> > > > + ""generate\_number\_parts" :true ,\n"  
> > > > + ""catenate\_words" :true ,\n"  
> > > > + ""catenate\_numbers" :false ,\n"  
> > > > + ""catenate\_all" :true ,\n"  
> > > > + ""split\_on\_case\_change" :true ,\n"  
> > > > + ""preserve\_original" :true ,\n"  
> > > > + ""split\_on\_numerics":true ,\n"  
> > > > + ""stem\_english\_possessive":true,\n"  
> > > > + ""protected\_words\_path " : "" +  
> > > > protectedwordfilepath + "",\n"  
> > > > + ""type\_table\_path " : "" + typetablefilepath +  
> > > > ""\n"  
> > > > + "}\n"  
> > > > + "}\n"  
> > > > + "}\n"  
> > > > + "}\n"  
> > > > + "}";
> > > > 
> > > > and my index are following.  
> > > > {id=4, Description=abc}  
> > > > {id=5, Description=xyz}  
> > > > {id=1, Description=Eloqua Automation}  
> > > > {id=3, Description=Oracle america}
> > > > 
> > > > when I search using "_oracle america_" i get '_abc_' and '_oracle  
> > > > america_'  
> > > > it's ok. but when searching using '_oracle_' i got only '_oracle  
> > > > america_'  
> > > > it's also ok.  
> > > > but I search using _'america_' I got "_oracle america_" and "\*abc" \*  
> > > > both.  
> > > > why both value are come when using america it my doubt.what it wrong.
> > > > 
> > > > Please help.
> > > > 
> > > > Thanks  
> > > > Rajesh

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![ppearcy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ppearcy/32/980_2.png) [@ppearcy](https://discuss.elastic.co/u/ppearcy)\
**Post date:** [March 5, 2013, 4:45pm UTC](https://discuss.elastic.co/t/synonym-multi-words-search/10964/7 "2013-03-05T16:45:33Z")

</div>

AFAIK, this may not be possible if you are processing fields with other  
text. The core issue is that oracle and america are two separate terms in  
your index.

If you used a keyword tokenizer on your field and a keyword tokenizer on  
your synonym file this would work with your example above, but only if your  
field contains just "oracle america" or just "abc". If your field became  
"Oracle America Company" it would not longer work with keyword analyzer.

The key thing to try to understand is that this is a term based search and  
that ends up having some conflicts with synonym phrases. Perhaps you should  
stick to single work synonyms?

Maybe others have further input, but I don't think there is much more I can  
do to help. I recommend playing around with the various options and seeing  
which tradeoffs are acceptable.

Thanks,  
Paul

On Monday, March 4, 2013 10:07:28 PM UTC-7, raj wrote:

> hi paul,
> 
> i want to use synonym "_abc_" to "_oracle america_" and "_oracle america_"  
> to "_abc_" only.  
> i don't want to match "_oracle_" or "_america_" to "_abc_" and reverse. _how  
> to define it in synonym file? please explain._  
> I am using synonym in only in index time not in search time according to  
> you.
> 
> Thanks  
> Rajesh
> 
> On Tuesday, March 5, 2013 2:30:30 AM UTC+5:30, ppearcy wrote:
> 
> > For a specific field's mappings, you have the ability to specify a  
> > index\_analyzer and a search\_analyzer. For, index time synonym expansion,  
> > you include the synonym token filter in the index analyzer and not the  
> > search analyzer.
> > 
> > Keep in mind the search engine is term based. So, if your synonym list is:  
> > abc, oracle america
> > 
> > Here is what happens to various data in the document:  
> > abc -\> abc oracle america  
> > oracle america -\> oracle america abc  
> > oracle -\> oracle  
> > america -\> america
> > 
> > Play around with index time synonym expansion and let us know if you have  
> > more specific questions and recreations of examples.
> > 
> > Thanks,  
> > Paul
> > 
> > On Monday, March 4, 2013 1:51:53 AM UTC-7, raj wrote:
> > 
> > > hi Paul,
> > > 
> > > Thankyou for reply,  
> > > I am new for Elasticsearch.  
> > > please may you explain "synonym expansion at index time"  
> > > and if I using "directional synonyms" then I not get appropriated result.  
> > > For example  
> > > when I using this  
> > > oracle america =\> abc  
> > > and search using "oracle america" I got only "oracle america"  
> > > i can't got abc.
> > > 
> > > "abc" is available in index document.  
> > > please help.
> > > 
> > > Thankyou  
> > > Rajesh
> > > 
> > > On Saturday, March 2, 2013 5:04:31 AM UTC+5:30, ppearcy wrote:
> > > 
> > > > I find the following approach effective if you are doing multi-word  
> > > > synonyms (synonym phrases):
> > > > 
> > > > - Only apply the synonym expansion at index time
> > > > - Don't have the synonym filter applied search
> > > > - Use directional synonyms where appropriate. You want to make sure  
> > > > that you're not injecting terms that are too general.
> > > > 
> > > > For example, you probably want:  
> > > > oracle america =\> abc
> > > > 
> > > > Otherwise, a more general term "america" will get injected when you see  
> > > > something specific.
> > > > 
> > > > If you provide a reproducible curl based gist with your current and  
> > > > expected behavior, I could provide more details.
> > > > 
> > > > Best Regards,  
> > > > Paul
> > > > 
> > > > On Friday, March 1, 2013 5:02:40 AM UTC-7, raj wrote:
> > > > 
> > > > > hi,
> > > > > 
> > > > > i am using Elasticsearch 0.20.5 version.
> > > > > 
> > > > > I have mapped in synonym file as follow
> > > > > 
> > > > > abc,oracle america  
> > > > > xyz,abc  
> > > > > xyz,eloqua automation
> > > > > 
> > > > > and my analyzer mapping as following
> > > > > 
> > > > > { \n"  
> > > > > + " "index" : {\n"  
> > > > > + ""analysis" : {\n"  
> > > > > + ""analyzer" : {\n"  
> > > > > + " "mainindexanalyzer" : {\n"  
> > > > > + ""type":"custom",\n" //whitespace standard  
> > > > > + ""tokenizer" : "standard",\n"  
> > > > > //,"myshingle" ,"mystopword" "length", "length",["lowercase"  
> > > > > ,"asciifolding","myworddelimiter","my\_snowball"]  
> > > > > + ""filter" :  
> > > > > ["lowercase","asciifolding","length","mystopword","my\_snowball","mystemmer","myshingle","mysynonym"],\n"  
> > > > > + ""char\_filter" :["html\_strip"]\n"  
> > > > > + " },\n"  
> > > > > + ""mainsearchanalyzer" : {\n"  
> > > > > + ""type":"custom",\n"  
> > > > > + " "tokenizer" :  
> > > > > "whitespace",\n"//["lowercase","asciifolding","mystopword","mysynonym","myworddelimiter","my\_snowball"]  
> > > > > + ""filter" :  
> > > > > ["lowercase","asciifolding","mystopword","mystemmer","myshingle","mysynonym","myworddelimiter"],\n"  
> > > > > + ""char\_filter" :["html\_strip"]\n"  
> > > > > + "}\n"  
> > > > > + "},\n"  
> > > > > + ""filter" : {\n"  
> > > > > + ""mystemmer":{\n"  
> > > > > + " "type" : "stemmer",\n"  
> > > > > + ""name" :"english"\n"  
> > > > > + "},\n"  
> > > > > + ""my\_snowball" : {\n"  
> > > > > + ""type" : "snowball",\n"  
> > > > > + ""language" : "English"\n"  
> > > > > + "},\n"  
> > > > > + ""mystopword": {\n"  
> > > > > + " "type" : "stop",\n"  
> > > > > + ""stopwords\_path" :"" + stopwordfilepath +  
> > > > > "" ,\n"  
> > > > > //+ ""stopwords\_path"  
> > > > > :"F:/resources/stopwordeng.txt" ,\n"  
> > > > > + ""ignore\_case":true\n"  
> > > > > + "},\n"  
> > > > > + ""myshingle":{\n"  
> > > > > + " "type" : "shingle",\n"  
> > > > > + ""max\_shingle\_size" :100,\n"  
> > > > > + ""min\_shingle\_size":2,\n"  
> > > > > + ""output\_unigrams":true \n"  
> > > > > + "},\n"  
> > > > > + ""mysynonym": {\n"  
> > > > > + " "type" : "synonym",\n"  
> > > > > + ""synonyms\_path" :"" + synonymfilepaths\_to\_p
> > > > > 
> > > > > - "" ,\n"  
> > > > > + ""ignore\_case":true,\n"  
> > > > > + ""expand":true\n"  
> > > > > + "},\n"  
> > > > > + ""myworddelimiter":{\n"  
> > > > > + " "type" : "word\_delimiter",\n"  
> > > > > + ""generate\_word\_parts" :true ,\n"  
> > > > > + ""generate\_number\_parts" :true ,\n"  
> > > > > + ""catenate\_words" :true ,\n"  
> > > > > + ""catenate\_numbers" :false ,\n"  
> > > > > + ""catenate\_all" :true ,\n"  
> > > > > + ""split\_on\_case\_change" :true ,\n"  
> > > > > + ""preserve\_original" :true ,\n"  
> > > > > + ""split\_on\_numerics":true ,\n"  
> > > > > + ""stem\_english\_possessive":true,\n"  
> > > > > + ""protected\_words\_path " : "" +  
> > > > > protectedwordfilepath + "",\n"  
> > > > > + ""type\_table\_path " : "" + typetablefilepath
> > > > > - ""\n"  
> > > > > + "}\n"  
> > > > > + "}\n"  
> > > > > + "}\n"  
> > > > > + "}\n"  
> > > > > + "}";
> > > > > 
> > > > > and my index are following.  
> > > > > {id=4, Description=abc}  
> > > > > {id=5, Description=xyz}  
> > > > > {id=1, Description=Eloqua Automation}  
> > > > > {id=3, Description=Oracle america}
> > > > > 
> > > > > when I search using "_oracle america_" i get '_abc_' and '_oracle  
> > > > > america_'  
> > > > > it's ok. but when searching using '_oracle_' i got only '_oracle  
> > > > > america_'  
> > > > > it's also ok.  
> > > > > but I search using _'america_' I got "_oracle america_" and "\*abc" \*  
> > > > > both.  
> > > > > why both value are come when using america it my doubt.what it wrong.
> > > > > 
> > > > > Please help.
> > > > > 
> > > > > Thanks  
> > > > > Rajesh

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:48am UTC](https://discuss.elastic.co/t/synonym-multi-words-search/10964/8 "2017-07-06T02:48:16Z")

</div>


