# Synonym expansion only during search time?

**URL:** <https://discuss.elastic.co/t/synonym-expansion-only-during-search-time/4659>\
**Category:** Elasticsearch\
**Created:** [June 20, 2011, 1:24pm UTC](https://discuss.elastic.co/t/synonym-expansion-only-during-search-time/4659 "2011-06-20T13:24:22Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Lukas\_Vlcek1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lukas_vlcek1/32/819_2.png) [@Lukas\_Vlcek1](https://discuss.elastic.co/u/Lukas_Vlcek1)\
**Post date:** [June 20, 2011, 1:24pm UTC](https://discuss.elastic.co/t/synonym-expansion-only-during-search-time/4659/1 "2011-06-20T13:24:22Z")

</div>

Hi,

I would like to make a synonym for "ws". It should translate to "web  
service". After some experiments I found out that using a synonym filter is  
not a good solution for this because if I use a rule like "ws =\> web  
service" or "ws, web service" then document having "ws" term gets injected  
two other terms "web" and "service" during indexing which means that it is  
relevant for search "web" (note the omitted "service" part of the query).  
The other problem is that if there is such a synonym rule then during search  
the query "ws" will be expanded to two other terms "web " and "service"  
which means that documents having only "web" in it can match no matter they  
do NOT contain "service" term.

Not sure if there is other solution but it seems to me that what I am  
looking for in this case is functionality that would expand Lucene query by  
adding proximity part like "web service"~1 (or something like that). Is  
something like that possible in ES? Or are there other better solutions how  
to handle "ws" - "web service" synonym and avoid false search hits?

Regards,  
Lukas

---

<div class="post-metadata">

**Author:** ![Jan\_Fiedler](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jan_fiedler/32/2518_2.png) [@Jan\_Fiedler](https://discuss.elastic.co/u/Jan_Fiedler)\
**Post date:** [June 20, 2011, 2:37pm UTC](https://discuss.elastic.co/t/synonym-expansion-only-during-search-time/4659/2 "2011-06-20T14:37:40Z")

</div>

I guess you are opening an interesting can of worms there. I have  
worked on similar issues on pure Lucene (2.x versions back then) and I  
am not aware are any build-in Lucene solutions (this may have changed  
though).

Back then I eventually ended up writing my own index time & query time  
analyzers to handle this. I think what you are looking for feature  
wise is called 'common phrases'. For example, 'baby doll' (a kind of  
special pyjama) is a classical common phrase that has a completely  
different meaning than the individual terms 'baby' and 'doll'. Usually  
you do not want 'baby doll' to match if someone searches for 'doll'.

To support something like this you basically have to detect certain  
sequences of tokens at indexing time and modify them such that they  
are not matched against the individual terms anymore (e.g. one can  
simply concatenate them with a separator). Then at query time, you  
would apply the same analyzing process such that someone searching for  
'baby doll' gets the correctly encoded keyword (e.g. 'baby-doll')  
while someone searching for 'doll for a baby' gets the classical  
tokens ('baby', 'doll' assuming stopword removal).

In your case, you actually want synonyms + common phrases (i.e. expand  
'ws' into 'web service' and handle 'web service' as a common phrase).  
It would be interesting to learn whether there any open-source common  
phrase solutions out there ...

On Jun 20, 3:24 pm, Lukáš Vlček [lukas.vl...@gmail.com](mailto:lukas.vl...@gmail.com) wrote:

> Hi,
> 
> I would like to make a synonym for "ws". It should translate to "web  
> service". After some experiments I found out that using a synonym filter is  
> not a good solution for this because if I use a rule like "ws =\> web  
> service" or "ws, web service" then document having "ws" term gets injected  
> two other terms "web" and "service" during indexing which means that it is  
> relevant for search "web" (note the omitted "service" part of the query).  
> The other problem is that if there is such a synonym rule then during search  
> the query "ws" will be expanded to two other terms "web " and "service"  
> which means that documents having only "web" in it can match no matter they  
> do NOT contain "service" term.
> 
> Not sure if there is other solution but it seems to me that what I am  
> looking for in this case is functionality that would expand Lucene query by  
> adding proximity part like "web service"~1 (or something like that). Is  
> something like that possible in ES? Or are there other better solutions how  
> to handle "ws" - "web service" synonym and avoid false search hits?
> 
> Regards,  
> Lukas

---

<div class="post-metadata">

**Author:** ![ppearcy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ppearcy/32/980_2.png) [@ppearcy](https://discuss.elastic.co/u/ppearcy)\
**Post date:** [July 27, 2011, 6:49am UTC](https://discuss.elastic.co/t/synonym-expansion-only-during-search-time/4659/3 "2011-07-27T06:49:04Z")

</div>

Hey Lukas,  
Jumping on this a little late, but I was just confronted with  
similar. For your "web service" example, I think things would work as  
expected with a synonym file that looks like this:  
ws, web service

Make sure to keep the single token synonym on the front, as that is  
the one that will go into the index.

Then with an analyzer with expand set to false that looks something  
like this:  
analysis :  
filter :  
english\_snowball:  
type : snowball  
language : English  
synonym :  
type : synonym  
synonyms\_path : analysis/synonym.txt  
ignore\_case : true  
expand : false  
analyzer :  
test\_analyzer :  
type : custom  
filter : [standard, lowercase, synonym, stop,  
english\_snowball]  
tokenizer : standard

When a user enters a query I am performing a phrase search and a  
normal term search, eg:  
"web service" OR (web service)

which should correctly match anything with ws. Where things got a  
little tricky for me is when I had multiple synonym phrases that did  
not have a single token synonym form. So, I ended up using a uuid for  
the single token. So, if I had a list like this:  
nicotine addiction, smoking cessation

I would transform it into something like this:  
7b4d385ab81911e0bc60001e0beaa9c4, nicotine addiction, smoking  
cessation

Would be curious to know if this approach works for your use case and  
if you see any holes with this technique. My testing so far looks  
good.

Thanks!  
Paul

On Jun 20, 8:37 am, Jan Fiedler [fiedler....@gmail.com](mailto:fiedler....@gmail.com) wrote:

> I guess you are opening an interesting can of worms there. I have  
> worked on similar issues on pure Lucene (2.x versions back then) and I  
> am not aware are any build-in Lucene solutions (this may have changed  
> though).
> 
> Back then I eventually ended up writing my own index time & query time  
> analyzers to handle this. I think what you are looking for feature  
> wise is called 'common phrases'. For example, 'baby doll' (a kind of  
> special pyjama) is a classical common phrase that has a completely  
> different meaning than the individual terms 'baby' and 'doll'. Usually  
> you do not want 'baby doll' to match if someone searches for 'doll'.
> 
> To support something like this you basically have to detect certain  
> sequences of tokens at indexing time and modify them such that they  
> are not matched against the individual terms anymore (e.g. one can  
> simply concatenate them with a separator). Then at query time, you  
> would apply the same analyzing process such that someone searching for  
> 'baby doll' gets the correctly encoded keyword (e.g. 'baby-doll')  
> while someone searching for 'doll for a baby' gets the classical  
> tokens ('baby', 'doll' assuming stopword removal).
> 
> In your case, you actually want synonyms + common phrases (i.e. expand  
> 'ws' into 'web service' and handle 'web service' as a common phrase).  
> It would be interesting to learn whether there any open-source common  
> phrase solutions out there ...
> 
> On Jun 20, 3:24 pm, Lukáš Vlček [lukas.vl...@gmail.com](mailto:lukas.vl...@gmail.com) wrote:
> 
> > Hi,
> 
> > I would like to make asynonymfor "ws". It should translate to "web  
> > service". After some experiments I found out that using asynonymfilter is  
> > not a good solution for this because if I use a rule like "ws =\> web  
> > service" or "ws, web service" then document having "ws" term gets injected  
> > two other terms "web" and "service" during indexing which means that it is  
> > relevant for search "web" (note the omitted "service" part of the query).  
> > The other problem is that if there is such asynonymrule then during search  
> > the query "ws" will be expanded to two other terms "web " and "service"  
> > which means that documents having only "web" in it can match no matter they  
> > do NOT contain "service" term.
> 
> > Not sure if there is other solution but it seems to me that what I am  
> > looking for in this case is functionality that would expand Lucene query by  
> > adding proximity part like "web service"~1 (or something like that). Is  
> > something like that possible in ES? Or are there other better solutions how  
> > to handle "ws" - "web service"synonymand avoid false search hits?
> 
> > Regards,  
> > Lukas

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:59am UTC](https://discuss.elastic.co/t/synonym-expansion-only-during-search-time/4659/4 "2017-07-06T03:59:06Z")

</div>


