# Problem with word-separators in bool search with standard tokenizer

**URL:** <https://discuss.elastic.co/t/problem-with-word-separators-in-bool-search-with-standard-tokenizer/19883>\
**Category:** Elasticsearch\
**Created:** [September 19, 2014, 3:05pm UTC](https://discuss.elastic.co/t/problem-with-word-separators-in-bool-search-with-standard-tokenizer/19883 "2014-09-19T15:05:59Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ankush\_Jhalani](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ankush_jhalani/32/875_2.png) [@Ankush\_Jhalani](https://discuss.elastic.co/u/Ankush_Jhalani)\
**Post date:** [September 19, 2014, 3:05pm UTC](https://discuss.elastic.co/t/problem-with-word-separators-in-bool-search-with-standard-tokenizer/19883/1 "2014-09-19T15:05:59Z")

</div>

In our search we have configured text with 2 analyzers, english and  
standard so we can match phrases on the standard-analyzer. We break the  
keywords by space, and create a bool query for each word.

This is working fine for all cases except where the query has standard  
word-separators like & (ampersand), ; (semi-colon), etc. As  
word-separators are stripped in index by analyzer, searching for them  
returns 0 results. Gist.

> <https://gist.github.com/ajhalani/3def3ea7caec5cd58490>

I don't want to use a whitespace analyzer because we do actually want to  
ignore word separators. I was thinking about hacky workarounds like  
removing all standalone non-alphanumeric characters, or moving them in  
"should" instead of default "must" (in case we do have analyzers in future  
that are whitespace).

Thanks in advance.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/f2abbc24-52d5-4567-afa3-66610956ce0b%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/f2abbc24-52d5-4567-afa3-66610956ce0b%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Ankush\_Jhalani](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ankush_jhalani/32/875_2.png) [@Ankush\_Jhalani](https://discuss.elastic.co/u/Ankush_Jhalani)\
**Post date:** [September 19, 2014, 3:12pm UTC](https://discuss.elastic.co/t/problem-with-word-separators-in-bool-search-with-standard-tokenizer/19883/2 "2014-09-19T15:12:37Z")

</div>

On other hand, If I use a single query\_string instead of bool of terms it  
works. Does ES/lucene determines not to use the word-separators by looking  
at the definition of the fields.

On Friday, September 19, 2014 11:05:59 AM UTC-4, Ankush Jhalani wrote:

> In our search we have configured text with 2 analyzers, english and  
> standard so we can match phrases on the standard-analyzer. We break the  
> keywords by space, and create a bool query for each word.
> 
> This is working fine for all cases except where the query has standard  
> word-separators like & (ampersand), ; (semi-colon), etc. As  
> word-separators are stripped in index by analyzer, searching for them  
> returns 0 results. Gist.  
> [elasticsearch - bool search - word separator issue · GitHub](https://gist.github.com/ajhalani/3def3ea7caec5cd58490)
> 
> I don't want to use a whitespace analyzer because we do actually want to  
> ignore word separators. I was thinking about hacky workarounds like  
> removing all standalone non-alphanumeric characters, or moving them in  
> "should" instead of default "must" (in case we do have analyzers in future  
> that are whitespace).
> 
> Thanks in advance.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/4b205133-eecd-490a-a028-9a53a3230973%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/4b205133-eecd-490a-a028-9a53a3230973%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Ankush\_Jhalani](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ankush_jhalani/32/875_2.png) [@Ankush\_Jhalani](https://discuss.elastic.co/u/Ankush_Jhalani)\
**Post date:** [September 22, 2014, 4:19pm UTC](https://discuss.elastic.co/t/problem-with-word-separators-in-bool-search-with-standard-tokenizer/19883/3 "2014-09-22T16:19:10Z")

</div>

just checking back if anyone has any ideas.. thanks!

On Friday, September 19, 2014 11:05:59 AM UTC-4, Ankush Jhalani wrote:

> In our search we have configured text with 2 analyzers, english and  
> standard so we can match phrases on the standard-analyzer. We break the  
> keywords by space, and create a bool query for each word.
> 
> This is working fine for all cases except where the query has standard  
> word-separators like & (ampersand), ; (semi-colon), etc. As  
> word-separators are stripped in index by analyzer, searching for them  
> returns 0 results. Gist.  
> [elasticsearch - bool search - word separator issue · GitHub](https://gist.github.com/ajhalani/3def3ea7caec5cd58490)
> 
> I don't want to use a whitespace analyzer because we do actually want to  
> ignore word separators. I was thinking about hacky workarounds like  
> removing all standalone non-alphanumeric characters, or moving them in  
> "should" instead of default "must" (in case we do have analyzers in future  
> that are whitespace).
> 
> Thanks in advance.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/e7dfb594-58c1-4127-8ae7-73f2c1f0adca%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/e7dfb594-58c1-4127-8ae7-73f2c1f0adca%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [September 23, 2014, 3:33am UTC](https://discuss.elastic.co/t/problem-with-word-separators-in-bool-search-with-standard-tokenizer/19883/4 "2014-09-23T03:33:15Z")

</div>

The query string query is working because the ampersand is also being  
stripped from the query.

Your best bet is to use the pattern tokenizer and explicitly define which  
characters to split the input text on.

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

Cheers,

Ivan

On Mon, Sep 22, 2014 at 9:19 AM, Ankush Jhalani [ankush.jhalani@gmail.com](mailto:ankush.jhalani@gmail.com)  
wrote:

> just checking back if anyone has any ideas.. thanks!
> 
> On Friday, September 19, 2014 11:05:59 AM UTC-4, Ankush Jhalani wrote:
> 
> > In our search we have configured text with 2 analyzers, english and  
> > standard so we can match phrases on the standard-analyzer. We break the  
> > keywords by space, and create a bool query for each word.
> > 
> > This is working fine for all cases except where the query has standard  
> > word-separators like & (ampersand), ; (semi-colon), etc. As  
> > word-separators are stripped in index by analyzer, searching for them  
> > returns 0 results. Gist. [https://gist.github.com/](https://gist.github.com/)  
> > ajhalani/3def3ea7caec5cd58490
> > 
> > I don't want to use a whitespace analyzer because we do actually want to  
> > ignore word separators. I was thinking about hacky workarounds like  
> > removing all standalone non-alphanumeric characters, or moving them in  
> > "should" instead of default "must" (in case we do have analyzers in future  
> > that are whitespace).
> > 
> > Thanks in advance.
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/e7dfb594-58c1-4127-8ae7-73f2c1f0adca%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/e7dfb594-58c1-4127-8ae7-73f2c1f0adca%40googlegroups.com)  
> > [https://groups.google.com/d/msgid/elasticsearch/e7dfb594-58c1-4127-8ae7-73f2c1f0adca%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/e7dfb594-58c1-4127-8ae7-73f2c1f0adca%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > .
> 
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQCqrVg8kWgArY\_t5paHSCeEG9LWAdv\_0Q2rm9vdcnPqeQ%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQCqrVg8kWgArY_t5paHSCeEG9LWAdv_0Q2rm9vdcnPqeQ%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Bryan\_Warner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bryan_warner/32/1225_2.png) [@Bryan\_Warner](https://discuss.elastic.co/u/Bryan_Warner)\
**Post date:** [September 23, 2014, 2:03pm UTC](https://discuss.elastic.co/t/problem-with-word-separators-in-bool-search-with-standard-tokenizer/19883/5 "2014-09-23T14:03:23Z")

</div>

Hi Ankush,

A few weeks ago I released an Elasticsearch plugin that allows you to  
override the default word boundary properties for Unicode characters as  
implemented by the StandardTokenizer algorithm. I had the same issue where  
I wanted to use the StandardTokenizer but override the word boundary  
properties for special characters like '#', '@', etc. (for example, treat  
them the same way as the '\_' , which is categorized as an extended  
num-letter)

Plugin: [GitHub - bbguitar77/elasticsearch-analysis-standardext](https://github.com/bbguitar77/elasticsearch-analysis-standardext)

I hope this helps solve your issue.

Thanks  
Bryan

On Monday, September 22, 2014 12:19:10 PM UTC-4, Ankush Jhalani wrote:

> just checking back if anyone has any ideas.. thanks!
> 
> On Friday, September 19, 2014 11:05:59 AM UTC-4, Ankush Jhalani wrote:
> 
> > In our search we have configured text with 2 analyzers, english and  
> > standard so we can match phrases on the standard-analyzer. We break the  
> > keywords by space, and create a bool query for each word.
> > 
> > This is working fine for all cases except where the query has standard  
> > word-separators like & (ampersand), ; (semi-colon), etc. As  
> > word-separators are stripped in index by analyzer, searching for them  
> > returns 0 results. Gist.  
> > [elasticsearch - bool search - word separator issue · GitHub](https://gist.github.com/ajhalani/3def3ea7caec5cd58490)
> > 
> > I don't want to use a whitespace analyzer because we do actually want to  
> > ignore word separators. I was thinking about hacky workarounds like  
> > removing all standalone non-alphanumeric characters, or moving them in  
> > "should" instead of default "must" (in case we do have analyzers in future  
> > that are whitespace).
> > 
> > Thanks in advance.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/6af13c45-93e5-4a8e-9520-88fdc14056f8%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/6af13c45-93e5-4a8e-9520-88fdc14056f8%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:00am UTC](https://discuss.elastic.co/t/problem-with-word-separators-in-bool-search-with-standard-tokenizer/19883/6 "2017-07-06T01:00:26Z")

</div>


