# Help with analyzer and mapping

**URL:** <https://discuss.elastic.co/t/help-with-analyzer-and-mapping/9373>\
**Category:** Elasticsearch\
**Created:** [October 16, 2012, 9:53am UTC](https://discuss.elastic.co/t/help-with-analyzer-and-mapping/9373 "2012-10-16T09:53:28Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![sujoysett](https://avatars.discourse-cdn.com/v4/letter/s/2acd7d/32.png) [@sujoysett](https://discuss.elastic.co/u/sujoysett)\
**Post date:** [October 16, 2012, 9:53am UTC](https://discuss.elastic.co/t/help-with-analyzer-and-mapping/9373/1 "2012-10-16T09:53:28Z")

</div>

Hi,

In an index I have a field "text", analyzed with _standard analyzer_. I now  
want to return the documents which has the _keyword "at&t" occurring in the  
field "text"_.

However, "at" is probably a member of Stop Token Filter, and "&" is again  
probably a member of the tokenizer used. (I am don't have much clarity on  
the exact logic here).  
I have tried using _match_, _text_, and \*query\_string \*in my query, and all  
returns quite a lot of junk documents in additional to the required  
documents.

I was thinking of custom analyzer here, but I want to use "at" and "&" as  
is for other search functions on this same field, and a custom analyzer  
might upset that.

Am I missing something simpler here? How to search for a text that probably  
includes stopwords and tokenizer characters?  
Is something like exact search irrespective of tokens (might be time  
consuming search, I accept) possible in elasticsearch?

Thanks in advance,  
-- Sujoy.

--

---

<div class="post-metadata">

**Author:** ![Tanguy1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tanguy1/32/1749_2.png) [@Tanguy1](https://discuss.elastic.co/u/Tanguy1)\
**Post date:** [October 16, 2012, 10:09am UTC](https://discuss.elastic.co/t/help-with-analyzer-and-mapping/9373/2 "2012-10-16T10:09:16Z")

</div>

Hi,

You can use the \_analyze API to understand the logic behind analyzers:  
[http://localhost:9200/\_analyze?pretty=true&analyzer=standard&text=The+at%26t+company](http://localhost:9200/_analyze?pretty=true&analyzer=standard&text=The+at%26t+company)

If you index "The at&t company" with the standard analyzer, the token that  
are really indexed are "t" and "company". The same logic applies when  
searching with match & query\_string queries and that explains the results  
you have.

There are many ways to get the expected results when searching for "at&t".  
Some suggestions:

- use a custom analyzer for the "text" field in mapping
- declare "text" as multi\_field and search for exact matches

Hope this helps,

-- Tanguy  
Twitter: @tlrx

> **[tlrx - Overview](https://github.com/tlrx)**
>
> tlrx has 23 repositories available. Follow their code on GitHub.

Le mardi 16 octobre 2012 11:53:28 UTC+2, Sujoy Sett a écrit :

> Hi,
> 
> In an index I have a field "text", analyzed with _standard analyzer_. I  
> now want to return the documents which has the _keyword  
> "at&t" occurring in the field "text"_.
> 
> However, "at" is probably a member of Stop Token Filter, and "&" is again  
> probably a member of the tokenizer used. (I am don't have much clarity on  
> the exact logic here).  
> I have tried using _match_, _text_, and \*query\_string \*in my query, and  
> all returns quite a lot of junk documents in additional to the required  
> documents.
> 
> I was thinking of custom analyzer here, but I want to use "at" and "&" as  
> is for other search functions on this same field, and a custom analyzer  
> might upset that.
> 
> Am I missing something simpler here? How to search for a text that  
> probably includes stopwords and tokenizer characters?  
> Is something like exact search irrespective of tokens (might be time  
> consuming search, I accept) possible in elasticsearch?
> 
> Thanks in advance,  
> -- Sujoy.

--

---

<div class="post-metadata">

**Author:** ![sujoysett](https://avatars.discourse-cdn.com/v4/letter/s/2acd7d/32.png) [@sujoysett](https://discuss.elastic.co/u/sujoysett)\
**Post date:** [October 16, 2012, 2:00pm UTC](https://discuss.elastic.co/t/help-with-analyzer-and-mapping/9373/3 "2012-10-16T14:00:15Z")

</div>

Thanks Tanguy,

I will surely try the multi-field.

-- Sujoy.

On Tuesday, October 16, 2012 3:39:16 PM UTC+5:30, Tanguy wrote:

> Hi,
> 
> You can use the \_analyze API to understand the logic behind analyzers:
> 
> [http://localhost:9200/\_analyze?pretty=true&analyzer=standard&text=The+at%26t+company](http://localhost:9200/_analyze?pretty=true&analyzer=standard&text=The+at%26t+company)
> 
> If you index "The at&t company" with the standard analyzer, the token that  
> are really indexed are "t" and "company". The same logic applies when  
> searching with match & query\_string queries and that explains the results  
> you have.
> 
> There are many ways to get the expected results when searching for "at&t".  
> Some suggestions:
> 
> - use a custom analyzer for the "text" field in mapping
> - declare "text" as multi\_field and search for exact matches
> 
> Hope this helps,
> 
> -- Tanguy  
> Twitter: @tlrx  
> [tlrx (Tanguy Leroux) · GitHub](https://github.com/tlrx)
> 
> Le mardi 16 octobre 2012 11:53:28 UTC+2, Sujoy Sett a écrit :
> 
> > Hi,
> > 
> > In an index I have a field "text", analyzed with _standard analyzer_. I  
> > now want to return the documents which has the _keyword  
> > "at&t" occurring in the field "text"_.
> > 
> > However, "at" is probably a member of Stop Token Filter, and "&" is again  
> > probably a member of the tokenizer used. (I am don't have much clarity on  
> > the exact logic here).  
> > I have tried using _match_, _text_, and \*query\_string \*in my query, and  
> > all returns quite a lot of junk documents in additional to the required  
> > documents.
> > 
> > I was thinking of custom analyzer here, but I want to use "at" and "&" as  
> > is for other search functions on this same field, and a custom analyzer  
> > might upset that.
> > 
> > Am I missing something simpler here? How to search for a text that  
> > probably includes stopwords and tokenizer characters?  
> > Is something like exact search irrespective of tokens (might be time  
> > consuming search, I accept) possible in elasticsearch?
> > 
> > Thanks in advance,  
> > -- Sujoy.

--

---

<div class="post-metadata">

**Author:** ![sujoysett](https://avatars.discourse-cdn.com/v4/letter/s/2acd7d/32.png) [@sujoysett](https://discuss.elastic.co/u/sujoysett)\
**Post date:** [October 24, 2012, 10:21am UTC](https://discuss.elastic.co/t/help-with-analyzer-and-mapping/9373/4 "2012-10-24T10:21:05Z")

</div>

Hi All,

Can anyone explain what does "type" mean for a token?

[http://localhost:9200/[index]/\_analyze?pretty=true&analyzer=custom-whitespace-lowercase&text=ipad](http://localhost:9200/%5Bindex%5D/_analyze?pretty=true&analyzer=custom-whitespace-lowercase&text=ipad)  
gives response

{  
"tokens": [  
{  
"token": "ipad",  
"start\_offset": 0,  
"end\_offset": 4,  
"type": "word",  
"position": 1  
}  
]  
}

whereas,

[http://localhost:9200/[index]/\_analyze?pretty=true&analyzer=standard&text=ipad](http://localhost:9200/%5Bindex%5D/_analyze?pretty=true&analyzer=standard&text=ipad)  
gives response

{  
"tokens": [  
{  
"token": "ipad",  
"start\_offset": 0,  
"end\_offset": 4,  
"type": "",  
"position": 1  
}  
]  
}

custom-whitespace-lowercase is a custom analyzer defined  
with whitespace tokenizer and lowercase filter.  
The purpose of defining this analyzer was to avoid the stop-word filter  
that comes by default in standard analyzer.

But this analyzer is creating a different problem by not identifying the  
term "ipad" while querying.  
Also, a bit of extra information, I don't know whether relevant or not, my  
mapping is as follows:  
properties: {

- text: {
  - type: multi\_field
  - fields: {
    - text: {
      - type: string  
}

    - text\_custom\_1: {
      - include\_in\_all: false
      - analyzer: custom-whitespace-lowercase
      - type: string  
}  
}  
}

Thanks,  
-- Sujoy.

On Tuesday, October 16, 2012 7:30:15 PM UTC+5:30, Sujoy Sett wrote:

> Thanks Tanguy,
> 
> I will surely try the multi-field.
> 
> -- Sujoy.
> 
> On Tuesday, October 16, 2012 3:39:16 PM UTC+5:30, Tanguy wrote:
> 
> > Hi,
> > 
> > You can use the \_analyze API to understand the logic behind analyzers:
> > 
> > [http://localhost:9200/\_analyze?pretty=true&analyzer=standard&text=The+at%26t+company](http://localhost:9200/_analyze?pretty=true&analyzer=standard&text=The+at%26t+company)
> > 
> > If you index "The at&t company" with the standard analyzer, the token  
> > that are really indexed are "t" and "company". The same logic applies when  
> > searching with match & query\_string queries and that explains the results  
> > you have.
> > 
> > There are many ways to get the expected results when searching for  
> > "at&t". Some suggestions:
> > 
> > - use a custom analyzer for the "text" field in mapping
> > - declare "text" as multi\_field and search for exact matches
> > 
> > Hope this helps,
> > 
> > -- Tanguy  
> > Twitter: @tlrx  
> > [tlrx (Tanguy Leroux) · GitHub](https://github.com/tlrx)
> > 
> > Le mardi 16 octobre 2012 11:53:28 UTC+2, Sujoy Sett a écrit :
> > 
> > > Hi,
> > > 
> > > In an index I have a field "text", analyzed with _standard analyzer_. I  
> > > now want to return the documents which has the _keyword  
> > > "at&t" occurring in the field "text"_.
> > > 
> > > However, "at" is probably a member of Stop Token Filter, and "&" is  
> > > again probably a member of the tokenizer used. (I am don't have much  
> > > clarity on the exact logic here).  
> > > I have tried using _match_, _text_, and \*query\_string \*in my query, and  
> > > all returns quite a lot of junk documents in additional to the required  
> > > documents.
> > > 
> > > I was thinking of custom analyzer here, but I want to use "at" and "&"  
> > > as is for other search functions on this same field, and a custom analyzer  
> > > might upset that.
> > > 
> > > Am I missing something simpler here? How to search for a text that  
> > > probably includes stopwords and tokenizer characters?  
> > > Is something like exact search irrespective of tokens (might be time  
> > > consuming search, I accept) possible in elasticsearch?
> > > 
> > > Thanks in advance,  
> > > -- Sujoy.

--

---

<div class="post-metadata">

**Author:** ![simonw\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/simonw_2/32/1130_2.png) [@simonw\_2](https://discuss.elastic.co/u/simonw_2)\
**Post date:** [October 24, 2012, 2:45pm UTC](https://discuss.elastic.co/t/help-with-analyzer-and-mapping/9373/5 "2012-10-24T14:45:07Z")

</div>

hey, the type is set by the tokenizer or token filter. the default type is  
"word". StandardTokenizer might set it to "alphanum", "url", "email" etc.  
other token filters like ShingleFilter set this to "shingle" to indicate  
what this 'token' is. if you want to use standard analyzer but without  
stopwords you can just compose it out of standard tokenizer, & lowercase

simon

On Wednesday, October 24, 2012 12:21:05 PM UTC+2, Sujoy Sett wrote:

> Hi All,
> 
> Can anyone explain what does "type" mean for a token?
> 
> [http://localhost:9200/[index]/\_analyze?pretty=true&analyzer=custom-whitespace-lowercase&text=ipad](http://localhost:9200/%5Bindex%5D/_analyze?pretty=true&analyzer=custom-whitespace-lowercase&text=ipad)  
> gives response
> 
> {  
> "tokens": [  
> {  
> "token": "ipad",  
> "start\_offset": 0,  
> "end\_offset": 4,  
> "type": "word",  
> "position": 1  
> }  
> ]  
> }
> 
> whereas,
> 
> [http://localhost:9200/[index]/\_analyze?pretty=true&analyzer=standard&text=ipad](http://localhost:9200/%5Bindex%5D/_analyze?pretty=true&analyzer=standard&text=ipad)  
> gives response
> 
> {  
> "tokens": [  
> {  
> "token": "ipad",  
> "start\_offset": 0,  
> "end\_offset": 4,  
> "type": "",  
> "position": 1  
> }  
> ]  
> }
> 
> custom-whitespace-lowercase is a custom analyzer defined  
> with whitespace tokenizer and lowercase filter.  
> The purpose of defining this analyzer was to avoid the stop-word filter  
> that comes by default in standard analyzer.
> 
> But this analyzer is creating a different problem by not identifying the  
> term "ipad" while querying.  
> Also, a bit of extra information, I don't know whether relevant or not, my  
> mapping is as follows:  
> properties: {
> 
> - text: {
> - type: multi\_field
> - fields: {
> - text: {
> - type: string  
> }
> 
> - text\_custom\_1: {
> - include\_in\_all: false
> - analyzer: custom-whitespace-lowercase
> - type: string  
> }  
> }  
> }
> 
> Thanks,  
> -- Sujoy.
> 
> On Tuesday, October 16, 2012 7:30:15 PM UTC+5:30, Sujoy Sett wrote:
> 
> > Thanks Tanguy,
> > 
> > I will surely try the multi-field.
> > 
> > -- Sujoy.
> > 
> > On Tuesday, October 16, 2012 3:39:16 PM UTC+5:30, Tanguy wrote:
> > 
> > > Hi,
> > > 
> > > You can use the \_analyze API to understand the logic behind analyzers:
> > > 
> > > [http://localhost:9200/\_analyze?pretty=true&analyzer=standard&text=The+at%26t+company](http://localhost:9200/_analyze?pretty=true&analyzer=standard&text=The+at%26t+company)
> > > 
> > > If you index "The at&t company" with the standard analyzer, the token  
> > > that are really indexed are "t" and "company". The same logic applies when  
> > > searching with match & query\_string queries and that explains the results  
> > > you have.
> > > 
> > > There are many ways to get the expected results when searching for  
> > > "at&t". Some suggestions:
> > > 
> > > - use a custom analyzer for the "text" field in mapping
> > > - declare "text" as multi\_field and search for exact matches
> > > 
> > > Hope this helps,
> > > 
> > > -- Tanguy  
> > > Twitter: @tlrx  
> > > [tlrx (Tanguy Leroux) · GitHub](https://github.com/tlrx)
> > > 
> > > Le mardi 16 octobre 2012 11:53:28 UTC+2, Sujoy Sett a écrit :
> > > 
> > > > Hi,
> > > > 
> > > > In an index I have a field "text", analyzed with _standard analyzer_. I  
> > > > now want to return the documents which has the _keyword  
> > > > "at&t" occurring in the field "text"_.
> > > > 
> > > > However, "at" is probably a member of Stop Token Filter, and "&" is  
> > > > again probably a member of the tokenizer used. (I am don't have much  
> > > > clarity on the exact logic here).  
> > > > I have tried using _match_, _text_, and \*query\_string \*in my query,  
> > > > and all returns quite a lot of junk documents in additional to the required  
> > > > documents.
> > > > 
> > > > I was thinking of custom analyzer here, but I want to use "at" and "&"  
> > > > as is for other search functions on this same field, and a custom analyzer  
> > > > might upset that.
> > > > 
> > > > Am I missing something simpler here? How to search for a text that  
> > > > probably includes stopwords and tokenizer characters?  
> > > > Is something like exact search irrespective of tokens (might be time  
> > > > consuming search, I accept) possible in elasticsearch?
> > > > 
> > > > Thanks in advance,  
> > > > -- Sujoy.

--

---

<div class="post-metadata">

**Author:** ![sujoysett](https://avatars.discourse-cdn.com/v4/letter/s/2acd7d/32.png) [@sujoysett](https://discuss.elastic.co/u/sujoysett)\
**Post date:** [October 24, 2012, 5:58pm UTC](https://discuss.elastic.co/t/help-with-analyzer-and-mapping/9373/6 "2012-10-24T17:58:57Z")

</div>

Thanks Simon.

Probably standard tokenizer + lowercase filter will not server my purpose,  
as I want words like AT&T as a single token, whereas, standard tokenizer  
breaks down text by special characters like '&'.

But that is different issue. What I am concerned with right now is that  
querying for the term 'ipad' on a multifield analyzed with this custom  
analyzer is not fetching me proper results. I am querying by a term query  
on 'ipad' within a boolean must\_not, and I am finding results with term  
'ipad' in it. But going by standard analyzer is fetching results as  
expected. Any hint to the cause of this behavior?

Thanks,  
-- Sujoy.

On Wednesday, October 24, 2012 8:15:07 PM UTC+5:30, simonw wrote:

> hey, the type is set by the tokenizer or token filter. the default type is  
> "word". StandardTokenizer might set it to "alphanum", "url", "email" etc.  
> other token filters like ShingleFilter set this to "shingle" to indicate  
> what this 'token' is. if you want to use standard analyzer but without  
> stopwords you can just compose it out of standard tokenizer, & lowercase
> 
> simon
> 
> On Wednesday, October 24, 2012 12:21:05 PM UTC+2, Sujoy Sett wrote:
> 
> > Hi All,
> > 
> > Can anyone explain what does "type" mean for a token?
> > 
> > [http://localhost:9200/[index]/\_analyze?pretty=true&analyzer=custom-whitespace-lowercase&text=ipad](http://localhost:9200/%5Bindex%5D/_analyze?pretty=true&analyzer=custom-whitespace-lowercase&text=ipad)  
> > gives response
> > 
> > {  
> > "tokens": [  
> > {  
> > "token": "ipad",  
> > "start\_offset": 0,  
> > "end\_offset": 4,  
> > "type": "word",  
> > "position": 1  
> > }  
> > ]  
> > }
> > 
> > whereas,
> > 
> > [http://localhost:9200/[index]/\_analyze?pretty=true&analyzer=standard&text=ipad](http://localhost:9200/%5Bindex%5D/_analyze?pretty=true&analyzer=standard&text=ipad)  
> > gives response
> > 
> > {  
> > "tokens": [  
> > {  
> > "token": "ipad",  
> > "start\_offset": 0,  
> > "end\_offset": 4,  
> > "type": "",  
> > "position": 1  
> > }  
> > ]  
> > }
> > 
> > custom-whitespace-lowercase is a custom analyzer defined  
> > with whitespace tokenizer and lowercase filter.  
> > The purpose of defining this analyzer was to avoid the stop-word filter  
> > that comes by default in standard analyzer.
> > 
> > But this analyzer is creating a different problem by not identifying the  
> > term "ipad" while querying.  
> > Also, a bit of extra information, I don't know whether relevant or not,  
> > my mapping is as follows:  
> > properties: {
> > 
> > - text: {
> > - type: multi\_field
> > - fields: {
> > - text: {
> > - type: string  
> > }
> > 
> > - text\_custom\_1: {
> > - include\_in\_all: false
> > - analyzer: custom-whitespace-lowercase
> > - type: string  
> > }  
> > }  
> > }
> > 
> > Thanks,  
> > -- Sujoy.
> > 
> > On Tuesday, October 16, 2012 7:30:15 PM UTC+5:30, Sujoy Sett wrote:
> > 
> > > Thanks Tanguy,
> > > 
> > > I will surely try the multi-field.
> > > 
> > > -- Sujoy.
> > > 
> > > On Tuesday, October 16, 2012 3:39:16 PM UTC+5:30, Tanguy wrote:
> > > 
> > > > Hi,
> > > > 
> > > > You can use the \_analyze API to understand the logic behind analyzers:
> > > > 
> > > > [http://localhost:9200/\_analyze?pretty=true&analyzer=standard&text=The+at%26t+company](http://localhost:9200/_analyze?pretty=true&analyzer=standard&text=The+at%26t+company)
> > > > 
> > > > If you index "The at&t company" with the standard analyzer, the token  
> > > > that are really indexed are "t" and "company". The same logic applies when  
> > > > searching with match & query\_string queries and that explains the results  
> > > > you have.
> > > > 
> > > > There are many ways to get the expected results when searching for  
> > > > "at&t". Some suggestions:
> > > > 
> > > > - use a custom analyzer for the "text" field in mapping
> > > > - declare "text" as multi\_field and search for exact matches
> > > > 
> > > > Hope this helps,
> > > > 
> > > > -- Tanguy  
> > > > Twitter: @tlrx  
> > > > [tlrx (Tanguy Leroux) · GitHub](https://github.com/tlrx)
> > > > 
> > > > Le mardi 16 octobre 2012 11:53:28 UTC+2, Sujoy Sett a écrit :
> > > > 
> > > > > Hi,
> > > > > 
> > > > > In an index I have a field "text", analyzed with _standard analyzer_. I  
> > > > > now want to return the documents which has the _keyword  
> > > > > "at&t" occurring in the field "text"_.
> > > > > 
> > > > > However, "at" is probably a member of Stop Token Filter, and "&" is  
> > > > > again probably a member of the tokenizer used. (I am don't have much  
> > > > > clarity on the exact logic here).  
> > > > > I have tried using _match_, _text_, and \*query\_string \*in my query,  
> > > > > and all returns quite a lot of junk documents in additional to the required  
> > > > > documents.
> > > > > 
> > > > > I was thinking of custom analyzer here, but I want to use "at" and "&"  
> > > > > as is for other search functions on this same field, and a custom analyzer  
> > > > > might upset that.
> > > > > 
> > > > > Am I missing something simpler here? How to search for a text that  
> > > > > probably includes stopwords and tokenizer characters?  
> > > > > Is something like exact search irrespective of tokens (might be time  
> > > > > consuming search, I accept) possible in elasticsearch?
> > > > > 
> > > > > Thanks in advance,  
> > > > > -- Sujoy.

--

---

<div class="post-metadata">

**Author:** ![simonw\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/simonw_2/32/1130_2.png) [@simonw\_2](https://discuss.elastic.co/u/simonw_2)\
**Post date:** [October 24, 2012, 7:04pm UTC](https://discuss.elastic.co/t/help-with-analyzer-and-mapping/9373/7 "2012-10-24T19:04:24Z")

</div>

On Wednesday, October 24, 2012 7:58:57 PM UTC+2, Sujoy Sett wrote:

> Thanks Simon.
> 
> Probably standard tokenizer + lowercase filter will not server my purpose,  
> as I want words like AT&T as a single token, whereas, standard tokenizer  
> breaks down text by special characters like '&'.
> 
> But that is different issue. What I am concerned with right now is that  
> querying for the term 'ipad' on a multifield analyzed with this custom  
> analyzer is not fetching me proper results. I am querying by a term query  
> on 'ipad' within a boolean must\_not, and I am finding results with term  
> 'ipad' in it. But going by standard analyzer is fetching results as  
> expected. Any hint to the cause of this behavior?

the documents that are returned, do they contain "I Pad" or "ipad" ? I mean  
are you sure the are analyzed correctly?

simon

> Thanks,  
> -- Sujoy.
> 
> On Wednesday, October 24, 2012 8:15:07 PM UTC+5:30, simonw wrote:
> 
> > hey, the type is set by the tokenizer or token filter. the default type  
> > is "word". StandardTokenizer might set it to "alphanum", "url", "email"  
> > etc. other token filters like ShingleFilter set this to "shingle" to  
> > indicate what this 'token' is. if you want to use standard analyzer but  
> > without stopwords you can just compose it out of standard tokenizer, &  
> > lowercase
> > 
> > simon
> > 
> > On Wednesday, October 24, 2012 12:21:05 PM UTC+2, Sujoy Sett wrote:
> > 
> > > Hi All,
> > > 
> > > Can anyone explain what does "type" mean for a token?
> > > 
> > > [http://localhost:9200/[index]/\_analyze?pretty=true&analyzer=custom-whitespace-lowercase&text=ipad](http://localhost:9200/%5Bindex%5D/_analyze?pretty=true&analyzer=custom-whitespace-lowercase&text=ipad)  
> > > gives response
> > > 
> > > {  
> > > "tokens": [  
> > > {  
> > > "token": "ipad",  
> > > "start\_offset": 0,  
> > > "end\_offset": 4,  
> > > "type": "word",  
> > > "position": 1  
> > > }  
> > > ]  
> > > }
> > > 
> > > whereas,
> > > 
> > > [http://localhost:9200/[index]/\_analyze?pretty=true&analyzer=standard&text=ipad](http://localhost:9200/%5Bindex%5D/_analyze?pretty=true&analyzer=standard&text=ipad)  
> > > gives response
> > > 
> > > {  
> > > "tokens": [  
> > > {  
> > > "token": "ipad",  
> > > "start\_offset": 0,  
> > > "end\_offset": 4,  
> > > "type": "",  
> > > "position": 1  
> > > }  
> > > ]  
> > > }
> > > 
> > > custom-whitespace-lowercase is a custom analyzer defined  
> > > with whitespace tokenizer and lowercase filter.  
> > > The purpose of defining this analyzer was to avoid the stop-word filter  
> > > that comes by default in standard analyzer.
> > > 
> > > But this analyzer is creating a different problem by not identifying the  
> > > term "ipad" while querying.  
> > > Also, a bit of extra information, I don't know whether relevant or not,  
> > > my mapping is as follows:  
> > > properties: {
> > > 
> > > - text: {
> > > - type: multi\_field
> > > - fields: {
> > > - text: {
> > > - type: string  
> > > }
> > > 
> > > - text\_custom\_1: {
> > > - include\_in\_all: false
> > > - analyzer: custom-whitespace-lowercase
> > > - type: string  
> > > }  
> > > }  
> > > }
> > > 
> > > Thanks,  
> > > -- Sujoy.
> > > 
> > > On Tuesday, October 16, 2012 7:30:15 PM UTC+5:30, Sujoy Sett wrote:
> > > 
> > > > Thanks Tanguy,
> > > > 
> > > > I will surely try the multi-field.
> > > > 
> > > > -- Sujoy.
> > > > 
> > > > On Tuesday, October 16, 2012 3:39:16 PM UTC+5:30, Tanguy wrote:
> > > > 
> > > > > Hi,
> > > > > 
> > > > > You can use the \_analyze API to understand the logic behind analyzers:
> > > > > 
> > > > > [http://localhost:9200/\_analyze?pretty=true&analyzer=standard&text=The+at%26t+company](http://localhost:9200/_analyze?pretty=true&analyzer=standard&text=The+at%26t+company)
> > > > > 
> > > > > If you index "The at&t company" with the standard analyzer, the token  
> > > > > that are really indexed are "t" and "company". The same logic applies when  
> > > > > searching with match & query\_string queries and that explains the results  
> > > > > you have.
> > > > > 
> > > > > There are many ways to get the expected results when searching for  
> > > > > "at&t". Some suggestions:
> > > > > 
> > > > > - use a custom analyzer for the "text" field in mapping
> > > > > - declare "text" as multi\_field and search for exact matches
> > > > > 
> > > > > Hope this helps,
> > > > > 
> > > > > -- Tanguy  
> > > > > Twitter: @tlrx  
> > > > > [tlrx (Tanguy Leroux) · GitHub](https://github.com/tlrx)
> > > > > 
> > > > > Le mardi 16 octobre 2012 11:53:28 UTC+2, Sujoy Sett a écrit :
> > > > > 
> > > > > > Hi,
> > > > > > 
> > > > > > In an index I have a field "text", analyzed with _standard analyzer_. I  
> > > > > > now want to return the documents which has the _keyword  
> > > > > > "at&t" occurring in the field "text"_.
> > > > > > 
> > > > > > However, "at" is probably a member of Stop Token Filter, and "&" is  
> > > > > > again probably a member of the tokenizer used. (I am don't have much  
> > > > > > clarity on the exact logic here).  
> > > > > > I have tried using _match_, _text_, and \*query\_string \*in my query,  
> > > > > > and all returns quite a lot of junk documents in additional to the required  
> > > > > > documents.
> > > > > > 
> > > > > > I was thinking of custom analyzer here, but I want to use "at" and  
> > > > > > "&" as is for other search functions on this same field, and a custom  
> > > > > > analyzer might upset that.
> > > > > > 
> > > > > > Am I missing something simpler here? How to search for a text that  
> > > > > > probably includes stopwords and tokenizer characters?  
> > > > > > Is something like exact search irrespective of tokens (might be time  
> > > > > > consuming search, I accept) possible in elasticsearch?
> > > > > > 
> > > > > > Thanks in advance,  
> > > > > > -- Sujoy.

--

---

<div class="post-metadata">

**Author:** ![Chris\_Male](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chris_male/32/2607_2.png) [@Chris\_Male](https://discuss.elastic.co/u/Chris_Male)\
**Post date:** [October 25, 2012, 3:31am UTC](https://discuss.elastic.co/t/help-with-analyzer-and-mapping/9373/8 "2012-10-25T03:31:02Z")

</div>

Are you able to provide the query you're using? Just so we can see which  
fields you're querying and what not.

On Thursday, October 25, 2012 6:58:57 AM UTC+13, Sujoy Sett wrote:

> Thanks Simon.
> 
> Probably standard tokenizer + lowercase filter will not server my purpose,  
> as I want words like AT&T as a single token, whereas, standard tokenizer  
> breaks down text by special characters like '&'.
> 
> But that is different issue. What I am concerned with right now is that  
> querying for the term 'ipad' on a multifield analyzed with this custom  
> analyzer is not fetching me proper results. I am querying by a term query  
> on 'ipad' within a boolean must\_not, and I am finding results with term  
> 'ipad' in it. But going by standard analyzer is fetching results as  
> expected. Any hint to the cause of this behavior?
> 
> Thanks,  
> -- Sujoy.
> 
> On Wednesday, October 24, 2012 8:15:07 PM UTC+5:30, simonw wrote:
> 
> > hey, the type is set by the tokenizer or token filter. the default type  
> > is "word". StandardTokenizer might set it to "alphanum", "url", "email"  
> > etc. other token filters like ShingleFilter set this to "shingle" to  
> > indicate what this 'token' is. if you want to use standard analyzer but  
> > without stopwords you can just compose it out of standard tokenizer, &  
> > lowercase
> > 
> > simon
> > 
> > On Wednesday, October 24, 2012 12:21:05 PM UTC+2, Sujoy Sett wrote:
> > 
> > > Hi All,
> > > 
> > > Can anyone explain what does "type" mean for a token?
> > > 
> > > [http://localhost:9200/[index]/\_analyze?pretty=true&analyzer=custom-whitespace-lowercase&text=ipad](http://localhost:9200/%5Bindex%5D/_analyze?pretty=true&analyzer=custom-whitespace-lowercase&text=ipad)  
> > > gives response
> > > 
> > > {  
> > > "tokens": [  
> > > {  
> > > "token": "ipad",  
> > > "start\_offset": 0,  
> > > "end\_offset": 4,  
> > > "type": "word",  
> > > "position": 1  
> > > }  
> > > ]  
> > > }
> > > 
> > > whereas,
> > > 
> > > [http://localhost:9200/[index]/\_analyze?pretty=true&analyzer=standard&text=ipad](http://localhost:9200/%5Bindex%5D/_analyze?pretty=true&analyzer=standard&text=ipad)  
> > > gives response
> > > 
> > > {  
> > > "tokens": [  
> > > {  
> > > "token": "ipad",  
> > > "start\_offset": 0,  
> > > "end\_offset": 4,  
> > > "type": "",  
> > > "position": 1  
> > > }  
> > > ]  
> > > }
> > > 
> > > custom-whitespace-lowercase is a custom analyzer defined  
> > > with whitespace tokenizer and lowercase filter.  
> > > The purpose of defining this analyzer was to avoid the stop-word filter  
> > > that comes by default in standard analyzer.
> > > 
> > > But this analyzer is creating a different problem by not identifying the  
> > > term "ipad" while querying.  
> > > Also, a bit of extra information, I don't know whether relevant or not,  
> > > my mapping is as follows:  
> > > properties: {
> > > 
> > > - text: {
> > > - type: multi\_field
> > > - fields: {
> > > - text: {
> > > - type: string  
> > > }
> > > 
> > > - text\_custom\_1: {
> > > - include\_in\_all: false
> > > - analyzer: custom-whitespace-lowercase
> > > - type: string  
> > > }  
> > > }  
> > > }
> > > 
> > > Thanks,  
> > > -- Sujoy.
> > > 
> > > On Tuesday, October 16, 2012 7:30:15 PM UTC+5:30, Sujoy Sett wrote:
> > > 
> > > > Thanks Tanguy,
> > > > 
> > > > I will surely try the multi-field.
> > > > 
> > > > -- Sujoy.
> > > > 
> > > > On Tuesday, October 16, 2012 3:39:16 PM UTC+5:30, Tanguy wrote:
> > > > 
> > > > > Hi,
> > > > > 
> > > > > You can use the \_analyze API to understand the logic behind analyzers:
> > > > > 
> > > > > [http://localhost:9200/\_analyze?pretty=true&analyzer=standard&text=The+at%26t+company](http://localhost:9200/_analyze?pretty=true&analyzer=standard&text=The+at%26t+company)
> > > > > 
> > > > > If you index "The at&t company" with the standard analyzer, the token  
> > > > > that are really indexed are "t" and "company". The same logic applies when  
> > > > > searching with match & query\_string queries and that explains the results  
> > > > > you have.
> > > > > 
> > > > > There are many ways to get the expected results when searching for  
> > > > > "at&t". Some suggestions:
> > > > > 
> > > > > - use a custom analyzer for the "text" field in mapping
> > > > > - declare "text" as multi\_field and search for exact matches
> > > > > 
> > > > > Hope this helps,
> > > > > 
> > > > > -- Tanguy  
> > > > > Twitter: @tlrx  
> > > > > [tlrx (Tanguy Leroux) · GitHub](https://github.com/tlrx)
> > > > > 
> > > > > Le mardi 16 octobre 2012 11:53:28 UTC+2, Sujoy Sett a écrit :
> > > > > 
> > > > > > Hi,
> > > > > > 
> > > > > > In an index I have a field "text", analyzed with _standard analyzer_. I  
> > > > > > now want to return the documents which has the _keyword  
> > > > > > "at&t" occurring in the field "text"_.
> > > > > > 
> > > > > > However, "at" is probably a member of Stop Token Filter, and "&" is  
> > > > > > again probably a member of the tokenizer used. (I am don't have much  
> > > > > > clarity on the exact logic here).  
> > > > > > I have tried using _match_, _text_, and \*query\_string \*in my query,  
> > > > > > and all returns quite a lot of junk documents in additional to the required  
> > > > > > documents.
> > > > > > 
> > > > > > I was thinking of custom analyzer here, but I want to use "at" and  
> > > > > > "&" as is for other search functions on this same field, and a custom  
> > > > > > analyzer might upset that.
> > > > > > 
> > > > > > Am I missing something simpler here? How to search for a text that  
> > > > > > probably includes stopwords and tokenizer characters?  
> > > > > > Is something like exact search irrespective of tokens (might be time  
> > > > > > consuming search, I accept) possible in elasticsearch?
> > > > > > 
> > > > > > Thanks in advance,  
> > > > > > -- Sujoy.

--

---

<div class="post-metadata">

**Author:** ![sujoysett](https://avatars.discourse-cdn.com/v4/letter/s/2acd7d/32.png) [@sujoysett](https://discuss.elastic.co/u/sujoysett)\
**Post date:** [October 25, 2012, 9:46am UTC](https://discuss.elastic.co/t/help-with-analyzer-and-mapping/9373/9 "2012-10-25T09:46:25Z")

</div>

Hi,

Tried to create a gist with small set of docs, but was not able to recreate  
the problem.  
Apparently, some docs missed the multi-field mapping while indexing and  
were the reason behind faulty responses from the search query being used -

{  
"size": 100,  
"query": {  
"bool": {  
"must\_not": [  
{  
"match": {  
"text.text\_custom\_1": {  
"query": "ipad",  
"type": "phrase"  
}  
}  
}  
]  
}  
}  
}

Applying proper filter with this query removed the faulty docs. It was  
really a silly fault. Thanks very much for all your help.

Thanks  
-- Sujoy.

On Thursday, October 25, 2012 9:01:02 AM UTC+5:30, Chris Male wrote:

> Are you able to provide the query you're using? Just so we can see which  
> fields you're querying and what not.
> 
> On Thursday, October 25, 2012 6:58:57 AM UTC+13, Sujoy Sett wrote:
> 
> > Thanks Simon.
> > 
> > Probably standard tokenizer + lowercase filter will not server my  
> > purpose, as I want words like AT&T as a single token, whereas, standard  
> > tokenizer breaks down text by special characters like '&'.
> > 
> > But that is different issue. What I am concerned with right now is that  
> > querying for the term 'ipad' on a multifield analyzed with this custom  
> > analyzer is not fetching me proper results. I am querying by a term query  
> > on 'ipad' within a boolean must\_not, and I am finding results with term  
> > 'ipad' in it. But going by standard analyzer is fetching results as  
> > expected. Any hint to the cause of this behavior?
> > 
> > Thanks,  
> > -- Sujoy.
> > 
> > On Wednesday, October 24, 2012 8:15:07 PM UTC+5:30, simonw wrote:
> > 
> > > hey, the type is set by the tokenizer or token filter. the default type  
> > > is "word". StandardTokenizer might set it to "alphanum", "url", "email"  
> > > etc. other token filters like ShingleFilter set this to "shingle" to  
> > > indicate what this 'token' is. if you want to use standard analyzer but  
> > > without stopwords you can just compose it out of standard tokenizer, &  
> > > lowercase
> > > 
> > > simon
> > > 
> > > On Wednesday, October 24, 2012 12:21:05 PM UTC+2, Sujoy Sett wrote:
> > > 
> > > > Hi All,
> > > > 
> > > > Can anyone explain what does "type" mean for a token?
> > > > 
> > > > [http://localhost:9200/[index]/\_analyze?pretty=true&analyzer=custom-whitespace-lowercase&text=ipad](http://localhost:9200/%5Bindex%5D/_analyze?pretty=true&analyzer=custom-whitespace-lowercase&text=ipad)  
> > > > gives response
> > > > 
> > > > {  
> > > > "tokens": [  
> > > > {  
> > > > "token": "ipad",  
> > > > "start\_offset": 0,  
> > > > "end\_offset": 4,  
> > > > "type": "word",  
> > > > "position": 1  
> > > > }  
> > > > ]  
> > > > }
> > > > 
> > > > whereas,
> > > > 
> > > > [http://localhost:9200/[index]/\_analyze?pretty=true&analyzer=standard&text=ipad](http://localhost:9200/%5Bindex%5D/_analyze?pretty=true&analyzer=standard&text=ipad)  
> > > > gives response
> > > > 
> > > > {  
> > > > "tokens": [  
> > > > {  
> > > > "token": "ipad",  
> > > > "start\_offset": 0,  
> > > > "end\_offset": 4,  
> > > > "type": "",  
> > > > "position": 1  
> > > > }  
> > > > ]  
> > > > }
> > > > 
> > > > custom-whitespace-lowercase is a custom analyzer defined  
> > > > with whitespace tokenizer and lowercase filter.  
> > > > The purpose of defining this analyzer was to avoid the stop-word filter  
> > > > that comes by default in standard analyzer.
> > > > 
> > > > But this analyzer is creating a different problem by not identifying  
> > > > the term "ipad" while querying.  
> > > > Also, a bit of extra information, I don't know whether relevant or not,  
> > > > my mapping is as follows:  
> > > > properties: {
> > > > 
> > > > - text: {
> > > > - type: multi\_field
> > > > - fields: {
> > > > - text: {
> > > > - type: string  
> > > > }
> > > > 
> > > > - text\_custom\_1: {
> > > > - include\_in\_all: false
> > > > - analyzer: custom-whitespace-lowercase
> > > > - type: string  
> > > > }  
> > > > }  
> > > > }
> > > > 
> > > > Thanks,  
> > > > -- Sujoy.
> > > > 
> > > > On Tuesday, October 16, 2012 7:30:15 PM UTC+5:30, Sujoy Sett wrote:
> > > > 
> > > > > Thanks Tanguy,
> > > > > 
> > > > > I will surely try the multi-field.
> > > > > 
> > > > > -- Sujoy.
> > > > > 
> > > > > On Tuesday, October 16, 2012 3:39:16 PM UTC+5:30, Tanguy wrote:
> > > > > 
> > > > > > Hi,
> > > > > > 
> > > > > > You can use the \_analyze API to understand the logic behind analyzers:
> > > > > > 
> > > > > > [http://localhost:9200/\_analyze?pretty=true&analyzer=standard&text=The+at%26t+company](http://localhost:9200/_analyze?pretty=true&analyzer=standard&text=The+at%26t+company)
> > > > > > 
> > > > > > If you index "The at&t company" with the standard analyzer, the token  
> > > > > > that are really indexed are "t" and "company". The same logic applies when  
> > > > > > searching with match & query\_string queries and that explains the results  
> > > > > > you have.
> > > > > > 
> > > > > > There are many ways to get the expected results when searching for  
> > > > > > "at&t". Some suggestions:
> > > > > > 
> > > > > > - use a custom analyzer for the "text" field in mapping
> > > > > > - declare "text" as multi\_field and search for exact matches
> > > > > > 
> > > > > > Hope this helps,
> > > > > > 
> > > > > > -- Tanguy  
> > > > > > Twitter: @tlrx  
> > > > > > [tlrx (Tanguy Leroux) · GitHub](https://github.com/tlrx)
> > > > > > 
> > > > > > Le mardi 16 octobre 2012 11:53:28 UTC+2, Sujoy Sett a écrit :
> > > > > > 
> > > > > > > Hi,
> > > > > > > 
> > > > > > > In an index I have a field "text", analyzed with _standard analyzer_. I  
> > > > > > > now want to return the documents which has the _keyword  
> > > > > > > "at&t" occurring in the field "text"_.
> > > > > > > 
> > > > > > > However, "at" is probably a member of Stop Token Filter, and "&" is  
> > > > > > > again probably a member of the tokenizer used. (I am don't have much  
> > > > > > > clarity on the exact logic here).  
> > > > > > > I have tried using _match_, _text_, and \*query\_string \*in my query,  
> > > > > > > and all returns quite a lot of junk documents in additional to the required  
> > > > > > > documents.
> > > > > > > 
> > > > > > > I was thinking of custom analyzer here, but I want to use "at" and  
> > > > > > > "&" as is for other search functions on this same field, and a custom  
> > > > > > > analyzer might upset that.
> > > > > > > 
> > > > > > > Am I missing something simpler here? How to search for a text that  
> > > > > > > probably includes stopwords and tokenizer characters?  
> > > > > > > Is something like exact search irrespective of tokens (might be time  
> > > > > > > consuming search, I accept) possible in elasticsearch?
> > > > > > > 
> > > > > > > Thanks in advance,  
> > > > > > > -- Sujoy.

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:07am UTC](https://discuss.elastic.co/t/help-with-analyzer-and-mapping/9373/10 "2017-07-06T03:07:15Z")

</div>


