# Whitespace tokenizer not working as I'd expect

**URL:** <https://discuss.elastic.co/t/whitespace-tokenizer-not-working-as-id-expect/22634>\
**Category:** Elasticsearch\
**Created:** [March 12, 2015, 2:41pm UTC](https://discuss.elastic.co/t/whitespace-tokenizer-not-working-as-id-expect/22634 "2015-03-12T14:41:19Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![cching](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cching/32/44800_2.png) [@cching](https://discuss.elastic.co/u/cching)\
**Post date:** [March 12, 2015, 2:41pm UTC](https://discuss.elastic.co/t/whitespace-tokenizer-not-working-as-id-expect/22634/1 "2015-03-12T14:41:19Z")

</div>

Hi all,

I'm trying to break up some strings to use in a full text search leaving  
the original field intact. I have created a "full\_text" field that is  
populated from a "name" field using "copy\_to" and an analyzer that looks  
like this:

```
"settings" : {
    "analysis": {
        "char_filter" : {
            "full_text_mapping" : {
                "type": "mapping",
                "mappings" : [".=>%20", "_=>%20"]
            }
        },
        "analyzer" : {
            "full_text_analyzer" : {
                "type" : "custom",
                "char_filter" : "full_text_mapping",
                "tokenizer" : "whitespace",
                "filter" : ["lowercase"]
            }
        }
    }
},

```

As you can see I'm trying to convert '.' and '\_' to ' ' before the  
whitespace tokenizer kicks in. It's my understanding that the char\_filter  
will replace those characters with whitespace that the whitespace tokenizer  
would then tokenize and then all components could be searchable. For  
instance, I would expect "GRIZZLY.BEAR" to be found using both "grizzly"  
and "bear". But with the whitespace tokenizer I am not able to find the  
document with either term. So what am I not understanding? Full script  
showing what I'm doing:

#!/bin/sh

ES=localhost:9200

echo "\>\>\> Deleting \_all"  
curl -XDELETE $ES/\_all

echo "\>\>\> Creating the index 'animals'"  
curl -XPUT $ES/animals -d'  
{  
"settings" : {  
"analysis": {  
"char\_filter" : {  
"full\_text\_mapping" : {  
"type": "mapping",  
"mappings" : [".=\>%20", "\_=\>%20"]  
}  
},  
"analyzer" : {  
"full\_text\_analyzer" : {  
"type" : "custom",  
"char\_filter" : "full\_text\_mapping",  
"tokenizer" : "whitespace",  
"filter" : ["lowercase"]  
}  
}  
}  
},  
"mappings" : {  
"bear" : {  
"properties" : {  
"suggest" : {  
"type" : "completion",  
"analyzer" : "simple",  
"payloads" : true  
},  
"full\_text" : {  
"type" : "string",  
"analyzer" : "full\_text\_analyzer"  
},  
"name" : {  
"type" : "string",  
"index" : "not\_analyzed",  
"copy\_to" : "full\_text"  
}  
}  
}  
}  
}' && echo

echo "\>\>\> Indexing the GRIZZLY.BEAR document"  
curl -XPOST $ES/animals/bear -d'  
{  
"name": "GRIZZLY.BEAR"  
}  
' && echo

curl -XPOST $ES/animals/\_flush && echo

# Search for the document using the name

echo  
echo "\>\>\> Searching for name:GRIZZLY.BEAR"  
echo  
curl $ES/animals/bear/\_search -d'  
{  
"query" : {  
"match" : {  
"name" : "GRIZZLY.BEAR"  
}  
}  
}  
' && echo

# Search for the document using a general term

echo  
echo "\>\>\> Searching for full\_text:grizzly"  
echo  
curl $ES/animals/bear/\_search -d'  
{  
"query" : {  
"match" : {  
"full\_text" : "grizzly"  
}  
}  
}  
' && echo

# Search for the document using a general term

echo  
echo "\>\>\> Searching for full\_text:bear"  
echo  
curl $ES/animals/bear/\_search -d'  
{  
"query" : {  
"match" : {  
"full\_text" : "bear"  
}  
}  
}  
' && echo

I appreciate any help with this!

Cheers,  
Craig

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/5fa2347f-3019-4973-9d67-7f18b3dfee9e%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/5fa2347f-3019-4973-9d67-7f18b3dfee9e%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [March 13, 2015, 9:47am UTC](https://discuss.elastic.co/t/whitespace-tokenizer-not-working-as-id-expect/22634/2 "2015-03-13T09:47:21Z")

</div>

From which source did you assume that %20 is a white space?

The mapping char filter understands \uXXXX notation (which is not  
documented in ES).

With curl, on bash, you have to escape the \u notation with double  
backslash like this

". =\> \u0020"

Here is a working example

> <https://gist.github.com/jprante/d7c839d47e9a9dc78311>

Jörg

On Thu, Mar 12, 2015 at 3:41 PM, Craig Ching [craigching@gmail.com](mailto:craigching@gmail.com) wrote:

> Hi all,
> 
> I'm trying to break up some strings to use in a full text search leaving  
> the original field intact. I have created a "full\_text" field that is  
> populated from a "name" field using "copy\_to" and an analyzer that looks  
> like this:
> 
> ```
> "settings" : {
> "analysis": {
> "char_filter" : {
> "full_text_mapping" : {
> "type": "mapping",
> "mappings" : [".=>%20", "_=>%20"]
> }
> },
> "analyzer" : {
> "full_text_analyzer" : {
> "type" : "custom",
> "char_filter" : "full_text_mapping",
> "tokenizer" : "whitespace",
> "filter" : ["lowercase"]
> }
> }
> }
> },
> 
> ```
> 
> As you can see I'm trying to convert '.' and '\_' to ' ' before the  
> whitespace tokenizer kicks in. It's my understanding that the char\_filter  
> will replace those characters with whitespace that the whitespace tokenizer  
> would then tokenize and then all components could be searchable. For  
> instance, I would expect "GRIZZLY.BEAR" to be found using both "grizzly"  
> and "bear". But with the whitespace tokenizer I am not able to find the  
> document with either term. So what am I not understanding? Full script  
> showing what I'm doing:
> 
> #!/bin/sh
> 
> ES=localhost:9200
> 
> echo "\>\>\> Deleting \_all"  
> curl -XDELETE $ES/\_all
> 
> echo "\>\>\> Creating the index 'animals'"  
> curl -XPUT $ES/animals -d'  
> {  
> "settings" : {  
> "analysis": {  
> "char\_filter" : {  
> "full\_text\_mapping" : {  
> "type": "mapping",  
> "mappings" : [".=\>%20", "\_=\>%20"]  
> }  
> },  
> "analyzer" : {  
> "full\_text\_analyzer" : {  
> "type" : "custom",  
> "char\_filter" : "full\_text\_mapping",  
> "tokenizer" : "whitespace",  
> "filter" : ["lowercase"]  
> }  
> }  
> }  
> },  
> "mappings" : {  
> "bear" : {  
> "properties" : {  
> "suggest" : {  
> "type" : "completion",  
> "analyzer" : "simple",  
> "payloads" : true  
> },  
> "full\_text" : {  
> "type" : "string",  
> "analyzer" : "full\_text\_analyzer"  
> },  
> "name" : {  
> "type" : "string",  
> "index" : "not\_analyzed",  
> "copy\_to" : "full\_text"  
> }  
> }  
> }  
> }  
> }' && echo
> 
> echo "\>\>\> Indexing the GRIZZLY.BEAR document"  
> curl -XPOST $ES/animals/bear -d'  
> {  
> "name": "GRIZZLY.BEAR"  
> }  
> ' && echo
> 
> curl -XPOST $ES/animals/\_flush && echo
> 
> # Search for the document using the name
> 
> echo  
> echo "\>\>\> Searching for name:GRIZZLY.BEAR"  
> echo  
> curl $ES/animals/bear/\_search -d'  
> {  
> "query" : {  
> "match" : {  
> "name" : "GRIZZLY.BEAR"  
> }  
> }  
> }  
> ' && echo
> 
> # Search for the document using a general term
> 
> echo  
> echo "\>\>\> Searching for full\_text:grizzly"  
> echo  
> curl $ES/animals/bear/\_search -d'  
> {  
> "query" : {  
> "match" : {  
> "full\_text" : "grizzly"  
> }  
> }  
> }  
> ' && echo
> 
> # Search for the document using a general term
> 
> echo  
> echo "\>\>\> Searching for full\_text:bear"  
> echo  
> curl $ES/animals/bear/\_search -d'  
> {  
> "query" : {  
> "match" : {  
> "full\_text" : "bear"  
> }  
> }  
> }  
> ' && echo
> 
> I appreciate any help with this!
> 
> Cheers,  
> Craig
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/5fa2347f-3019-4973-9d67-7f18b3dfee9e%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/5fa2347f-3019-4973-9d67-7f18b3dfee9e%40googlegroups.com)  
> [https://groups.google.com/d/msgid/elasticsearch/5fa2347f-3019-4973-9d67-7f18b3dfee9e%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/5fa2347f-3019-4973-9d67-7f18b3dfee9e%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoHgoBCQjHMgWUVHDrWG%3DmD8SiCo52%3DVQSaLzt%3D-V%3DTe%2BA%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoHgoBCQjHMgWUVHDrWG%3DmD8SiCo52%3DVQSaLzt%3D-V%3DTe%2BA%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![cching](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cching/32/44800_2.png) [@cching](https://discuss.elastic.co/u/cching)\
**Post date:** [March 16, 2015, 1:41pm UTC](https://discuss.elastic.co/t/whitespace-tokenizer-not-working-as-id-expect/22634/3 "2015-03-16T13:41:43Z")

</div>

On Friday, March 13, 2015 at 4:47:31 AM UTC-5, Jörg Prante wrote:

> From which source did you assume that %20 is a white space?

It was just a guess since, as you say, it's not documented 😉 After using  
%20, it _did_ appear to tokenize differently, though I couldn't figure out  
how to prove that it had worked and I just assumed it did I guess.

> The mapping char filter understands \uXXXX notation (which is not  
> documented in ES).
> 
> With curl, on bash, you have to escape the \u notation with double  
> backslash like this
> 
> ". =\> \u0020"
> 
> Here is a working example
> 
> [Char filter demo · GitHub](https://gist.github.com/jprante/d7c839d47e9a9dc78311)

Awesome, thanks very much for that, works like a charm!

> Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/d72e8b10-e025-429d-8edf-1a1ab3776fbd%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/d72e8b10-e025-429d-8edf-1a1ab3776fbd%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:26am UTC](https://discuss.elastic.co/t/whitespace-tokenizer-not-working-as-id-expect/22634/4 "2017-07-06T00:26:34Z")

</div>


