# Whitespace analyzer (char-filter And token-filter)

**URL:** <https://discuss.elastic.co/t/whitespace-analyzer-char-filter-and-token-filter/205664>\
**Category:** Elasticsearch\
**Created:** [October 29, 2019, 12:33pm UTC](https://discuss.elastic.co/t/whitespace-analyzer-char-filter-and-token-filter/205664 "2019-10-29T12:33:41Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Irakli](https://avatars.discourse-cdn.com/v4/letter/i/ecc23a/32.png) [@Irakli](https://discuss.elastic.co/u/Irakli)\
**Post date:** [October 29, 2019, 12:33pm UTC](https://discuss.elastic.co/t/whitespace-analyzer-char-filter-and-token-filter/205664/1 "2019-10-29T12:33:42Z")

</div>

Hi there,

As I know we have 3 steps on text field when document with text field is indexed in elasticsearch.  
First - Char-filter process,  
Second - Tokenizing process,  
Third - Token filtering process

I have the following mapping in my index:  
"mappings": {  
"properties": {  
"field1": {  
"type": "text",  
"analyzer": "whitespace"  
},  
}  
}

So I am wandering what the Whitespace analyzer exactly does on text field on this exact situation?  
Is there any Char-filter process by default on field1? As I know, if I don't set it, there will not be any char-filter and first step will be tokenizing process by default,Then tokenizing process only split text by space, After that if I don't set token filter, it will be default just lowercase filter. Is it correct or not ?

---

<div class="post-metadata">

**Author:** ![abdon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/abdon/32/9195_2.png) [@abdon](https://discuss.elastic.co/u/abdon)\
**Post date:** [October 29, 2019, 1:00pm UTC](https://discuss.elastic.co/t/whitespace-analyzer-char-filter-and-token-filter/205664/2 "2019-10-29T13:00:58Z")

</div>

Yes, you are exactly right. You can see how the whitespace analyzer has been defined [in the documentation](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-whitespace-analyzer.html#_definition_7). The whitespace analyzer has no character filters, so the first step is the [whitespace tokenizer](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-whitespace-tokenizer.html), which breaks strings on whitespace. Finally, there are no token filters. So that's really all that this analyzer does: it breaks strings on whitespace.

---

<div class="post-metadata">

**Author:** ![Irakli](https://avatars.discourse-cdn.com/v4/letter/i/ecc23a/32.png) [@Irakli](https://discuss.elastic.co/u/Irakli)\
**Post date:** [October 29, 2019, 1:04pm UTC](https://discuss.elastic.co/t/whitespace-analyzer-char-filter-and-token-filter/205664/3 "2019-10-29T13:04:15Z")

</div>

Thanks such a quick answer, But what about token-filter ? isn't there any lowercase token filter after tokenizing process ?

---

<div class="post-metadata">

**Author:** ![abdon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/abdon/32/9195_2.png) [@abdon](https://discuss.elastic.co/u/abdon)\
**Post date:** [October 29, 2019, 1:06pm UTC](https://discuss.elastic.co/t/whitespace-analyzer-char-filter-and-token-filter/205664/4 "2019-10-29T13:06:28Z")

</div>

No, there are no token filters in this analyzer. If you would like to use the whitespace tokenizer in combination with the lowercase token filter, you would have to create a [custom analyzer](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-custom-analyzer.html) that combines these two.

---

<div class="post-metadata">

**Author:** ![Irakli](https://avatars.discourse-cdn.com/v4/letter/i/ecc23a/32.png) [@Irakli](https://discuss.elastic.co/u/Irakli)\
**Post date:** [October 29, 2019, 1:09pm UTC](https://discuss.elastic.co/t/whitespace-analyzer-char-filter-and-token-filter/205664/5 "2019-10-29T13:09:47Z")

</div>

Thank you Abdon

---

<div class="post-metadata">

**Author:** ![Irakli](https://avatars.discourse-cdn.com/v4/letter/i/ecc23a/32.png) [@Irakli](https://discuss.elastic.co/u/Irakli)\
**Post date:** [October 29, 2019, 4:07pm UTC](https://discuss.elastic.co/t/whitespace-analyzer-char-filter-and-token-filter/205664/6 "2019-10-29T16:07:14Z")

</div>

Abdon one more question please, It is index time analyzer in my example above, right ? And at the search time, is it necessary to reference which analyzer could be used with that field above?

---

<div class="post-metadata">

**Author:** ![abdon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/abdon/32/9195_2.png) [@abdon](https://discuss.elastic.co/u/abdon)\
**Post date:** [October 30, 2019, 12:12pm UTC](https://discuss.elastic.co/t/whitespace-analyzer-char-filter-and-token-filter/205664/7 "2019-10-30T12:12:56Z")

</div>

By default, when you query a field, Elasticsearch will apply the analyzer that's defined in the mapping to the query terms. There is no need to specify the analyzer at search time.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 27, 2019, 12:12pm UTC](https://discuss.elastic.co/t/whitespace-analyzer-char-filter-and-token-filter/205664/8 "2019-11-27T12:12:59Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
