# Whitespace analyzer

**URL:** <https://discuss.elastic.co/t/whitespace-analyzer/11860>\
**Category:** Elasticsearch\
**Created:** [May 7, 2013, 6:33pm UTC](https://discuss.elastic.co/t/whitespace-analyzer/11860 "2013-05-07T18:33:15Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mohit\_Anchlia](https://avatars.discourse-cdn.com/v4/letter/m/848f3c/32.png) [@Mohit\_Anchlia](https://discuss.elastic.co/u/Mohit_Anchlia)\
**Post date:** [May 7, 2013, 6:33pm UTC](https://discuss.elastic.co/t/whitespace-analyzer/11860/1 "2013-05-07T18:33:15Z")

</div>

I am using whitespace analyzer and I have a text of format "p1-\>p2-\>p3". My  
assumption is that when using whitespace analyzer this text will not be  
broken down into terms p1,p2 and p3. Is that correct?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [May 7, 2013, 6:37pm UTC](https://discuss.elastic.co/t/whitespace-analyzer/11860/2 "2013-05-07T18:37:40Z")

</div>

Hey,

you are right, as a whitespace analyzer splits by whitespace, which is not  
included here. It is easy for you to verify (you can try the standard  
analyzer to get a different behaviour):

curl -X POST 'localhost:9200/\_analyze?analyzer=whitespace&pretty' -d  
'p1-\>p2-\>p3'

{  
"tokens" : [ {  
"token" : "p1-\>p2-\>p3",  
"start\_offset" : 0,  
"end\_offset" : 10,  
"type" : "word",  
"position" : 1  
} ]  
}

As you can se with the analyze API, your input remains one token..

--Alex

On Tue, May 7, 2013 at 8:33 PM, Mohit Anchlia [mohitanchlia@gmail.com](mailto:mohitanchlia@gmail.com)wrote:

> I am using whitespace analyzer and I have a text of format "p1-\>p2-\>p3".  
> My assumption is that when using whitespace analyzer this text will not be  
> broken down into terms p1,p2 and p3. Is that correct?
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Mohit\_Anchlia](https://avatars.discourse-cdn.com/v4/letter/m/848f3c/32.png) [@Mohit\_Anchlia](https://discuss.elastic.co/u/Mohit_Anchlia)\
**Post date:** [May 7, 2013, 6:41pm UTC](https://discuss.elastic.co/t/whitespace-analyzer/11860/3 "2013-05-07T18:41:16Z")

</div>

Thanks! Is it possible to run \_analyze on top of existing index? I had an  
index with standard tokenizer which I converted to whitespace but it  
doesn't seem to be working as expected:

curl -XPOST 'localhost:9200/pflow1/\_close'

curl -XPUT 'localhost:9200/pflow1/\_settings' -d '{  
"analysis" : {  
"analyzer":{  
"content":{  
"type":"whitespace",  
"tokenizer":"whitespace"  
}  
}  
}  
}'

curl -XPOST 'localhost:9200/pflow1/\_open'

On Tue, May 7, 2013 at 11:37 AM, Alexander Reelsen [alr@spinscale.de](mailto:alr@spinscale.de) wrote:

> Hey,
> 
> you are right, as a whitespace analyzer splits by whitespace, which is not  
> included here. It is easy for you to verify (you can try the standard  
> analyzer to get a different behaviour):
> 
> curl -X POST 'localhost:9200/\_analyze?analyzer=whitespace&pretty' -d  
> 'p1-\>p2-\>p3'
> 
> {  
> "tokens" : [ {  
> "token" : "p1-\>p2-\>p3",  
> "start\_offset" : 0,  
> "end\_offset" : 10,  
> "type" : "word",  
> "position" : 1  
> } ]  
> }
> 
> As you can se with the analyze API, your input remains one token..
> 
> --Alex
> 
> On Tue, May 7, 2013 at 8:33 PM, Mohit Anchlia [mohitanchlia@gmail.com](mailto:mohitanchlia@gmail.com)wrote:
> 
> > I am using whitespace analyzer and I have a text of format "p1-\>p2-\>p3".  
> > My assumption is that when using whitespace analyzer this text will not be  
> > broken down into terms p1,p2 and p3. Is that correct?
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [May 7, 2013, 8:51pm UTC](https://discuss.elastic.co/t/whitespace-analyzer/11860/4 "2013-05-07T20:51:30Z")

</div>

Hey,

you can use any analyzer you want by specifying it, see

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

--Alex

On Tue, May 7, 2013 at 8:41 PM, Mohit Anchlia [mohitanchlia@gmail.com](mailto:mohitanchlia@gmail.com)wrote:

> Thanks! Is it possible to run \_analyze on top of existing index? I had an  
> index with standard tokenizer which I converted to whitespace but it  
> doesn't seem to be working as expected:
> 
> curl -XPOST 'localhost:9200/pflow1/\_close'
> 
> curl -XPUT 'localhost:9200/pflow1/\_settings' -d '{  
> "analysis" : {  
> "analyzer":{  
> "content":{  
> "type":"whitespace",  
> "tokenizer":"whitespace"  
> }  
> }  
> }  
> }'
> 
> curl -XPOST 'localhost:9200/pflow1/\_open'
> 
> On Tue, May 7, 2013 at 11:37 AM, Alexander Reelsen [alr@spinscale.de](mailto:alr@spinscale.de)wrote:
> 
> > Hey,
> > 
> > you are right, as a whitespace analyzer splits by whitespace, which is  
> > not included here. It is easy for you to verify (you can try the standard  
> > analyzer to get a different behaviour):
> > 
> > curl -X POST 'localhost:9200/\_analyze?analyzer=whitespace&pretty' -d  
> > 'p1-\>p2-\>p3'
> > 
> > {  
> > "tokens" : [ {  
> > "token" : "p1-\>p2-\>p3",  
> > "start\_offset" : 0,  
> > "end\_offset" : 10,  
> > "type" : "word",  
> > "position" : 1  
> > } ]  
> > }
> > 
> > As you can se with the analyze API, your input remains one token..
> > 
> > --Alex
> > 
> > On Tue, May 7, 2013 at 8:33 PM, Mohit Anchlia [mohitanchlia@gmail.com](mailto:mohitanchlia@gmail.com)wrote:
> > 
> > > I am using whitespace analyzer and I have a text of format "p1-\>p2-\>p3".  
> > > My assumption is that when using whitespace analyzer this text will not be  
> > > broken down into terms p1,p2 and p3. Is that correct?
> > > 
> > > --  
> > > You received this message because you are subscribed to the Google  
> > > Groups "elasticsearch" group.  
> > > To unsubscribe from this group and stop receiving emails from it, send  
> > > an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:37am UTC](https://discuss.elastic.co/t/whitespace-analyzer/11860/5 "2017-07-06T02:37:49Z")

</div>


