# Custom analyzer without a tokenizer

**URL:** <https://discuss.elastic.co/t/custom-analyzer-without-a-tokenizer/22496>\
**Category:** Elasticsearch\
**Created:** [March 3, 2015, 6:02pm UTC](https://discuss.elastic.co/t/custom-analyzer-without-a-tokenizer/22496 "2015-03-03T18:02:30Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Sagar\_Shah](https://avatars.discourse-cdn.com/v4/letter/s/a9a28c/32.png) [@Sagar\_Shah](https://discuss.elastic.co/u/Sagar_Shah)\
**Post date:** [March 3, 2015, 6:02pm UTC](https://discuss.elastic.co/t/custom-analyzer-without-a-tokenizer/22496/1 "2015-03-03T18:02:30Z")

</div>

Hello everyone,  
I am working on a defining a mapping in elastic search, which can have few  
fields on the fly. I can define the types & index using dynamic templates,  
but I would like to know the difference between following two and which one  
is preferred over the other.  
I do not want to break down the string into tokens but use it in single  
complete string

Option 1. Field Index : not\_analyzed  
Option 2. Field Index: customer analyzer with no tokenizer

Are there any performance differences for the above two approaches?  
How does the 2nd option works (A custom analyzer with no tokenizer)? and  
how can I create mapping for the same?

Appreciate your inputs!

Regards,  
Sagar Shah

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/fadbe843-aaba-4e9e-9f7f-a4d20b0a4569%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/fadbe843-aaba-4e9e-9f7f-a4d20b0a4569%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [March 3, 2015, 6:36pm UTC](https://discuss.elastic.co/t/custom-analyzer-without-a-tokenizer/22496/2 "2015-03-03T18:36:42Z")

</div>

On Tue, Mar 3, 2015 at 1:02 PM, Sagar Shah [sagarshah1983@gmail.com](mailto:sagarshah1983@gmail.com) wrote:

Hello everyone,

> I am working on a defining a mapping in Elasticsearch, which can have few  
> fields on the fly. I can define the types & index using dynamic templates,  
> but I would like to know the difference between following two and which one  
> is preferred over the other.  
> I do not want to break down the string into tokens but use it in single  
> complete string
> 
> Option 1. Field Index : not\_analyzed  
> Option 2. Field Index: customer analyzer with no tokenizer
> 
> Are there any performance differences for the above two approaches?  
> How does the 2nd option works (A custom analyzer with no tokenizer)? and  
> how can I create mapping for the same?

I believe if you leave the tokenizer out you get the StandardTokenizer.  
Its very different from not\_analyzed. not\_analyzed is like "don't break  
this up - don't change it at all - I'm going to search for it _exactly_ how  
I send it to you". I believe the standard tokenizer does icu word  
segmentation: [UAX #29: Unicode Text Segmentation](http://unicode.org/reports/tr29/) .

Nik

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAPmjWd3fx1hhX7RE9w\_jK1YZK47URUiaUjK%3D2yXzyPQGPMd4BA%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAPmjWd3fx1hhX7RE9w_jK1YZK47URUiaUjK%3D2yXzyPQGPMd4BA%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Sagar\_Shah](https://avatars.discourse-cdn.com/v4/letter/s/a9a28c/32.png) [@Sagar\_Shah](https://discuss.elastic.co/u/Sagar_Shah)\
**Post date:** [March 4, 2015, 12:43am UTC](https://discuss.elastic.co/t/custom-analyzer-without-a-tokenizer/22496/3 "2015-03-04T00:43:48Z")

</div>

Thanks Nikolas for clarification.  
Do you think there would be any difference in performance between the two,  
given that I would always be searching full text or matching phrase using  
match\_phrase for the partial matching?

Please advise.

On Tue, Mar 3, 2015 at 12:36 PM, Nikolas Everett [nik9000@gmail.com](mailto:nik9000@gmail.com) wrote:

> On Tue, Mar 3, 2015 at 1:02 PM, Sagar Shah [sagarshah1983@gmail.com](mailto:sagarshah1983@gmail.com)  
> wrote:
> 
> Hello everyone,
> 
> > I am working on a defining a mapping in Elasticsearch, which can have  
> > few fields on the fly. I can define the types & index using dynamic  
> > templates, but I would like to know the difference between following two  
> > and which one is preferred over the other.  
> > I do not want to break down the string into tokens but use it in single  
> > complete string
> > 
> > Option 1. Field Index : not\_analyzed  
> > Option 2. Field Index: customer analyzer with no tokenizer
> > 
> > Are there any performance differences for the above two approaches?  
> > How does the 2nd option works (A custom analyzer with no tokenizer)? and  
> > how can I create mapping for the same?
> 
> I believe if you leave the tokenizer out you get the StandardTokenizer.  
> Its very different from not\_analyzed. not\_analyzed is like "don't break  
> this up - don't change it at all - I'm going to search for it _exactly_ how  
> I send it to you". I believe the standard tokenizer does icu word  
> segmentation: [UAX #29: Unicode Text Segmentation](http://unicode.org/reports/tr29/) .
> 
> Nik
> 
> --  
> You received this message because you are subscribed to a topic in the  
> Google Groups "elasticsearch" group.  
> To unsubscribe from this topic, visit  
> [https://groups.google.com/d/topic/elasticsearch/fvbTw4j3\_iw/unsubscribe](https://groups.google.com/d/topic/elasticsearch/fvbTw4j3_iw/unsubscribe).  
> To unsubscribe from this group and all its topics, send an email to  
> [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/CAPmjWd3fx1hhX7RE9w\_jK1YZK47URUiaUjK%3D2yXzyPQGPMd4BA%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAPmjWd3fx1hhX7RE9w_jK1YZK47URUiaUjK%3D2yXzyPQGPMd4BA%40mail.gmail.com)  
> [https://groups.google.com/d/msgid/elasticsearch/CAPmjWd3fx1hhX7RE9w\_jK1YZK47URUiaUjK%3D2yXzyPQGPMd4BA%40mail.gmail.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/CAPmjWd3fx1hhX7RE9w_jK1YZK47URUiaUjK%3D2yXzyPQGPMd4BA%40mail.gmail.com?utm_medium=email&utm_source=footer)  
> .
> 
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--

Regards,  
Sagar Shah

Too many people think more of security instead of opportunity. They seem  
more afraid of life than death!!!

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CANZnv5hFf\_eXc%3DG3TqP-0dUV9%3DDEATnh7hBoHYG2W2VcUTxUcQ%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CANZnv5hFf_eXc%3DG3TqP-0dUV9%3DDEATnh7hBoHYG2W2VcUTxUcQ%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:28am UTC](https://discuss.elastic.co/t/custom-analyzer-without-a-tokenizer/22496/4 "2017-07-06T00:28:43Z")

</div>


