# Differences with standard analyzer in 16.2 vs 14.2

**URL:** <https://discuss.elastic.co/t/differences-with-standard-analyzer-in-16-2-vs-14-2/4686>\
**Category:** Elasticsearch\
**Created:** [June 23, 2011, 12:05am UTC](https://discuss.elastic.co/t/differences-with-standard-analyzer-in-16-2-vs-14-2/4686 "2011-06-23T00:05:15Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![ppearcy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ppearcy/32/980_2.png) [@ppearcy](https://discuss.elastic.co/u/ppearcy)\
**Post date:** [June 23, 2011, 12:05am UTC](https://discuss.elastic.co/t/differences-with-standard-analyzer-in-16-2-vs-14-2/4686/1 "2011-06-23T00:05:15Z")

</div>

Hey,  
Probably documented in the release notes, but I wanted to point out  
a change with the standard analyzer between 14.2 and 16.2 that doesn't  
give great results. This is definitely an edge case and I have ~8  
other examples where the results are better, so a win overall.

The old behavior of standard analyzer correctly recognizes AT&T:

curl -XGET 'localhost:9200/index19/\_analyze?analyzer=standard' -d  
'AT&T'  
{"tokens":[{"token":"at&t","start\_offset":0,"end\_offset":  
4,"type":"","position":1}]}

The new behavior ends up removing AT as a stop word:

curl -XGET 'localhost:9200/index19/\_analyze?analyzer=standard' -d  
'AT&T'  
{"tokens":[{"token":"t","start\_offset":3,"end\_offset":  
4,"type":"","position":2}]}

Just wanted to point this out.

Thanks!  
Paul

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [June 23, 2011, 3:53am UTC](https://discuss.elastic.co/t/differences-with-standard-analyzer-in-16-2-vs-14-2/4686/2 "2011-06-23T03:53:47Z")

</div>

Yes, standard tokenizer was changed in Lucene 3.1 (  
[[LUCENE-2167] Implement StandardTokenizer with the UAX#29 Standard - ASF JIRA](https://issues.apache.org/jira/browse/LUCENE-2167)). I think you can bring  
old behavior back by specifying analyzer version in config file:

index.analysis.analyzer.standard.type: standard  
index.analysis.analyzer.standard.version: 3.0

On Wed, Jun 22, 2011 at 8:05 PM, Paul [ppearcy@gmail.com](mailto:ppearcy@gmail.com) wrote:

> Hey,  
> Probably documented in the release notes, but I wanted to point out  
> a change with the standard analyzer between 14.2 and 16.2 that doesn't  
> give great results. This is definitely an edge case and I have ~8  
> other examples where the results are better, so a win overall.
> 
> The old behavior of standard analyzer correctly recognizes AT&T:
> 
> curl -XGET 'localhost:9200/index19/\_analyze?analyzer=standard' -d  
> 'AT&T'  
> {"tokens":[{"token":"at&t","start\_offset":0,"end\_offset":  
> 4,"type":"","position":1}]}
> 
> The new behavior ends up removing AT as a stop word:
> 
> curl -XGET 'localhost:9200/index19/\_analyze?analyzer=standard' -d  
> 'AT&T'  
> {"tokens":[{"token":"t","start\_offset":3,"end\_offset":  
> 4,"type":"","position":2}]}
> 
> Just wanted to point this out.
> 
> Thanks!  
> Paul

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:02am UTC](https://discuss.elastic.co/t/differences-with-standard-analyzer-in-16-2-vs-14-2/4686/3 "2017-07-06T04:02:54Z")

</div>


