# Difference in analyzer between 1.3.4 and 0.20.2

**URL:** <https://discuss.elastic.co/t/difference-in-analyzer-between-1-3-4-and-0-20-2/20605>\
**Category:** Elasticsearch\
**Created:** [November 6, 2014, 12:31pm UTC](https://discuss.elastic.co/t/difference-in-analyzer-between-1-3-4-and-0-20-2/20605 "2014-11-06T12:31:22Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ben\_George](https://avatars.discourse-cdn.com/v4/letter/b/58956e/32.png) [@Ben\_George](https://discuss.elastic.co/u/Ben_George)\
**Post date:** [November 6, 2014, 12:31pm UTC](https://discuss.elastic.co/t/difference-in-analyzer-between-1-3-4-and-0-20-2/20605/1 "2014-11-06T12:31:22Z")

</div>

I am in process of upgrading ES from 0.20.2 to 1.3.4. Below are two  
requests to test an analyzer / filter, and although the mapping files are  
semantically the same the results are slightly different.

Can anyone provide some insight as to why the differ (the start\_offest,  
end\_offset and position) ? Also does it matter ? The reason I noticed  
this is because I'm trying to debug some unexpected behaviour with a query  
where the result set for "a" are same for "aa" or even "axxxxxxxxxx".

The filter config is:

```
            "filter_edge_ngram_front": {
                "type": "edgeNGram",
                "max_gram": "20",
                "min_gram": "1",
                "side": "front"
            }

```

v.20.2/\_analyze?text=aa+b&filters=filter\_edge\_ngram\_front&tokenizer=standard

{

- 
## tokens: [

## { - token: "a", - start\_offset: 0, - end\_offset: 1, - type: "word", - position: 1 },

## { - token: "aa", - start\_offset: 0, - end\_offset: 2, - type: "word", - position: 2 },
{  
- token: "b",  
- start\_offset: 3,  
- end\_offset: 4,  
- type: "word",  
- position: 3  
}  
]

}

v1.3.4/\_analyze?text=aa+b&filters=filter\_edge\_ngram\_front&tokenizer=standard

{

- 
## tokens: [

## { - token: "a", - start\_offset: 0, - end\_offset: 2, - type: "word", - position: 1 },

## { - token: "aa", - start\_offset: 0, - end\_offset: 2, - type: "word", - position: 1 },
{  
- token: "b",  
- start\_offset: 3,  
- end\_offset: 4,  
- type: "word",  
- position: 2  
}  
]

}

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/2e58239a-0091-4d8b-872a-e5b5414b72ad%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/2e58239a-0091-4d8b-872a-e5b5414b72ad%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![simonw\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/simonw_2/32/1130_2.png) [@simonw\_2](https://discuss.elastic.co/u/simonw_2)\
**Post date:** [November 6, 2014, 7:18pm UTC](https://discuss.elastic.co/t/difference-in-analyzer-between-1-3-4-and-0-20-2/20605/2 "2014-11-06T19:18:43Z")

</div>

We fixed EdgeNGram tokenizer / filter in the 1.x series but don't ask me  
when exactly I think it was lucene 4.4 or so. Those offsets are now correct  
while they where broken before.  
not sure if this helps you to debug your problem

On Thursday, November 6, 2014 1:31:22 PM UTC+1, Ben George wrote:

> I am in process of upgrading ES from 0.20.2 to 1.3.4. Below are two  
> requests to test an analyzer / filter, and although the mapping files are  
> semantically the same the results are slightly different.
> 
> Can anyone provide some insight as to why the differ (the start\_offest,  
> end\_offset and position) ? Also does it matter ? The reason I noticed  
> this is because I'm trying to debug some unexpected behaviour with a query  
> where the result set for "a" are same for "aa" or even "axxxxxxxxxx".
> 
> The filter config is:
> 
> ```
> "filter_edge_ngram_front": {
> "type": "edgeNGram",
> "max_gram": "20",
> "min_gram": "1",
> "side": "front"
> }
> 
> ```
> 
> v.20.2/\_analyze?text=aa+b&filters=filter\_edge\_ngram\_front&tokenizer=standard
> 
> {
> 
> - 
> ## tokens: [
> 
> ## { - token: "a", - start\_offset: 0, - end\_offset: 1, - type: "word", - position: 1 },
> 
> ## { - token: "aa", - start\_offset: 0, - end\_offset: 2, - type: "word", - position: 2 },
> {  
> - token: "b",  
> - start\_offset: 3,  
> - end\_offset: 4,  
> - type: "word",  
> - position: 3  
> }  
> ]
> 
> }
> 
> v1.3.4/\_analyze?text=aa+b&filters=filter\_edge\_ngram\_front&tokenizer=standard
> 
> {
> 
> - 
> ## tokens: [
> 
> ## { - token: "a", - start\_offset: 0, - end\_offset: 2, - type: "word", - position: 1 },
> 
> ## { - token: "aa", - start\_offset: 0, - end\_offset: 2, - type: "word", - position: 1 },
> {  
> - token: "b",  
> - start\_offset: 3,  
> - end\_offset: 4,  
> - type: "word",  
> - position: 2  
> }  
> ]
> 
> }

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/ab062e84-b429-40d7-bb8b-bb94e9ec9316%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ab062e84-b429-40d7-bb8b-bb94e9ec9316%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:51am UTC](https://discuss.elastic.co/t/difference-in-analyzer-between-1-3-4-and-0-20-2/20605/3 "2017-07-06T00:51:40Z")

</div>


