# My custom analyzer is registered but not used during indexing

**URL:** https://discuss.elastic.co/t/my-custom-analyzer-is-registered-but-not-used-during-indexing/19362
**Category:** Elasticsearch
**Created:** [August 20, 2014, 10:50am UTC](https://discuss.elastic.co/t/my-custom-analyzer-is-registered-but-not-used-during-indexing/19362 "2014-08-20T10:50:10Z")
**Posts on this page:** 2
**Page:** 1

<div class="post-metadata">

### Author: ![Frederic\_Esnault](https://avatars.discourse-cdn.com/v4/letter/f/4af34b/32.png) [@Frederic\_Esnault](https://discuss.elastic.co/u/Frederic_Esnault)
#### Post date: [August 20, 2014, 10:50am UTC](https://discuss.elastic.co/t/my-custom-analyzer-is-registered-but-not-used-during-indexing/19362/1 "2014-08-20T10:50:10Z")

</div>

Hi everyone,

I'm facing a curious problem.

I defined an analyzer in my settings, this way :

{  
"_index_":{  
"cluster.name":"test-cluster",  
"client.transport.sniff":true,  
"_analysis_":{  
"_filter_":{  
"french\_elision":{  
"type":"elision",  
"articles":[  
...skipped...  
]  
},  
"french\_stop":{  
"type":"stop",  
"stopwords":"_french_",  
"ignore\_case":true  
},  
"snowball":{  
"type":"snowball",  
"language":"french"  
}  
},  
"_analyzer_":{  
"_my\_french_":{  
"type":"custom",  
"tokenizer":"standard",  
"filter":[  
"french\_elision",  
"lowercase",  
"french\_stop",  
"snowball"  
]  
},  
"_lower\_analyzer_":{  
"type":"custom",  
"tokenizer":"keyword",  
"filter":"lowercase"  
},  
"_token\_analyzer_":{  
"type":"custom",  
"tokenizer":"whitespace"  
}  
}  
}  
}  
}

And in my mapping, i have two text fields. On one of them i specify  
explicitly the analyzer to use to my\_french, and on the other i let the  
global mapping analyzer (also set to my\_french) kick in automatically. Here  
is the mapping.

{  
"_record_":{  
"_\_all_":{  
"enabled":false  
},  
_"analyzer":"my\_french",_  
"_properties_":{  
"\_uuid":{  
"type":"string",  
"store":"yes",  
"index":"not\_analyzed"  
},  
"_a_":{  
"_type_":"multi\_field",  
"_fields_":{  
_"a"_:{  
"type":"string",  
"store":"yes",  
"index":"analyzed",  
_"analyzer":"my\_french"_  
},  
"_raw_":{  
"type":"string",  
"store":"no",  
"index":"not\_analyzed"  
},  
"_tokens_":{  
"type":"string",  
"store":"no",  
"index":"analyzed",  
"analyzer":"token\_analyzer"  
},  
"_lower_":{  
"type":"string",  
"store":"no",  
"index":"analyzed",  
"analyzer":"lower\_analyzer"  
}  
}  
},  
"g\_r":{  
"type":"string",  
"store":"yes",  
"index":"analyzed"  
}  
}  
}  
}

When i try to analyze with a REST query using my analyzer, the result is  
correct :

$ curl -XGET 'localhost:9200/test-index/\_analyze?analyzer=_my\_french_&pretty=true'  
-d "_j'aime les chevaux_"  
{  
"_tokens_" : [ {  
"token" : "_aim_",  
"start\_offset" : 0,  
"end\_offset" : 6,  
"type" : "",  
"position" : 1  
}, {  
"_token_" : "_cheval_",  
"start\_offset" : 11,  
"end\_offset" : 18,  
"type" : "",  
"position" : 3  
} ]  
}

But when i index my data and search on it, i don't get the expected result.  
So with a facet, i wanted to know what tokens have been stored in the  
index, and here is the result :

$ curl -X POST "[http://localhost:9200/test-index/\_search?pretty=true](http://localhost:9200/test-index/_search?pretty=true)" -d  
'{"query": {"match": {"\_id": "12"}},"facets": {"tokens": {"terms":  
{"field": "a"}}}}'  
{  
"took" : 2,  
"timed\_out" : false,  
"\_shards" : {  
"total" : 5,  
"successful" : 5,  
"failed" : 0  
},  
"_hits_" : {  
"total" : 1,  
"max\_score" : 1.0,  
"hits" : [ {  
"\_index" : "test-index",  
"\_type" : "record",  
"\_id" : "12",  
"\_score" : 1.0,  
"\_source":{"\_uuid":"12","a\_t":false,"a\_n":false,"a":"J'aime les  
chevaux","b\_r":null,"b\_t":false,"b\_n":false,"b":1407664800000,"c\_r":null,"c\_t":false,"c\_n":false,"c":2,"d\_r":"m3","d\_t":true,"d\_n":false,"d":null,"e\_r":null,"e\_t":false,"e\_n":true,"e":12,"f\_r":null,"f\_t":false,"f\_n":false,"f":true,"g\_r":"J'aime  
les chevaux","g\_t":false,"g\_n":false,"g":12.0}  
} ]  
},  
"_facets_" : {  
"_tokens_" : {  
_"\_type" : "terms",_  
"missing" : 0,  
"total" : 2,  
"other" : 0,  
"_terms_" : [ {  
_"term" : "j'aim",_  
"count" : 1  
}, {  
_"term" : "cheval",_  
"count" : 1  
} ]  
}  
}  
}

As you can see, the tokens are not correct. I expect 'aim' and 'cheval', as  
resulting from my analysis, but i got 'j'aim' and 'cheval'.

Anyone could tell me what i missed or what i'm doing wrong here ?

Thanks !

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/c7f4de41-59e4-4b74-96e6-b1157e1d7a8f%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/c7f4de41-59e4-4b74-96e6-b1157e1d7a8f%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 6, 2017, 1:07am UTC](https://discuss.elastic.co/t/my-custom-analyzer-is-registered-but-not-used-during-indexing/19362/2 "2017-07-06T01:07:28Z")

</div>


