# Override built-in analyzer

**URL:** <https://discuss.elastic.co/t/override-built-in-analyzer/14687>\
**Category:** Elasticsearch\
**Created:** [December 3, 2013, 8:04pm UTC](https://discuss.elastic.co/t/override-built-in-analyzer/14687 "2013-12-03T20:04:18Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![sashao](https://avatars.discourse-cdn.com/v4/letter/s/41988e/32.png) [@sashao](https://discuss.elastic.co/u/sashao)\
**Post date:** [December 3, 2013, 8:04pm UTC](https://discuss.elastic.co/t/override-built-in-analyzer/14687/1 "2013-12-03T20:04:18Z")

</div>

Hello friends,

I'm trying to preserve specific characters during tokenization using word\_delimiter filter by defining the type\_table (as described in [http://www.fullscale.co/blog/2013/03/04/preserving\_specific\_characters\_during\_tokenizing\_in\_elasticsearch.html](http://www.fullscale.co/blog/2013/03/04/preserving_specific_characters_during_tokenizing_in_elasticsearch.html)).  
Actually my idea is to override the built-in English analyzer by including custom configured "word\_delimiter" ("type\_table": ["# =\> ALPHA", "@ =\> ALPHA"]) filter, but I cannot find any way to do it.  
I also tried to create a custom english analyzer but still getting next problems:

1. I don't actually know the default settings of the built-in english analyzer (But I really want to preserve it)
2. While trying to set "tokenizer": english getting an error on creating index, saying that english tokenizer is not found.  
I'm using 0.90.5 ES

Hope for your kind help!  
Sasha

---

<div class="post-metadata">

**Author:** ![sashao](https://avatars.discourse-cdn.com/v4/letter/s/41988e/32.png) [@sashao](https://discuss.elastic.co/u/sashao)\
**Post date:** [December 4, 2013, 10:03am UTC](https://discuss.elastic.co/t/override-built-in-analyzer/14687/2 "2013-12-04T10:03:54Z")

</div>

**This is my configuration (using Sense plugin for chrome):**  
POST \_template/temp1  
{  
"template": "_",  
"order": "5",  
"settings": {  
"index": {  
"analysis": {  
"filter": {  
"word\_delimiter\_filter": {  
"type": "word\_delimiter",  
"generate\_word\_parts": false,  
"catenate\_words": true,  
"split\_on\_numerics": false,  
"preserve\_original": true,  
"type\_table": [  
"# =\> ALPHA",  
"@ =\> ALPHA",  
"% =\> ALPHA",  
"$ =\> ALPHA",  
"% =\> ALPHA"  
]  
},  
"stop\_english": {  
"type": "stop",  
"stopwords": [  
"english"  
]  
},  
"english\_stemmer": {  
"type": "stemmer",  
"name": "english"  
}  
},  
"analyzer": {  
"english": { **//HERE I'M TRYING TO OVERRIDE THE BUILT IN english ANALYZER**  
"filter": [  
"word\_delimiter\_filter"  
]  
},  
"english2": { **//HERE I'M TRYING TO CONFIG MY OWN english ANALYZER THAT WOULD BEHAVE LIKE** THE BUILT IN  
"type": "custom",  
"tokenizer": "english",  
"filter": [  
"lowercase",  
"word\_delimiter\_filter",  
"english\_stemmer",  
"stop\_english"  
]  
}  
}  
}  
}  
},  
"mappings": {  
"default": {  
"dynamic\_templates": [  
{  
"template\_textEnglish": {  
"match": "text.English._",  
"mapping": {  
"type": "string",  
"store": "yes",  
"index": "analyzed",  
"analyzer": "english",  
"term\_vector": "with\_positions\_offsets"  
}  
}  
},  
{  
"template\_textEnglish": {  
"match": "text.English2.\*",  
"mapping": {  
"type": "string",  
"store": "yes",  
"index": "analyzed",  
"analyzer": "english2",  
"term\_vector": "with\_positions\_offsets"  
}  
}  
}  
]  
}  
}  
}

**and this is the error I get trying to create a new index:**  
{  
"error": "IndexCreationException[[test1] failed to create index]; nested: ElasticSearchIllegalArgumentException[failed to find analyzer type [null] or tokenizer for [english]]; ",  
"status": 400  
}

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:03am UTC](https://discuss.elastic.co/t/override-built-in-analyzer/14687/3 "2017-07-06T02:03:26Z")

</div>


