# Extend built-in analyzers

**URL:** <https://discuss.elastic.co/t/extend-built-in-analyzers/134778>\
**Category:** Elasticsearch\
**Created:** [June 6, 2018, 10:01am UTC](https://discuss.elastic.co/t/extend-built-in-analyzers/134778 "2018-06-06T10:01:31Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![yansal](https://avatars.discourse-cdn.com/v4/letter/y/dc4da7/32.png) [@yansal](https://discuss.elastic.co/u/yansal)\
**Post date:** [June 6, 2018, 10:01am UTC](https://discuss.elastic.co/t/extend-built-in-analyzers/134778/1 "2018-06-06T10:01:31Z")

</div>

My use case is to add the [html\_strip char\_filter](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-htmlstrip-charfilter.html) to an existing language analyzer.

For example, I would like to create an index like this one:

```auto
PUT /myindex
{
    "settings": {
        "analysis": {
            "analyzer": {
                "english_html_strip": {
                    "type": "english",
                    "char_filter": [
                        "html_strip"
                    ]
                }
            }
        }
    },
    "mappings": {
        "_doc": {
            "properties": {
                "description_english_html": {
                    "type": "text",
                    "analyzer": "english_html_strip"
                }
            }
        }
    }
}

```

I then expect the following `_analyze` request to both use the english analyzer and to strip html tags.

```auto
POST /myindex/_analyze
{
    "field": "description_english_html",
    "text": "<h1>A header</h1><p>A paragraph</p>"
}

```

However it doesn't strip html tags, see the output:

```auto
{
    "tokens": [
        {
            "token": "h1",
            "start_offset": 1,
            "end_offset": 3,
            "type": "<ALPHANUM>",
            "position": 0
        },
        {
            "token": "header",
            "start_offset": 6,
            "end_offset": 12,
            "type": "<ALPHANUM>",
            "position": 2
        },
        {
            "token": "h1",
            "start_offset": 14,
            "end_offset": 16,
            "type": "<ALPHANUM>",
            "position": 3
        },
        {
            "token": "p",
            "start_offset": 18,
            "end_offset": 19,
            "type": "<ALPHANUM>",
            "position": 4
        },
        {
            "token": "paragraph",
            "start_offset": 22,
            "end_offset": 31,
            "type": "<ALPHANUM>",
            "position": 6
        },
        {
            "token": "p",
            "start_offset": 33,
            "end_offset": 34,
            "type": "<ALPHANUM>",
            "position": 7
        }
    ]
}

```

Is there a solution, besides rebuilding the language analyzer from scratch?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [June 6, 2018, 10:14am UTC](https://discuss.elastic.co/t/extend-built-in-analyzers/134778/2 "2018-06-06T10:14:02Z")

</div>

I'd follow this: [https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-lang-analyzer.html#english-analyzer](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-lang-analyzer.html#english-analyzer)

---

<div class="post-metadata">

**Author:** ![tdasch](https://avatars.discourse-cdn.com/v4/letter/t/58f4c7/32.png) [@tdasch](https://discuss.elastic.co/u/tdasch)\
**Post date:** [June 6, 2018, 4:02pm UTC](https://discuss.elastic.co/t/extend-built-in-analyzers/134778/3 "2018-06-06T16:02:16Z")

</div>

I have no idea if this is correct, but it appears to work. I feel like i've gotten it wrong but maybe it can help you reach the right answer. The documentation was confusing on adding a character filter to a language analyzer.

```
> PUT test23
> {
> "settings": {
> "analysis": {
> "analyzer": {
> "english": {
> "type": "custom",
> "char_filter": ["html_strip"],
> "tokenizer": "standard"
> }
> }
> }
> } 
> }

> POST test23/_analyze
> {
> "analyzer": "english",
> "text": "<h1>A header</h1><p>A paragraph</p>"
> }

> {
> "tokens": [
> {
> "token": "A",
> "start_offset": 4,
> "end_offset": 5,
> "type": "<ALPHANUM>",
> "position": 0
> },
> {
> "token": "header",
> "start_offset": 6,
> "end_offset": 12,
> "type": "<ALPHANUM>",
> "position": 1
> },
> {
> "token": "A",
> "start_offset": 20,
> "end_offset": 21,
> "type": "<ALPHANUM>",
> "position": 2
> },
> {
> "token": "paragraph",
> "start_offset": 22,
> "end_offset": 31,
> "type": "<ALPHANUM>",
> "position": 3
> }
> ]
> }
```

---

<div class="post-metadata">

**Author:** ![yansal](https://avatars.discourse-cdn.com/v4/letter/y/dc4da7/32.png) [@yansal](https://discuss.elastic.co/u/yansal)\
**Post date:** [June 7, 2018, 8:54am UTC](https://discuss.elastic.co/t/extend-built-in-analyzers/134778/4 "2018-06-07T08:54:22Z")

</div>

It doesn't work. As you can see it just erases the "english" analyzer. "A" appears as a token twice and shouldn't because it's an english article.

---

<div class="post-metadata">

**Author:** ![tdasch](https://avatars.discourse-cdn.com/v4/letter/t/58f4c7/32.png) [@tdasch](https://discuss.elastic.co/u/tdasch)\
**Post date:** [June 7, 2018, 12:12pm UTC](https://discuss.elastic.co/t/extend-built-in-analyzers/134778/5 "2018-06-07T12:12:22Z")

</div>

Yeah I see that! I slept on it and did some more looking this morning with fresh eyes. Mr. Pilato was on the money it would seem but it took me a while to figure out why.

This is my final product, which copies the English analyzer custom example from the doc Mr. Pilato linked - customization will come from your needs.

```
> PUT test23
> {
> "settings": {
> "analysis": {
> "filter": {
> "english_stop": {
> "type": "stop",
> "stopwords": "_english_" 
> },
> "english_keywords": {
> "type": "keyword_marker",
> "keywords": ["example"] 
> },
> "english_stemmer": {
> "type": "stemmer",
> "language": "english"
> },
> "english_possessive_stemmer": {
> "type": "stemmer",
> "language": "possessive_english"
> }
> },
> "analyzer": {
> "english": {
> "tokenizer": "standard",
> "filter": [
> "english_possessive_stemmer",
> "lowercase",
> "english_stop",
> "english_keywords",
> "english_stemmer"
> ],
> "char_filter": ["html_strip"]
> }
> }
> }
> }
> }

```

My Test:

```
> POST test23/_analyze
> {
> "analyzer": "english",
> "text": "<h1>A header</h1><p>A paragraph</p>"
> }

```

My Result:

```
> {
> "tokens": [
> {
> "token": "header",
> "start_offset": 6,
> "end_offset": 12,
> "type": "<ALPHANUM>",
> "position": 1
> },
> {
> "token": "paragraph",
> "start_offset": 22,
> "end_offset": 31,
> "type": "<ALPHANUM>",
> "position": 3
> }
> ]
> }

```

I wanted to lay out my thought process incase someone wanted to provide further insight or correct my error(s). `Settings` is customizing the token filters of the english analyzer (this is the part that would be tailored for your needs).`Analyzer` is selecting the english analyzer, setting the standard tokenizer, setting the filters customized in `settings`, and adding the char\_filter `html_strip`. I hope this helps?

---

<div class="post-metadata">

**Author:** ![yansal](https://avatars.discourse-cdn.com/v4/letter/y/dc4da7/32.png) [@yansal](https://discuss.elastic.co/u/yansal)\
**Post date:** [June 7, 2018, 1:10pm UTC](https://discuss.elastic.co/t/extend-built-in-analyzers/134778/6 "2018-06-07T13:10:04Z")

</div>

It works by rebuilding the language analyzer from scratch. The question is is there a solution to extend an existing analyzer without rebuilding it from scratch.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [June 7, 2018, 1:51pm UTC](https://discuss.elastic.co/t/extend-built-in-analyzers/134778/7 "2018-06-07T13:51:22Z")

</div>

No. There's not apart the options documented for each analyzer if any.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [June 7, 2018, 2:19pm UTC](https://discuss.elastic.co/t/extend-built-in-analyzers/134778/8 "2018-06-07T14:19:52Z")

</div>

I'm not sure I agree with the fact you opened

> <https://github.com/elastic/elasticsearch/issues/31174>

I mean that the workaround is super easy as the analyzer is documented and I don't see much value of implementing this ^^^.

But let's see what the team is saying.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2018, 2:19pm UTC](https://discuss.elastic.co/t/extend-built-in-analyzers/134778/9 "2018-07-05T14:19:55Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
