# Strip\_html

**URL:** <https://discuss.elastic.co/t/strip-html/3777>\
**Category:** Elasticsearch\
**Created:** [January 14, 2011, 9:58am UTC](https://discuss.elastic.co/t/strip-html/3777 "2011-01-14T09:58:04Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Andrew\_Degtiariov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrew_degtiariov/32/3223_2.png) [@Andrew\_Degtiariov](https://discuss.elastic.co/u/Andrew_Degtiariov)\
**Post date:** [January 14, 2011, 9:58am UTC](https://discuss.elastic.co/t/strip-html/3777/1 "2011-01-14T09:58:04Z")

</div>

Hi!

I'm using html\_strip for striping all HTML tags from some fields. But when I  
search messages with "img" in its body then ES finds messages with ![]() tag  
too.  
Where I'm wrong?

There is the query:

{'query': {  
'filtered': {  
'filter': {'term': {'owner\_id': '4d07646affc84a6717000011'}  
},  
'query': {'query\_string': {'query': u'img\*'}}  
}  
}

Here is part of my ES config:

index:  
analysis:  
analyzer:  
message\_content:  
tokenizer: standard  
char\_filter: [html\_strip]  
read\_ahead: 1024

Part of the mapping of index "messages":

```
    "message": {
        "_source" : {"enabled": false},
        "properties" : {
            "body": {
                 "type": "string",
                 "index_analyzer": "message_content"
            }
       }
  }

```

--  
Andrew Degtiariov  
DA-RIPE

---

<div class="post-metadata">

**Author:** ![searchersteve](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/searchersteve/32/3008_2.png) [@searchersteve](https://discuss.elastic.co/u/searchersteve)\
**Post date:** [January 14, 2011, 10:36pm UTC](https://discuss.elastic.co/t/strip-html/3777/2 "2011-01-14T22:36:51Z")

</div>

I'm afraid I'm not writing with a solution but a related question...

My reading of the ES manual suggests that the analyzer will run html\_strip by default -- in other words, there's no need to declare it in the configuration. But am I correct? See below:

[http://www.elasticsearch.com/docs/elasticsearch/index\_modules/analysis/charfilter/](http://www.elasticsearch.com/docs/elasticsearch/index_modules/analysis/charfilter/)

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [January 15, 2011, 1:11pm UTC](https://discuss.elastic.co/t/strip-html/3777/3 "2011-01-15T13:11:12Z")

</div>

Hi Andrew

> I'm using html\_strip for striping all HTML tags from some fields. But  
> when I search messages with "img" in its body then ES finds messages  
> with ![]() tag too.  
> Where I'm wrong?

I created a small HTML strip test here

> <https://gist.github.com/clintongormley/780895>

This creates an index with two custom analyzers:

test\_1: {  
"tokenizer" : "standard",  
"char\_filter" : ["html\_strip"]  
}

test\_2: {  
"tokenizer" : "standard",  
"char\_filter" : ["html\_strip"],  
"filter" : ["standard","lowercase","stop","asciifolding"],  
}

Then I use the analyze call to compare the results of the standard,  
test\_1 and test\_2 analyzers when indexing the text:

"the **quick** brÃ¶wn ![]() "jumped""

The results show that this is working correctly.

Are you sure that you added your mapping with the message\_content  
analyzer before you indexed all of your docs?

Can you create a curl recreation of the issue that you are seeing?

thanks

Clint

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [January 15, 2011, 1:11pm UTC](https://discuss.elastic.co/t/strip-html/3777/4 "2011-01-15T13:11:49Z")

</div>

Hiya

On Fri, 2011-01-14 at 14:36 -0800, searchersteve wrote:

> I'm afraid I'm not writing with a solution but a related question...
> 
> My reading of the ES manual suggests that the analyzer will run html\_strip  
> by default -- in other words, there's no need to declare it in the  
> configuration. But am I correct? See below:

No, it is available by default, but not enabled by default.

See [HTML Strip charfilter test for ElasticSearch · GitHub](https://gist.github.com/780895) for an example of how to use it.

clint

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:13am UTC](https://discuss.elastic.co/t/strip-html/3777/5 "2017-07-06T04:13:52Z")

</div>


