# How to use html\_strip Char filter?

**URL:** <https://discuss.elastic.co/t/how-to-use-html-strip-char-filter/5696>\
**Category:** Elasticsearch\
**Created:** [October 26, 2011, 6:57pm UTC](https://discuss.elastic.co/t/how-to-use-html-strip-char-filter/5696 "2011-10-26T18:57:22Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mauricio\_Alarcon](https://avatars.discourse-cdn.com/v4/letter/m/cab0a1/32.png) [@Mauricio\_Alarcon](https://discuss.elastic.co/u/Mauricio_Alarcon)\
**Post date:** [October 26, 2011, 6:57pm UTC](https://discuss.elastic.co/t/how-to-use-html-strip-char-filter/5696/1 "2011-10-26T18:57:22Z")

</div>

Guys, I'm in need of remove all html from a specific field on my  
documents corpus. I based my configuration on  
[http://www.elasticsearch.org/guide/reference/index-modules/analysis/custom-analyzer.html](http://www.elasticsearch.org/guide/reference/index-modules/analysis/custom-analyzer.html)

and ended with this inside my elasticsearch.yml

11 index :  
12 analysis :  
13 analyzer:  
14 descriptionAnalyzer:  
15 type: custom  
16 tokenizer: standard  
17 filter: standard  
18 char\_filter: html\_strip

And in the mappings I pointed the field that I wanted to this  
analyzer  
"description" : { "type" : "string", "index" : "analyzed",  
"analyzer" : "descriptionAnalyzer" }

I confirmed that it was used after few indexed docs

```
 "description" : {
      "analyzer" : "descriptionAnalyzer",
      "type" : "string"
    },

```

But I'm still getting html in my search results.

What am I doing wrong?

Cheers

~M

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [October 26, 2011, 8:48pm UTC](https://discuss.elastic.co/t/how-to-use-html-strip-char-filter/5696/2 "2011-10-26T20:48:28Z")

</div>

What do you mean that you get HTML in your search results? You get them as  
part of the \_source? If so, then it makes sense, since the \_source is just  
the document you indexed.

On Wed, Oct 26, 2011 at 8:57 PM, maverick [mauricio.alarcon@gmail.com](mailto:mauricio.alarcon@gmail.com)wrote:

> Guys, I'm in need of remove all html from a specific field on my  
> documents corpus. I based my configuration on
> 
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/custom-analyzer.html)
> 
> and ended with this inside my elasticsearch.yml
> 
> 11 index :  
> 12 analysis :  
> 13 analyzer:  
> 14 descriptionAnalyzer:  
> 15 type: custom  
> 16 tokenizer: standard  
> 17 filter: standard  
> 18 char\_filter: html\_strip
> 
> And in the mappings I pointed the field that I wanted to this  
> analyzer  
> "description" : { "type" : "string", "index" : "analyzed",  
> "analyzer" : "descriptionAnalyzer" }
> 
> I confirmed that it was used after few indexed docs
> 
> ```
> "description" : {
> "analyzer" : "descriptionAnalyzer",
> "type" : "string"
> },
> 
> ```
> 
> But I'm still getting html in my search results.
> 
> What am I doing wrong?
> 
> Cheers
> 
> ~M

---

<div class="post-metadata">

**Author:** ![phobos182](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/phobos182/32/3011_2.png) [@phobos182](https://discuss.elastic.co/u/phobos182)\
**Post date:** [October 27, 2011, 1:33pm UTC](https://discuss.elastic.co/t/how-to-use-html-strip-char-filter/5696/3 "2011-10-27T13:33:48Z")

</div>

I submitted something like this a few months back. The HTML script character filter just removes the items from the index, but not from the stored \_source / value.

We use JSoup to remove HTML entries before indexing on the client side.

---

<div class="post-metadata">

**Author:** ![Mauricio\_Alarcon](https://avatars.discourse-cdn.com/v4/letter/m/cab0a1/32.png) [@Mauricio\_Alarcon](https://discuss.elastic.co/u/Mauricio_Alarcon)\
**Post date:** [October 27, 2011, 1:40pm UTC](https://discuss.elastic.co/t/how-to-use-html-strip-char-filter/5696/4 "2011-10-27T13:40:03Z")

</div>

Thanks Shay,

I spoke too soon, seems that I had a brain fart. After posting the  
message I started playing with the analyzer and indeed the analyzer  
does its job just right. I gisted it here for the record

> <https://gist.github.com/mauricioalarcon/1319548>

For a moment I thought that playing with the analyzer + setting  
store=yes will also get rid off the html on my source, which was a  
simple dirty way to remove unwanted formatting for my view. But indeed  
this doesn't make any sense

Sorry for the false alarm

Cheers

~M

On Oct 26, 4:48 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:

> What do you mean that you get HTML in your search results? You get them as  
> part of the \_source? If so, then it makes sense, since the \_source is just  
> the document you indexed.
> 
> On Wed, Oct 26, 2011 at 8:57 PM, maverick [mauricio.alar...@gmail.com](mailto:mauricio.alar...@gmail.com)wrote:
> 
> > Guys, I'm in need of remove all html from a specific field on my  
> > documents corpus. I based my configuration on
> 
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/c)...
> 
> > and ended with this inside my elasticsearch.yml
> 
> > 11 index :  
> > 12 analysis :  
> > 13 analyzer:  
> > 14 descriptionAnalyzer:  
> > 15 type: custom  
> > 16 tokenizer: standard  
> > 17 filter: standard  
> > 18 char\_filter: html\_strip
> 
> > And in the mappings I pointed the field that I wanted to this  
> > analyzer  
> > "description" : { "type" : "string", "index" : "analyzed",  
> > "analyzer" : "descriptionAnalyzer" }
> 
> > I confirmed that it was used after few indexed docs
> 
> > ```
> > "description" : {
> > "analyzer" : "descriptionAnalyzer",
> > "type" : "string"
> > },
> > 
> > ```
> 
> > But I'm still getting html in my search results.
> 
> > What am I doing wrong?
> 
> > Cheers
> 
> > ~M

---

<div class="post-metadata">

**Author:** ![Mauricio\_Alarcon](https://avatars.discourse-cdn.com/v4/letter/m/cab0a1/32.png) [@Mauricio\_Alarcon](https://discuss.elastic.co/u/Mauricio_Alarcon)\
**Post date:** [October 27, 2011, 1:41pm UTC](https://discuss.elastic.co/t/how-to-use-html-strip-char-filter/5696/5 "2011-10-27T13:41:12Z")

</div>

Thanks for the tip phobos182 this is just exactly what I was looking  
for

~M

On Oct 27, 9:33 am, phobos182 [phobos...@gmail.com](mailto:phobos...@gmail.com) wrote:

> I submitted something like this a few months back. The HTML script character  
> filter just removes the items from the index, but not from the stored  
> \_source / value.
> 
> We use JSoup to remove HTML entries before indexing on the client side.
> 
> --  
> View this message in context:[http://elasticsearch-users.115913.n3.nabble.com/How-to-use-html-strip](http://elasticsearch-users.115913.n3.nabble.com/How-to-use-html-strip)...  
> Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:50am UTC](https://discuss.elastic.co/t/how-to-use-html-strip-char-filter/5696/6 "2017-07-06T03:50:40Z")

</div>


