# Problem with standard\_html\_strip

**URL:** <https://discuss.elastic.co/t/problem-with-standard-html-strip/8206>\
**Category:** Elasticsearch\
**Created:** [June 24, 2012, 3:44am UTC](https://discuss.elastic.co/t/problem-with-standard-html-strip/8206 "2012-06-24T03:44:30Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![M11](https://avatars.discourse-cdn.com/v4/letter/m/85f322/32.png) [@M11](https://discuss.elastic.co/u/M11)\
**Post date:** [June 24, 2012, 3:44am UTC](https://discuss.elastic.co/t/problem-with-standard-html-strip/8206/1 "2012-06-24T03:44:30Z")

</div>

Hi,

I'd like to index words from HTML source. I'm having same problem with  
this ( [https://gist.github.com/1478233](https://gist.github.com/1478233) )  
It appears that ElasticSearch indexes HTML tags along with words.

(copy from the URL)  
$ curl -XPUT [http://localhost:9200/foo](http://localhost:9200/foo)

{"ok":true,"acknowledged":true}

$ curl -XPUT [http://localhost:9200/foo/bar/\_mapping](http://localhost:9200/foo/bar/_mapping) -d '{  
"properties" : {  
"body": {"type":"string", "analyzer":"standard\_html\_strip"}  
}  
}'

{"ok":true,"acknowledged":true}

$ curl -XPUT [http://localhost:9200/foo/bar/1](http://localhost:9200/foo/bar/1) -d '{  
"body": "

This is a **test**.

"  
}'

{"ok":true,"\_index":"foo","\_type":"bar","\_id":"1","\_version":1}

$ curl -XGET [http://localhost:9200/foo/bar/\_search?q=body:strong](http://localhost:9200/foo/bar/_search?q=body:strong)

{"took":35,"timed\_out":false,"\_shards":{"total":5,"successful":  
5,"failed":0},"hits":{"total":1,"max\_score":0.18985549,"hits":  
[{"\_index":"foo","\_type":"bar","\_id":"1","\_score":0.18985549,  
"\_source" : {  
"body": "

This is a **test**.

"  
}}]}}

I assume that the last query should not yield anything because  
"strong" is a tag name. Am I doing something wrong or is it just a  
Lucene bug?

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [June 24, 2012, 8:32am UTC](https://discuss.elastic.co/t/problem-with-standard-html-strip/8206/2 "2012-06-24T08:32:25Z")

</div>

Hi

there are two bugs in your configuration. You are treating the html\_strip  
filter as an analyzer, which does not work and you are indexing the mapping  
wrong.

Put this in your elasticsearch.yml:

index:  
analysis:  
analyzer:  
default:  
type: standard

```
  strip_html_analyzer:
    type: custom
    tokenizer: standard
    filter: [standard]
    char_filter: html_strip

```

Then correct setting the mapping (you need to add the type as well as  
setting the right analyzer:)

curl -XPUT [http://localhost:9200/foo/bar/\_mapping](http://localhost:9200/foo/bar/_mapping) -d '{  
"bar":  
{ "properties" : {  
"body": {"type":"string", "analyzer":"strip\_html\_analyzer" }  
}  
}  
}'

By checking the mapping with a GET on the above URL you will also see that  
it is empty with your current configuration, so the mapping was not applied  
(without getting back an error message, which is not too nice..)

Now index your document again and start searching for "strong" and "test"  
and it should work.

--Alexander

---

<div class="post-metadata">

**Author:** ![M11](https://avatars.discourse-cdn.com/v4/letter/m/85f322/32.png) [@M11](https://discuss.elastic.co/u/M11)\
**Post date:** [June 24, 2012, 6:20pm UTC](https://discuss.elastic.co/t/problem-with-standard-html-strip/8206/3 "2012-06-24T18:20:29Z")

</div>

Hi Alex,

html\_strip is a char filter and I was mentioning standard\_html\_strip which  
I believe it is an analyzer with html\_strip.

Anyway, I tried your suggestion.

$ curl -XPUT [http://localhost:9200/foo/bar/\_mapping](http://localhost:9200/foo/bar/_mapping) -d '{  
"bar":  
{ "properties" : {  
"body": {"type":"string", "analyzer":"strip\_html\_analyzer" }  
}  
}  
}'  
{"ok":true,"acknowledged":true}

$ curl -XPUT localhost:9200/foo/bar/1 -d '{

> "body" : "
> 
> hello world 
> 
> there it is
> 
>  
> "  
> }  
> '  
> {"ok":true,"\_index":"foo","\_type":"bar","\_id":"1","\_version":1}

$ curl -XGET localhost:9200/foo/bar/\_search?q=color  
{"took":3,"timed\_out":false,"\_shards":{"total":5,"successful":5,"failed":0},"hits":{"total":1,"max\_score":0.06780553,"hits":[{"\_index":"foo","\_type":"bar","\_id":"1","\_score":0.06780553,  
"\_source" : {  
"body" : "

hello world 

there it is

 
"  
}  
}]}}

Calling to \_analyze handler directly seems to be working fine.

$ curl -XPOST localhost:9200/foo/\_analyze?analyzer=strip\_html\_analyzer -d  
'

 hello world 
'  
{"tokens":[{"token":"hello","start\_offset":17,"end\_offset":22,"type":"","position":1},{"token":"world","start\_offset":23,"end\_offset":28,"type":"","position":2}]}

The list of token does not include HTML tags.

Am I still doing something wrong?

Thanks,

2012년 6월 24일 일요일 오전 4시 32분 25초 UTC-4, Alexander Reelsen 님의 말:

> Hi
> 
> there are two bugs in your configuration. You are treating the html\_strip  
> filter as an analyzer, which does not work and you are indexing the mapping  
> wrong.
> 
> Put this in your elasticsearch.yml:
> 
> index:  
> analysis:  
> analyzer:  
> default:  
> type: standard
> 
> ```
> strip_html_analyzer:
> type: custom
> tokenizer: standard
> filter: [standard]
> char_filter: html_strip
> 
> ```
> 
> Then correct setting the mapping (you need to add the type as well as  
> setting the right analyzer:)
> 
> curl -XPUT [http://localhost:9200/foo/bar/\_mapping](http://localhost:9200/foo/bar/_mapping) -d '{  
> "bar":  
> { "properties" : {  
> "body": {"type":"string", "analyzer":"strip\_html\_analyzer" }  
> }  
> }  
> }'
> 
> By checking the mapping with a GET on the above URL you will also see that  
> it is empty with your current configuration, so the mapping was not applied  
> (without getting back an error message, which is not too nice..)
> 
> Now index your document again and start searching for "strong" and "test"  
> and it should work.
> 
> --Alexander

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [June 25, 2012, 7:55pm UTC](https://discuss.elastic.co/t/problem-with-standard-html-strip/8206/4 "2012-06-25T19:55:46Z")

</div>

Your search query is searching the \_all field[http://www.elasticsearch.org/guide/reference/mapping/all-field.html](http://www.elasticsearch.org/guide/reference/mapping/all-field.html),  
that includes content of all fields but is indexed using default analyzer.  
To search only the body field and use body-specific analyzer, you need to  
specify this field in your query:

curl -XGET "localhost:9200/foo/bar/\_search?q=body:color"

On Sunday, June 24, 2012 2:20:29 PM UTC-4, M wrote:

> Hi Alex,
> 
> html\_strip is a char filter and I was mentioning standard\_html\_strip which  
> I believe it is an analyzer with html\_strip.
> 
> Anyway, I tried your suggestion.
> 
> $ curl -XPUT [http://localhost:9200/foo/bar/\_mapping](http://localhost:9200/foo/bar/_mapping) -d '{  
> "bar":  
> { "properties" : {  
> "body": {"type":"string", "analyzer":"strip\_html\_analyzer" }  
> }  
> }  
> }'  
> {"ok":true,"acknowledged":true}
> 
> $ curl -XPUT localhost:9200/foo/bar/1 -d '{
> 
> > "body" : "
> > 
> > hello world 
> > 
> > there it is
> > 
> >  
> > "  
> > }  
> > '  
> > {"ok":true,"\_index":"foo","\_type":"bar","\_id":"1","\_version":1}
> 
> $ curl -XGET localhost:9200/foo/bar/\_search?q=color  
> {"took":3,"timed\_out":false,"\_shards":{"total":5,"successful":5,"failed":0},"hits":{"total":1,"max\_score":0.06780553,"hits":[{"\_index":"foo","\_type":"bar","\_id":"1","\_score":0.06780553,  
> "\_source" : {  
> "body" : "
> 
> hello world 
> 
> there it is
> 
>  
> "  
> }  
> }]}}
> 
> Calling to \_analyze handler directly seems to be working fine.
> 
> $ curl -XPOST localhost:9200/foo/\_analyze?analyzer=strip\_html\_analyzer -d  
> '
> 
> hello world 
> '
> 
> {"tokens":[{"token":"hello","start\_offset":17,"end\_offset":22,"type":"","position":1},{"token":"world","start\_offset":23,"end\_offset":28,"type":"","position":2}]}
> 
> The list of token does not include HTML tags.
> 
> Am I still doing something wrong?
> 
> Thanks,
> 
> 2012년 6월 24일 일요일 오전 4시 32분 25초 UTC-4, Alexander Reelsen 님의 말:
> 
> > Hi
> > 
> > there are two bugs in your configuration. You are treating the html\_strip  
> > filter as an analyzer, which does not work and you are indexing the mapping  
> > wrong.
> > 
> > Put this in your elasticsearch.yml:
> > 
> > index:  
> > analysis:  
> > analyzer:  
> > default:  
> > type: standard
> > 
> > ```
> > strip_html_analyzer:
> > type: custom
> > tokenizer: standard
> > filter: [standard]
> > char_filter: html_strip
> > 
> > ```
> > 
> > Then correct setting the mapping (you need to add the type as well as  
> > setting the right analyzer:)
> > 
> > curl -XPUT [http://localhost:9200/foo/bar/\_mapping](http://localhost:9200/foo/bar/_mapping) -d '{  
> > "bar":  
> > { "properties" : {  
> > "body": {"type":"string", "analyzer":"strip\_html\_analyzer" }  
> > }  
> > }  
> > }'
> > 
> > By checking the mapping with a GET on the above URL you will also see  
> > that it is empty with your current configuration, so the mapping was not  
> > applied (without getting back an error message, which is not too nice..)
> > 
> > Now index your document again and start searching for "strong" and "test"  
> > and it should work.
> > 
> > --Alexander

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:22am UTC](https://discuss.elastic.co/t/problem-with-standard-html-strip/8206/5 "2017-07-06T03:22:34Z")

</div>


