# Simple question about html stripping

**URL:** <https://discuss.elastic.co/t/simple-question-about-html-stripping/7251>\
**Category:** Elasticsearch\
**Created:** [April 6, 2012, 3:24pm UTC](https://discuss.elastic.co/t/simple-question-about-html-stripping/7251 "2012-04-06T15:24:53Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Evgeniy\_Galkin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/evgeniy_galkin/32/2741_2.png) [@Evgeniy\_Galkin](https://discuss.elastic.co/u/Evgeniy_Galkin)\
**Post date:** [April 6, 2012, 3:24pm UTC](https://discuss.elastic.co/t/simple-question-about-html-stripping/7251/1 "2012-04-06T15:24:53Z")

</div>

Hello.  
I'm trying to begin use the elasticsearch.  
So, i have my custom analyzer with next config (part of the  
elasticsearch.yml):

analyzer :  
russian\_html :  
type : custom  
tokenizer : standard  
filter : [standard, lowercase, rus\_stem]  
char\_filter : [html\_strip]  
filter :  
rus\_stem :  
type : stemmer  
name : russian

Search is working ok, except next thing:  
Next query - {"query":{"text":{"body":"style"}}} - returns result in  
which word "style" is part of (for example)

some  
things

, although in this case the keyword is an html attribute. Is  
it normal behaviour or it's a bug/my fault in configuration?

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [April 8, 2012, 6:18pm UTC](https://discuss.elastic.co/t/simple-question-about-html-stripping/7251/2 "2012-04-08T18:18:56Z")

</div>

It should be stripped, can you test it with a sample doc and using the  
analyze API endpoint? See if its stripeed.

On Fri, Apr 6, 2012 at 6:24 PM, Evgeniy Galkin [evgeniy@parkflyer.ru](mailto:evgeniy@parkflyer.ru) wrote:

> Hello.  
> I'm trying to begin use the elasticsearch.  
> So, i have my custom analyzer with next config (part of the  
> elasticsearch.yml):
> 
> analyzer :  
> russian\_html :  
> type : custom  
> tokenizer : standard  
> filter : [standard, lowercase, rus\_stem]  
> char\_filter : [html\_strip]  
> filter :  
> rus\_stem :  
> type : stemmer  
> name : russian
> 
> Search is working ok, except next thing:  
> Next query - {"query":{"text":{"body":"style"}}} - returns result in  
> which word "style" is part of (for example)
> 
> some  
> things
> 
> , although in this case the keyword is an html attribute. Is  
> it normal behaviour or it's a bug/my fault in configuration?

---

<div class="post-metadata">

**Author:** ![Evgeniy\_Galkin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/evgeniy_galkin/32/2741_2.png) [@Evgeniy\_Galkin](https://discuss.elastic.co/u/Evgeniy_Galkin)\
**Post date:** [April 9, 2012, 6:37am UTC](https://discuss.elastic.co/t/simple-question-about-html-stripping/7251/3 "2012-04-09T06:37:17Z")

</div>

Thank you for hint about analyze API.  
Probably, I've found a bug.  
The problem appears when html tag contains an attribute which starts  
with '\_' sign (underscore).  
Example:  
$ curl -XGET [http://localhost:9200/comments/\_analyze?field=comment.body](http://localhost:9200/comments/_analyze?field=comment.body)  
-d "\<span style="color: rgb(0, 0, 0); background-color: transparent;  
" sometag="someval"\>"  
{"tokens":}

$ curl -XGET [http://localhost:9200/comments/\_analyze?field=comment.body](http://localhost:9200/comments/_analyze?field=comment.body)  
-d "\<span style="color: rgb(0, 0, 0); background-color: transparent;  
" \_sometag="someval"\>"  
{"tokens":[{"token":"span","start\_offset":1,"end\_offset":  
5,"type":"","position":1},{"token":"style","start\_offset":  
6,"end\_offset":11,"type":"","position":2},  
{"token":"color","start\_offset":13,"end\_offset":  
18,"type":"","position":3},{"token":"rgb","start\_offset":  
20,"end\_offset":23,"type":"","position":4},  
{"token":"0","start\_offset":24,"end\_offset":  
25,"type":"","position":5},{"token":"0","start\_offset":  
27,"end\_offset":28,"type":"","position":6},  
{"token":"0","start\_offset":30,"end\_offset":  
31,"type":"","position":7},{"token":"background","start\_offset":  
34,"end\_offset":44,"type":"","position":8},  
{"token":"color","start\_offset":45,"end\_offset":  
50,"type":"","position":9},  
{"token":"transparent","start\_offset":52,"end\_offset":  
63,"type":"","position":10},  
{"token":"\_sometag","start\_offset":66,"end\_offset":  
74,"type":"","position":11},  
{"token":"someval","start\_offset":76,"end\_offset":  
83,"type":"","position":12}]}

Underscore sign in '\_sometag' is the only difference.

So, should i create bug report on github?

On 9 апр, 00:18, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:

> It should be stripped, can you test it with a sample doc and using the  
> analyze API endpoint? See if its stripeed.
> 
> On Fri, Apr 6, 2012 at 6:24 PM, Evgeniy Galkin [evge...@parkflyer.ru](mailto:evge...@parkflyer.ru) wrote:
> 
> > Hello.  
> > I'm trying to begin use the elasticsearch.  
> > So, i have my custom analyzer with next config (part of the  
> > elasticsearch.yml):
> 
> > analyzer :  
> > russian\_html :  
> > type : custom  
> > tokenizer : standard  
> > filter : [standard, lowercase, rus\_stem]  
> > char\_filter : [html\_strip]  
> > filter :  
> > rus\_stem :  
> > type : stemmer  
> > name : russian
> 
> > Search is working ok, except next thing:  
> > Next query - {"query":{"text":{"body":"style"}}} - returns result in  
> > which word "style" is part of (for example)
> > 
> > some  
> > things
> > 
> > , although in this case the keyword is an html attribute. Is  
> > it normal behaviour or it's a bug/my fault in configuration?

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [April 11, 2012, 9:28am UTC](https://discuss.elastic.co/t/simple-question-about-html-stripping/7251/4 "2012-04-11T09:28:24Z")

</div>

The html stripping is part of Lucene, so the bug is probably there. I can  
try and chase it and open a recreation on the Lucene level.

On Mon, Apr 9, 2012 at 9:37 AM, Evgeniy Galkin [evgeniy@parkflyer.ru](mailto:evgeniy@parkflyer.ru) wrote:

> Thank you for hint about analyze API.  
> Probably, I've found a bug.  
> The problem appears when html tag contains an attribute which starts  
> with '\_' sign (underscore).  
> Example:  
> $ curl -XGET [http://localhost:9200/comments/\_analyze?field=comment.body](http://localhost:9200/comments/_analyze?field=comment.body)  
> -d "\<span style="color: rgb(0, 0, 0); background-color: transparent;  
> " sometag="someval"\>"  
> {"tokens":}
> 
> $ curl -XGET [http://localhost:9200/comments/\_analyze?field=comment.body](http://localhost:9200/comments/_analyze?field=comment.body)  
> -d "\<span style="color: rgb(0, 0, 0); background-color: transparent;  
> " \_sometag="someval"\>"  
> {"tokens":[{"token":"span","start\_offset":1,"end\_offset":  
> 5,"type":"","position":1},{"token":"style","start\_offset":  
> 6,"end\_offset":11,"type":"","position":2},  
> {"token":"color","start\_offset":13,"end\_offset":  
> 18,"type":"","position":3},{"token":"rgb","start\_offset":  
> 20,"end\_offset":23,"type":"","position":4},  
> {"token":"0","start\_offset":24,"end\_offset":  
> 25,"type":"","position":5},{"token":"0","start\_offset":  
> 27,"end\_offset":28,"type":"","position":6},  
> {"token":"0","start\_offset":30,"end\_offset":  
> 31,"type":"","position":7},{"token":"background","start\_offset":  
> 34,"end\_offset":44,"type":"","position":8},  
> {"token":"color","start\_offset":45,"end\_offset":  
> 50,"type":"","position":9},  
> {"token":"transparent","start\_offset":52,"end\_offset":  
> 63,"type":"","position":10},  
> {"token":"\_sometag","start\_offset":66,"end\_offset":  
> 74,"type":"","position":11},  
> {"token":"someval","start\_offset":76,"end\_offset":  
> 83,"type":"","position":12}]}
> 
> Underscore sign in '\_sometag' is the only difference.
> 
> So, should i create bug report on github?
> 
> On 9 апр, 00:18, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> 
> > It should be stripped, can you test it with a sample doc and using the  
> > analyze API endpoint? See if its stripeed.
> > 
> > On Fri, Apr 6, 2012 at 6:24 PM, Evgeniy Galkin [evge...@parkflyer.ru](mailto:evge...@parkflyer.ru)  
> > wrote:
> > 
> > > Hello.  
> > > I'm trying to begin use the elasticsearch.  
> > > So, i have my custom analyzer with next config (part of the  
> > > elasticsearch.yml):
> > 
> > > analyzer :  
> > > russian\_html :  
> > > type : custom  
> > > tokenizer : standard  
> > > filter : [standard, lowercase, rus\_stem]  
> > > char\_filter : [html\_strip]  
> > > filter :  
> > > rus\_stem :  
> > > type : stemmer  
> > > name : russian
> > 
> > > Search is working ok, except next thing:  
> > > Next query - {"query":{"text":{"body":"style"}}} - returns result in  
> > > which word "style" is part of (for example)
> > > 
> > > some  
> > > things
> > > 
> > > , although in this case the keyword is an html attribute. Is  
> > > it normal behaviour or it's a bug/my fault in configuration?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:32am UTC](https://discuss.elastic.co/t/simple-question-about-html-stripping/7251/5 "2017-07-06T03:32:59Z")

</div>


