# Tokenizing HTML

**URL:** <https://discuss.elastic.co/t/tokenizing-html/8648>\
**Category:** Elasticsearch\
**Created:** [August 6, 2012, 3:42pm UTC](https://discuss.elastic.co/t/tokenizing-html/8648 "2012-08-06T15:42:07Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![David\_James](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/david_james/32/2768_2.png) [@David\_James](https://discuss.elastic.co/u/David_James)\
**Post date:** [August 6, 2012, 3:42pm UTC](https://discuss.elastic.co/t/tokenizing-html/8648/1 "2012-08-06T15:42:07Z")

</div>

Converting to plain text before tokenizing can be a problem if it loses the  
natural delimiters that HTML tags provide. For example:

- New  
York
- City of Chicago
 should not be tokenized as New York City,  
of, Chicago but that could happen it you convert it to New York City of  
Chicago first.

Converting tag breaks to \n would probably help, but you have to be  
careful: do you do it with

? Probably. _? Probably not._

Is there already an analyzer that does the above or close to it?

My application is in Ruby -- I'm happy (prefer actually) to write my custom  
HTML parser in Ruby with Nokogiri. If I do that, what approaches should I  
consider when I connect to ES?

Thanks!

---

<div class="post-metadata">

**Author:** ![vineeth\_mohan](https://avatars.discourse-cdn.com/v4/letter/v/bc79bd/32.png) [@vineeth\_mohan](https://discuss.elastic.co/u/vineeth_mohan)\
**Post date:** [August 6, 2012, 3:51pm UTC](https://discuss.elastic.co/t/tokenizing-html/8648/2 "2012-08-06T15:51:25Z")

</div>

You can use html\_strip character filter -

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

This way you can preserve the html and at the same time make only stripped  
text search-able.

Thanks  
Vineeth

On Mon, Aug 6, 2012 at 9:12 PM, David James [davidcjames@gmail.com](mailto:davidcjames@gmail.com) wrote:

> Converting to plain text before tokenizing can be a problem if it loses  
> the natural delimiters that HTML tags provide. For example:
> 
> - New  
> York
> - City of Chicago
> should not be tokenized as New York City  
> , of, Chicago but that could happen it you convert it to New York City of  
> Chicago first.
> 
> Converting tag breaks to \n would probably help, but you have to be  
> careful: do you do it with
> 
> ? Probably. _? Probably not._
> 
>  
> 
> _Is there already an analyzer that does the above or close to it?_
> 
>  
> 
> _My application is in Ruby -- I'm happy (prefer actually) to write my  
> custom HTML parser in Ruby with Nokogiri. If I do that, what approaches  
> should I consider when I connect to ES?_
> 
>  
> 
> _Thanks!_

---

<div class="post-metadata">

**Author:** ![David\_James](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/david_james/32/2768_2.png) [@David\_James](https://discuss.elastic.co/u/David_James)\
**Post date:** [August 6, 2012, 6:55pm UTC](https://discuss.elastic.co/t/tokenizing-html/8648/3 "2012-08-06T18:55:33Z")

</div>

A filter can only run after a tokenizer. Right? (This is what reading  
[http://www.manning.com/hatcher3/](http://www.manning.com/hatcher3/) and  
[http://www.elasticsearch.org/guide/reference/index-modules/analysis/](http://www.elasticsearch.org/guide/reference/index-modules/analysis/) leads  
me to believe.) So, if the HTMLStripCharFilter just removes tags, it cannot  
possibly adjust the tokenization. Right? Therefore, it could not possibly  
handle the situation I described above:

- New York
- City of  
Chicago
. Right?

Only stripping tags -- without thinking about tokenization as well -- is  
naive in my opinion.

I'm new to ElasticSearch and Lucene, but I've read quite a bit. I would be  
surprised if the situation I described above is **not** handled properly...  
somewhere.

P.S. I just did a code skim of HTMLStripCharFilter and it scared the crap  
out of me in the "I didn't know Java could be this bad" kind of way. Then I  
realized it was generated by JFlex.

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [August 6, 2012, 8:44pm UTC](https://discuss.elastic.co/t/tokenizing-html/8648/4 "2012-08-06T20:44:28Z")

</div>

On Mon, 2012-08-06 at 11:55 -0700, David James wrote:

> A filter can only run after a tokenizer.

A token filter runs after a tokenizer. But it can transform the tokens  
which are finally indexed.

That said, this is a character filter, not a token filter, and it runs  
before the tokenizer.

clint

> Right? (This is what reading [Lucene in Action, Second Edition](http://www.manning.com/hatcher3/) and  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/)  
> leads me to believe.) So, if the HTMLStripCharFilter just removes  
> tags, it cannot possibly adjust the tokenization. Right? Therefore, it  
> could not possibly handle the situation I described above:
> 
> - New  
> York
> - City of Chicago
> . Right?
> 
> Only stripping tags -- without thinking about tokenization as well --  
> is naive in my opinion.
> 
> I'm new to Elasticsearch and Lucene, but I've read quite a bit. I  
> would be surprised if the situation I described above is **not**  
> handled properly... somewhere.
> 
> P.S. I just did a code skim of HTMLStripCharFilter and it scared the  
> crap out of me in the "I didn't know Java could be this bad" kind of  
> way. Then I realized it was generated by JFlex.

---

<div class="post-metadata">

**Author:** ![David\_James](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/david_james/32/2768_2.png) [@David\_James](https://discuss.elastic.co/u/David_James)\
**Post date:** [August 7, 2012, 4:06pm UTC](https://discuss.elastic.co/t/tokenizing-html/8648/5 "2012-08-07T16:06:47Z")

</div>

Thanks for clearing that up!

I am currently using Nokogiri (a Ruby library) to parse the HTML and remove  
the tags "intelligently". I wrote some custom logic that:

- liberally adds newlines ("\n") in the resulting plain text where  
appropriate, so that 
- New York
- \<City of Chicago
 becomes New  
York\n\nCity of Chicago\n\n- I want to go  
_fast_

Downstream (in Elasticsearch), I can pay attention to the double newline to  
help tokenization.

It does this by paying attention to three kinds of HTML elements:

1. elements that _separate_ tokens (e.g. article aside blockquote body  
button canvas caption center col colgroup dd dir div dl dt embed fieldset  
figcaption figure footer form h1 h2 h3 h4 h5 h6 header hgroup hr iframe  
label legend li menu nav noscript ol optgroup option output p pre progress  
samp section select table tbody td textarea tfoot th thead title tr ul video  
)
2. elements that _do not separate_ tokens (e.g. a abbr acronym address b  
bdo big br cite code dfn em font i input ins kbd nobr q script small span  
strong style time tt u var wbr)
3. elements that should be _skipped_ (e.g. applet area base basefont del  
img link map meta object param s strike sub sup) (Note that this ignores  
strikethrough text.)

These lists roughly correspond to block-level and inline elements. That  
said, I tweaked them based on documents I'm seeing. So, the lists are  
subjective. They do not understand changes effected by CSS -- such as  
changing a div to be an inline element (that is so evil!).

This is not fancy, but I wanted to share in case it is helpful and/or crazy.

-David

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:17am UTC](https://discuss.elastic.co/t/tokenizing-html/8648/6 "2017-07-06T03:17:28Z")

</div>


