# Indexing of HTML content

**URL:** <https://discuss.elastic.co/t/indexing-of-html-content/3190>\
**Category:** Elasticsearch\
**Created:** [August 5, 2010, 5:50pm UTC](https://discuss.elastic.co/t/indexing-of-html-content/3190 "2010-08-05T17:50:03Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![James\_Cook](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@James\_Cook](https://discuss.elastic.co/u/James_Cook)\
**Post date:** [August 5, 2010, 5:50pm UTC](https://discuss.elastic.co/t/indexing-of-html-content/3190/1 "2010-08-05T17:50:03Z")

</div>

I see that the attachments plugin uses Tika under the hood to intelligently  
index text content embedded in other formats.

We have a situation where we are storing a field in our JSON object that can  
contain HTML content. (Think the content of a discussion thread.)

Is there a strategy of technique to make sure the HTML tags are not indexed?  
An existing mapping type?

---

<div class="post-metadata">

**Author:** ![Lukas\_Vlcek1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lukas_vlcek1/32/819_2.png) [@Lukas\_Vlcek1](https://discuss.elastic.co/u/Lukas_Vlcek1)\
**Post date:** [August 5, 2010, 6:27pm UTC](https://discuss.elastic.co/t/indexing-of-html-content/3190/2 "2010-08-05T18:27:19Z")

</div>

Hi,

AFAIK Tika is using TagSoup for parsing HTML documents. So you can either  
use directly TagSoup or you can try Tika's Swing GUI to test output of your  
documents first.  
See [Apache Tika – Getting Started with Apache Tika](http://tika.apache.org/0.7/gettingstarted.html) where in "Using Tika as a  
command line utility" you can find option -g or --gui

But if you have a json document that can contain HTML inside (as a value of  
some property) then I think you will have to do this manually. The way  
attachment plugin works is that it assumes that whole input document is of  
the same content-type, there is no support for parsing documents with nested  
content-types.

Regards,  
Lukas

On Thu, Aug 5, 2010 at 7:50 PM, James Cook [jcook@tracermedia.com](mailto:jcook@tracermedia.com) wrote:

> I see that the attachments plugin uses Tika under the hood to intelligently  
> index text content embedded in other formats.
> 
> We have a situation where we are storing a field in our JSON object that  
> can contain HTML content. (Think the content of a discussion thread.)
> 
> Is there a strategy of technique to make sure the HTML tags are not  
> indexed? An existing mapping type?

---

<div class="post-metadata">

**Author:** ![James\_Cook](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@James\_Cook](https://discuss.elastic.co/u/James_Cook)\
**Post date:** [August 6, 2010, 7:38pm UTC](https://discuss.elastic.co/t/indexing-of-html-content/3190/3 "2010-08-06T19:38:59Z")

</div>

It would be cool to add tagsoup as a mapping type. Would that be the most  
appropriate place for this functionality to reside?

-- jim

On Thu, Aug 5, 2010 at 2:27 PM, Lukáš Vlček [lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com) wrote:

> Hi,
> 
> AFAIK Tika is using TagSoup for parsing HTML documents. So you can either  
> use directly TagSoup or you can try Tika's Swing GUI to test output of your  
> documents first.  
> See [Apache Tika – Getting Started with Apache Tika](http://tika.apache.org/0.7/gettingstarted.html) where in "Using Tika as  
> a command line utility" you can find option -g or --gui
> 
> But if you have a json document that can contain HTML inside (as a value of  
> some property) then I think you will have to do this manually. The way  
> attachment plugin works is that it assumes that whole input document is of  
> the same content-type, there is no support for parsing documents with nested  
> content-types.
> 
> Regards,  
> Lukas
> 
> On Thu, Aug 5, 2010 at 7:50 PM, James Cook [jcook@tracermedia.com](mailto:jcook@tracermedia.com) wrote:
> 
> > I see that the attachments plugin uses Tika under the hood to  
> > intelligently index text content embedded in other formats.
> > 
> > We have a situation where we are storing a field in our JSON object that  
> > can contain HTML content. (Think the content of a discussion thread.)
> > 
> > Is there a strategy of technique to make sure the HTML tags are not  
> > indexed? An existing mapping type?

---

<div class="post-metadata">

**Author:** ![Lukas\_Vlcek1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lukas_vlcek1/32/819_2.png) [@Lukas\_Vlcek1](https://discuss.elastic.co/u/Lukas_Vlcek1)\
**Post date:** [August 6, 2010, 7:47pm UTC](https://discuss.elastic.co/t/indexing-of-html-content/3190/4 "2010-08-06T19:47:29Z")

</div>

Wouldn't it be easier to do it on the client side? I think doing it on the  
server side can be too complex: consider that if you add HTML as a type to  
mapping then you have to specify and handle a lot of other things because  
every html has some structure, it has head and body, the head usually  
contains a lot of metadata, body contains paragraphs, divs, headlines,  
links, ... etc.

Can you give some example how you would like to use it?

Regards,  
Lukas

On Fri, Aug 6, 2010 at 9:38 PM, James Cook [jcook@tracermedia.com](mailto:jcook@tracermedia.com) wrote:

> It would be cool to add tagsoup as a mapping type. Would that be the most  
> appropriate place for this functionality to reside?
> 
> -- jim
> 
> On Thu, Aug 5, 2010 at 2:27 PM, Lukáš Vlček [lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com) wrote:
> 
> > Hi,
> > 
> > AFAIK Tika is using TagSoup for parsing HTML documents. So you can either  
> > use directly TagSoup or you can try Tika's Swing GUI to test output of your  
> > documents first.  
> > See [Apache Tika – Getting Started with Apache Tika](http://tika.apache.org/0.7/gettingstarted.html) where in "Using Tika  
> > as a command line utility" you can find option -g or --gui
> > 
> > But if you have a json document that can contain HTML inside (as a value  
> > of some property) then I think you will have to do this manually. The way  
> > attachment plugin works is that it assumes that whole input document is of  
> > the same content-type, there is no support for parsing documents with nested  
> > content-types.
> > 
> > Regards,  
> > Lukas
> > 
> > On Thu, Aug 5, 2010 at 7:50 PM, James Cook [jcook@tracermedia.com](mailto:jcook@tracermedia.com) wrote:
> > 
> > > I see that the attachments plugin uses Tika under the hood to  
> > > intelligently index text content embedded in other formats.
> > > 
> > > We have a situation where we are storing a field in our JSON object that  
> > > can contain HTML content. (Think the content of a discussion thread.)
> > > 
> > > Is there a strategy of technique to make sure the HTML tags are not  
> > > indexed? An existing mapping type?

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [August 8, 2010, 3:00pm UTC](https://discuss.elastic.co/t/indexing-of-html-content/3190/5 "2010-08-08T15:00:19Z")

</div>

On Thu, 2010-08-05 at 13:50 -0400, James Cook wrote:

> I see that the attachments plugin uses Tika under the hood to  
> intelligently index text content embedded in other formats.

> We have a situation where we are storing a field in our JSON object  
> that can contain HTML content. (Think the content of a discussion  
> thread.)

> Is there a strategy of technique to make sure the HTML tags are not  
> indexed? An existing mapping type?

I think that this is an important use case, and I've opened this issue:

> <https://github.com/elastic/elasticsearch/issues/301>
>
> The need to index fields containing HTML is a frequent use case. I think that it… is important to have a built-in tokenizer which can handle HTML (ie remove tags and decode entities).
> 
> At the moment, in my client code I do the following:
> 
> \`\`\`
> $value =~ s/\<\[^\>\]+\>/ /g; # replace any \<.....\> extents with a single space
> $value =~ s/\\s+/ /g; # replace multiple spaces with a single space
> $value =~ s/^ //; # trim leading whitespace
> $value =~ s/ $//; # trim trailing whitespace
> decode\_entities($value); # translate all HTML entities to the equiv UTF-8 char
> \`\`\`
> 
> This is sufficient to convert HTML to text suitable for indexing by the default analyzer - doesn't need to do any more than this.
> 
> Any chance of getting this built in?

> 

clint

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 8, 2010, 9:23pm UTC](https://discuss.elastic.co/t/indexing-of-html-content/3190/6 "2010-08-08T21:23:13Z")

</div>

Agreed. In any case, I am leaning towards more native "attachment" like  
support for specific types, without using Tika. Planning to tackle this post  
0.9.1.

On Sun, Aug 8, 2010 at 6:00 PM, Clinton Gormley [clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)wrote:

> On Thu, 2010-08-05 at 13:50 -0400, James Cook wrote:
> 
> > I see that the attachments plugin uses Tika under the hood to  
> > intelligently index text content embedded in other formats.
> 
> > We have a situation where we are storing a field in our JSON object  
> > that can contain HTML content. (Think the content of a discussion  
> > thread.)
> 
> > Is there a strategy of technique to make sure the HTML tags are not  
> > indexed? An existing mapping type?
> 
> I think that this is an important use case, and I've opened this issue:
> 
> [HTML tokenizer · Issue #301 · elastic/elasticsearch · GitHub](http://github.com/elasticsearch/elasticsearch/issues/issue/301)
> 
> > 
> 
> clint

---

<div class="post-metadata">

**Author:** ![James\_Cook](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@James\_Cook](https://discuss.elastic.co/u/James_Cook)\
**Post date:** [August 12, 2010, 12:16pm UTC](https://discuss.elastic.co/t/indexing-of-html-content/3190/7 "2010-08-12T12:16:19Z")

</div>

Thanks for opening the feature request. We could implement this on our  
client side, but it would cause a large increase in our storage needs to  
store the HTML and also provide a field that can be tokenized containing the  
raw text content from that same HTML field.

Is it feasible for someone with little knowledge of ES internals to add a  
built-in tokenizer like this? (Not sure if tokenizer is the right  
terminology. Is it a mapping or an analyzer?)

On Sun, Aug 8, 2010 at 5:23 PM, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:

> Agreed. In any case, I am leaning towards more native "attachment" like  
> support for specific types, without using Tika. Planning to tackle this post  
> 0.9.1.
> 
> On Sun, Aug 8, 2010 at 6:00 PM, Clinton Gormley [clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)wrote:
> 
> > On Thu, 2010-08-05 at 13:50 -0400, James Cook wrote:
> > 
> > > I see that the attachments plugin uses Tika under the hood to  
> > > intelligently index text content embedded in other formats.
> > 
> > > We have a situation where we are storing a field in our JSON object  
> > > that can contain HTML content. (Think the content of a discussion  
> > > thread.)
> > 
> > > Is there a strategy of technique to make sure the HTML tags are not  
> > > indexed? An existing mapping type?
> > 
> > I think that this is an important use case, and I've opened this issue:
> > 
> > [HTML tokenizer · Issue #301 · elastic/elasticsearch · GitHub](http://github.com/elasticsearch/elasticsearch/issues/issue/301)
> > 
> > > 
> > 
> > clint

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 12, 2010, 12:32pm UTC](https://discuss.elastic.co/t/indexing-of-html-content/3190/8 "2010-08-12T12:32:09Z")

</div>

Do you just want to strip out the html characters, or also, as a result of  
the parsing of the html, add properties automatically like title, tags and  
so on (on top of the default body level text).

-shay.banon

On Thu, Aug 12, 2010 at 3:16 PM, James Cook [jcook@tracermedia.com](mailto:jcook@tracermedia.com) wrote:

> Thanks for opening the feature request. We could implement this on our  
> client side, but it would cause a large increase in our storage needs to  
> store the HTML and also provide a field that can be tokenized containing the  
> raw text content from that same HTML field.
> 
> Is it feasible for someone with little knowledge of ES internals to add a  
> built-in tokenizer like this? (Not sure if tokenizer is the right  
> terminology. Is it a mapping or an analyzer?)
> 
> On Sun, Aug 8, 2010 at 5:23 PM, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:
> 
> > Agreed. In any case, I am leaning towards more native "attachment" like  
> > support for specific types, without using Tika. Planning to tackle this post  
> > 0.9.1.
> > 
> > On Sun, Aug 8, 2010 at 6:00 PM, Clinton Gormley \<[clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)
> > 
> > > wrote:
> > 
> > > On Thu, 2010-08-05 at 13:50 -0400, James Cook wrote:
> > > 
> > > > I see that the attachments plugin uses Tika under the hood to  
> > > > intelligently index text content embedded in other formats.
> > > 
> > > > We have a situation where we are storing a field in our JSON object  
> > > > that can contain HTML content. (Think the content of a discussion  
> > > > thread.)
> > > 
> > > > Is there a strategy of technique to make sure the HTML tags are not  
> > > > indexed? An existing mapping type?
> > > 
> > > I think that this is an important use case, and I've opened this issue:
> > > 
> > > [HTML tokenizer · Issue #301 · elastic/elasticsearch · GitHub](http://github.com/elasticsearch/elasticsearch/issues/issue/301)
> > > 
> > > > 
> > > 
> > > clint

---

<div class="post-metadata">

**Author:** ![Lukas\_Vlcek1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lukas_vlcek1/32/819_2.png) [@Lukas\_Vlcek1](https://discuss.elastic.co/u/Lukas_Vlcek1)\
**Post date:** [August 12, 2010, 12:39pm UTC](https://discuss.elastic.co/t/indexing-of-html-content/3190/9 "2010-08-12T12:39:09Z")

</div>

Hi,

just a note, if the later is required then I think this can get more  
complex. Especially when you realize that HTML5 is adding a lot of new (and  
useful) stuff: [http://diveintohtml5.org/semantics.html#new-elements](http://diveintohtml5.org/semantics.html#new-elements)

[http://diveintohtml5.org/semantics.html#new-elements](http://diveintohtml5.org/semantics.html#new-elements)Lukas

On Thu, Aug 12, 2010 at 2:32 PM, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:

> Do you just want to strip out the html characters, or also, as a result of  
> the parsing of the html, add properties automatically like title, tags and  
> so on (on top of the default body level text).
> 
> -shay.banon
> 
> On Thu, Aug 12, 2010 at 3:16 PM, James Cook [jcook@tracermedia.com](mailto:jcook@tracermedia.com) wrote:
> 
> > Thanks for opening the feature request. We could implement this on our  
> > client side, but it would cause a large increase in our storage needs to  
> > store the HTML and also provide a field that can be tokenized containing the  
> > raw text content from that same HTML field.
> > 
> > Is it feasible for someone with little knowledge of ES internals to add a  
> > built-in tokenizer like this? (Not sure if tokenizer is the right  
> > terminology. Is it a mapping or an analyzer?)
> > 
> > On Sun, Aug 8, 2010 at 5:23 PM, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:
> > 
> > > Agreed. In any case, I am leaning towards more native "attachment" like  
> > > support for specific types, without using Tika. Planning to tackle this post  
> > > 0.9.1.
> > > 
> > > On Sun, Aug 8, 2010 at 6:00 PM, Clinton Gormley \<  
> > > [clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)\> wrote:
> > > 
> > > > On Thu, 2010-08-05 at 13:50 -0400, James Cook wrote:
> > > > 
> > > > > I see that the attachments plugin uses Tika under the hood to  
> > > > > intelligently index text content embedded in other formats.
> > > > 
> > > > > We have a situation where we are storing a field in our JSON object  
> > > > > that can contain HTML content. (Think the content of a discussion  
> > > > > thread.)
> > > > 
> > > > > Is there a strategy of technique to make sure the HTML tags are not  
> > > > > indexed? An existing mapping type?
> > > > 
> > > > I think that this is an important use case, and I've opened this issue:
> > > > 
> > > > [HTML tokenizer · Issue #301 · elastic/elasticsearch · GitHub](http://github.com/elasticsearch/elasticsearch/issues/issue/301)
> > > > 
> > > > > 
> > > > 
> > > > clint

---

<div class="post-metadata">

**Author:** ![James\_Cook](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@James\_Cook](https://discuss.elastic.co/u/James_Cook)\
**Post date:** [August 12, 2010, 12:41pm UTC](https://discuss.elastic.co/t/indexing-of-html-content/3190/10 "2010-08-12T12:41:34Z")

</div>

I was hoping for just the stripping of tags. I'm indexing html fragments  
that a user creates using a CMS. So they are editing using something like  
TinyMCE.

On Thu, Aug 12, 2010 at 8:32 AM, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:

> Do you just want to strip out the html characters, or also, as a result of  
> the parsing of the html, add properties automatically like title, tags and  
> so on (on top of the default body level text).
> 
> -shay.banon
> 
> On Thu, Aug 12, 2010 at 3:16 PM, James Cook [jcook@tracermedia.com](mailto:jcook@tracermedia.com) wrote:
> 
> > Thanks for opening the feature request. We could implement this on our  
> > client side, but it would cause a large increase in our storage needs to  
> > store the HTML and also provide a field that can be tokenized containing the  
> > raw text content from that same HTML field.
> > 
> > Is it feasible for someone with little knowledge of ES internals to add a  
> > built-in tokenizer like this? (Not sure if tokenizer is the right  
> > terminology. Is it a mapping or an analyzer?)
> > 
> > On Sun, Aug 8, 2010 at 5:23 PM, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:
> > 
> > > Agreed. In any case, I am leaning towards more native "attachment" like  
> > > support for specific types, without using Tika. Planning to tackle this post  
> > > 0.9.1.
> > > 
> > > On Sun, Aug 8, 2010 at 6:00 PM, Clinton Gormley \<  
> > > [clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)\> wrote:
> > > 
> > > > On Thu, 2010-08-05 at 13:50 -0400, James Cook wrote:
> > > > 
> > > > > I see that the attachments plugin uses Tika under the hood to  
> > > > > intelligently index text content embedded in other formats.
> > > > 
> > > > > We have a situation where we are storing a field in our JSON object  
> > > > > that can contain HTML content. (Think the content of a discussion  
> > > > > thread.)
> > > > 
> > > > > Is there a strategy of technique to make sure the HTML tags are not  
> > > > > indexed? An existing mapping type?
> > > > 
> > > > I think that this is an important use case, and I've opened this issue:
> > > > 
> > > > [HTML tokenizer · Issue #301 · elastic/elasticsearch · GitHub](http://github.com/elasticsearch/elasticsearch/issues/issue/301)
> > > > 
> > > > > 
> > > > 
> > > > clint

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 12, 2010, 3:23pm UTC](https://discuss.elastic.co/t/indexing-of-html-content/3190/11 "2010-08-12T15:23:57Z")

</div>

Done: [Analysis: Add `char_filter` on top of `tokenizer`, `filter`, and `analyzer`. Add an `html_strip` char filter and `standard_html_strip` analyzer · Issue #315 · elastic/elasticsearch · GitHub](http://github.com/elasticsearch/elasticsearch/issues/issue/315). Note,  
there is a new "class" of filters, called `char_filter`. You will need to  
create a custom analyzer, with the appropriate tokenizer and filters, and  
add the html\_strip as a char filter. The standard analyzer, for example, is  
composed of standard tokenizer, standard filter, lowercase filter, and stop  
filter.

-shay.banon

On Thu, Aug 12, 2010 at 3:41 PM, James Cook [jcook@tracermedia.com](mailto:jcook@tracermedia.com) wrote:

> I was hoping for just the stripping of tags. I'm indexing html fragments  
> that a user creates using a CMS. So they are editing using something like  
> TinyMCE.
> 
> On Thu, Aug 12, 2010 at 8:32 AM, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:
> 
> > Do you just want to strip out the html characters, or also, as a result of  
> > the parsing of the html, add properties automatically like title, tags and  
> > so on (on top of the default body level text).
> > 
> > -shay.banon
> > 
> > On Thu, Aug 12, 2010 at 3:16 PM, James Cook [jcook@tracermedia.com](mailto:jcook@tracermedia.com)wrote:
> > 
> > > Thanks for opening the feature request. We could implement this on our  
> > > client side, but it would cause a large increase in our storage needs to  
> > > store the HTML and also provide a field that can be tokenized containing the  
> > > raw text content from that same HTML field.
> > > 
> > > Is it feasible for someone with little knowledge of ES internals to add a  
> > > built-in tokenizer like this? (Not sure if tokenizer is the right  
> > > terminology. Is it a mapping or an analyzer?)
> > > 
> > > On Sun, Aug 8, 2010 at 5:23 PM, Shay Banon \<[shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)
> > > 
> > > > wrote:
> > > 
> > > > Agreed. In any case, I am leaning towards more native "attachment" like  
> > > > support for specific types, without using Tika. Planning to tackle this post  
> > > > 0.9.1.
> > > > 
> > > > On Sun, Aug 8, 2010 at 6:00 PM, Clinton Gormley \<  
> > > > [clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)\> wrote:
> > > > 
> > > > > On Thu, 2010-08-05 at 13:50 -0400, James Cook wrote:
> > > > > 
> > > > > > I see that the attachments plugin uses Tika under the hood to  
> > > > > > intelligently index text content embedded in other formats.
> > > > > 
> > > > > > We have a situation where we are storing a field in our JSON object  
> > > > > > that can contain HTML content. (Think the content of a discussion  
> > > > > > thread.)
> > > > > 
> > > > > > Is there a strategy of technique to make sure the HTML tags are not  
> > > > > > indexed? An existing mapping type?
> > > > > 
> > > > > I think that this is an important use case, and I've opened this issue:
> > > > > 
> > > > > [HTML tokenizer · Issue #301 · elastic/elasticsearch · GitHub](http://github.com/elasticsearch/elasticsearch/issues/issue/301)
> > > > > 
> > > > > > 
> > > > > 
> > > > > clint

---

<div class="post-metadata">

**Author:** ![James\_Cook](https://avatars.discourse-cdn.com/v4/letter/j/898d66/32.png) [@James\_Cook](https://discuss.elastic.co/u/James_Cook)\
**Post date:** [August 13, 2010, 1:10am UTC](https://discuss.elastic.co/t/indexing-of-html-content/3190/12 "2010-08-13T01:10:10Z")

</div>

Great! Thanks for jumping on this. It looks like you are stripping tags  
during streaming so the memory use is very efficient. Very nice.

-- jim

On Thu, Aug 12, 2010 at 11:23 AM, Shay Banon  
[shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:

> Done: [Analysis: Add `char_filter` on top of `tokenizer`, `filter`, and `analyzer`. Add an `html_strip` char filter and `standard_html_strip` analyzer · Issue #315 · elastic/elasticsearch · GitHub](http://github.com/elasticsearch/elasticsearch/issues/issue/315).  
> Note, there is a new "class" of filters, called `char_filter`. You will need  
> to create a custom analyzer, with the appropriate tokenizer and filters, and  
> add the html\_strip as a char filter. The standard analyzer, for example, is  
> composed of standard tokenizer, standard filter, lowercase filter, and stop  
> filter.
> 
> -shay.banon
> 
> On Thu, Aug 12, 2010 at 3:41 PM, James Cook [jcook@tracermedia.com](mailto:jcook@tracermedia.com) wrote:
> 
> > I was hoping for just the stripping of tags. I'm indexing html fragments  
> > that a user creates using a CMS. So they are editing using something like  
> > TinyMCE.
> > 
> > On Thu, Aug 12, 2010 at 8:32 AM, Shay Banon \<[shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)
> > 
> > > wrote:
> > 
> > > Do you just want to strip out the html characters, or also, as a result  
> > > of the parsing of the html, add properties automatically like title, tags  
> > > and so on (on top of the default body level text).
> > > 
> > > -shay.banon
> > > 
> > > On Thu, Aug 12, 2010 at 3:16 PM, James Cook [jcook@tracermedia.com](mailto:jcook@tracermedia.com)wrote:
> > > 
> > > > Thanks for opening the feature request. We could implement this on our  
> > > > client side, but it would cause a large increase in our storage needs to  
> > > > store the HTML and also provide a field that can be tokenized containing the  
> > > > raw text content from that same HTML field.
> > > > 
> > > > Is it feasible for someone with little knowledge of ES internals to add  
> > > > a built-in tokenizer like this? (Not sure if tokenizer is the right  
> > > > terminology. Is it a mapping or an analyzer?)
> > > > 
> > > > On Sun, Aug 8, 2010 at 5:23 PM, Shay Banon \<  
> > > > [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)\> wrote:
> > > > 
> > > > > Agreed. In any case, I am leaning towards more native "attachment" like  
> > > > > support for specific types, without using Tika. Planning to tackle this post  
> > > > > 0.9.1.
> > > > > 
> > > > > On Sun, Aug 8, 2010 at 6:00 PM, Clinton Gormley \<  
> > > > > [clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)\> wrote:
> > > > > 
> > > > > > On Thu, 2010-08-05 at 13:50 -0400, James Cook wrote:
> > > > > > 
> > > > > > > I see that the attachments plugin uses Tika under the hood to  
> > > > > > > intelligently index text content embedded in other formats.
> > > > > > 
> > > > > > > We have a situation where we are storing a field in our JSON object  
> > > > > > > that can contain HTML content. (Think the content of a discussion  
> > > > > > > thread.)
> > > > > > 
> > > > > > > Is there a strategy of technique to make sure the HTML tags are not  
> > > > > > > indexed? An existing mapping type?
> > > > > > 
> > > > > > I think that this is an important use case, and I've opened this  
> > > > > > issue:
> > > > > > 
> > > > > > [HTML tokenizer · Issue #301 · elastic/elasticsearch · GitHub](http://github.com/elasticsearch/elasticsearch/issues/issue/301)
> > > > > > 
> > > > > > > 
> > > > > > 
> > > > > > clint

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:20am UTC](https://discuss.elastic.co/t/indexing-of-html-content/3190/13 "2017-07-06T04:20:51Z")

</div>


