# Indexing HTML

**URL:** <https://discuss.elastic.co/t/indexing-html/11661>\
**Category:** Elasticsearch\
**Created:** [April 22, 2013, 9:24pm UTC](https://discuss.elastic.co/t/indexing-html/11661 "2013-04-22T21:24:33Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Nathan\_Akeris](https://avatars.discourse-cdn.com/v4/letter/n/77aa72/32.png) [@Nathan\_Akeris](https://discuss.elastic.co/u/Nathan_Akeris)\
**Post date:** [April 22, 2013, 9:24pm UTC](https://discuss.elastic.co/t/indexing-html/11661/1 "2013-04-22T21:24:33Z")

</div>

Hi all,

Been trying to figure this out and seem to be having no luck, so I thought  
I'd throw my question up here.

Goal:

- I want to be able to store HTML in a document, but not have it indexed.  
For example, I don't want searches for "\<span style" to return results,  
but I want the raw HTML saved.

The use case is basically: We have a bunch of HTML documents, a user  
performs a search, these documents are rendered in some research results in  
their native HTML. I'd like to highlight the results but at this point I'd  
be content just not screwing up the search results.

Things I've tried:

Field mapping, changing the default analyzer, using html\_strip, setting  
include\_in\_all to false, setting index to no and store to yes, etc.

What ends up happening in most cases is ES just ignores what I'm doing  
(despite seeing the mapping correctly configured in the index), letting  
searches like \_all:\<span work when that field shouldn't even be indexed nor  
included in the \_all field. In some cases I'd lose the  
default analyzer and gain case sensitivity, so I stopped trying to do  
this globally.

This seems like something that should be straight forward but I've wasted a  
day on it. Any ideas? I've used the normal documentation, googled various  
groups, etc. and just can't seem to get it to work.

Thanks for your time all.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [April 22, 2013, 9:28pm UTC](https://discuss.elastic.co/t/indexing-html/11661/2 "2013-04-22T21:28:08Z")

</div>

I would probably store it as a binary BASE64 encoded content.  
That way, it won't be touched anyhow by Elasticsearch.

My 2 cents.

--  
David Pilato | Technical Advocate | [Elasticsearch.com](http://Elasticsearch.com)  
@dadoonet | @elasticsearchfr | @scrutmydocs

Le 22 avr. 2013 à 23:24, Nathan Akeris [nathan.akeris@gmail.com](mailto:nathan.akeris@gmail.com) a écrit :

> Hi all,
> 
> Been trying to figure this out and seem to be having no luck, so I thought I'd throw my question up here.
> 
> Goal:
> 
> - I want to be able to store HTML in a document, but not have it indexed. For example, I don't want searches for "\<span style" to return results, but I want the raw HTML saved.
> 
> The use case is basically: We have a bunch of HTML documents, a user performs a search, these documents are rendered in some research results in their native HTML. I'd like to highlight the results but at this point I'd be content just not screwing up the search results.
> 
> Things I've tried:
> 
> Field mapping, changing the default analyzer, using html\_strip, setting include\_in\_all to false, setting index to no and store to yes, etc.
> 
> What ends up happening in most cases is ES just ignores what I'm doing (despite seeing the mapping correctly configured in the index), letting searches like \_all:\<span work when that field shouldn't even be indexed nor included in the \_all field. In some cases I'd lose the default analyzer and gain case sensitivity, so I stopped trying to do this globally.
> 
> This seems like something that should be straight forward but I've wasted a day on it. Any ideas? I've used the normal documentation, googled various groups, etc. and just can't seem to get it to work.
> 
> Thanks for your time all.
> 
> --  
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Nathan\_Akeris](https://avatars.discourse-cdn.com/v4/letter/n/77aa72/32.png) [@Nathan\_Akeris](https://discuss.elastic.co/u/Nathan_Akeris)\
**Post date:** [April 22, 2013, 9:56pm UTC](https://discuss.elastic.co/t/indexing-html/11661/3 "2013-04-22T21:56:00Z")

</div>

Thanks, that worked well. I guess I'm still confused as to what went wrong  
before though - I'd like to use ES the way it's intended. Do these  
features not work as described, or is there some sort of trick to getting  
it to work? Has anyone else had these issues or is this type of request  
normally straight forward?

Thanks again.

On Monday, April 22, 2013 5:28:37 PM UTC-4, David Pilato wrote:

> I would probably store it as a binary BASE64 encoded content.  
> That way, it won't be touched anyhow by Elasticsearch.
> 
> My 2 cents.
> 
> --  
> _David Pilato_ | _Technical Advocate_ | _[Elasticsearch.com](http://Elasticsearch.com)_  
> @dadoonet [https://twitter.com/dadoonet](https://twitter.com/dadoonet) | @elasticsearchfr[https://twitter.com/elasticsearchfr](https://twitter.com/elasticsearchfr)  
> | @scrutmydocs [https://twitter.com/scrutmydocs](https://twitter.com/scrutmydocs)
> 
> Le 22 avr. 2013 à 23:24, Nathan Akeris \<[nathan...@gmail.com](mailto:nathan...@gmail.com) \<javascript:\>\>  
> a écrit :
> 
> Hi all,
> 
> Been trying to figure this out and seem to be having no luck, so I thought  
> I'd throw my question up here.
> 
> Goal:
> 
> - I want to be able to store HTML in a document, but not have it indexed.  
> For example, I don't want searches for "\<span style" to return results,  
> but I want the raw HTML saved.
> 
> The use case is basically: We have a bunch of HTML documents, a user  
> performs a search, these documents are rendered in some research results in  
> their native HTML. I'd like to highlight the results but at this point I'd  
> be content just not screwing up the search results.
> 
> Things I've tried:
> 
> Field mapping, changing the default analyzer, using html\_strip, setting  
> include\_in\_all to false, setting index to no and store to yes, etc.
> 
> What ends up happening in most cases is ES just ignores what I'm doing  
> (despite seeing the mapping correctly configured in the index), letting  
> searches like \_all:\<span work when that field shouldn't even be indexed nor  
> included in the \_all field. In some cases I'd lose the  
> default analyzer and gain case sensitivity, so I stopped trying to do  
> this globally.
> 
> This seems like something that should be straight forward but I've wasted  
> a day on it. Any ideas? I've used the normal documentation, googled  
> various groups, etc. and just can't seem to get it to work.
> 
> Thanks for your time all.
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![egaumer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/egaumer/32/2365_2.png) [@egaumer](https://discuss.elastic.co/u/egaumer)\
**Post date:** [April 22, 2013, 11:40pm UTC](https://discuss.elastic.co/t/indexing-html/11661/4 "2013-04-22T23:40:58Z")

</div>

If you just want to store the HTML content and not search on it (i.e., not  
indexed) then just set index=no and include\_in\_all=false in you mapping  
properties.

-Eric

On Monday, April 22, 2013 5:56:00 PM UTC-4, Nathan Akeris wrote:

> Thanks, that worked well. I guess I'm still confused as to what went  
> wrong before though - I'd like to use ES the way it's intended. Do these  
> features not work as described, or is there some sort of trick to getting  
> it to work? Has anyone else had these issues or is this type of request  
> normally straight forward?
> 
> Thanks again.
> 
> On Monday, April 22, 2013 5:28:37 PM UTC-4, David Pilato wrote:
> 
> > I would probably store it as a binary BASE64 encoded content.  
> > That way, it won't be touched anyhow by Elasticsearch.
> > 
> > My 2 cents.
> > 
> > --  
> > _David Pilato_ | _Technical Advocate_ | _[Elasticsearch.com](http://Elasticsearch.com)_  
> > @dadoonet [https://twitter.com/dadoonet](https://twitter.com/dadoonet) | @elasticsearchfr[https://twitter.com/elasticsearchfr](https://twitter.com/elasticsearchfr)  
> > | @scrutmydocs [https://twitter.com/scrutmydocs](https://twitter.com/scrutmydocs)
> > 
> > Le 22 avr. 2013 à 23:24, Nathan Akeris [nathan...@gmail.com](mailto:nathan...@gmail.com) a écrit :
> > 
> > Hi all,
> > 
> > Been trying to figure this out and seem to be having no luck, so I  
> > thought I'd throw my question up here.
> > 
> > Goal:
> > 
> > - I want to be able to store HTML in a document, but not have it  
> > indexed. For example, I don't want searches for "\<span style" to return  
> > results, but I want the raw HTML saved.
> > 
> > The use case is basically: We have a bunch of HTML documents, a user  
> > performs a search, these documents are rendered in some research results in  
> > their native HTML. I'd like to highlight the results but at this point I'd  
> > be content just not screwing up the search results.
> > 
> > Things I've tried:
> > 
> > Field mapping, changing the default analyzer, using html\_strip, setting  
> > include\_in\_all to false, setting index to no and store to yes, etc.
> > 
> > What ends up happening in most cases is ES just ignores what I'm doing  
> > (despite seeing the mapping correctly configured in the index), letting  
> > searches like \_all:\<span work when that field shouldn't even be indexed nor  
> > included in the \_all field. In some cases I'd lose the  
> > default analyzer and gain case sensitivity, so I stopped trying to do  
> > this globally.
> > 
> > This seems like something that should be straight forward but I've wasted  
> > a day on it. Any ideas? I've used the normal documentation, googled  
> > various groups, etc. and just can't seem to get it to work.
> > 
> > Thanks for your time all.
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Greg\_Brown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/greg_brown/32/1373_2.png) [@Greg\_Brown](https://discuss.elastic.co/u/Greg_Brown)\
**Post date:** [April 23, 2013, 1:09am UTC](https://discuss.elastic.co/t/indexing-html/11661/5 "2013-04-23T01:09:18Z")

</div>

ES isn't very aware of html tags other than the char filter being able to  
strip them.

You're going to run into problems with highlighting that I don't think can  
be worked around. The highlighting doesn't pay attention to existing HTML  
structures and so when highlighting tags are inserted they end up producing  
invalid HTML nesting. Stuff like **Match1Match2** NoMatch

The way I deal with this is all html I index gets run through:  
$clean\_content = html\_entity\_decode( strip\_tags( $content ) );

Then I use IDs attached to the ES documents in the search results to  
retrieve the original html. When I'm doing highlighting I ignore the  
original html structure. Not perfect, but it works pretty well and ensures  
that html tags never cause a problem.

-Greg

On Monday, April 22, 2013 3:24:33 PM UTC-6, Nathan Akeris wrote:

> Hi all,
> 
> Been trying to figure this out and seem to be having no luck, so I thought  
> I'd throw my question up here.
> 
> Goal:
> 
> - I want to be able to store HTML in a document, but not have it indexed.  
> For example, I don't want searches for "\<span style" to return results,  
> but I want the raw HTML saved.
> 
> The use case is basically: We have a bunch of HTML documents, a user  
> performs a search, these documents are rendered in some research results in  
> their native HTML. I'd like to highlight the results but at this point I'd  
> be content just not screwing up the search results.
> 
> Things I've tried:
> 
> Field mapping, changing the default analyzer, using html\_strip, setting  
> include\_in\_all to false, setting index to no and store to yes, etc.
> 
> What ends up happening in most cases is ES just ignores what I'm doing  
> (despite seeing the mapping correctly configured in the index), letting  
> searches like \_all:\<span work when that field shouldn't even be indexed nor  
> included in the \_all field. In some cases I'd lose the  
> default analyzer and gain case sensitivity, so I stopped trying to do  
> this globally.
> 
> This seems like something that should be straight forward but I've wasted  
> a day on it. Any ideas? I've used the normal documentation, googled  
> various groups, etc. and just can't seem to get it to work.
> 
> Thanks for your time all.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:40am UTC](https://discuss.elastic.co/t/indexing-html/11661/6 "2017-07-06T02:40:00Z")

</div>


