# Strip\_HTML on indexing does not store results?

**URL:** <https://discuss.elastic.co/t/strip-html-on-indexing-does-not-store-results/4577>\
**Category:** Elasticsearch\
**Created:** [June 8, 2011, 3:47pm UTC](https://discuss.elastic.co/t/strip-html-on-indexing-does-not-store-results/4577 "2011-06-08T15:47:16Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![phobos182](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/phobos182/32/3011_2.png) [@phobos182](https://discuss.elastic.co/u/phobos182)\
**Post date:** [June 8, 2011, 3:47pm UTC](https://discuss.elastic.co/t/strip-html-on-indexing-does-not-store-results/4577/1 "2011-06-08T15:47:16Z")

</div>

Here is a copy of my analyzer which includes the strip\_html character filter. When retrieving documents from a field stored with this analyzer, it looks like the HTML codes are still in the document. Does the strip\_html just stip the text for term indexing, or does it strip it from the content before it is stored? I was expecting the document field to be retrieved without any HTML markup.

Thanks,

# -- Config --

curl -XPOST '[http://localhost:9200/test](http://localhost:9200/test)' -d '  
{  
"settings": {  
"index": {  
"analysis": {  
"analyzer": {  
"test\_analyzer": {  
"type": "custom",  
"char\_filter": [  
"scrub\_html"  
],  
"tokenizer": "standard",  
"filter": [  
"standard",  
"lowercase"  
]  
}  
},  
"char\_filter": {  
"scrub\_html": {  
"type": "html\_strip",  
"read\_ahead": 4096  
}  
}  
}  
}  
},  
"mappings": {  
"media": {  
"\_source": {  
"compress": true  
},  
"\_size": {  
"enabled": true,  
"store": "yes"  
},  
"properties": {  
"content": {  
"include\_in\_all": true,  
"omit\_norms": true,  
"store": "yes",  
"null\_value": "na",  
"analyzer": "test\_analyzer",  
"term\_vector": "with\_positions\_offsets",  
"type": "string"  
}  
}  
}  
}  
}  
'  
curl -XPOST '[http://localhost:9200/test/media](http://localhost:9200/test/media)' -d '  
{  
"content": "

Thank you!

"  
}  
'  
curl -XGET '[http://localhost:9200/test/\_search?q=\*&fields=content](http://localhost:9200/test/_search?q=*&fields=content)'  
{  
"took": 2,  
"timed\_out": false,  
"\_shards": {  
"total": 8,  
"successful": 8,  
"failed": 0  
},  
"hits": {  
"total": 1,  
"max\_score": 1,  
"hits": [  
{  
"\_index": "test",  
"\_type": "media",  
"\_id": "INyEgpcISFOOb8QrBUuUKQ",  
"\_score": 1,  
"fields": {  
"content": "

Thank you!

"  
}  
}  
]  
}  
}

---

<div class="post-metadata">

**Author:** ![Greg\_Brown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/greg_brown/32/1373_2.png) [@Greg\_Brown](https://discuss.elastic.co/u/Greg_Brown)\
**Post date:** [June 8, 2011, 11:01pm UTC](https://discuss.elastic.co/t/strip-html-on-indexing-does-not-store-results/4577/2 "2011-06-08T23:01:39Z")

</div>

I had run into the same problem before. The strip\_html does not save  
the document with the tags stripped out so any highlights you do from  
the index (for instance) will still contain the original html. I  
solved it by stripping html tags myself before indexing the document.  
-Greg

On Jun 8, 9:47 am, phobos182 [phobos...@gmail.com](mailto:phobos...@gmail.com) wrote:

> Here is a copy of my analyzer which includes the strip\_html character filter.  
> When retrieving documents from a field stored with this analyzer, it looks  
> like the HTML codes are still in the document. Does the strip\_html just stip  
> the text for term indexing, or does it strip it from the content before it  
> is stored? I was expecting the document field to be retrieved without any  
> HTML markup.
> 
> Thanks,
> 
> # -- Config --
> 
> curl -XPOST '[http://localhost:9200/test'-d](http://localhost:9200/test'-d) '  
> {  
> "settings": {  
> "index": {  
> "analysis": {  
> "analyzer": {  
> "test\_analyzer": {  
> "type": "custom",  
> "char\_filter": [  
> "scrub\_html"  
> ],  
> "tokenizer": "standard",  
> "filter": [  
> "standard",  
> "lowercase"  
> ]  
> }  
> },  
> "char\_filter": {  
> "scrub\_html": {  
> "type": "html\_strip",  
> "read\_ahead": 4096  
> }  
> }  
> }  
> }  
> },  
> "mappings": {  
> "media": {  
> "\_source": {  
> "compress": true  
> },  
> "\_size": {  
> "enabled": true,  
> "store": "yes"  
> },  
> "properties": {  
> "content": {  
> "include\_in\_all": true,  
> "omit\_norms": true,  
> "store": "yes",  
> "null\_value": "na",  
> "analyzer": "test\_analyzer",  
> "term\_vector": "with\_positions\_offsets",  
> "type": "string"  
> }  
> }  
> }  
> }}
> 
> '  
> curl -XPOST '[http://localhost:9200/test/media'-d](http://localhost:9200/test/media'-d) '  
> {  
> "content": "
> 
> Thank you!
> 
> "}
> 
> '  
> curl -XGET '[http://localhost:9200/test/\_search?q=\*&fields=content](http://localhost:9200/test/_search?q=*&fields=content)'  
> {  
> "took": 2,  
> "timed\_out": false,  
> "\_shards": {  
> "total": 8,  
> "successful": 8,  
> "failed": 0  
> },  
> "hits": {  
> "total": 1,  
> "max\_score": 1,  
> "hits": [  
> {  
> "\_index": "test",  
> "\_type": "media",  
> "\_id": "INyEgpcISFOOb8QrBUuUKQ",  
> "\_score": 1,  
> "fields": {  
> "content": "
> 
> Thank you!
> 
> "  
> }  
> }  
> ]  
> }
> 
> }
> 
> --  
> View this message in context:[http://elasticsearch-users.115913.n3.nabble.com/Strip-HTML-on-indexin](http://elasticsearch-users.115913.n3.nabble.com/Strip-HTML-on-indexin)...  
> Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).

---

<div class="post-metadata">

**Author:** ![fashionalwallet](https://avatars.discourse-cdn.com/v4/letter/f/839c29/32.png) [@fashionalwallet](https://discuss.elastic.co/u/fashionalwallet)\
**Post date:** [June 9, 2011, 9:05am UTC](https://discuss.elastic.co/t/strip-html-on-indexing-does-not-store-results/4577/3 "2011-06-09T09:05:32Z")

</div>

- deleted -

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [June 9, 2011, 6:57pm UTC](https://discuss.elastic.co/t/strip-html-on-indexing-does-not-store-results/4577/4 "2011-06-09T18:57:37Z")

</div>

Btw, we can have an html "type", which will strip the content and store it as is (and index it as well).

On Thursday, June 9, 2011 at 2:01 AM, Greg B wrote:

> I had run into the same problem before. The strip\_html does not save  
> the document with the tags stripped out so any highlights you do from  
> the index (for instance) will still contain the original html. I  
> solved it by stripping html tags myself before indexing the document.  
> -Greg
> 
> On Jun 8, 9:47 am, phobos182 \<[phobos...@gmail.com](mailto:phobos...@gmail.com) ([http://gmail.com](http://gmail.com))\> wrote:
> 
> > Here is a copy of my analyzer which includes the strip\_html character filter.  
> > When retrieving documents from a field stored with this analyzer, it looks  
> > like the HTML codes are still in the document. Does the strip\_html just stip  
> > the text for term indexing, or does it strip it from the content before it  
> > is stored? I was expecting the document field to be retrieved without any  
> > HTML markup.
> > 
> > Thanks,
> > 
> > # -- Config --
> > 
> > curl -XPOST '[http://localhost:9200/test'-d](http://localhost:9200/test'-d) '  
> > {  
> > "settings": {  
> > "index": {  
> > "analysis": {  
> > "analyzer": {  
> > "test\_analyzer": {  
> > "type": "custom",  
> > "char\_filter": [  
> > "scrub\_html"  
> > ],  
> > "tokenizer": "standard",  
> > "filter": [  
> > "standard",  
> > "lowercase"  
> > ]  
> > }  
> > },  
> > "char\_filter": {  
> > "scrub\_html": {  
> > "type": "html\_strip",  
> > "read\_ahead": 4096  
> > }  
> > }  
> > }  
> > }  
> > },  
> > "mappings": {  
> > "media": {  
> > "\_source": {  
> > "compress": true  
> > },  
> > "\_size": {  
> > "enabled": true,  
> > "store": "yes"  
> > },  
> > "properties": {  
> > "content": {  
> > "include\_in\_all": true,  
> > "omit\_norms": true,  
> > "store": "yes",  
> > "null\_value": "na",  
> > "analyzer": "test\_analyzer",  
> > "term\_vector": "with\_positions\_offsets",  
> > "type": "string"  
> > }  
> > }  
> > }  
> > }}
> > 
> > '  
> > curl -XPOST '[http://localhost:9200/test/media'-d](http://localhost:9200/test/media'-d) '  
> > {  
> > "content": "
> > 
> > Thank you!
> > 
> > "}
> > 
> > '  
> > curl -XGET '[http://localhost:9200/test/\_search?q=\*&fields=content](http://localhost:9200/test/_search?q=*&fields=content)'  
> > {  
> > "took": 2,  
> > "timed\_out": false,  
> > "\_shards": {  
> > "total": 8,  
> > "successful": 8,  
> > "failed": 0  
> > },  
> > "hits": {  
> > "total": 1,  
> > "max\_score": 1,  
> > "hits": [  
> > {  
> > "\_index": "test",  
> > "\_type": "media",  
> > "\_id": "INyEgpcISFOOb8QrBUuUKQ",  
> > "\_score": 1,  
> > "fields": {  
> > "content": "
> > 
> > Thank you!
> > 
> > "  
> > }  
> > }  
> > ]  
> > }
> > 
> > }
> > 
> > --  
> > View this message in context:[http://elasticsearch-users.115913.n3.nabble.com/Strip-HTML-on-indexin](http://elasticsearch-users.115913.n3.nabble.com/Strip-HTML-on-indexin)...  
> > Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com) ([http://Nabble.com](http://Nabble.com)).

---

<div class="post-metadata">

**Author:** ![phobos182](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/phobos182/32/3011_2.png) [@phobos182](https://discuss.elastic.co/u/phobos182)\
**Post date:** [June 9, 2011, 7:34pm UTC](https://discuss.elastic.co/t/strip-html-on-indexing-does-not-store-results/4577/5 "2011-06-09T19:34:16Z")

</div>

Having a core type of "html" would be a big convenience factor. The \_source field could contain the raw document (with markup), and leave the fields as scrubbed and stripped. I would get the best of both worlds by having my terms not contain markup for tag clouds, and the stored body not having markup for highlighting.

---

<div class="post-metadata">

**Author:** ![karmi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/karmi/32/44951_2.png) [@karmi](https://discuss.elastic.co/u/karmi)\
**Post date:** [June 10, 2011, 10:06am UTC](https://discuss.elastic.co/t/strip-html-on-indexing-does-not-store-results/4577/6 "2011-06-10T10:06:53Z")

</div>

That's a great idea, I've talked to many people who would seriously  
enjoy this.

On Jun 9, 8:57 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> Btw, we can have an html "type", which will strip the content and store it as is (and index it as well).

---

<div class="post-metadata">

**Author:** ![Lukas\_Vlcek1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lukas_vlcek1/32/819_2.png) [@Lukas\_Vlcek1](https://discuss.elastic.co/u/Lukas_Vlcek1)\
**Post date:** [June 10, 2011, 11:21am UTC](https://discuss.elastic.co/t/strip-html-on-indexing-does-not-store-results/4577/7 "2011-06-10T11:21:57Z")

</div>

+1!

It would be good to have some configuration options here however. Having an  
option to tell the html cleaner which html tags to remove and which to keep  
when storing the original html field content could be very useful (it can be  
handy for document preview).

I think that jsoup could be used for this. It has a nice API for cleaning  
HTML and allows to specify tag set to be remove (can be also customized).

Check  
[Jsoup: jsoup HTML Parser Documentation](http://jsoup.org/apidocs/org/jsoup/Jsoup.html#clean)(java.lang.String,  
org.jsoup.safety.Whitelist)  
[Safelist: jsoup HTML Parser Documentation](http://jsoup.org/apidocs/org/jsoup/safety/Whitelist.html)

Just my cents.  
Lukas

On Fri, Jun 10, 2011 at 12:06 PM, Karel Minarik [karel.minarik@gmail.com](mailto:karel.minarik@gmail.com)wrote:

> That's a great idea, I've talked to many people who would seriously  
> enjoy this.
> 
> On Jun 9, 8:57 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > Btw, we can have an html "type", which will strip the content and store  
> > it as is (and index it as well).

---

<div class="post-metadata">

**Author:** ![Administrator\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/administrator_2/32/3211_2.png) [@Administrator\_2](https://discuss.elastic.co/u/Administrator_2)\
**Post date:** [June 10, 2011, 11:29am UTC](https://discuss.elastic.co/t/strip-html-on-indexing-does-not-store-results/4577/8 "2011-06-10T11:29:17Z")

</div>

Definitely; while .NET has some great support for HTML in the form of the HTML Agility Pack (great for stripping documents) it would be great to have ES have intimate knowledge of this document type.

I assume storage of the original document would be provided on top of the parses version?

- Nick

On Jun 10, 2011, at 7:21 AM, Lukáš Vlček [lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com) wrote:

> +1!
> 
> It would be good to have some configuration options here however. Having an option to tell the html cleaner which html tags to remove and which to keep when storing the original html field content could be very useful (it can be handy for document preview).
> 
> I think that jsoup could be used for this. It has a nice API for cleaning HTML and allows to specify tag set to be remove (can be also customized).
> 
> Check  
> [Jsoup: jsoup HTML Parser Documentation](http://jsoup.org/apidocs/org/jsoup/Jsoup.html#clean)(java.lang.String, org.jsoup.safety.Whitelist)  
> [Safelist: jsoup HTML Parser Documentation](http://jsoup.org/apidocs/org/jsoup/safety/Whitelist.html)
> 
> Just my cents.  
> Lukas
> 
> On Fri, Jun 10, 2011 at 12:06 PM, Karel Minarik [karel.minarik@gmail.com](mailto:karel.minarik@gmail.com) wrote:  
> That's a great idea, I've talked to many people who would seriously  
> enjoy this.
> 
> On Jun 9, 8:57 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > Btw, we can have an html "type", which will strip the content and store it as is (and index it as well).

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [June 10, 2011, 10:32pm UTC](https://discuss.elastic.co/t/strip-html-on-indexing-does-not-store-results/4577/9 "2011-06-10T22:32:18Z")

</div>

Make sense, lets open an issue so we can keep track of this. It should be pretty simple to add an html type (even as a plugin, similar to the attachments one). If someone is up for the challenge, I am here to help!

On Friday, June 10, 2011 at 2:29 PM, administrator wrote:

> Definitely; while .NET has some great support for HTML in the form of the HTML Agility Pack (great for stripping documents) it would be great to have ES have intimate knowledge of this document type.
> 
> I assume storage of the original document would be provided on top of the parses version?
> 
> - Nick
> 
> On Jun 10, 2011, at 7:21 AM, Lukáš Vlček \<[lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com) ([mailto:lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com))\> wrote:
> 
> > +1!
> > 
> > It would be good to have some configuration options here however. Having an option to tell the html cleaner which html tags to remove and which to keep when storing the original html field content could be very useful (it can be handy for document preview).
> > 
> > I think that jsoup could be used for this. It has a nice API for cleaning HTML and allows to specify tag set to be remove (can be also customized).
> > 
> > Check  
> > [Jsoup: jsoup HTML Parser Documentation](http://jsoup.org/apidocs/org/jsoup/Jsoup.html#clean)(java.lang.String, org.jsoup.safety.Whitelist)  
> > [Safelist: jsoup HTML Parser Documentation](http://jsoup.org/apidocs/org/jsoup/safety/Whitelist.html)
> > 
> > Just my cents.  
> > Lukas
> > 
> > On Fri, Jun 10, 2011 at 12:06 PM, Karel Minarik \<[karel.minarik@gmail.com](mailto:karel.minarik@gmail.com) ([mailto:karel.minarik@gmail.com](mailto:karel.minarik@gmail.com))\> wrote:
> > 
> > > That's a great idea, I've talked to many people who would seriously  
> > > enjoy this.
> > > 
> > > On Jun 9, 8:57 pm, Shay Banon \<[shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) ([mailto:shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com))\> wrote:
> > > 
> > > > Btw, we can have an html "type", which will strip the content and store it as is (and index it as well).

---

<div class="post-metadata">

**Author:** ![Lukas\_Vlcek1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lukas_vlcek1/32/819_2.png) [@Lukas\_Vlcek1](https://discuss.elastic.co/u/Lukas_Vlcek1)\
**Post date:** [June 13, 2011, 7:43am UTC](https://discuss.elastic.co/t/strip-html-on-indexing-does-not-store-results/4577/10 "2011-06-13T07:43:48Z")

</div>

Just created a ticket for it:

> <https://github.com/elastic/elasticsearch/issues/1026>
>
> Provide out of the box support for fields with HTML content. The goal would be t…o allow to store both original and cleaned content (for highlighting or document preview). HTML cleaner should be configurable.
> 
> More detailed discussion can be found in ML here: http://elasticsearch-users.115913.n3.nabble.com/Strip-HTML-on-indexing-does-not-store-results-td3039614.html

On Sat, Jun 11, 2011 at 12:32 AM, Shay Banon  
[shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:

> Make sense, lets open an issue so we can keep track of this. It should be  
> pretty simple to add an html type (even as a plugin, similar to the  
> attachments one). If someone is up for the challenge, I am here to help!
> 
> On Friday, June 10, 2011 at 2:29 PM, administrator wrote:
> 
> Definitely; while .NET has some great support for HTML in the form of the  
> HTML Agility Pack (great for stripping documents) it would be great to have  
> ES have intimate knowledge of this document type.
> 
> I assume storage of the original document would be provided on top of the  
> parses version?
> 
> - Nick
> 
> On Jun 10, 2011, at 7:21 AM, Lukáš Vlček [lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com) wrote:
> 
> +1!
> 
> It would be good to have some configuration options here however. Having an  
> option to tell the html cleaner which html tags to remove and which to keep  
> when storing the original html field content could be very useful (it can be  
> handy for document preview).
> 
> I think that jsoup could be used for this. It has a nice API for cleaning  
> HTML and allows to specify tag set to be remove (can be also customized).
> 
> Check
> 
> [http://jsoup.org/apidocs/org/jsoup/Jsoup.html#clean(java.lang.String,%20org.jsoup.safety.Whitelist)](http://jsoup.org/apidocs/org/jsoup/Jsoup.html#clean(java.lang.String,%20org.jsoup.safety.Whitelist))  
> [Jsoup: jsoup HTML Parser Documentation](http://jsoup.org/apidocs/org/jsoup/Jsoup.html#clean)(java.lang.String,  
> org.jsoup.safety.Whitelist)  
> [http://jsoup.org/apidocs/org/jsoup/safety/Whitelist.html](http://jsoup.org/apidocs/org/jsoup/safety/Whitelist.html)  
> [Safelist: jsoup HTML Parser Documentation](http://jsoup.org/apidocs/org/jsoup/safety/Whitelist.html)
> 
> Just my cents.  
> Lukas
> 
> On Fri, Jun 10, 2011 at 12:06 PM, Karel Minarik \<[karel.minarik@gmail.com](mailto:karel.minarik@gmail.com)  
> [karel.minarik@gmail.com](mailto:karel.minarik@gmail.com)\> wrote:
> 
> That's a great idea, I've talked to many people who would seriously  
> enjoy this.
> 
> On Jun 9, 8:57 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > Btw, we can have an html "type", which will strip the content and store  
> > it as is (and index it as well).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:03am UTC](https://discuss.elastic.co/t/strip-html-on-indexing-does-not-store-results/4577/11 "2017-07-06T04:03:44Z")

</div>


