# Help stripping HTML tags

**URL:** <https://discuss.elastic.co/t/help-stripping-html-tags/4528>\
**Category:** Elasticsearch\
**Created:** [June 1, 2011, 11:32pm UTC](https://discuss.elastic.co/t/help-stripping-html-tags/4528 "2011-06-01T23:32:54Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Greg\_Brown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/greg_brown/32/1373_2.png) [@Greg\_Brown](https://discuss.elastic.co/u/Greg_Brown)\
**Post date:** [June 1, 2011, 11:32pm UTC](https://discuss.elastic.co/t/help-stripping-html-tags/4528/1 "2011-06-01T23:32:54Z")

</div>

I've configured a default custom analyzer as follows:

index :  
analysis :  
filter :  
snowball :  
type : snowball  
language : English  
wd\_filter :  
type : word\_delimiter  
generate\_word\_parts : true  
generate\_number\_parts : true  
catenate\_words : true  
split\_on\_case\_change : true  
preserve\_original : true  
split\_on\_numerics: true  
analyzer :  
default :  
type : custom  
tokenizer : uax\_url\_email  
filter : [lowercase,snowball,wd\_filter]  
char\_filter : [html\_strip]

But when I index a doc into 'content' and then examine the indexed  
files with:  
curl -XGET '[http://localhost:9200/mgs/p2/385?fields=content&pretty](http://localhost:9200/mgs/p2/385?fields=content&pretty)'

I still am seeing all of the html tags in the text. Reading back the  
\_mapping for the "content" field it is:  
"content" : {  
"store" : "yes",  
"type" : "string"  
},  
'index' defaults to 'analyzed', so this appears correct.

Further complicating is that if I run:  
curl -XGET '[http://localhost:9200/mgs/\_analyze](http://localhost:9200/mgs/_analyze)' -d '

this is  
a tests

'

Then I get:  
{"tokens":[{"token":"this","start\_offset":3,"end\_offset":  
7,"type":"","position":1},{"token":"is","start\_offset":  
11,"end\_offset":13,"type":"","position":2},  
{"token":"a","start\_offset":18,"end\_offset":  
19,"type":"","position":3},{"token":"test","start\_offset":  
20,"end\_offset":25,"type":"","position":4}]}

Which appears correct as the tags have been stripped out by using the  
default analyzer, and the word 'tests' has been stemmed to 'test'.

So the analyzer seems to be working correctly except for when I  
actually add a document. I am using Elastica for adding documents,  
but I tried

> curl -XPUT '[http://localhost:9200/mgs/p2/0](http://localhost:9200/mgs/p2/0)' -d '{"content" : "
> 
> This is a tests
> 
> " }'

Then:

> curl -XGET '[http://localhost:9200/mgs/p2/0?fields=content](http://localhost:9200/mgs/p2/0?fields=content)'  
> Gives:  
> {"\_index":"mgs","\_type":"p2","\_id":"0","\_version":2,"fields":  
> {"content":"
> 
> This is a tests
> 
> "}}

So adding the document does not seem to be causing the analyzer to be  
run on the added document. Any ideas on what I am missing/doing  
wrong?

Thanks  
-Greg

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [June 2, 2011, 7:42am UTC](https://discuss.elastic.co/t/help-stripping-html-tags/4528/2 "2011-06-02T07:42:55Z")

</div>

The stored version of a field stores the content as is, and not the analyzed form. The analysis process only controls how the text is broken down into terms and indexed.

On Thursday, June 2, 2011 at 2:32 AM, Greg B wrote:

> I've configured a default custom analyzer as follows:
> 
> index :  
> analysis :  
> filter :  
> snowball :  
> type : snowball  
> language : English  
> wd\_\_filter :  
> type : word\_delimiter  
> generate\_word\_parts : true  
> generate\_number\_parts : true  
> catenate\_\_words : true  
> split\_on\_case\_change : true  
> preserve\_original : true  
> split\_on\_numerics: true  
> analyzer :  
> default :  
> type : custom  
> tokenizer : uax\_\_url\_email  
> filter : [lowercase,snowball,wd\_filter]  
> char\_filter : [html\_strip]
> 
> But when I index a doc into 'content' and then examine the indexed  
> files with:  
> curl -XGET '[http://localhost:9200/mgs/p2/385?fields=content&pretty](http://localhost:9200/mgs/p2/385?fields=content&pretty)'
> 
> I still am seeing all of the html tags in the text. Reading back the  
> \_mapping for the "content" field it is:  
> "content" : {  
> "store" : "yes",  
> ""type" : "string"  
> },  
> 'index' defaults to 'analyzed', so this appears correct.
> 
> Further complicating is that if I run:  
> curl -XGET '[http://localhost:9200/mgs/\_analyze](http://localhost:9200/mgs/_analyze)' -d '
> 
> this is  
> a tests
> 
> '
> 
> Then I get:  
> {"tokens":[{"token":"this","start\_offset":3,"end\_offset":  
> 7,"type":"","position":1},{"token":"is","start\_offset":  
> 11,"end\_offset":13,"type":"","position":2},  
> {"token":"a","start\_offset":18,"end\_offset":  
> 19,"type":"","position":3},{"token":"test","start\_offset":  
> 20,"end\_offset":25,"type":"","position":4}]}
> 
> Which appears correct as the tags have been stripped out by using the  
> default analyzer, and the word 'tests' has been stemmed to 'test'.
> 
> So the analyzer seems to be working correctly except for when I  
> actually add a document. I am using Elastica for adding documents,  
> but I tried
> 
> > curl -XPUT '[http://localhost:9200/mgs/p2/0](http://localhost:9200/mgs/p2/0)' -d '{"content" : "
> > 
> > This is a tests
> > 
> > " }'
> 
> Then:
> 
> > curl -XGET '[http://localhost:9200/mgs/p2/0?fields=content](http://localhost:9200/mgs/p2/0?fields=content)'  
> > Gives:  
> > {"\_index":"mgs","\_type":"p2","\_id":"0","\_version":2,"fields":  
> > {"content":"
> > 
> > This is a tests
> > 
> > "}}
> 
> So adding the document does not seem to be causing the analyzer to be  
> run on the added document. Any ideas on what I am missing/doing  
> wrong?
> 
> Thanks  
> -Greg

---

<div class="post-metadata">

**Author:** ![Greg\_Brown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/greg_brown/32/1373_2.png) [@Greg\_Brown](https://discuss.elastic.co/u/Greg_Brown)\
**Post date:** [June 2, 2011, 3:56pm UTC](https://discuss.elastic.co/t/help-stripping-html-tags/4528/3 "2011-06-02T15:56:15Z")

</div>

Hi Shay, thanks for the fast response.

Is there a way to store the version with the html stripped, or do I  
need to implement my own stripping to remove the html tags. The tags  
are particularly a problem when trying to do highlighting in my search  
results where the extraneous tags rather thoroughly screw up my  
results page.

Thanks  
-Greg

On Jun 2, 1:42 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> The stored version of a field stores the content as is, and not the analyzed form. The analysis process only controls how the text is broken down into terms and indexed.
> 
> On Thursday, June 2, 2011 at 2:32 AM, Greg B wrote:
> 
> > I've configured a default custom analyzer as follows:
> 
> > index :  
> > analysis :  
> > filter :  
> > snowball :  
> > type : snowball  
> > language : English  
> > wd\_\_filter :  
> > type : word\_delimiter  
> > generate\_word\_parts : true  
> > generate\_number\_parts : true  
> > catenate\_\_words : true  
> > split\_on\_case\_change : true  
> > preserve\_original : true  
> > split\_on\_numerics: true  
> > analyzer :  
> > default :  
> > type : custom  
> > tokenizer : uax\_\_url\_email  
> > filter : [lowercase,snowball,wd\_filter]  
> > char\_filter : [html\_strip]
> 
> > But when I index a doc into 'content' and then examine the indexed  
> > files with:  
> > curl -XGET '[http://localhost:9200/mgs/p2/385?fields=content&pretty](http://localhost:9200/mgs/p2/385?fields=content&pretty)'
> 
> > I still am seeing all of the html tags in the text. Reading back the  
> > \_mapping for the "content" field it is:  
> > "content" : {  
> > "store" : "yes",  
> > ""type" : "string"  
> > },  
> > 'index' defaults to 'analyzed', so this appears correct.
> 
> > Further complicating is that if I run:  
> > curl -XGET '[http://localhost:9200/mgs/\_analyze'-d](http://localhost:9200/mgs/_analyze'-d) '
> > 
> > this is  
> > a tests
> > 
> > '
> 
> > Then I get:  
> > {"tokens":[{"token":"this","start\_offset":3,"end\_offset":  
> > 7,"type":"","position":1},{"token":"is","start\_offset":  
> > 11,"end\_offset":13,"type":"","position":2},  
> > {"token":"a","start\_offset":18,"end\_offset":  
> > 19,"type":"","position":3},{"token":"test","start\_offset":  
> > 20,"end\_offset":25,"type":"","position":4}]}
> 
> > Which appears correct as the tags have been stripped out by using the  
> > default analyzer, and the word 'tests' has been stemmed to 'test'.
> 
> > So the analyzer seems to be working correctly except for when I  
> > actually add a document. I am using Elastica for adding documents,  
> > but I tried
> > 
> > > curl -XPUT '[http://localhost:9200/mgs/p2/0'-d](http://localhost:9200/mgs/p2/0'-d) '{"content" : "
> > > 
> > > This is a tests
> > > 
> > > " }'
> 
> > Then:
> > 
> > > curl -XGET '[http://localhost:9200/mgs/p2/0?fields=content](http://localhost:9200/mgs/p2/0?fields=content)'  
> > > Gives:  
> > > {"\_index":"mgs","\_type":"p2","\_id":"0","\_version":2,"fields":  
> > > {"content":"
> > > 
> > > This is a tests
> > > 
> > > "}}
> 
> > So adding the document does not seem to be causing the analyzer to be  
> > run on the added document. Any ideas on what I am missing/doing  
> > wrong?
> 
> > Thanks  
> > -Greg

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [June 2, 2011, 4:04pm UTC](https://discuss.elastic.co/t/help-stripping-html-tags/4528/4 "2011-06-02T16:04:52Z")

</div>

In this case, you will need to do your own stripping.

On Thursday, June 2, 2011 at 6:56 PM, Greg B wrote:

> Hi Shay, thanks for the fast response.
> 
> Is there a way to store the version with the html stripped, or do I  
> need to implement my own stripping to remove the html tags. The tags  
> are particularly a problem when trying to do highlighting in my search  
> results where the extraneous tags rather thoroughly screw up my  
> results page.
> 
> Thanks  
> -Greg
> 
> On Jun 2, 1:42 am, Shay Banon \<[shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) ([http://elasticsearch.com](http://elasticsearch.com))\> wrote:
> 
> > The stored version of a field stores the content as is, and not the analyzed form. The analysis process only controls how the text is broken down into terms and indexed.
> > 
> > On Thursday, June 2, 2011 at 2:32 AM, Greg B wrote:
> > 
> > > I've configured a default custom analyzer as follows:
> > 
> > > index :  
> > > analysis :  
> > > filter :  
> > > snowball :  
> > > type : snowball  
> > > language : English  
> > > wd\_\_filter :  
> > > type : word\_delimiter  
> > > generate\_word\_parts : true  
> > > generate\_number\_parts : true  
> > > catenate\_\_words : true  
> > > split\_on\_case\_change : true  
> > > preserve\_original : true  
> > > split\_on\_numerics: true  
> > > analyzer :  
> > > default :  
> > > type : custom  
> > > tokenizer : uax\_\_url\_email  
> > > filter : [lowercase,snowball,wd\_filter]  
> > > char\_filter : [html\_strip]
> > 
> > > But when I index a doc into 'content' and then examine the indexed  
> > > files with:  
> > > curl -XGET '[http://localhost:9200/mgs/p2/385?fields=content&pretty](http://localhost:9200/mgs/p2/385?fields=content&pretty)'
> > 
> > > I still am seeing all of the html tags in the text. Reading back the  
> > > \_mapping for the "content" field it is:  
> > > "content" : {  
> > > "store" : "yes",  
> > > ""type" : "string"  
> > > },  
> > > 'index' defaults to 'analyzed', so this appears correct.
> > 
> > > Further complicating is that if I run:  
> > > curl -XGET '[http://localhost:9200/mgs/\_analyze'-d](http://localhost:9200/mgs/_analyze'-d) '
> > > 
> > > this is  
> > > a tests
> > > 
> > > '
> > 
> > > Then I get:  
> > > {"tokens":[{"token":"this","start\_offset":3,"end\_offset":  
> > > 7,"type":"","position":1},{"token":"is","start\_offset":  
> > > 11,"end\_offset":13,"type":"","position":2},  
> > > {"token":"a","start\_offset":18,"end\_offset":  
> > > 19,"type":"","position":3},{"token":"test","start\_offset":  
> > > 20,"end\_offset":25,"type":"","position":4}]}
> > 
> > > Which appears correct as the tags have been stripped out by using the  
> > > default analyzer, and the word 'tests' has been stemmed to 'test'.
> > 
> > > So the analyzer seems to be working correctly except for when I  
> > > actually add a document. I am using Elastica for adding documents,  
> > > but I tried
> > > 
> > > > curl -XPUT '[http://localhost:9200/mgs/p2/0'-d](http://localhost:9200/mgs/p2/0'-d) '{"content" : "
> > > > 
> > > > This is a tests
> > > > 
> > > > " }'
> > 
> > > Then:
> > > 
> > > > curl -XGET '[http://localhost:9200/mgs/p2/0?fields=content](http://localhost:9200/mgs/p2/0?fields=content)'  
> > > > Gives:  
> > > > {"\_index":"mgs","\_type":"p2","\_id":"0","\_version":2,"fields":  
> > > > {"content":"
> > > > 
> > > > This is a tests
> > > > 
> > > > "}}
> > 
> > > So adding the document does not seem to be causing the analyzer to be  
> > > run on the added document. Any ideas on what I am missing/doing  
> > > wrong?
> > 
> > > Thanks  
> > > -Greg

---

<div class="post-metadata">

**Author:** ![Administrator\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/administrator_2/32/3211_2.png) [@Administrator\_2](https://discuss.elastic.co/u/Administrator_2)\
**Post date:** [June 2, 2011, 4:06pm UTC](https://discuss.elastic.co/t/help-stripping-html-tags/4528/5 "2011-06-02T16:06:48Z")

</div>

Greg,

You have to do the stripping yourself.

- Nick

-----Original Message-----  
From: Greg B [[mailto:gbrown5878@gmail.com](mailto:gbrown5878@gmail.com)]  
Sent: Thursday, June 02, 2011 11:56 AM  
To: users  
Subject: SPAM-LOW: Re: Help stripping HTML tags

Hi Shay, thanks for the fast response.

Is there a way to store the version with the html stripped, or do I  
need to implement my own stripping to remove the html tags. The tags  
are particularly a problem when trying to do highlighting in my search  
results where the extraneous tags rather thoroughly screw up my  
results page.

Thanks  
-Greg

On Jun 2, 1:42 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> The stored version of a field stores the content as is, and not the  
> analyzed form. The analysis process only controls how the text is broken  
> down into terms and indexed.
> 
> On Thursday, June 2, 2011 at 2:32 AM, Greg B wrote:
> 
> > I've configured a default custom analyzer as follows:
> 
> > index :  
> > analysis :  
> > filter :  
> > snowball :  
> > type : snowball  
> > language : English  
> > wd\_\_filter :  
> > type : word\_delimiter  
> > generate\_word\_parts : true  
> > generate\_number\_parts : true  
> > catenate\_\_words : true  
> > split\_on\_case\_change : true  
> > preserve\_original : true  
> > split\_on\_numerics: true  
> > analyzer :  
> > default :  
> > type : custom  
> > tokenizer : uax\_\_url\_email  
> > filter : [lowercase,snowball,wd\_filter]  
> > char\_filter : [html\_strip]
> 
> > But when I index a doc into 'content' and then examine the indexed  
> > files with:  
> > curl -XGET '[http://localhost:9200/mgs/p2/385?fields=content&pretty](http://localhost:9200/mgs/p2/385?fields=content&pretty)'
> 
> > I still am seeing all of the html tags in the text. Reading back the  
> > \_mapping for the "content" field it is:  
> > "content" : {  
> > "store" : "yes",  
> > ""type" : "string"  
> > },  
> > 'index' defaults to 'analyzed', so this appears correct.
> 
> > Further complicating is that if I run:  
> > curl -XGET '[http://localhost:9200/mgs/\_analyze'-d](http://localhost:9200/mgs/_analyze'-d) '
> > 
> > this is  
> > a tests
> > 
> > '
> 
> > Then I get:  
> > {"tokens":[{"token":"this","start\_offset":3,"end\_offset":  
> > 7,"type":"","position":1},{"token":"is","start\_offset":  
> > 11,"end\_offset":13,"type":"","position":2},  
> > {"token":"a","start\_offset":18,"end\_offset":  
> > 19,"type":"","position":3},{"token":"test","start\_offset":  
> > 20,"end\_offset":25,"type":"","position":4}]}
> 
> > Which appears correct as the tags have been stripped out by using the  
> > default analyzer, and the word 'tests' has been stemmed to 'test'.
> 
> > So the analyzer seems to be working correctly except for when I  
> > actually add a document. I am using Elastica for adding documents,  
> > but I tried
> > 
> > > curl -XPUT '[http://localhost:9200/mgs/p2/0'-d](http://localhost:9200/mgs/p2/0'-d) '{"content" : "
> > > 
> > > This  
> > > is a tests
> > > 
> > > " }'
> 
> > Then:
> > 
> > > curl -XGET '[http://localhost:9200/mgs/p2/0?fields=content](http://localhost:9200/mgs/p2/0?fields=content)'  
> > > Gives:  
> > > {"\_index":"mgs","\_type":"p2","\_id":"0","\_version":2,"fields":  
> > > {"content":"
> > > 
> > > This is a tests
> > > 
> > > "}}
> 
> > So adding the document does not seem to be causing the analyzer to be  
> > run on the added document. Any ideas on what I am missing/doing  
> > wrong?
> 
> > Thanks  
> > -Greg

---

<div class="post-metadata">

**Author:** ![Greg\_Brown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/greg_brown/32/1373_2.png) [@Greg\_Brown](https://discuss.elastic.co/u/Greg_Brown)\
**Post date:** [June 2, 2011, 4:18pm UTC](https://discuss.elastic.co/t/help-stripping-html-tags/4528/6 "2011-06-02T16:18:30Z")

</div>

OK, thanks.  
-Greg

On Jun 2, 10:06 am, "Administrator" [ad...@sf4answers.com](mailto:ad...@sf4answers.com) wrote:

> Greg,
> 
> You have to do the stripping yourself.
> 
> - Nick
> 
> -----Original Message-----  
> From: Greg B [[mailto:gbrown5...@gmail.com](mailto:gbrown5...@gmail.com)]  
> Sent: Thursday, June 02, 2011 11:56 AM  
> To: users  
> Subject: SPAM-LOW: Re: Help stripping HTML tags
> 
> Hi Shay, thanks for the fast response.
> 
> Is there a way to store the version with the html stripped, or do I  
> need to implement my own stripping to remove the html tags. The tags  
> are particularly a problem when trying to do highlighting in my search  
> results where the extraneous tags rather thoroughly screw up my  
> results page.
> 
> Thanks  
> -Greg
> 
> On Jun 2, 1:42 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > The stored version of a field stores the content as is, and not the  
> > analyzed form. The analysis process only controls how the text is broken  
> > down into terms and indexed.
> 
> > On Thursday, June 2, 2011 at 2:32 AM, Greg B wrote:
> > 
> > > I've configured a default custom analyzer as follows:
> 
> > > index :  
> > > analysis :  
> > > filter :  
> > > snowball :  
> > > type : snowball  
> > > language : English  
> > > wd\_\_filter :  
> > > type : word\_delimiter  
> > > generate\_word\_parts : true  
> > > generate\_number\_parts : true  
> > > catenate\_\_words : true  
> > > split\_on\_case\_change : true  
> > > preserve\_original : true  
> > > split\_on\_numerics: true  
> > > analyzer :  
> > > default :  
> > > type : custom  
> > > tokenizer : uax\_\_url\_email  
> > > filter : [lowercase,snowball,wd\_filter]  
> > > char\_filter : [html\_strip]
> 
> > > But when I index a doc into 'content' and then examine the indexed  
> > > files with:  
> > > curl -XGET '[http://localhost:9200/mgs/p2/385?fields=content&pretty](http://localhost:9200/mgs/p2/385?fields=content&pretty)'
> 
> > > I still am seeing all of the html tags in the text. Reading back the  
> > > \_mapping for the "content" field it is:  
> > > "content" : {  
> > > "store" : "yes",  
> > > ""type" : "string"  
> > > },  
> > > 'index' defaults to 'analyzed', so this appears correct.
> 
> > > Further complicating is that if I run:  
> > > curl -XGET '[http://localhost:9200/mgs/\_analyze'-d'](http://localhost:9200/mgs/_analyze'-d')
> > > 
> > > this is  
> > > a tests
> > > 
> > > '
> 
> > > Then I get:  
> > > {"tokens":[{"token":"this","start\_offset":3,"end\_offset":  
> > > 7,"type":"","position":1},{"token":"is","start\_offset":  
> > > 11,"end\_offset":13,"type":"","position":2},  
> > > {"token":"a","start\_offset":18,"end\_offset":  
> > > 19,"type":"","position":3},{"token":"test","start\_offset":  
> > > 20,"end\_offset":25,"type":"","position":4}]}
> 
> > > Which appears correct as the tags have been stripped out by using the  
> > > default analyzer, and the word 'tests' has been stemmed to 'test'.
> 
> > > So the analyzer seems to be working correctly except for when I  
> > > actually add a document. I am using Elastica for adding documents,  
> > > but I tried
> > > 
> > > > curl -XPUT '[http://localhost:9200/mgs/p2/0'-d'](http://localhost:9200/mgs/p2/0'-d'){"content" : "
> > > > 
> > > > This  
> > > > is a tests
> > > > 
> > > > " }'
> 
> > > Then:
> > > 
> > > > curl -XGET '[http://localhost:9200/mgs/p2/0?fields=content](http://localhost:9200/mgs/p2/0?fields=content)'  
> > > > Gives:  
> > > > {"\_index":"mgs","\_type":"p2","\_id":"0","\_version":2,"fields":  
> > > > {"content":"
> > > > 
> > > > This is a tests
> > > > 
> > > > "}}
> 
> > > So adding the document does not seem to be causing the analyzer to be  
> > > run on the added document. Any ideas on what I am missing/doing  
> > > wrong?
> 
> > > Thanks  
> > > -Greg

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:04am UTC](https://discuss.elastic.co/t/help-stripping-html-tags/4528/7 "2017-07-06T04:04:52Z")

</div>


