# Indexing HTML documents, problems with JSON

**URL:** <https://discuss.elastic.co/t/indexing-html-documents-problems-with-json/3473>\
**Category:** Elasticsearch\
**Created:** [October 25, 2010, 2:29pm UTC](https://discuss.elastic.co/t/indexing-html-documents-problems-with-json/3473 "2010-10-25T14:29:04Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Albin\_Stigo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/albin_stigo/32/3227_2.png) [@Albin\_Stigo](https://discuss.elastic.co/u/Albin_Stigo)\
**Post date:** [October 25, 2010, 2:29pm UTC](https://discuss.elastic.co/t/indexing-html-documents-problems-with-json/3473/1 "2010-10-25T14:29:04Z")

</div>

Hi,

I have a bunch of html documents that I would like to index (around  
3000, so not so many). I put the title as well as some other metadata  
in separate properties but I would like to make the content searchable  
as well, and I would also like to be able to display the orignal  
document... and I would like to do this over JSON... But:

"JSON does not look like XML, so HTML text fed to a JSON parser will  
produce an error."

So im having problem parsing my hits back so...

How do you guys solve this... do you strip out the html out of the  
document and only index the plain text content and then pull the  
original from another database (based on an indexed id) or are there  
other ways?

Sorry for the rather long post.

--Albin

---

<div class="post-metadata">

**Author:** ![Lukas\_Vlcek1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lukas_vlcek1/32/819_2.png) [@Lukas\_Vlcek1](https://discuss.elastic.co/u/Lukas_Vlcek1)\
**Post date:** [October 25, 2010, 6:46pm UTC](https://discuss.elastic.co/t/indexing-html-documents-problems-with-json/3473/2 "2010-10-25T18:46:43Z")

</div>

Hi,

you should check attachment type:  
[http://www.elasticsearch.com/docs/elasticsearch/mapping/attachment/](http://www.elasticsearch.com/docs/elasticsearch/mapping/attachment/)  
Note that as of 0.12.0 the plugin needs to be installed in extracted form.  
The best option is to use bin/plugin script to install plugins (or you can  
do it manually, just create a "plugins" directory in ES HOME, unpack  
particular plugin into this folder and you are done ... start ES).

But this way you will not get raw HTML in \_source (it will be kept in base64  
form). So either you can try decode it from result hits on client side or  
you need to extract raw HTML before indexing, then escaping it to make it  
JSON valid (shouldn't be that hard) and using html\_strip filter (see  
[http://www.elasticsearch.com/docs/elasticsearch/index\_modules/analysis/charfilter/](http://www.elasticsearch.com/docs/elasticsearch/index_modules/analysis/charfilter/)  
,  
for more details:  
[Analysis: Add `char_filter` on top of `tokenizer`, `filter`, and `analyzer`. Add an `html_strip` char filter and `standard_html_strip` analyzer · Issue #315 · elastic/elasticsearch · GitHub](http://github.com/elasticsearch/elasticsearch/issues/issue/315)). However, I  
did not try it myself yet.

Regards,  
Lukas

On Mon, Oct 25, 2010 at 4:29 PM, Albin Stigo [albin.stigo@gmail.com](mailto:albin.stigo@gmail.com) wrote:

> Hi,
> 
> I have a bunch of html documents that I would like to index (around  
> 3000, so not so many). I put the title as well as some other metadata  
> in separate properties but I would like to make the content searchable  
> as well, and I would also like to be able to display the orignal  
> document... and I would like to do this over JSON... But:

> "JSON does not look like XML, so HTML text fed to a JSON parser will  
> produce an error."
> 
> So im having problem parsing my hits back so...
> 
> How do you guys solve this... do you strip out the html out of the  
> document and only index the plain text content and then pull the  
> original from another database (based on an indexed id) or are there  
> other ways?
> 
> Sorry for the rather long post.
> 
> --Albin

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [October 26, 2010, 10:10pm UTC](https://discuss.elastic.co/t/indexing-html-documents-problems-with-json/3473/3 "2010-10-26T22:10:53Z")

</div>

Regarding feeding html within the json itself, not sure how you generate the  
json, but most "to\_json" utils/libs also escape relevant characters.

On Mon, Oct 25, 2010 at 8:46 PM, Lukáš Vlček [lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com) wrote:

> Hi,
> 
> you should check attachment type:  
> [http://www.elasticsearch.com/docs/elasticsearch/mapping/attachment/](http://www.elasticsearch.com/docs/elasticsearch/mapping/attachment/)  
> Note that as of 0.12.0 the plugin needs to be installed in extracted form.  
> The best option is to use bin/plugin script to install plugins (or you can  
> do it manually, just create a "plugins" directory in ES HOME, unpack  
> particular plugin into this folder and you are done ... start ES).
> 
> But this way you will not get raw HTML in \_source (it will be kept in  
> base64 form). So either you can try decode it from result hits on client  
> side or you need to extract raw HTML before indexing, then escaping it to  
> make it JSON valid (shouldn't be that hard) and using html\_strip filter  
> (see  
> [http://www.elasticsearch.com/docs/elasticsearch/index\_modules/analysis/charfilter/](http://www.elasticsearch.com/docs/elasticsearch/index_modules/analysis/charfilter/) ,  
> for more details:  
> [Analysis: Add `char_filter` on top of `tokenizer`, `filter`, and `analyzer`. Add an `html_strip` char filter and `standard_html_strip` analyzer · Issue #315 · elastic/elasticsearch · GitHub](http://github.com/elasticsearch/elasticsearch/issues/issue/315)). However,  
> I did not try it myself yet.
> 
> Regards,  
> Lukas
> 
> On Mon, Oct 25, 2010 at 4:29 PM, Albin Stigo [albin.stigo@gmail.com](mailto:albin.stigo@gmail.com)wrote:
> 
> > Hi,
> > 
> > I have a bunch of html documents that I would like to index (around  
> > 3000, so not so many). I put the title as well as some other metadata  
> > in separate properties but I would like to make the content searchable  
> > as well, and I would also like to be able to display the orignal  
> > document... and I would like to do this over JSON... But:
> 
> > "JSON does not look like XML, so HTML text fed to a JSON parser will  
> > produce an error."
> > 
> > So im having problem parsing my hits back so...
> > 
> > How do you guys solve this... do you strip out the html out of the  
> > document and only index the plain text content and then pull the  
> > original from another database (based on an indexed id) or are there  
> > other ways?
> > 
> > Sorry for the rather long post.
> > 
> > --Albin

---

<div class="post-metadata">

**Author:** ![Albin\_Stigo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/albin_stigo/32/3227_2.png) [@Albin\_Stigo](https://discuss.elastic.co/u/Albin_Stigo)\
**Post date:** [October 27, 2010, 8:05am UTC](https://discuss.elastic.co/t/indexing-html-documents-problems-with-json/3473/4 "2010-10-27T08:05:21Z")

</div>

Ok!

I'm using node.js which as you know is running on V8 and  
JSON.parse(str) actually returns an error (on some hits) when trying  
to parse back to JSON. But of course I managed to put it into ES using  
JSON in the first place, using a python script with the json package  
and that escaped everything fine.

Maybe it's a bug in node.js..

--Albin

On Wed, Oct 27, 2010 at 12:10 AM, Shay Banon  
[shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com) wrote:

> Regarding feeding html within the json itself, not sure how you generate the  
> json, but most "to\_json" utils/libs also escape relevant characters.
> 
> On Mon, Oct 25, 2010 at 8:46 PM, Lukáš Vlček [lukas.vlcek@gmail.com](mailto:lukas.vlcek@gmail.com) wrote:
> 
> > Hi,  
> > you should check attachment  
> > type: [http://www.elasticsearch.com/docs/elasticsearch/mapping/attachment/](http://www.elasticsearch.com/docs/elasticsearch/mapping/attachment/)  
> > Note that as of 0.12.0 the plugin needs to be installed in extracted form.  
> > The best option is to use bin/plugin script to install plugins (or you can  
> > do it manually, just create a "plugins" directory in ES HOME, unpack  
> > particular plugin into this folder and you are done ... start ES).  
> > But this way you will not get raw HTML in \_source (it will be kept in  
> > base64 form). So either you can try decode it from result hits on client  
> > side or you need to extract raw HTML before indexing, then escaping it to  
> > make it JSON valid (shouldn't be that hard) and using html\_strip filter  
> > (see [http://www.elasticsearch.com/docs/elasticsearch/index\_modules/analysis/charfilter/](http://www.elasticsearch.com/docs/elasticsearch/index_modules/analysis/charfilter/) ,  
> > for more  
> > details: [Analysis: Add `char_filter` on top of `tokenizer`, `filter`, and `analyzer`. Add an `html_strip` char filter and `standard_html_strip` analyzer · Issue #315 · elastic/elasticsearch · GitHub](http://github.com/elasticsearch/elasticsearch/issues/issue/315)).  
> > However, I did not try it myself yet.  
> > Regards,  
> > Lukas  
> > On Mon, Oct 25, 2010 at 4:29 PM, Albin Stigo [albin.stigo@gmail.com](mailto:albin.stigo@gmail.com)  
> > wrote:
> > 
> > > Hi,
> > > 
> > > I have a bunch of html documents that I would like to index (around  
> > > 3000, so not so many). I put the title as well as some other metadata  
> > > in separate properties but I would like to make the content searchable  
> > > as well, and I would also like to be able to display the orignal  
> > > document... and I would like to do this over JSON... But:
> > > 
> > > "JSON does not look like XML, so HTML text fed to a JSON parser will  
> > > produce an error."
> > > 
> > > So im having problem parsing my hits back so...
> > > 
> > > How do you guys solve this... do you strip out the html out of the  
> > > document and only index the plain text content and then pull the  
> > > original from another database (based on an indexed id) or are there  
> > > other ways?
> > > 
> > > Sorry for the rather long post.
> > > 
> > > --Albin

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [October 27, 2010, 8:59am UTC](https://discuss.elastic.co/t/indexing-html-documents-problems-with-json/3473/5 "2010-10-27T08:59:37Z")

</div>

On Wed, 2010-10-27 at 10:05 +0200, Albin Stigo wrote:

> Ok!
> 
> I'm using node.js which as you know is running on V8 and  
> JSON.parse(str) actually returns an error (on some hits) when trying  
> to parse back to JSON. But of course I managed to put it into ES using  
> JSON in the first place, using a python script with the json package  
> and that escaped everything fine.

Request the doc that is throwing the error in node.js using curl from  
the command line. Have a look at what is in the \_source field - try to  
decode it from JSON using python and node.js

It's more likely that there was a bug putting the \_source INTO ES, than  
the other way around

clint

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:17am UTC](https://discuss.elastic.co/t/indexing-html-documents-problems-with-json/3473/6 "2017-07-06T04:17:32Z")

</div>


