# How to encode content of web page or file attachment with elasticsearch-river-mongodb

**URL:** <https://discuss.elastic.co/t/how-to-encode-content-of-web-page-or-file-attachment-with-elasticsearch-river-mongodb/14588>\
**Category:** Elasticsearch\
**Created:** [November 25, 2013, 11:15pm UTC](https://discuss.elastic.co/t/how-to-encode-content-of-web-page-or-file-attachment-with-elasticsearch-river-mongodb/14588 "2013-11-25T23:15:14Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Zoran\_Jeremic](https://avatars.discourse-cdn.com/v4/letter/z/7ba0ec/32.png) [@Zoran\_Jeremic](https://discuss.elastic.co/u/Zoran_Jeremic)\
**Post date:** [November 25, 2013, 11:15pm UTC](https://discuss.elastic.co/t/how-to-encode-content-of-web-page-or-file-attachment-with-elasticsearch-river-mongodb/14588/1 "2013-11-25T23:15:14Z")

</div>

Hi,

I'm using elasticsearch-river-mongodb to index data from mongodb and make  
it possible to search in elasticsearch. Document stored in mongodb could  
contain different types of fields, and one of fields is attachment where  
content of web pages of other files should be stored. Mappings I created  
looks like this:

{  
"document": {  
"properties": {  
"engine\_id": {  
"store": "yes",  
"type": "string"  
},  
"fields": {  
"type": "nested",  
"properties": {  
"text\_value": {  
"type": "string",  
"analyzer": "simple"  
},  
"name\_value": {  
"type": "string",  
"analyzer": "simple"  
},  
"float\_value": {  
"type": "double"  
},  
"key": {  
"index": "not\_analyzed",  
"type": "string",  
"index\_options": "docs",  
"omit\_norms": true  
},  
"file\_value": {  
"type": "attachment",  
"file\_value":{  
"term\_vector":"with\_positions\_offsets",  
"store":"yes"  
}  
}  
}  
}  
}  
}  
}  
}

Field "file\_value" stores content of web page. I tried to store it in  
several different ways e.g.:

byte[] encodedContent = org.elasticsearch.common.io.Streams.copyToByteArray(  
inputStream);  
String encodedContent = org.elasticsearch.common.Base64.encodeFromFile(  
"test.html");

However, encoded value seems to be treated as regular string in  
elasticsearch and I can search it only if I use encoded value in search  
query. If I insert real query, I don't have any results. This used to work  
fine when I have direct inserts into the elasticsearch, but with mongodb  
river it doesn't work or I'm making some mistake. The only solution I have  
at the moment to store the whole web page content (with html including) and  
store it or to use pre-processing of web page to extract the content and  
store as a string.

This is a sample of document stored in mongodb:

{  
"\_id" : ObjectId("5293cf6a2318b3b53ca5694d"),  
"engine\_id" : "engineid1234",  
"fields" : [  
{  
"key" : "title",  
"text\_value" : "Healthcare in India"  
},  
{  
"key" : "file",  
"file\_value" :  
"em9yYW4gamVyZW1pYyBsb2dpdGVjaCBzZWFyY2ggZWxhc3RpY3NlYXJjaAo="  
}  
]  
}

I hope that some of you guys could give me idea what's wrong here.

Thanks,  
Zoran

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 26, 2013, 8:36am UTC](https://discuss.elastic.co/t/how-to-encode-content-of-web-page-or-file-attachment-with-elasticsearch-river-mongodb/14588/2 "2013-11-26T08:36:38Z")

</div>

Could you copy a document as elasticsearch has indexed it? I mean a [http://localhost:9200/yourindex/yourtype/anyid/\_source](http://localhost:9200/yourindex/yourtype/anyid/_source)

--  
David Pilato | Technical Advocate | [Elasticsearch.com](http://Elasticsearch.com)  
@dadoonet | @elasticsearchfr

Le 26 novembre 2013 at 00:15:17, Zoran Jeremic ([zoran.jeremic@gmail.com](mailto:zoran.jeremic@gmail.com)) a écrit:

Hi,

I'm using elasticsearch-river-mongodb to index data from mongodb and make it possible to search in elasticsearch. Document stored in mongodb could contain different types of fields, and one of fields is attachment where content of web pages of other files should be stored. Mappings I created looks like this:

{  
"document": {  
"properties": {  
"engine\_id": {  
"store": "yes",  
"type": "string"  
},  
"fields": {  
"type": "nested",  
"properties": {  
"text\_value": {  
"type": "string",  
"analyzer": "simple"  
},  
"name\_value": {  
"type": "string",  
"analyzer": "simple"  
},  
"float\_value": {  
"type": "double"  
},  
"key": {  
"index": "not\_analyzed",  
"type": "string",  
"index\_options": "docs",  
"omit\_norms": true  
},  
"file\_value": {  
"type": "attachment",  
"file\_value":{  
"term\_vector":"with\_positions\_offsets",  
"store":"yes"  
}  
}  
}  
}  
}  
}  
}  
}

Field "file\_value" stores content of web page. I tried to store it in several different ways e.g.:

byte[] encodedContent = org.elasticsearch.common.io.Streams.copyToByteArray(inputStream);  
String encodedContent = org.elasticsearch.common.Base64.encodeFromFile("test.html");

However, encoded value seems to be treated as regular string in elasticsearch and I can search it only if I use encoded value in search query. If I insert real query, I don't have any results. This used to work fine when I have direct inserts into the elasticsearch, but with mongodb river it doesn't work or I'm making some mistake. The only solution I have at the moment to store the whole web page content (with html including) and store it or to use pre-processing of web page to extract the content and store as a string.

This is a sample of document stored in mongodb:

{  
"\_id" : ObjectId("5293cf6a2318b3b53ca5694d"),  
"engine\_id" : "engineid1234",  
"fields" : [  
{  
"key" : "title",  
"text\_value" : "Healthcare in India"  
},  
{  
"key" : "file",  
"file\_value" : "em9yYW4gamVyZW1pYyBsb2dpdGVjaCBzZWFyY2ggZWxhc3RpY3NlYXJjaAo="  
}  
]  
}

I hope that some of you guys could give me idea what's wrong here.

Thanks,  
Zoran

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Zoran\_Jeremic](https://avatars.discourse-cdn.com/v4/letter/z/7ba0ec/32.png) [@Zoran\_Jeremic](https://discuss.elastic.co/u/Zoran_Jeremic)\
**Post date:** [November 26, 2013, 6:47pm UTC](https://discuss.elastic.co/t/how-to-encode-content-of-web-page-or-file-attachment-with-elasticsearch-river-mongodb/14588/3 "2013-11-26T18:47:28Z")

</div>

Hi David,

Thank you for your quick response. Actually, after checking index mapping  
after inserting document, I realized that there was mistake in index name  
provided for mapping (documents) and index name I set in river (document),  
so ES created default mapping instead of using the one I created, and  
probably that was the reason why attachment was treated as index. However,  
after fixing this issue, I got into the another. I can't search any of the  
nested fields. I'm tried the following query:

{  
"query": {  
"bool": {  
"should": {  
"match": {"fields.text\_value":"Healthcare in India"}  
}  
}  
}  
}

The document is stored in ES index as follows:

{ "\_index": "inextweb\_documents",  
"\_type": "documents",  
"\_id": "5294e14023189da221c66101",  
"\_version": 1,  
"exists": true,  
"\_source": {  
"\_id": "5294e14023189da221c66101",  
"engine\_id": "engineid1234",  
"fields": [  
{  
"key": "title",  
"text\_value": "Healthcare in India"  
},  
{  
"key": "domain",  
"text\_value": "[wikipedia.org](http://wikipedia.org)"  
},  
{  
"key": "jobId",  
"text\_value": "jobid123"  
},  
{  
"key": "file",  
"file\_value":  
"SG93IHRvIGVuY29kZSBjb250ZW50IG9mIHdlYiBwYWdlIG9yIGZpbGUgYXR0YWNobWVudCB3aXRoIGVsYXN0aWNzZWFyY2gtcml2ZXItbW9uZ29kYgo="  
}  
]  
}  
}

Zoran

On Tuesday, November 26, 2013 12:36:38 AM UTC-8, David Pilato wrote:

> Could you copy a document as elasticsearch has indexed it? I mean a  
> [http://localhost:9200/yourindex/yourtype/anyid/\_source](http://localhost:9200/yourindex/yourtype/anyid/_source)[http://www.google.com/url?q=http%3A%2F%2Flocalhost%3A9200%2Fyourindex%2Fyourtype%2Fanyid%2F\_source&sa=D&sntz=1&usg=AFQjCNEHe7plMwQFK90OLpiGqvm-4JlKTQ](http://www.google.com/url?q=http%3A%2F%2Flocalhost%3A9200%2Fyourindex%2Fyourtype%2Fanyid%2F_source&sa=D&sntz=1&usg=AFQjCNEHe7plMwQFK90OLpiGqvm-4JlKTQ)
> 
> --  
> _David Pilato_ | _Technical Advocate_ | _[Elasticsearch.com](http://Elasticsearch.com)_  
> @dadoonet[https://www.google.com/url?q=https%3A%2F%2Ftwitter.com%2Fdadoonet&sa=D&sntz=1&usg=AFQjCNE-DMC3YEu3X\_lhRIhUzuSZGsaSqA](https://www.google.com/url?q=https%3A%2F%2Ftwitter.com%2Fdadoonet&sa=D&sntz=1&usg=AFQjCNE-DMC3YEu3X_lhRIhUzuSZGsaSqA)  
> | @elasticsearchfr[https://www.google.com/url?q=https%3A%2F%2Ftwitter.com%2Felasticsearchfr&sa=D&sntz=1&usg=AFQjCNGfXdQ98RWFMJXdiqpKnZb5GMg0zA](https://www.google.com/url?q=https%3A%2F%2Ftwitter.com%2Felasticsearchfr&sa=D&sntz=1&usg=AFQjCNGfXdQ98RWFMJXdiqpKnZb5GMg0zA)
> 
> Le 26 novembre 2013 at 00:15:17, Zoran Jeremic ([zoran....@gmail.com](mailto:zoran....@gmail.com)\<javascript:\>)  
> a écrit:
> 
> Hi,
> 
> I'm using elasticsearch-river-mongodb to index data from mongodb and make  
> it possible to search in elasticsearch. Document stored in mongodb could  
> contain different types of fields, and one of fields is attachment where  
> content of web pages of other files should be stored. Mappings I created  
> looks like this:
> 
> {  
> "document": {  
> "properties": {  
> "engine\_id": {  
> "store": "yes",  
> "type": "string"  
> },  
> "fields": {  
> "type": "nested",  
> "properties": {  
> "text\_value": {  
> "type": "string",  
> "analyzer": "simple"  
> },  
> "name\_value": {  
> "type": "string",  
> "analyzer": "simple"  
> },  
> "float\_value": {  
> "type": "double"  
> },  
> "key": {  
> "index": "not\_analyzed",  
> "type": "string",  
> "index\_options": "docs",  
> "omit\_norms": true  
> },  
> "file\_value": {  
> "type": "attachment",  
> "file\_value":{  
> "term\_vector":"with\_positions\_offsets",  
> "store":"yes"  
> }  
> }  
> }  
> }  
> }  
> }  
> }  
> }
> 
> Field "file\_value" stores content of web page. I tried to store it in  
> several different ways e.g.:
> 
> byte encodedContent = org.elasticsearch.common.io.Streams.  
> copyToByteArray(inputStream);  
> String encodedContent = org.elasticsearch.common.Base64.encodeFromFile(  
> "test.html");
> 
> However, encoded value seems to be treated as regular string in  
> elasticsearch and I can search it only if I use encoded value in search  
> query. If I insert real query, I don't have any results. This used to work  
> fine when I have direct inserts into the elasticsearch, but with mongodb  
> river it doesn't work or I'm making some mistake. The only solution I have  
> at the moment to store the whole web page content (with html including) and  
> store it or to use pre-processing of web page to extract the content and  
> store as a string.
> 
> This is a sample of document stored in mongodb:
> 
> {  
> "\_id" : ObjectId("5293cf6a2318b3b53ca5694d"),  
> "engine\_id" : "engineid1234",  
> "fields" : [  
> {  
> "key" : "title",  
> "text\_value" : "Healthcare in India"  
> },  
> {  
> "key" : "file",  
> "file\_value" :  
> "em9yYW4gamVyZW1pYyBsb2dpdGVjaCBzZWFyY2ggZWxhc3RpY3NlYXJjaAo="  
> }  
> ]  
> }
> 
> I hope that some of you guys could give me idea what's wrong here.
> 
> Thanks,  
> Zoran
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/4a716181-a2c6-4901-b11c-bf7b4c6a4539%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/4a716181-a2c6-4901-b11c-bf7b4c6a4539%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 26, 2013, 7:49pm UTC](https://discuss.elastic.co/t/how-to-encode-content-of-web-page-or-file-attachment-with-elasticsearch-river-mongodb/14588/4 "2013-11-26T19:49:25Z")

</div>

I guess you need to use nested query: [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/query-dsl-nested-query.html)

--  
David 😉  
Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

Le 26 nov. 2013 à 19:47, Zoran Jeremic [zoran.jeremic@gmail.com](mailto:zoran.jeremic@gmail.com) a écrit :

Hi David,

Thank you for your quick response. Actually, after checking index mapping after inserting document, I realized that there was mistake in index name provided for mapping (documents) and index name I set in river (document), so ES created default mapping instead of using the one I created, and probably that was the reason why attachment was treated as index. However, after fixing this issue, I got into the another. I can't search any of the nested fields. I'm tried the following query:

{  
"query": {  
"bool": {  
"should": {  
"match": {"fields.text\_value":"Healthcare in India"}  
}  
}  
}  
}

The document is stored in ES index as follows:

{ "\_index": "inextweb\_documents",  
"\_type": "documents",  
"\_id": "5294e14023189da221c66101",  
"\_version": 1,  
"exists": true,  
"\_source": {  
"\_id": "5294e14023189da221c66101",  
"engine\_id": "engineid1234",  
"fields": [  
{  
"key": "title",  
"text\_value": "Healthcare in India"  
},  
{  
"key": "domain",  
"text\_value": "[wikipedia.org](http://wikipedia.org)"  
},  
{  
"key": "jobId",  
"text\_value": "jobid123"  
},  
{  
"key": "file",  
"file\_value": "SG93IHRvIGVuY29kZSBjb250ZW50IG9mIHdlYiBwYWdlIG9yIGZpbGUgYXR0YWNobWVudCB3aXRoIGVsYXN0aWNzZWFyY2gtcml2ZXItbW9uZ29kYgo="  
}  
]  
}  
}

Zoran

> On Tuesday, November 26, 2013 12:36:38 AM UTC-8, David Pilato wrote:  
> Could you copy a document as elasticsearch has indexed it? I mean a [http://localhost:9200/yourindex/yourtype/anyid/\_source](http://localhost:9200/yourindex/yourtype/anyid/_source)
> 
> --  
> David Pilato | Technical Advocate | [Elasticsearch.com](http://Elasticsearch.com)  
> @dadoonet | @elasticsearchfr
> 
> Le 26 novembre 2013 at 00:15:17, Zoran Jeremic ([zoran....@gmail.com](mailto:zoran....@gmail.com)) a écrit:
> 
> > Hi,
> > 
> > I'm using elasticsearch-river-mongodb to index data from mongodb and make it possible to search in elasticsearch. Document stored in mongodb could contain different types of fields, and one of fields is attachment where content of web pages of other files should be stored. Mappings I created looks like this:
> > 
> > {  
> > "document": {  
> > "properties": {  
> > "engine\_id": {  
> > "store": "yes",  
> > "type": "string"  
> > },  
> > "fields": {  
> > "type": "nested",  
> > "properties": {  
> > "text\_value": {  
> > "type": "string",  
> > "analyzer": "simple"  
> > },  
> > "name\_value": {  
> > "type": "string",  
> > "analyzer": "simple"  
> > },  
> > "float\_value": {  
> > "type": "double"  
> > },  
> > "key": {  
> > "index": "not\_analyzed",  
> > "type": "string",  
> > "index\_options": "docs",  
> > "omit\_norms": true  
> > },  
> > "file\_value": {  
> > "type": "attachment",  
> > "file\_value":{  
> > "term\_vector":"with\_positions\_offsets",  
> > "store":"yes"  
> > }  
> > }  
> > }  
> > }  
> > }  
> > }  
> > }  
> > }
> > 
> > Field "file\_value" stores content of web page. I tried to store it in several different ways e.g.:
> > 
> > byte encodedContent = org.elasticsearch.common.io.Streams.copyToByteArray(inputStream);  
> > String encodedContent = org.elasticsearch.common.Base64.encodeFromFile("test.html");
> > 
> > However, encoded value seems to be treated as regular string in elasticsearch and I can search it only if I use encoded value in search query. If I insert real query, I don't have any results. This used to work fine when I have direct inserts into the elasticsearch, but with mongodb river it doesn't work or I'm making some mistake. The only solution I have at the moment to store the whole web page content (with html including) and store it or to use pre-processing of web page to extract the content and store as a string.
> > 
> > This is a sample of document stored in mongodb:
> > 
> > {  
> > "\_id" : ObjectId("5293cf6a2318b3b53ca5694d"),  
> > "engine\_id" : "engineid1234",  
> > "fields" : [  
> > {  
> > "key" : "title",  
> > "text\_value" : "Healthcare in India"  
> > },  
> > {  
> > "key" : "file",  
> > "file\_value" : "em9yYW4gamVyZW1pYyBsb2dpdGVjaCBzZWFyY2ggZWxhc3RpY3NlYXJjaAo="  
> > }  
> > ]  
> > }
> > 
> > I hope that some of you guys could give me idea what's wrong here.
> > 
> > Thanks,  
> > Zoran
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/4a716181-a2c6-4901-b11c-bf7b4c6a4539%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/4a716181-a2c6-4901-b11c-bf7b4c6a4539%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/BC62803D-059F-41F8-B863-9FCA594B81ED%40pilato.fr](https://groups.google.com/d/msgid/elasticsearch/BC62803D-059F-41F8-B863-9FCA594B81ED%40pilato.fr).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Zoran\_Jeremic](https://avatars.discourse-cdn.com/v4/letter/z/7ba0ec/32.png) [@Zoran\_Jeremic](https://discuss.elastic.co/u/Zoran_Jeremic)\
**Post date:** [November 26, 2013, 9:58pm UTC](https://discuss.elastic.co/t/how-to-encode-content-of-web-page-or-file-attachment-with-elasticsearch-river-mongodb/14588/5 "2013-11-26T21:58:50Z")

</div>

Yes. That's it. It works now 🙂

One additional question. I see that highlighting is not supported with  
nested queries. Is there any workaround it?

Thanks.  
Zoran

On Tuesday, November 26, 2013 11:49:25 AM UTC-8, David Pilato wrote:

> I guess you need to use nested query:  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/query-dsl-nested-query.html)
> 
> --  
> David 😉  
> Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs
> 
> Le 26 nov. 2013 à 19:47, Zoran Jeremic \<[zoran....@gmail.com](mailto:zoran....@gmail.com) \<javascript:\>\>  
> a écrit :
> 
> Hi David,
> 
> Thank you for your quick response. Actually, after checking index mapping  
> after inserting document, I realized that there was mistake in index name  
> provided for mapping (documents) and index name I set in river (document),  
> so ES created default mapping instead of using the one I created, and  
> probably that was the reason why attachment was treated as index. However,  
> after fixing this issue, I got into the another. I can't search any of the  
> nested fields. I'm tried the following query:
> 
> {  
> "query": {  
> "bool": {  
> "should": {  
> "match": {"fields.text\_value":"Healthcare in India"}  
> }  
> }  
> }  
> }
> 
> The document is stored in ES index as follows:
> 
> { "\_index": "inextweb\_documents",  
> "\_type": "documents",  
> "\_id": "5294e14023189da221c66101",  
> "\_version": 1,  
> "exists": true,  
> "\_source": {  
> "\_id": "5294e14023189da221c66101",  
> "engine\_id": "engineid1234",  
> "fields": [  
> {  
> "key": "title",  
> "text\_value": "Healthcare in India"  
> },  
> {  
> "key": "domain",  
> "text\_value": "[wikipedia.org](http://wikipedia.org)"  
> },  
> {  
> "key": "jobId",  
> "text\_value": "jobid123"  
> },  
> {  
> "key": "file",  
> "file\_value":  
> "SG93IHRvIGVuY29kZSBjb250ZW50IG9mIHdlYiBwYWdlIG9yIGZpbGUgYXR0YWNobWVudCB3aXRoIGVsYXN0aWNzZWFyY2gtcml2ZXItbW9uZ29kYgo="  
> }  
> ]  
> }  
> }
> 
> Zoran
> 
> On Tuesday, November 26, 2013 12:36:38 AM UTC-8, David Pilato wrote:
> 
> > Could you copy a document as elasticsearch has indexed it? I mean a  
> > [http://localhost:9200/yourindex/yourtype/anyid/\_source](http://localhost:9200/yourindex/yourtype/anyid/_source)[http://www.google.com/url?q=http%3A%2F%2Flocalhost%3A9200%2Fyourindex%2Fyourtype%2Fanyid%2F\_source&sa=D&sntz=1&usg=AFQjCNEHe7plMwQFK90OLpiGqvm-4JlKTQ](http://www.google.com/url?q=http%3A%2F%2Flocalhost%3A9200%2Fyourindex%2Fyourtype%2Fanyid%2F_source&sa=D&sntz=1&usg=AFQjCNEHe7plMwQFK90OLpiGqvm-4JlKTQ)
> > 
> > --  
> > _David Pilato_ | _Technical Advocate_ | _[Elasticsearch.com](http://Elasticsearch.com)  
> > [http://Elasticsearch.com](http://Elasticsearch.com)_  
> > @dadoonet[https://www.google.com/url?q=https%3A%2F%2Ftwitter.com%2Fdadoonet&sa=D&sntz=1&usg=AFQjCNE-DMC3YEu3X\_lhRIhUzuSZGsaSqA](https://www.google.com/url?q=https%3A%2F%2Ftwitter.com%2Fdadoonet&sa=D&sntz=1&usg=AFQjCNE-DMC3YEu3X_lhRIhUzuSZGsaSqA)  
> > | @elasticsearchfr[https://www.google.com/url?q=https%3A%2F%2Ftwitter.com%2Felasticsearchfr&sa=D&sntz=1&usg=AFQjCNGfXdQ98RWFMJXdiqpKnZb5GMg0zA](https://www.google.com/url?q=https%3A%2F%2Ftwitter.com%2Felasticsearchfr&sa=D&sntz=1&usg=AFQjCNGfXdQ98RWFMJXdiqpKnZb5GMg0zA)
> > 
> > Le 26 novembre 2013 at 00:15:17, Zoran Jeremic ([zoran....@gmail.com](mailto:zoran....@gmail.com)) a  
> > écrit:
> > 
> > Hi,
> > 
> > I'm using elasticsearch-river-mongodb to index data from mongodb and make  
> > it possible to search in elasticsearch. Document stored in mongodb could  
> > contain different types of fields, and one of fields is attachment where  
> > content of web pages of other files should be stored. Mappings I created  
> > looks like this:
> > 
> > {  
> > "document": {  
> > "properties": {  
> > "engine\_id": {  
> > "store": "yes",  
> > "type": "string"  
> > },  
> > "fields": {  
> > "type": "nested",  
> > "properties": {  
> > "text\_value": {  
> > "type": "string",  
> > "analyzer": "simple"  
> > },  
> > "name\_value": {  
> > "type": "string",  
> > "analyzer": "simple"  
> > },  
> > "float\_value": {  
> > "type": "double"  
> > },  
> > "key": {  
> > "index": "not\_analyzed",  
> > "type": "string",  
> > "index\_options": "docs",  
> > "omit\_norms": true  
> > },  
> > "file\_value": {  
> > "type": "attachment",  
> > "file\_value":{  
> > "term\_vector":"with\_positions\_offsets",  
> > "store":"yes"  
> > }  
> > }  
> > }  
> > }  
> > }  
> > }  
> > }  
> > }
> > 
> > Field "file\_value" stores content of web page. I tried to store it in  
> > several different ways e.g.:
> > 
> > byte encodedContent = org.elasticsearch.common.io.Streams.  
> > copyToByteArray(inputStream);  
> > String encodedContent = org.elasticsearch.common.Base64.encodeFromFile(  
> > "test.html");
> > 
> > However, encoded value seems to be treated as regular string in  
> > elasticsearch and I can search it only if I use encoded value in search  
> > query. If I insert real query, I don't have any results. This used to work  
> > fine when I have direct inserts into the elasticsearch, but with mongodb  
> > river it doesn't work or I'm making some mistake. The only solution I have  
> > at the moment to store the whole web page content (with html including) and  
> > store it or to use pre-processing of web page to extract the content and  
> > store as a string.
> > 
> > This is a sample of document stored in mongodb:
> > 
> > {  
> > "\_id" : ObjectId("5293cf6a2318b3b53ca5694d"),  
> > "engine\_id" : "engineid1234",  
> > "fields" : [  
> > {  
> > "key" : "title",  
> > "text\_value" : "Healthcare in India"  
> > },  
> > {  
> > "key" : "file",  
> > "file\_value" :  
> > "em9yYW4gamVyZW1pYyBsb2dpdGVjaCBzZWFyY2ggZWxhc3RpY3NlYXJjaAo="  
> > }  
> > ]  
> > }
> > 
> > I hope that some of you guys could give me idea what's wrong here.
> > 
> > Thanks,  
> > Zoran
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/4a716181-a2c6-4901-b11c-bf7b4c6a4539%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/4a716181-a2c6-4901-b11c-bf7b4c6a4539%40googlegroups.com)  
> > .  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/c0a2ade4-6c96-4de7-812d-6a7c6200fe6f%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/c0a2ade4-6c96-4de7-812d-6a7c6200fe6f%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:04am UTC](https://discuss.elastic.co/t/how-to-encode-content-of-web-page-or-file-attachment-with-elasticsearch-river-mongodb/14588/6 "2017-07-06T02:04:42Z")

</div>


