# JSON parsing and Payload injection

**URL:** <https://discuss.elastic.co/t/json-parsing-and-payload-injection/13445>\
**Category:** Elasticsearch\
**Created:** [September 3, 2013, 1:06pm UTC](https://discuss.elastic.co/t/json-parsing-and-payload-injection/13445 "2013-09-03T13:06:51Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Aurelien\_3](https://avatars.discourse-cdn.com/v4/letter/a/ce7236/32.png) [@Aurelien\_3](https://discuss.elastic.co/u/Aurelien_3)\
**Post date:** [September 3, 2013, 1:06pm UTC](https://discuss.elastic.co/t/json-parsing-and-payload-injection/13445/1 "2013-09-03T13:06:51Z")

</div>

Hi ESearchers,

I'm investigating into elastic search in order to inject payload  
information associated to each word. One interesting idea would be to use  
the JSON to inject objects holding the data and payload data together.

For example, my document would be :  
{  
"content": { "data" : "keyword1", "payload" : "2.0"},{"data":"keyword2",  
"payload":"3.0"}  
}

Thought to inject payload, I need to have access to several fields of the  
object.

As I understand from the documentation and code, objects are dynamically  
parsed and considered as nested fields. So the request parser will launch  
my analyzer for data and payload independantly.

I there a way to peronalize this behaviour to have access at the analyzer  
(and tokenizer level) to a whole chunk of JSON ?

This question was already asked before and "kimchy" was suggesting  
([https://groups.google.com/d/topic/elasticsearch/cO8J5i39cUE/discussion](https://groups.google.com/d/topic/elasticsearch/cO8J5i39cUE/discussion))  
to use scripts. I don't really undertand how scripts could solve this.

Any comment or ideas to dig are welcome.

Thanks by the way for the awesome product. It has much future ahead, I'm  
sure.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![javanna](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javanna/32/4698_2.png) [@javanna](https://discuss.elastic.co/u/javanna)\
**Post date:** [September 3, 2013, 3:24pm UTC](https://discuss.elastic.co/t/json-parsing-and-payload-injection/13445/2 "2013-09-03T15:24:50Z")

</div>

Hi,  
what Shay meant in that thread is that there is no easy way to read  
payloads, thus you would need to write a native script (Java), where you  
have access to the doc id and to the lucene reader, so that you can read  
the payload information by yourself from within the script.

It's now possible to plug in a custom similarity[http://www.elasticsearch.org/guide/reference/index-modules/similarity/](http://www.elasticsearch.org/guide/reference/index-modules/similarity/)per field, which is usually how you read payloads and you use them to score  
documents depending on your domain. Have a look at  
[GitHub - tlrx/elasticsearch-custom-similarity-provider: A custom SimilarityProvider example for Elasticsearch](https://github.com/tlrx/elasticsearch-custom-similarity-provider) for an  
example of custom similarity provider.

Not sure what your usecase is, but your questions refer more to the  
indexing process, the part that stored the payloads in the index. You would  
need to use a custom token filter that does something similar to what  
lucene DelimitedPayloadTokenFilter[http://lucene.apache.org/core/4\_4\_0/analyzers-common/org/apache/lucene/analysis/payloads/DelimitedPayloadTokenFilter.html](http://lucene.apache.org/core/4_4_0/analyzers-common/org/apache/lucene/analysis/payloads/DelimitedPayloadTokenFilter.html)does, but you need to specify the payload in the same field and use a  
proper delimiter (e.g. "keyword1|2.0").

I'm not aware of any other way to index payloads that could allow to pass  
it in as a separate field.

Hope this helps  
Luca

On Tuesday, September 3, 2013 3:06:51 PM UTC+2, Aurélien wrote:

> Hi ESearchers,
> 
> I'm investigating into Elasticsearch in order to inject payload  
> information associated to each word. One interesting idea would be to use  
> the JSON to inject objects holding the data and payload data together.
> 
> For example, my document would be :  
> {  
> "content": { "data" : "keyword1", "payload" : "2.0"},{"data":"keyword2",  
> "payload":"3.0"}  
> }
> 
> Thought to inject payload, I need to have access to several fields of the  
> object.
> 
> As I understand from the documentation and code, objects are dynamically  
> parsed and considered as nested fields. So the request parser will launch  
> my analyzer for data and payload independantly.
> 
> I there a way to peronalize this behaviour to have access at the analyzer  
> (and tokenizer level) to a whole chunk of JSON ?
> 
> This question was already asked before and "kimchy" was suggesting (  
> [https://groups.google.com/d/topic/elasticsearch/cO8J5i39cUE/discussion](https://groups.google.com/d/topic/elasticsearch/cO8J5i39cUE/discussion))  
> to use scripts. I don't really undertand how scripts could solve this.
> 
> Any comment or ideas to dig are welcome.
> 
> Thanks by the way for the awesome product. It has much future ahead, I'm  
> sure.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Aurelien\_3](https://avatars.discourse-cdn.com/v4/letter/a/ce7236/32.png) [@Aurelien\_3](https://discuss.elastic.co/u/Aurelien_3)\
**Post date:** [September 6, 2013, 2:46pm UTC](https://discuss.elastic.co/t/json-parsing-and-payload-injection/13445/3 "2013-09-06T14:46:59Z")

</div>

Thanks a lot.

Well I went to register DelimiterPayloadTokenFilter. For the moment seems  
simpler. I'll anyway try to find some time to experiment with Native Script.

Le mardi 3 septembre 2013 18:24:50 UTC+3, Luca Cavanna a écrit :

> Hi,  
> what Shay meant in that thread is that there is no easy way to read  
> payloads, thus you would need to write a native script (Java), where you  
> have access to the doc id and to the lucene reader, so that you can read  
> the payload information by yourself from within the script.
> 
> It's now possible to plug in a custom similarity[http://www.elasticsearch.org/guide/reference/index-modules/similarity/](http://www.elasticsearch.org/guide/reference/index-modules/similarity/)per field, which is usually how you read payloads and you use them to score  
> documents depending on your domain. Have a look at  
> [GitHub - tlrx/elasticsearch-custom-similarity-provider: A custom SimilarityProvider example for Elasticsearch](https://github.com/tlrx/elasticsearch-custom-similarity-provider) for an  
> example of custom similarity provider.
> 
> Not sure what your usecase is, but your questions refer more to the  
> indexing process, the part that stored the payloads in the index. You would  
> need to use a custom token filter that does something similar to what  
> lucene DelimitedPayloadTokenFilter[http://lucene.apache.org/core/4\_4\_0/analyzers-common/org/apache/lucene/analysis/payloads/DelimitedPayloadTokenFilter.html](http://lucene.apache.org/core/4_4_0/analyzers-common/org/apache/lucene/analysis/payloads/DelimitedPayloadTokenFilter.html)does, but you need to specify the payload in the same field and use a  
> proper delimiter (e.g. "keyword1|2.0").
> 
> I'm not aware of any other way to index payloads that could allow to pass  
> it in as a separate field.
> 
> Hope this helps  
> Luca
> 
> On Tuesday, September 3, 2013 3:06:51 PM UTC+2, Aurélien wrote:
> 
> > Hi ESearchers,
> > 
> > I'm investigating into Elasticsearch in order to inject payload  
> > information associated to each word. One interesting idea would be to use  
> > the JSON to inject objects holding the data and payload data together.
> > 
> > For example, my document would be :  
> > {  
> > "content": { "data" : "keyword1", "payload" :  
> > "2.0"},{"data":"keyword2", "payload":"3.0"}  
> > }
> > 
> > Thought to inject payload, I need to have access to several fields of the  
> > object.
> > 
> > As I understand from the documentation and code, objects are dynamically  
> > parsed and considered as nested fields. So the request parser will launch  
> > my analyzer for data and payload independantly.
> > 
> > I there a way to peronalize this behaviour to have access at the analyzer  
> > (and tokenizer level) to a whole chunk of JSON ?
> > 
> > This question was already asked before and "kimchy" was suggesting (  
> > [https://groups.google.com/d/topic/elasticsearch/cO8J5i39cUE/discussion](https://groups.google.com/d/topic/elasticsearch/cO8J5i39cUE/discussion))  
> > to use scripts. I don't really undertand how scripts could solve this.
> > 
> > Any comment or ideas to dig are welcome.
> > 
> > Thanks by the way for the awesome product. It has much future ahead, I'm  
> > sure.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:17am UTC](https://discuss.elastic.co/t/json-parsing-and-payload-injection/13445/4 "2017-07-06T02:17:48Z")

</div>


