# Convert into JSON for PERL module

**URL:** <https://discuss.elastic.co/t/convert-into-json-for-perl-module/7701>\
**Category:** Elasticsearch\
**Created:** [May 15, 2012, 8:32am UTC](https://discuss.elastic.co/t/convert-into-json-for-perl-module/7701 "2012-05-15T08:32:37Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jerome](https://avatars.discourse-cdn.com/v4/letter/j/a587f6/32.png) [@Jerome](https://discuss.elastic.co/u/Jerome)\
**Post date:** [May 15, 2012, 8:32am UTC](https://discuss.elastic.co/t/convert-into-json-for-perl-module/7701/1 "2012-05-15T08:32:37Z")

</div>

Hi !

I've Genbank Flat file (example : [http://www.ncbi.nlm.nih.gov/Sitemap/samplerecord.html](http://www.ncbi.nlm.nih.gov/Sitemap/samplerecord.html))  
and i want index it.  
My problem is the FEATURES, as you can see they're on 2 levels and  
they're différent with the file used. So i can't use the index  
function.

I used it for indexing all data which are present in any document :

$result = $es-\>index(  
index =\> $index,  
type =\> $type,  
id =\> $num\_acc,  
data =\> {  
ACC =\> $num\_acc,  
DESC =\> $desc,  
VERSION =\> $version,  
GI =\> $gi,  
ORGANISME =\> $espece,  
CLASSIFICATION =\> $classification,  
SEQUENCE =\> $seq  
},  
);

For the FEATURE i use the update funtion :

loop\_for\_the\_primaryTag {  
loop\_for\_the\_tag\_and\_value\_from\_this\_primaryTag {  
$result = $es-\>update(  
index =\> $index,  
type =\> $type,  
id =\> $num\_acc,

```
		script => "if(ctx._source.$primary_tag){if(ctx._source.$primary_tag\

```

["$tag"] == null){ctx.\_source.$primary\_tag["$tag"] = "$value"}  
else {ctx.\_source.$primary\_tag["$tag"] += " $value"}}else  
{ctx.\_source.$primary\_tag["$tag"] = "$value"}",  
);  
}  
}

I wrote you the script at the bottom.  
The objective is :  
{  
primary\_tag : {  
tag : value,  
tag : value  
},  
primary\_tag2 : {  
tag : value  
}  
}

But the script doesn't works. My question is : Where are the errors in  
the script ? Is it possible to do that ?

I think i can use the JSON to index it without this problem but i try  
and failed to convert my file in JSON. My second and most important  
question is : how can i convert this file in JSON (i read the doc on  
CPAN but...) and how can i index it in PERL ?

* * *

The pretty script

* * *

if(ctx.\_source.$primary\_tag){  
if(ctx.\_source.$primary\_tag["$tag"] == null){  
ctx.\_source.$primary\_tag["$tag"] = "$value"  
}  
else {  
ctx.\_source.$primary\_tag["$tag"] += " $value"  
}  
}  
else {  
ctx.\_source.$primary\_tag["$tag"] = "$value"  
}

Thanks for help !

I'll send an SOS to the World  
I hope that someone gets my  
Message in a forum

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [May 15, 2012, 9:41am UTC](https://discuss.elastic.co/t/convert-into-json-for-perl-module/7701/2 "2012-05-15T09:41:58Z")

</div>

Hi Jerome

> I've Genbank Flat file (example : [GenBank Sample Record](http://www.ncbi.nlm.nih.gov/Sitemap/samplerecord.html))  
> and i want index it.  
> My problem is the FEATURES, as you can see they're on 2 levels and  
> they're diffÃ©rent with the file used. So i can't use the index  
> function.

I'm afraid I don't really understand your post - rather difficult  
without real data.

But... I think what you're trying to do may be a lot easier using Perl  
directly, rather than trying to use the update() method.

I think you may be under the impression that you can only update docs in  
Elasticsearch via update(). This is incorrect. You can retrieve the  
doc using get(), change it, and reindex it using index().

You could just do something like:

my %doc = get\_next\_doc\_from\_flat\_file\_or\_from\_elasticsearch();  
while (my ($tag,$value) = get\_next\_tag()) {  
push @{$doc{$tag}},$value  
}  
$es-\>index( id =\> $num\_acc, data =\> %doc );

(this is of course pseudo-code, because I have no idea how you read the  
raw data).

clint

> I used it for indexing all data which are present in any document :
> 
> $result = $es-\>index(  
> index =\> $index,  
> type =\> $type,  
> id =\> $num\_acc,  
> data =\> {  
> ACC =\> $num\_acc,  
> DESC =\> $desc,  
> VERSION =\> $version,  
> GI =\> $gi,  
> ORGANISME =\> $espece,  
> CLASSIFICATION =\> $classification,  
> SEQUENCE =\> $seq  
> },  
> );
> 
> For the FEATURE i use the update funtion :
> 
> loop\_for\_the\_primaryTag {  
> loop\_for\_the\_tag\_and\_value\_from\_this\_primaryTag {  
> $result = $es-\>update(  
> index =\> $index,  
> type =\> $type,  
> id =\> $num\_acc,
> 
> ```
> script => "if(ctx._source.$primary_tag){if(ctx._source.$primary_tag\
> 
> ```
> 
> ["$tag"] == null){ctx.\_source.$primary\_tag["$tag"] = "$value"}  
> else {ctx.\_source.$primary\_tag["$tag"] += " $value"}}else  
> {ctx.\_source.$primary\_tag["$tag"] = "$value"}",  
> );  
> }  
> }
> 
> I wrote you the script at the bottom.  
> The objective is :  
> {  
> primary\_tag : {  
> tag : value,  
> tag : value  
> },  
> primary\_tag2 : {  
> tag : value  
> }  
> }
> 
> But the script doesn't works. My question is : Where are the errors in  
> the script ? Is it possible to do that ?
> 
> I think i can use the JSON to index it without this problem but i try  
> and failed to convert my file in JSON. My second and most important  
> question is : how can i convert this file in JSON (i read the doc on  
> CPAN but...) and how can i index it in PERL ?
> 
> * * *
> 
> The pretty script
> 
> * * *
> 
> if(ctx.\_source.$primary\_tag){  
> if(ctx.\_source.$primary\_tag["$tag"] == null){  
> ctx.\_source.$primary\_tag["$tag"] = "$value"  
> }  
> else {  
> ctx.\_source.$primary\_tag["$tag"] += " $value"  
> }  
> }  
> else {  
> ctx.\_source.$primary\_tag["$tag"] = "$value"  
> }
> 
> Thanks for help !
> 
> I'll send an SOS to the World  
> I hope that someone gets my  
> Message in a forum

---

<div class="post-metadata">

**Author:** ![Jerome](https://avatars.discourse-cdn.com/v4/letter/j/a587f6/32.png) [@Jerome](https://discuss.elastic.co/u/Jerome)\
**Post date:** [May 15, 2012, 10:32am UTC](https://discuss.elastic.co/t/convert-into-json-for-perl-module/7701/3 "2012-05-15T10:32:09Z")

</div>

Ah sorry for the difficult post.

You said i can update my docs using "get" and "index".  
If I must re-index my documents, I must necessarily provide ALL the  
document data or can I just provide only the new data?

This is why I used the update () function, it allows to update by  
entering only the new data.

In your example, if I understand correctly, you add in a hash the  
document present in Elasticsearch. Then you get couples tags /  
values ​​and add them before reindex the hash.

I think a hash with only one level of imbrication and my document has  
more than one. And my real problem is this second imbrication level, I  
managed to be indexed by putting everything on one level but that does  
not make practical research. I also looking for a way to automatically  
index, data on several levels, this datas are different each time  
making it impossible to define them by hand.

You can see the script here if you need it :

> **[DL.FREE.FR](http://dl.free.fr/getfile.pl?file=%2FjyVhHesXY)**
>
> Jusqu’à 28 Méga, 10Go d’espace disque, WiFi-MiMo, Ligne téléphonique, Appels illimités vers 70 destinations, 250 chaînes de télévision, Vidéo à la Demande

On 15 mai, 11:41, Clinton Gormley [cl...@traveljury.com](mailto:cl...@traveljury.com) wrote:

> Hi Jerome
> 
> > I've Genbank Flat file (example :[GenBank Sample Record](http://www.ncbi.nlm.nih.gov/Sitemap/samplerecord.html))  
> > and i want index it.  
> > My problem is the FEATURES, as you can see they're on 2 levels and  
> > they're différent with the file used. So i can't use the index  
> > function.
> 
> I'm afraid I don't really understand your post - rather difficult  
> without real data.
> 
> But... I think what you're trying to do may be a lot easier using Perl  
> directly, rather than trying to use the update() method.
> 
> I think you may be under the impression that you can only update docs in  
> Elasticsearch via update(). This is incorrect. You can retrieve the  
> doc using get(), change it, and reindex it using index().
> 
> You could just do something like:
> 
> my %doc = get\_next\_doc\_from\_flat\_file\_or\_from\_elasticsearch();  
> while (my ($tag,$value) = get\_next\_tag()) {  
> push @{$doc{$tag}},$value  
> }  
> $es-\>index( id =\> $num\_acc, data =\> %doc );
> 
> (this is of course pseudo-code, because I have no idea how you read the  
> raw data).
> 
> clint
> 
> > I used it for indexing all data which are present in any document :
> 
> > $result = $es-\>index(  
> > index =\> $index,  
> > type =\> $type,  
> > id =\> $num\_acc,  
> > data =\> {  
> > ACC =\> $num\_acc,  
> > DESC =\> $desc,  
> > VERSION =\> $version,  
> > GI =\> $gi,  
> > ORGANISME =\> $espece,  
> > CLASSIFICATION =\> $classification,  
> > SEQUENCE =\> $seq  
> > },  
> > );
> 
> > For the FEATURE i use the update funtion :
> 
> > loop\_for\_the\_primaryTag {  
> > loop\_for\_the\_tag\_and\_value\_from\_this\_primaryTag {  
> > $result = $es-\>update(  
> > index =\> $index,  
> > type =\> $type,  
> > id =\> $num\_acc,
> 
> > ```
> > script => "if(ctx._source.$primary_tag){if(ctx._source.$primary_tag\
> > 
> > ```
> > 
> > ["$tag"] == null){ctx.\_source.$primary\_tag["$tag"] = "$value"}  
> > else {ctx.\_source.$primary\_tag["$tag"] += " $value"}}else  
> > {ctx.\_source.$primary\_tag["$tag"] = "$value"}",  
> > );  
> > }  
> > }
> 
> > I wrote you the script at the bottom.  
> > The objective is :  
> > {  
> > primary\_tag : {  
> > tag : value,  
> > tag : value  
> > },  
> > primary\_tag2 : {  
> > tag : value  
> > }  
> > }
> 
> > But the script doesn't works. My question is : Where are the errors in  
> > the script ? Is it possible to do that ?
> 
> > I think i can use the JSON to index it without this problem but i try  
> > and failed to convert my file in JSON. My second and most important  
> > question is : how can i convert this file in JSON (i read the doc on  
> > CPAN but...) and how can i index it in PERL ?
> 
> > * * *
> > 
> > The pretty script
> > 
> > * * *
> 
> > if(ctx.\_source.$primary\_tag){  
> > if(ctx.\_source.$primary\_tag["$tag"] == null){  
> > ctx.\_source.$primary\_tag["$tag"] = "$value"  
> > }  
> > else {  
> > ctx.\_source.$primary\_tag["$tag"] += " $value"  
> > }  
> > }  
> > else {  
> > ctx.\_source.$primary\_tag["$tag"] = "$value"  
> > }
> 
> > Thanks for help !
> 
> > I'll send an SOS to the World  
> > I hope that someone gets my  
> > Message in a forum

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [May 15, 2012, 11:30am UTC](https://discuss.elastic.co/t/convert-into-json-for-perl-module/7701/4 "2012-05-15T11:30:59Z")

</div>

Hi Jerome

> You said i can update my docs using "get" and "index".  
> If I must re-index my documents, I must necessarily provide ALL the  
> document data or can I just provide only the new data?
> 
> This is why I used the update () function, it allows to update by  
> entering only the new data.

You can get() your existing doc, which will include all of the data  
already in the doc, then add your new data, and reindex it.

This is essentially the same thing that update() does internally.

The advantage of of using get() plus index() is that you can do it all  
in a language you are familiar with, as opposed to trying to debug mvel.

> In your example, if I understand correctly, you add in a hash the  
> document present in Elasticsearch. Then you get couples tags /  
> values ââand add them before reindex the hash.
> 
> I think a hash with only one level of imbrication and my document has  
> more than one. And my real problem is this second imbrication level, I  
> managed to be indexed by putting everything on one level but that does  
> not make practical research. I also looking for a way to automatically  
> index, data on several levels, this datas are different each time  
> making it impossible to define them by hand.

I'm not sure what imbrication means, but I assume you're talking about a  
structure like this:

$doc = {  
name =\> 'Foo',  
one =\> {  
two =\> {  
three =\> {  
tags =\> ['foo','bar','baz'],  
}  
}  
}  
}

This is easy to do in Perl. For instance, I could add a new tag to  
'tags' with:

```
push @{ $doc->{one}{two}{three} }, $new_tag;

```

You don't need to pre-create that structure. If your $doc looked like  
this:

$doc = { name =\> 'Foo' }

and you did this:

```
push @{ $doc->{one}{two}{three} }, $new_tag;

```

then you'd end up with this:

$doc = {  
name =\> 'Foo',  
one =\> {  
two =\> {  
three =\> {  
tags =\> ['foo','bar','baz'],  
}  
}  
}  
}

But this has nothing to do with Elasticsearch - it's basic Perl  
references. Perhaps you should read 'perlreftut':

[http://perldoc.perl.org/perlreftut.html](http://perldoc.perl.org/perlreftut.html)

clint

---

<div class="post-metadata">

**Author:** ![Jerome](https://avatars.discourse-cdn.com/v4/letter/j/a587f6/32.png) [@Jerome](https://discuss.elastic.co/u/Jerome)\
**Post date:** [May 15, 2012, 12:43pm UTC](https://discuss.elastic.co/t/convert-into-json-for-perl-module/7701/5 "2012-05-15T12:43:38Z")

</div>

Thanks.

> I'm not sure what imbrication means, but I assume you're talking about a  
> structure

Yes i was talking about structure.

I never use hash like that, i learn something today. ^^

But i tried it and it's works !

![](https://us1.discourse-cdn.com/elastic/original/3X/2/d/2d1811979bb5a2f8be94d32489c209601735982a.png)  
Thank you very much for this help !

On 15 mai, 13:30, Clinton Gormley [cl...@traveljury.com](mailto:cl...@traveljury.com) wrote:

> Hi Jerome
> 
> > You said i can update my docs using "get" and "index".  
> > If I must re-index my documents, I must necessarily provide ALL the  
> > document data or can I just provide only the new data?
> 
> > This is why I used the update () function, it allows to update by  
> > entering only the new data.
> 
> You can get() your existing doc, which will include all of the data  
> already in the doc, then add your new data, and reindex it.
> 
> This is essentially the same thing that update() does internally.
> 
> The advantage of of using get() plus index() is that you can do it all  
> in a language you are familiar with, as opposed to trying to debug mvel.
> 
> > In your example, if I understand correctly, you add in a hash the  
> > document present in Elasticsearch. Then you get couples tags /  
> > values ​​and add them before reindex the hash.
> 
> > I think a hash with only one level of imbrication and my document has  
> > more than one. And my real problem is this second imbrication level, I  
> > managed to be indexed by putting everything on one level but that does  
> > not make practical research. I also looking for a way to automatically  
> > index, data on several levels, this datas are different each time  
> > making it impossible to define them by hand.
> 
> I'm not sure what imbrication means, but I assume you're talking about a  
> structure like this:
> 
> $doc = {  
> name =\> 'Foo',  
> one =\> {  
> two =\> {  
> three =\> {  
> tags =\> ['foo','bar','baz'],  
> }  
> }  
> }  
> }
> 
> This is easy to do in Perl. For instance, I could add a new tag to  
> 'tags' with:
> 
> ```
> push @{ $doc->{one}{two}{three} }, $new_tag;
> 
> ```
> 
> You don't need to pre-create that structure. If your $doc looked like  
> this:
> 
> $doc = { name =\> 'Foo' }
> 
> and you did this:
> 
> ```
> push @{ $doc->{one}{two}{three} }, $new_tag;
> 
> ```
> 
> then you'd end up with this:
> 
> $doc = {  
> name =\> 'Foo',  
> one =\> {  
> two =\> {  
> three =\> {  
> tags =\> ['foo','bar','baz'],  
> }  
> }  
> }  
> }
> 
> But this has nothing to do with Elasticsearch - it's basic Perl  
> references. Perhaps you should read 'perlreftut':
> 
> [perlreftut - Mark's very short tutorial about references - Perldoc Browser](http://perldoc.perl.org/perlreftut.html)
> 
> clint

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:28am UTC](https://discuss.elastic.co/t/convert-into-json-for-perl-module/7701/6 "2017-07-06T03:28:51Z")

</div>


