# Elasticsearch Transforms

**URL:** <https://discuss.elastic.co/t/elasticsearch-transforms/265874>\
**Category:** Elasticsearch\
**Tags:** transforms\
**Created:** [March 1, 2021, 9:28pm UTC](https://discuss.elastic.co/t/elasticsearch-transforms/265874 "2021-03-01T21:28:00Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![karanp](https://avatars.discourse-cdn.com/v4/letter/k/b5e925/32.png) [@karanp](https://discuss.elastic.co/u/karanp)\
**Post date:** [March 1, 2021, 9:28pm UTC](https://discuss.elastic.co/t/elasticsearch-transforms/265874/1 "2021-03-01T21:28:00Z")

</div>

Hello, So recently I started using es pivot transform. At every stage I am getting some incomplete data and using scripted metric I am completing it and sending it to the new index. Currently what this does is that every time a new document with same group by id comes in it updates the data in the new index. My question is that is there a way that instead of updating the document in the new index can I create a new document everytime?  
For eg:  
"""  
doc1 :  
"state" : "create",  
"id" : "01",  
"time" : "xyz" ,  
"started" : True

```
  doc2 : 
           "state" : "update",
           "id" : "01",
            "time" : "abc",
             "data" : "added new data"

```

So, currently in the new index there will only be one document like this:  
doc :  
"state" : "update",  
"id" : "01",  
"time" : "abc" ,  
"started" : True,  
"data" : "added new data"

As you can see the data gets updated after the second docx comes in. what I would like to see in the new index is this:  
doc1 :  
"state" : "create",  
"id" : "01",  
"time" : "xyz" ,  
"started" : True

```
     doc2 : 
             "state" : "update",
           "id" : "01",
            "time" : "abc" ,
            "started" : True,
             "data" : "added new data"

```

Would like to know if there is a way to do this with transforms.

---

<div class="post-metadata">

**Author:** ![przemekwitek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/przemekwitek/32/79526_2.png) [@przemekwitek](https://discuss.elastic.co/u/przemekwitek)\
**Post date:** [March 2, 2021, 6:29am UTC](https://discuss.elastic.co/t/elasticsearch-transforms/265874/2 "2021-03-02T06:29:48Z")

</div>

Hi,  
It is by design that there is exactly one document corresponding to the given group-by id in the destination index. So if you use `id` field in group-by section on transform pivot config, the transform will do just this: take all the documents with the given `id` and summarize it in one document in the destination index.  
If you'd like to have 1:1 relation between documents in source and destination, you should be using an id that is unique among documents, something like "event id" or maybe even "timestamp".

If you do not need continuous updates functionality, you could also try using reindex + scripted field. There is another ticket in which this is discussed:

> [@Reindex data with a new Field](https://discuss.elastic.co/t/reindex-data-with-a-new-field/115073/2):
>
> There is an example of using reindex to modify the document in the docs [here](https://www.elastic.co/guide/en/elasticsearch/reference/current/docs-reindex.html#docs-reindex-change-name). It changes the name of a field but you can use the same script construct to do what you want to do. I suspect the script looks something like "script": "ctx.\_source['timestamp\_hour'] = ctx.\_source['@timestamp'].getHourOfDay(). You are right that it will be much better at search time to use this new field. Watch the time zone, btw. I believe the hour that you get in this case is UTC. You could also do this with \_update…

---

<div class="post-metadata">

**Author:** ![przemekwitek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/przemekwitek/32/79526_2.png) [@przemekwitek](https://discuss.elastic.co/u/przemekwitek)\
**Post date:** [March 2, 2021, 8:44am UTC](https://discuss.elastic.co/t/elasticsearch-transforms/265874/3 "2021-03-02T08:44:57Z")

</div>

[Update]  
I'm not sure about your particular use-case, but I could suggest one more option you might be interested in: ingest pipeline + script processor.  
The script defined in the pipeline would be executed on each document without any grouping (so it won't override existing docs).

See documentation: [Script processor | Elasticsearch Reference [7.11] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/script-processor.html)

If you are still unsure which of these options (reindex+script, ingest+script, transform) is suitable for you, please share more about your use-case.

---

<div class="post-metadata">

**Author:** ![karanp](https://avatars.discourse-cdn.com/v4/letter/k/b5e925/32.png) [@karanp](https://discuss.elastic.co/u/karanp)\
**Post date:** [March 2, 2021, 9:00pm UTC](https://discuss.elastic.co/t/elasticsearch-transforms/265874/4 "2021-03-02T21:00:06Z")

</div>

@przemekwitek Thank you for the reply. Let me elaborate a bit more on my use case. So, for my case a single event can have around 5 or more states(Unsure about the total number of events but it will be very high). Now, in the first state lets say it comes with 10 fields. After this state it will always come with fields which are either added or updated. So, the 2nd, 3rd and 4th states lets say comes with 4 fields and 5th with 6 field. My task is that every time a state comes in I have to create a docx with all the fields that have come in till that point and update the fields if any update has come in. So after every state there will be an extra docx in the new index.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 30, 2021, 9:00pm UTC](https://discuss.elastic.co/t/elasticsearch-transforms/265874/5 "2021-03-30T21:00:42Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
