# How to add field that is only present in some documents?

**URL:** <https://discuss.elastic.co/t/how-to-add-field-that-is-only-present-in-some-documents/295283>\
**Category:** Kibana\
**Tags:** transforms\
**Created:** [January 24, 2022, 8:40pm UTC](https://discuss.elastic.co/t/how-to-add-field-that-is-only-present-in-some-documents/295283 "2022-01-24T20:40:28Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![katja1](https://avatars.discourse-cdn.com/v4/letter/k/b487fb/32.png) [@katja1](https://discuss.elastic.co/u/katja1)\
**Post date:** [January 24, 2022, 8:40pm UTC](https://discuss.elastic.co/t/how-to-add-field-that-is-only-present-in-some-documents/295283/1 "2022-01-24T20:40:28Z")

</div>

Hi,

I have the following situation: Logs with the same ID [grouping ID] and different fields.

```auto

grouping ID: 1 organization: xyz message: abc
grouping ID: 1 organization: xyz message: abc
grouping ID: 1 person: name

```

I would like to use Transforms (or anything you suggest, but unfortunately not Logstash), to group all the logs with same ID and have the "enriched" log at the end, meaning it contains all the fields with information:

```auto
grouping ID: 1 organization: xyz message: abc person: name                                                           

```

So far, I only had situations, where I could use 'group\_by' in order to add fields. However, here this is not possible.

I would appreciate your help!

---

<div class="post-metadata">

**Author:** ![Tomo\_M](https://avatars.discourse-cdn.com/v4/letter/t/848f3c/32.png) [@Tomo\_M](https://discuss.elastic.co/u/Tomo_M)\
**Post date:** [January 25, 2022, 12:06am UTC](https://discuss.elastic.co/t/how-to-add-field-that-is-only-present-in-some-documents/295283/2 "2022-01-25T00:06:44Z")

</div>

What is your current transform configuration and its result?

---

<div class="post-metadata">

**Author:** ![katja1](https://avatars.discourse-cdn.com/v4/letter/k/b487fb/32.png) [@katja1](https://discuss.elastic.co/u/katja1)\
**Post date:** [January 25, 2022, 7:39am UTC](https://discuss.elastic.co/t/how-to-add-field-that-is-only-present-in-some-documents/295283/3 "2022-01-25T07:39:33Z")

</div>

```auto
  POST _transform/_preview
    {
 "source": {
   "index": [
        "[my_index]"
    ],
     "query": {
      "bool": {
        "must": [
          {
            "exists": {"field":"grouping_ID" }   
          }
        ]
        
      }
    }
  },
  "pivot": {
    "group_by": {
      "grouping_ID": {
        "terms": {
          "field": "grouping_ID"
       }
     }
   },
   
  "aggregations": {
        
      "organization":{
        "terms": {
          "field": "organization.keyword"
        }
      },
      "message":{
       "terms":{
        "field": "message.keyword"
       }
     },
      "person":{
       "terms":{
        "field": "person.keyword"
       }
     }
    }
        
  },
  "description": "test",
  "dest": {
    "index": "[my_new_index]"
    },
  "frequency": "1m",
  "sync": {
    "time": {
      "field": "time.iso8601",
      "delay": "60s"
      }
    }
  }

```

This is the closes I got, meaning the destination log includes all the needed fields, but I am aware that the "terms" aggregation is not right for the task.  
In my previous work with transforms, I could simply use "group\_by" to add the needed fields because all my logs included this field, but if I would do this here (for example, if I would additionally do group\_by organization), all the logs not having this field would be lost. I hope you understand what I'm trying to explain.

Probably I will need a scripted metric aggregation, but I didn't get far with it.

---

<div class="post-metadata">

**Author:** ![przemekwitek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/przemekwitek/32/79526_2.png) [@przemekwitek](https://discuss.elastic.co/u/przemekwitek)\
**Post date:** [January 25, 2022, 2:01pm UTC](https://discuss.elastic.co/t/how-to-add-field-that-is-only-present-in-some-documents/295283/4 "2022-01-25T14:01:51Z")

</div>

> [@katja1](#):
>
> This is the closes I got, meaning the destination log includes all the needed fields, but I am aware that the "terms" aggregation is not right for the task.

Well, it may be ok if it works for your use-case.

> [@katja1](#):
>
> In my previous work with transforms, I could simply use "group\_by" to add the needed fields because all my logs included this field, but if I would do this here (for example, if I would additionally do group\_by organization), all the logs not having this field would be lost. I hope you understand what I'm trying to explain.

Yes, I see the problem. When you group by a fixed set of fields, the transform requires them to be present. This is working as intended.

> [@katja1](#):
>
> Probably I will need a scripted metric aggregation, but I didn't get far with it.

If `terms` don't work for you for any reason, then I think scripted metric aggregation is the way to go. It's a bit more involved, especially if you didn't use the scripting language before, but gives the greatest flexibility.

---

<div class="post-metadata">

**Author:** ![katja1](https://avatars.discourse-cdn.com/v4/letter/k/b487fb/32.png) [@katja1](https://discuss.elastic.co/u/katja1)\
**Post date:** [January 25, 2022, 2:31pm UTC](https://discuss.elastic.co/t/how-to-add-field-that-is-only-present-in-some-documents/295283/5 "2022-01-25T14:31:22Z")

</div>

It is not working for my use-case - with terms aggregation I was only able to bring them in the destination log, but the field type is unknown, and the values of the field are written in {}, for example { "organizationNameA" : 3}, with the number 3 representing the count of logs containing that field. So this is a bit unexpected.

---

<div class="post-metadata">

**Author:** ![Tomo\_M](https://avatars.discourse-cdn.com/v4/letter/t/848f3c/32.png) [@Tomo\_M](https://discuss.elastic.co/u/Tomo_M)\
**Post date:** [January 25, 2022, 4:24pm UTC](https://discuss.elastic.co/t/how-to-add-field-that-is-only-present-in-some-documents/295283/6 "2022-01-25T16:24:06Z")

</div>

If your problem caused by that some documents miss the field to group\_by and grouping such missing documents together is acceptable, one workaround could be to set ingest pipeline to fill such missing field by 'NULL' value.

```auto
PUT _ingest/pipeline/set_NULL
{
  "description": "set 'NULL' for missing fields",
  "processors": [
    {"set":{
      "field":"organization",
      "value": "NULL",
      "if":"!ctx.containsKey('organization')"}}
  ]
}

PUT /your_index/_settings
{
  "index": {
    "default_pipeline": "set_NULL"
  }
}

# apply ingest_pipeline to exiting documents.
POST your_index/_update_by_query
{
  "query":{
    "match_all": {}
  }
}

```

> [@katja1](#):
>
> However, here this is not possible.

I think that sharing not only sample data that worked well, but also **sample data that didn't exactly work well** , and presenting what the desired output would be, will advance the discussion.

> [@katja1](#):
>
> with the number 3 representing the count of logs containing that field

I made a sample `scripted metric aggregation` to pick up unique values as an array, something like named "unique values aggregation".  
(This script is inspired from [this post](https://discuss.elastic.co/t/how-to-aggregate-data-on-elastic-sent-by-logstash/205140/2).)

```auto
"unique_organization":{
  "scripted_metric": {
   "init_script": "state.set = new HashSet()",
    "map_script": "if (params['_source'].containsKey(params.field)) {state.set.add(params['_source'][params.field])}",
    "combine_script": "return state.set",
    "reduce_script": "def ret = new HashSet(); for (s in states) {for (k in s) {ret.add(k);}} return ret",
    "params":{
      "field": "organization"
    }
  }
}

```

---

<div class="post-metadata">

**Author:** ![katja1](https://avatars.discourse-cdn.com/v4/letter/k/b487fb/32.png) [@katja1](https://discuss.elastic.co/u/katja1)\
**Post date:** [January 26, 2022, 9:33am UTC](https://discuss.elastic.co/t/how-to-add-field-that-is-only-present-in-some-documents/295283/7 "2022-01-26T09:33:44Z")

</div>

Thank you for your reply. I will try to apply your suggestions and will report the update.

Regarding the following:

> [@Tomo\_M](#):
>
> I think that sharing not only sample data that worked well, but also **sample data that didn't exactly work well** , and presenting what the desired output would be, will advance the discussion.

I'm sorry if I wasn't clear enough. What I meant was that I cannot use 'group\_by' groupingID, organization, message and person, because not all logs contain all fields. All I can do is use group\_by on groupingID and then the other fields I need to "bring" to the final log in some other way. My desired output would be as I wrote in the post - a log that contains all the fields with all the information.

---

<div class="post-metadata">

**Author:** ![przemekwitek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/przemekwitek/32/79526_2.png) [@przemekwitek](https://discuss.elastic.co/u/przemekwitek)\
**Post date:** [January 26, 2022, 12:30pm UTC](https://discuss.elastic.co/t/how-to-add-field-that-is-only-present-in-some-documents/295283/8 "2022-01-26T12:30:47Z")

</div>

IMO there are 2 more features that you should take into account when designing your solution:

1. specify `person` in the `group_by` clause and set `missing_bucket` to `true` so that the buckets **without** a `person` field are also returned

You can read more about `missing_bucket` in the docs: [Composite aggregation | Reference](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-composite-aggregation.html#_missing_bucket)

1. specify `person` in the `aggregations` clause but use `top_metrics` aggregation instead of `terms`.

You can read more about `top_metrics` in this blog:

> [@Dec 11th, 2021: \[en\] On a road trip with Transform](https://discuss.elastic.co/t/dec-11th-2021-en-on-a-road-trip-with-transform/290532):
>
> Are you already tired of hearing Christmas songs? One of the classics is "Driving home for Christmas". The lyrics of that song have not much content, Chris Rea basically tells about a long boring journey in his car driving home. In this post we accompany this song and spice it with new bits and pieces about [Transforms](https://www.elastic.co/guide/en/elasticsearch/reference/current/transforms.html). Transforms provide an easy way to summarize data. And it's been so long The aim of a search engine - like Elasticsearch - is to provide relevant results quickly. For your family…

---

<div class="post-metadata">

**Author:** ![katja1](https://avatars.discourse-cdn.com/v4/letter/k/b487fb/32.png) [@katja1](https://discuss.elastic.co/u/katja1)\
**Post date:** [January 31, 2022, 12:38pm UTC](https://discuss.elastic.co/t/how-to-add-field-that-is-only-present-in-some-documents/295283/9 "2022-01-31T12:38:47Z")

</div>

Thanks for the reply. I'm still playing around with it, but your reply brings valuable tips.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 28, 2022, 12:39pm UTC](https://discuss.elastic.co/t/how-to-add-field-that-is-only-present-in-some-documents/295283/10 "2022-02-28T12:39:08Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
