# Filtering a Pivot Transform

**URL:** <https://discuss.elastic.co/t/filtering-a-pivot-transform/299371>\
**Category:** Elasticsearch\
**Tags:** transforms\
**Created:** [March 10, 2022, 9:52pm UTC](https://discuss.elastic.co/t/filtering-a-pivot-transform/299371 "2022-03-10T21:52:23Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![iamtheschmitzer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/iamtheschmitzer/32/96988_2.png) [@iamtheschmitzer](https://discuss.elastic.co/u/iamtheschmitzer)\
**Post date:** [March 10, 2022, 9:52pm UTC](https://discuss.elastic.co/t/filtering-a-pivot-transform/299371/1 "2022-03-10T21:52:23Z")

</div>

I have a pivot transform that calculates a duration between two timestamped documents by grouping them, matching like documents, then aggregating the min timestamp, the max timestamp and using a script to calculate the duration between those two:

```auto
PUT _transform/transform-name
{
  "source": {
    "index": "source-index-name"
  },
  "dest" : { 
    "index" : "dest-index-name"
  },
  "pivot": {
    "group_by": { 
      "linkId": { "terms": { "field": "linkId" }},
      "string_metadata_hash": { "terms": { "field": "string_metadata_hash" }},
      "string_calculation_job_type": {
        "terms": {
          "script": {
            "source": """
              def m = /^calculation_job_(.*)_([^_]+)$/.matcher(doc['eventType'].value);
              if (m.matches()) {
                return m.group(1);
              }
            """
          }
        }
      }
    },
    "aggregations": {
      "metadata": {
        "scripted_metric": {
          "init_script": "state.doc = new HashMap()",
          "map_script": "if (state.doc.isEmpty()){state.doc = new HashMap(params['_source'].metadata)}",
          "combine_script": "return state.doc",
          "reduce_script": "return states[0]"
        }
      },
      "start": {
        "min": {
          "field": "timestamp"
        }
      },
      "complete": {
        "max": {
          "field": "timestamp"
        }
      },
      "duration_sec": {
        "bucket_script": {
          "buckets_path": {
            "start": "start.value",
            "complete": "complete.value"
          },
          "script": "return (params.complete - params.start)/1000;"
        }
      }
    }
  },
   "sync": {
    "time": { 
      "field": "timestamp",
      "delay": "60s"
    }
  }
}

```

This matches event Type `calculation_job_foo_start` with `calculation_job_foo_complete` to calculate the duration between the two.

However, it also results in a document when there is only a `calculation_job_foo_start` and no `calculation_job_foo_complete` - which results in a `start` and `complete` times that are the same (both start times) and a `duration_sec` of 0

Is there a way to prevent output to dest index until there is both a start and complete event type? In other words, filter the output?

---

<div class="post-metadata">

**Author:** ![Hendrik\_Muhs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hendrik_muhs/32/25802_2.png) [@Hendrik\_Muhs](https://discuss.elastic.co/u/Hendrik_Muhs)\
**Post date:** [March 14, 2022, 6:47am UTC](https://discuss.elastic.co/t/filtering-a-pivot-transform/299371/2 "2022-03-14T06:47:53Z")

</div>

You can define an ingest pipeline for `dest`. In the ingest pipeline you can use a [drop processor](https://www.elastic.co/guide/en/elasticsearch/reference/current/drop-processor.html) to prevent documents from getting indexed.

---

<div class="post-metadata">

**Author:** ![iamtheschmitzer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/iamtheschmitzer/32/96988_2.png) [@iamtheschmitzer](https://discuss.elastic.co/u/iamtheschmitzer)\
**Post date:** [March 18, 2022, 6:55pm UTC](https://discuss.elastic.co/t/filtering-a-pivot-transform/299371/3 "2022-03-18T18:55:03Z")

</div>

Thanks @Hendrik_Muhs did not end up needing the ingest pipeline here, but did use the concept in other ways. The premise for my question ended up being a false one.

The duration\_sec of 0 is temporary, and not a problem, as long as the `delay` is long enough to find all the documents needed.

---

<div class="post-metadata">

**Author:** ![iamtheschmitzer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/iamtheschmitzer/32/96988_2.png) [@iamtheschmitzer](https://discuss.elastic.co/u/iamtheschmitzer)\
**Post date:** [March 24, 2022, 7:11pm UTC](https://discuss.elastic.co/t/filtering-a-pivot-transform/299371/4 "2022-03-24T19:11:39Z")

</div>

So I did try this approach but cannot get it to work @Hendrik_Muhs

getting a compile error on `doc` within the `if`

```auto
PUT _ingest/pipeline/name-here
{
  "processors" : [
    {
      "drop": {
        "if": "doc['duration_sec'] == 0.0"
      }
    },
    {
      "set": {
        "field": "timestamp",
        "value": "{{_ingest.timestamp}}"
      }
    }
  ]
}

```

---

<div class="post-metadata">

**Author:** ![Hendrik\_Muhs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hendrik_muhs/32/25802_2.png) [@Hendrik\_Muhs](https://discuss.elastic.co/u/Hendrik_Muhs)\
**Post date:** [March 29, 2022, 9:35am UTC](https://discuss.elastic.co/t/filtering-a-pivot-transform/299371/5 "2022-03-29T09:35:37Z")

</div>

When inserted into the ingest pipeline it isn't a doc yet, you need to use the [ingest context](https://www.elastic.co/guide/en/elasticsearch/painless/8.1/painless-ingest-processor-context.html).

Try replacing `doc` with `ctx`.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 26, 2022, 9:36am UTC](https://discuss.elastic.co/t/filtering-a-pivot-transform/299371/6 "2022-04-26T09:36:29Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
