# How to (pre)filter data used in a visualization?

**URL:** <https://discuss.elastic.co/t/how-to-pre-filter-data-used-in-a-visualization/218325>\
**Category:** Kibana\
**Created:** [February 7, 2020, 10:51am UTC](https://discuss.elastic.co/t/how-to-pre-filter-data-used-in-a-visualization/218325 "2020-02-07T10:51:42Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![AxelR](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/axelr/32/61269_2.png) [@AxelR](https://discuss.elastic.co/u/AxelR)\
**Post date:** [February 7, 2020, 10:51am UTC](https://discuss.elastic.co/t/how-to-pre-filter-data-used-in-a-visualization/218325/1 "2020-02-07T10:51:43Z")

</div>

Hello,

I'm trying to build an histogram which must count only the first occurrences (chronologically speaking) of all recorded events for the corresponding period (in my data, a specific event can occur several times with a different outcome each time). From what I gathered so far, this might be done by using data aggregations.

However, I'm having trouble finding examples in Kibana on how to give an aggregation as in input to filter the elements being counted...

I'm not sure that I have been clear enough, do not hesitate to ask for further info.

Thanks in advance everybody 🙂

---

<div class="post-metadata">

**Author:** ![flash1293](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/flash1293/32/41227_2.png) [@flash1293](https://discuss.elastic.co/u/flash1293)\
**Post date:** [February 12, 2020, 12:52pm UTC](https://discuss.elastic.co/t/how-to-pre-filter-data-used-in-a-visualization/218325/2 "2020-02-12T12:52:27Z")

</div>

Hi @AxelR,

I'm not aware of a way to do this during query time - you have to make sure your data is already indexed in a "de-duped" way.

One way to do this is to set up a transform job as described here: [https://www.elastic.co/guide/en/elasticsearch/reference/7.6/put-transform.html](https://www.elastic.co/guide/en/elasticsearch/reference/7.6/put-transform.html)  
This will continuously run aggregations on your data and store the pre-aggregated results so you can build visualizations (like a histogram) on top of it

Let's say you detect an event being duplicated by its `event_id` field. By grouping by `event_id` and adding the min of your `timestamp` field to the result, the result index will only contain one document per event id (with the first occurence as its timestamp)

```auto
PUT _transform/first_event_transform
{
  "source": {
    "index": "all_events",
  },
  "pivot": {
    "group_by": {
      "event": {
        "terms": {
          "field": "event_id"
        }
      }
    },
    "aggregations": {
      "first_occurrence": {
        "min": {
          "field": "timestamp"
        }
      }
    }
  },
  // ...
}

```

Based on this index you can create your histogram aggregation as usual.

---

<div class="post-metadata">

**Author:** ![AxelR](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/axelr/32/61269_2.png) [@AxelR](https://discuss.elastic.co/u/AxelR)\
**Post date:** [February 12, 2020, 3:13pm UTC](https://discuss.elastic.co/t/how-to-pre-filter-data-used-in-a-visualization/218325/3 "2020-02-12T15:13:26Z")

</div>

Thanks, I think I understand the general idea!  
Since the transform job creates a new index, I guess I also have to store the outcome (correct/incorrect) of the first test, do I?

---

<div class="post-metadata">

**Author:** ![flash1293](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/flash1293/32/41227_2.png) [@flash1293](https://discuss.elastic.co/u/flash1293)\
**Post date:** [February 12, 2020, 4:26pm UTC](https://discuss.elastic.co/t/how-to-pre-filter-data-used-in-a-visualization/218325/4 "2020-02-12T16:26:10Z")

</div>

If you want to visualize it, yes - it's kind of similar to a sql query grouping by the event id - all fields you want to access to have to define together with the aggregation (because there could be multiple documents within each group).

For fetching the outcome of the first event, you probably have to resort to a scripted metric: [https://www.elastic.co/guide/en/elasticsearch/reference/master/search-aggregations-metrics-scripted-metric-aggregation.html](https://www.elastic.co/guide/en/elasticsearch/reference/master/search-aggregations-metrics-scripted-metric-aggregation.html)

If you are just after all fields of the first document, a solution using a logstash pipeline is probably a better fit: [https://www.elastic.co/blog/how-to-find-and-remove-duplicate-documents-in-elasticsearch](https://www.elastic.co/blog/how-to-find-and-remove-duplicate-documents-in-elasticsearch)

You need an additional service (Logstash) to process the data, but it's more straight forward for this kind of thing. Transforms are better suited if you just want to access aggregations of the groups.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 11, 2020, 4:26pm UTC](https://discuss.elastic.co/t/how-to-pre-filter-data-used-in-a-visualization/218325/5 "2020-03-11T16:26:17Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
