# Distinct count with filter

**URL:** <https://discuss.elastic.co/t/distinct-count-with-filter/155795>\
**Category:** Elasticsearch\
**Created:** [November 7, 2018, 9:03pm UTC](https://discuss.elastic.co/t/distinct-count-with-filter/155795 "2018-11-07T21:03:50Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![winder](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/winder/32/54090_2.png) [@winder](https://discuss.elastic.co/u/winder)\
**Post date:** [November 7, 2018, 9:03pm UTC](https://discuss.elastic.co/t/distinct-count-with-filter/155795/1 "2018-11-07T21:03:50Z")

</div>

Hoping to get some guidance here.

I'm trying to correlate session-id's from two different events, a **heartbeat** sent every 10 minutes and a **disconnect** which could be sent any time. The goal is to get the number of active sessions for the last 10 minutes in a kibana visualization.

I don't think this is possible with the raw events in Elasticsearch, is that correct?

Would a logstash pipeline be the general approach here?

It seems like I should be able to use something [like this Aggregate Filter example](https://www.elastic.co/guide/en/logstash-versioned-plugins/current/v2.9.0-plugins-filters-aggregate.html#v2.9.0-plugins-filters-aggregate-example3), using the **heartbeat** to add session-ids to a periodic "active sessions" event and the **shutdown** to remove the session-id. Does this seem like a reasonable approach or is there something simpler?

If this is the approach, can I initialize the next "active sessions" array from the most recent active-sessions event?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [November 9, 2018, 10:34am UTC](https://discuss.elastic.co/t/distinct-count-with-filter/155795/2 "2018-11-09T10:34:08Z")

</div>

As you indicate, Logstash has some features to join related documents in the ingest stream.

Another general approach is to land the events in the index first and then use a job to periodically (every few seconds?) update a separate "session" index with the latest recorded activities in the event index.  
See: "[entity-centric indexing](https://twitter.com/elasticmark/status/1009380268409610240)".

---

<div class="post-metadata">

**Author:** ![winder](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/winder/32/54090_2.png) [@winder](https://discuss.elastic.co/u/winder)\
**Post date:** [November 9, 2018, 2:49pm UTC](https://discuss.elastic.co/t/distinct-count-with-filter/155795/3 "2018-11-09T14:49:52Z")

</div>

Thanks Mark, this entity-centric indexing technique looks like exactly what I'm aiming for!

I'm still curious if such an index could be created in logstash, rather than introducing an extra set of scripts. It looks like it might be possible to recreate with the following:

1. Group events together and apply a timeout to make sure they are updated at the required interval: [like this example](https://www.elastic.co/guide/en/logstash/current/plugins-filters-aggregate.html#plugins-filters-aggregate-example5).
2. Using an update script with the ES Output Plugin [appears to be supported](https://www.elastic.co/guide/en/logstash/current/plugins-outputs-elasticsearch.html#plugins-outputs-elasticsearch-script).

I'll probably run my tests using your `ESEntityCentricIndexing` script and see if it can be rolled into Logstash for production.

For now I'll mark this as solved and create any further questions in the Logstash section

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [November 9, 2018, 3:05pm UTC](https://discuss.elastic.co/t/distinct-count-with-filter/155795/4 "2018-11-09T15:05:04Z")

</div>

> [@winder](#):
>
> I'm still curious if such an index could be created in logstash,

I'm not a logstash expert but I'd suggest checking what happens when you use a system that relies on joining things together using a transient blob of memory. The questions I would have are:

1. Do I have to route all related events through the same logstash process?
2. Does the memory grow endlessly waiting for "start" events to match their equivalent "end" event?
3. How much memory do I need to hold a window big enough to tally all starts with ends?
4. What happens to in-flight sessions when the power goes off?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 7, 2018, 3:19pm UTC](https://discuss.elastic.co/t/distinct-count-with-filter/155795/5 "2018-12-07T15:19:40Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
