# Transforms: do I need to filter source for time-series data?

**URL:** <https://discuss.elastic.co/t/transforms-do-i-need-to-filter-source-for-time-series-data/276057>\
**Category:** Elasticsearch\
**Tags:** transforms\
**Created:** [June 15, 2021, 9:08pm UTC](https://discuss.elastic.co/t/transforms-do-i-need-to-filter-source-for-time-series-data/276057 "2021-06-15T21:08:01Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![rokcarl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rokcarl/32/53239_2.png) [@rokcarl](https://discuss.elastic.co/u/rokcarl)\
**Post date:** [June 15, 2021, 9:08pm UTC](https://discuss.elastic.co/t/transforms-do-i-need-to-filter-source-for-time-series-data/276057/1 "2021-06-15T21:08:01Z")

</div>

I'm doing a transformation of time-series/events data that does not change for the past.

I've talked to an Elastic engineer Ben on Slack and it seems like I _might_ have stumbled upon a bug.

I'm using this [index config](https://gist.github.com/rokcarl/42c9d06a29c7d6db830c1840be1c49a8) and this [transform config](https://gist.github.com/rokcarl/8fa90cdf58d836d2ad2f3fe9cb43b649).

The problem is that each iteration seems to be processing all the data. I've been running the transform on our indices (`stats-[date]`) and here are the stats:

- 204B documents processed
- 14 min total indexing time
- 108 hours total search time
- 25 seconds total processing time
- "average" documents processed: 221M
- "average" checkpoint duration: 7 min

We use curator to delete old indices and so we always have 6 indices, each with under 50M docs, if I sum them up, it currently comes at 237M, right around the average of each run. So it looks like it processes all documents each time.

Ben says that there have been some improvements since 7.7.0, but I'm on 7.9.1.

Here's my CPU utilization and load across this week of turning it on.

 ![Screenshot 2021-06-15 at 07.47.43](https://us1.discourse-cdn.com/elastic/original/3X/7/d/7d98e3ed2e4af25e1cb3fd69d960cc2faa4ac8d1.png)

Since then, I have modified my transform to do a filter query on the `timestamp` field, here's the [new transform](https://gist.github.com/4c599ea06a73bc60305577808f0eba9d) with a `range` query in the `source`. This time, however, it seems like it's running fine. Is this expected behaviour?

---

<div class="post-metadata">

**Author:** ![Hendrik\_Muhs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hendrik_muhs/32/25802_2.png) [@Hendrik\_Muhs](https://discuss.elastic.co/u/Hendrik_Muhs)\
**Post date:** [June 16, 2021, 6:52am UTC](https://discuss.elastic.co/t/transforms-do-i-need-to-filter-source-for-time-series-data/276057/2 "2021-06-16T06:52:42Z")

</div>

Transform minimizes the amount of data that needs to be re-processed as explained in the [docs](https://www.elastic.co/guide/en/elasticsearch/reference/current/transform-checkpoints.html).

In your case transform looks for customers that have changed and _time buckets_ that have changed. Here is the problem: You have a fixed interval of `1d`, so that's a rather large bucket. On the other side you run the transform every `10m`. Transform will recalculate the intermediate bucket of "today" every 10 minutes, which means it will rewrite this bucket 144 times until it decides to ignore it. To reduce the rewrites you can set `frequency` to a higher value, meaning running less often.

Do you need the intermediate bucket results at a 10 minute frequency?

Or you change the `date_histogram` to less than a day, this will produce more documents, but avoid a lot of rewrites. To get back to daily results you can use aggregations on the transform destination index.

In your new configuration you use a range query to limit the amount of data transform reads, however you still group daily but only allow a 3 hour window back in time. This will produce wrong results, because the last bucket will _not_ contain all data of the day. Combining `delay` and the range query, it will only contain `180m-90m=90m` of data.

Please check your results in the transform index. It runs faster but I bet, it does not create correct results.

If you are curious what transform does, which queries it sends, you can enable debug logging to see the queries it sends:

```auto
PUT /_cluster/settings
{
   "transient": {
      "logger.org.elasticsearch.xpack.transform.transforms": "debug"
   }
}

```

You should see given your 1st transform that transform adds a range query to not re-process buckets from previous days. (Note your day starts at 01:30, because you set `delay` to `90m`, which means that transform expects that data can be up to `90m` late, so a data point from yesterday can arrive e.g. at `01:30` today).

Given your 2nd transform you should see 2 range queries combined.

Don't forget to set the log level back to e.g. _info_ or _warning_, whatever your preferences are.

---

<div class="post-metadata">

**Author:** ![rokcarl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rokcarl/32/53239_2.png) [@rokcarl](https://discuss.elastic.co/u/rokcarl)\
**Post date:** [June 16, 2021, 1:38pm UTC](https://discuss.elastic.co/t/transforms-do-i-need-to-filter-source-for-time-series-data/276057/3 "2021-06-16T13:38:15Z")

</div>

> [@](#):
>
> To reduce the rewrites you can set `frequency` to a higher value, meaning running less often.

Yes, that's why I set it to 1 hour in my new config. That is still too low for me because, as you pointed out, I'm going to be calculating the same day many times.

> [@](#):
>
> Do you need the intermediate bucket results at a 10 minute frequency?

Not sure what you mean here. What's an intermediate bucket result? I need the transformed data in 1-day buckets and I don't care that much about the delay and don't care if it comes in with latency.

> [@](#):
>
> Or you change the `date_histogram` to less than a day

Possibly, but that's not what we currently want as that would increase the query complexity by doing aggregations and increase the storage needed, but we'd like to retain the data for a long time.

> [@](#):
>
> In your new configuration you use a range query to limit the amount of data transform reads, however you still group daily but only allow a 3 hour window back in time.

Ah you're right, so I would need, if we do find out that transforms search all the docs al the time, set this to something like 26 hours (24 hours for one bucket + 1.5h for the delay).

I'll turn on debug logging in my staging cluster and check things out. But the main thing that I don't think you touched on is the fact that if I don't do a filter, Elastic will do a scan of all the documents every time and whether that's expected?

---

<div class="post-metadata">

**Author:** ![Hendrik\_Muhs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hendrik_muhs/32/25802_2.png) [@Hendrik\_Muhs](https://discuss.elastic.co/u/Hendrik_Muhs)\
**Post date:** [June 16, 2021, 2:33pm UTC](https://discuss.elastic.co/t/transforms-do-i-need-to-filter-source-for-time-series-data/276057/4 "2021-06-16T14:33:11Z")

</div>

> [@rokcarl](#):
>
> What's an intermediate bucket result?

That last bucket, the bucket of today, remains incomplete until the day is over, so it is an "intermediate result". I got you aren't looking for those and already set `frequency`  
to the maximum possible value. A bucket of today worst-case still gets re-calculated 24 times a day, but at least not 144 times as before.

We have plans to improve scheduling, which will let you run the transform at certain times a day and e.g. only once a day.

> [@rokcarl](#):
>
> But the main thing that I don't think you touched on is the fact that if I don't do a filter, Elastic will do a scan of all the documents every time and whether that's expected?

That's not correct, transform will _not_ do a full scan of all documents. Have you checked the link in my 1st post? The [docs](https://www.elastic.co/guide/en/elasticsearch/reference/current/transform-checkpoints.html) explain it, checkout the section about [change detection heuristics](https://www.elastic.co/guide/en/elasticsearch/reference/current/transform-checkpoints.html#ml-transform-checkpoint-heuristics).

If you prefer a more technical deep-dive, this [pull request](https://github.com/elastic/elasticsearch/issues/54254#issuecomment-604551359) has some numbers. In another [pull request](https://github.com/elastic/elasticsearch/pull/63315) you find a simple drawing about how change detection works.

Nevertheless, you should be able to confirm/see it in action by turning on debug logging.

---

<div class="post-metadata">

**Author:** ![rokcarl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rokcarl/32/53239_2.png) [@rokcarl](https://discuss.elastic.co/u/rokcarl)\
**Post date:** [June 16, 2021, 7:36pm UTC](https://discuss.elastic.co/t/transforms-do-i-need-to-filter-source-for-time-series-data/276057/5 "2021-06-16T19:36:00Z")

</div>

I have read the docs. Twice. They're pretty terse, so didn't take me much time. Have you read my original post? Did you see my numbers? It looks like, with a high degree of certainty, it's parsing every document.

Okay, so I now switched to our staging cluster to enable debugging as you suggested. Here's my [list of slightly modified commands](https://gist.github.com/1350074a4ed4da1b69e0b01bfc1d2958) for staging. Notable changes are that I run it every minute with a one minute delay and without a filter. Then I ran the following to get the logs that mention the docs processed: `docker logs stats_elasticsearch_1 | grep docs_processed | jq ".message"`.

The latest number is 2144822, you can check more of [docker logs here](https://gist.github.com/f387f049aef5aa4ed634160e6d4117a4). Check my indices:

 ![Screenshot 2021-06-16 at 21.33.15](https://us1.discourse-cdn.com/elastic/original/3X/a/7/a7560f531a5e97af80fe9532830651cff11c5186.png)

Clearly it adds up:  
`31601+201998+162229+400426+1350592 = 2146846 ~= 2144822`

---

<div class="post-metadata">

**Author:** ![Hendrik\_Muhs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hendrik_muhs/32/25802_2.png) [@Hendrik\_Muhs](https://discuss.elastic.co/u/Hendrik_Muhs)\
**Post date:** [June 16, 2021, 8:53pm UTC](https://discuss.elastic.co/t/transforms-do-i-need-to-filter-source-for-time-series-data/276057/6 "2021-06-16T20:53:48Z")

</div>

I checked the code again and realized that for `7.9.1` the query is only logged using trace logging.

That's the [line of interest](https://github.com/elastic/elasticsearch/blob/v7.9.1/x-pack/plugin/transform/src/main/java/org/elasticsearch/xpack/transform/transforms/TransformIndexer.java#L761).

So in order to get this line you need:

```auto
PUT /_cluster/settings
{
   "transient": {
      "logger.org.elasticsearch.xpack.transform.transforms": "trace"
   }
}

```

It would be great if you could re-run the test and check the logs for this line, you should see a range query with bucket boundaries and your filter query if you have one.

---

<div class="post-metadata">

**Author:** ![rokcarl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rokcarl/32/53239_2.png) [@rokcarl](https://discuss.elastic.co/u/rokcarl)\
**Post date:** [June 16, 2021, 9:18pm UTC](https://discuss.elastic.co/t/transforms-do-i-need-to-filter-source-for-time-series-data/276057/7 "2021-06-16T21:18:09Z")

</div>

Tracing logs [here](https://gist.github.com/5ce03b20c90950067958cc2433184cf7).

I've parsed some of the logs and I see two queries, one is called the _changes query_, this one looks fine:

```auto
{"size":0,"query":{"bool":{"filter":[{"match_all":{"boost":1.0}},{"range":{"timestamp":{"from":1623877564261,"to":1623877624260,"include_lower":true,"include_upper":false,"format":"epoch_millis","boost":1.0}}}],"adjust_pure_negative":true,"boost":1.0}},"aggregations":{"_transform":{"composite":{"size":500,"sources":[{"customer":{"terms":{"field":"customer","missing_bucket":false,"order":"asc"}}}]}}}}

```

You can see the `range` with a from-to. Good. But then there's a _query_ query:

```auto
{"size":0,"query":{"bool":{"filter":[{"match_all":{"boost":1.0}},{"range":{"timestamp":{"from":null,"to":1623877624260,"include_lower":true,"include_upper":false,"format":"epoch_millis","boost":1.0}}},{"bool":{"filter":[{"terms":{"customer":["_xyz_"],"boost":1.0}}],"adjust_pure_negative":true,"boost":1.0}}],"adjust_pure_negative":true,"boost":1.0}},"aggregations":{"_transform":{"composite":{"size":500,"sources":[{"day":{"date_histogram":{"field":"timestamp","missing_bucket":false,"value_type":"date","order":"asc","fixed_interval":"1d"}}},{"customer":{"terms":{"field":"customer","missing_bucket":false,"order":"asc"}}}]},"aggregations":{"credits":{"sum":{"field":"credits"}},"requests":{"value_count":{"field":"timestamp"}},"error_codes":{"scripted_metric":{"init_script":{"source":"state.responses = [:]","lang":"painless"},"map_script":{"source":" def key_name = 'missing'; if (doc['error.code'].size() != 0) { def error_code = doc['error.code'].value; if (error_code >= 400 && error_code <= 499) { key_name = '4xx'; } else if (error_code >= 500 && error_code <= 599) { key_name = '5xx'; } else { key_name = error_code; } }||| if (!state.responses.containsKey(key_name)) { state.responses[key_name] = 0; } state.responses[key_name] += 1; ","lang":"painless"},"combine_script":{"source":"state.responses","lang":"painless"},"reduce_script":{"source":" def counts = [:]; for (responses in states) { for (key in responses.keySet()) { if (key == 'missing') { continue; } if (!counts.containsKey(key)) { counts[key] = 0; } counts[key] += responses[key]; } } return counts; ","lang":"painless"}}}}}}}

```

This one is the one that also runs my painless code and I think is the one that gets the actual data. Notice how `from` is `null`?

---

<div class="post-metadata">

**Author:** ![Hendrik\_Muhs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hendrik_muhs/32/25802_2.png) [@Hendrik\_Muhs](https://discuss.elastic.co/u/Hendrik_Muhs)\
**Post date:** [June 17, 2021, 8:05am UTC](https://discuss.elastic.co/t/transforms-do-i-need-to-filter-source-for-time-series-data/276057/8 "2021-06-17T08:05:01Z")

</div>

Thank you! I found it.

I updated this [issue](https://github.com/elastic/elasticsearch/issues/60590), for some reason this fix did not make it into `7.9`.

Your options are

- update to `>7.10`
- I verified that if you change your config to use the same input and output field name, the optimization works:

```auto
        "group_by": {
            "timestamp": {"date_histogram": {"field": "timestamp", "fixed_interval": "1d"}},
...

```

TL/DR:

A query that applies the optimization contains something like this:

```auto
...
        {
          "range": {
            "timestamp": {
              "from": null,
              "to": 1623912178447,
              "include_lower": true,
              "include_upper": false,
              "format": "epoch_millis",
              "boost": 1
            }
          }
        },
        {
          "bool": {
            "filter": [
              {
                "range": {
                  "timestamp": {
                    "from": 1623888000000,
                    "to": null,
                    "include_lower": true,
                    "include_upper": true,
                    "format": "epoch_millis",
                    "boost": 1
                  }
                }
              },
              {
                "terms": {
                  "customer": [
                    "xyz"
                  ],
                  "boost": 1
                }
              }
...

```

The 1st range query sets the upper bound of the checkpoint, the other 2 filters are the change queries, one for each `group_by`. The range query - that's the one you miss - rounds down the lower bound to the bucket boundary: `1623888000000`.

I am sorry this slipped through the release process, the issue applies only to the `7.9` series.

---

<div class="post-metadata">

**Author:** ![rokcarl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rokcarl/32/53239_2.png) [@rokcarl](https://discuss.elastic.co/u/rokcarl)\
**Post date:** [June 17, 2021, 10:26am UTC](https://discuss.elastic.co/t/transforms-do-i-need-to-filter-source-for-time-series-data/276057/9 "2021-06-17T10:26:45Z")

</div>

Nice, modifying the grouping seems to be better. One downside that remains is that, even when I'll set it at one hour, it will process the whole day. But I think I can live with that. Great, thanks for the help.

---

<div class="post-metadata">

**Author:** ![Hendrik\_Muhs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hendrik_muhs/32/25802_2.png) [@Hendrik\_Muhs](https://discuss.elastic.co/u/Hendrik_Muhs)\
**Post date:** [June 18, 2021, 8:18am UTC](https://discuss.elastic.co/t/transforms-do-i-need-to-filter-source-for-time-series-data/276057/10 "2021-06-18T08:18:42Z")

</div>

Great, this helps. Further performance improvements are coming.

BTW. Your implementation for aggregating response codes can also be implemented using `filter` aggregations like in this [example](https://www.elastic.co/guide/en/elasticsearch/reference/current/transform-examples.html#example-clientips). I don't think it matters much in terms of performance, however it might simplify your configuration.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 16, 2021, 8:18am UTC](https://discuss.elastic.co/t/transforms-do-i-need-to-filter-source-for-time-series-data/276057/11 "2021-07-16T08:18:58Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
