I'm upgrading from v7 to 9.5 - lots of great new features! For my application, TSDS look ideal, and downsampling will be a major improvement.
There appears to be one flaw with downsampling for my use cases. Keyword fields not marked as time_series_dimension are downsampled with the latest value in each bucket. That will be highly misleading for end-user queries.
My datatream mappings are quite large, with many keywords and metrics. Many of those are only useful for detailed analytics, not for historical reporting. Thus many keyword and metric values should be dropped when creating summarised data for longer time range reporting and queries. In TSDS terms, the _tsid should only contain a subset of the available keywords.
For (highly simplified) example:
{
"mappings": {
"properties": {
"@timestamp": {
"type": "date"
},
"TechnologyID": {
"type": "keyword",
"time_series_dimension": true
},
"OperatorID": {
"type": "keyword",
"time_series_dimension": true
},
"DeviceID": {
"type": "keyword",
},
"PowerState": {
"type": "keyword",
},
"Latency": {
"type": "half_float",
"time_series_metric": "gauge"
},
"SignalStrength": {
"type": "half_float",
"time_series_metric": "gauge"
}
}
}
}
These measurements are taken every few minutes, and would be downsampled initially to one or more hour. But DeviceID and PowerState should be downsampled to null, because a query analysing latency by power state across both detailed and downsampled time range would give incorrect results, as the downsampled latency would represent many power states, not just the last one.
Unless I've missed something, there does not seem to be an option to drop keyword (or metric) values that are not marked as time series dimensions (or metrics) when downsampling. Altrhough I've not looked at the code yet, it feels like a fairly easy extension.
Have I understood this correctly? If so, do you plan to add downsample to null as an option?
Appreciate there are a number of workarounds, but all seem rather clumsy and resource intensive compared to the substantial improvement of TSDS + downsampling:
- Use the API to scan for backing datastream indexes which have crossed end time but not yet downsampled - mass update non-dimension keyword fields to null.
- Reindex downsampled indexes, then delete them - run queries across detail TSDS + reindexed history.
- Rollups, transforms...