# Number or date type for time serial data?

**URL:** <https://discuss.elastic.co/t/number-or-date-type-for-time-serial-data/29882>\
**Category:** Elasticsearch\
**Created:** [September 24, 2015, 1:09am UTC](https://discuss.elastic.co/t/number-or-date-type-for-time-serial-data/29882 "2015-09-24T01:09:07Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![limac](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/limac/32/3981_2.png) [@limac](https://discuss.elastic.co/u/limac)\
**Post date:** [September 24, 2015, 1:09am UTC](https://discuss.elastic.co/t/number-or-date-type-for-time-serial-data/29882/1 "2015-09-24T01:09:07Z")

</div>

hi, I am using elasticsearch to manage a lot of time serial data. It's about 500G(2000M events) for each day. the mapping looks like this:  
"mappings" : {  
"_default_" : {  
"\_all" : {"enabled" : false},  
"properties" : {  
"@version": { "index": "analyzed", "type": "integer" },  
"@timestamp": { "index": "analyzed", "type": "date" },  
"date\_time":{"index":"not analyzed", "type":"integer"}  
"netflow": {  
"dynamic": true,  
"type": "object",  
"properties": {  
"version": { "index": "analyzed", "type": "integer" },  
"flow\_seq\_num": { "index": "not\_analyzed", "type": "long" },  
"engine\_type": { "index": "not\_analyzed", "type": "integer" },  
"engine\_id": { "index": "not\_analyzed", "type": "integer" },  
"sampling\_algorithm": { "index": "not\_analyzed", "type": "integer" },  
"sampling\_interval": { "index": "not\_analyzed", "type": "integer" },  
"flow\_records": { "index": "not\_analyzed", "type": "integer" },  
"ipv4\_src\_addr": { "index": "analyzed", "type": "ip" },  
"ipv4\_dst\_addr": { "index": "analyzed", "type": "ip" },  
"ipv4\_next\_hop": { "index": "analyzed", "type": "ip" },  
"input\_snmp": { "index": "not\_analyzed", "type": "long" },  
"output\_snmp": { "index": "not\_analyzed", "type": "long" },  
"in\_pkts": { "index": "analyzed", "type": "long" },  
"in\_bytes": { "index": "analyzed", "type": "long" },  
"first\_switched": { "index": "not\_analyzed", "type": "date" },  
"last\_switched": { "index": "not\_analyzed", "type": "date" },  
"l4\_src\_port": { "index": "analyzed", "type": "long" },  
"l4\_dst\_port": { "index": "analyzed", "type": "long" },  
"tcp\_flags": { "index": "analyzed", "type": "integer" },  
"protocol": { "index": "analyzed", "type": "integer" },  
"src\_tos": { "index": "analyzed", "type": "integer" },  
"src\_as": { "index": "analyzed", "type": "integer" },  
"dst\_as": { "index": "analyzed", "type": "integer" },  
"src\_mask": { "index": "analyzed", "type": "integer" },  
"dst\_mask": { "index": "analyzed", "type": "integer" }  
}  
}  
}  
}  
}

the "date\_time"field is the unix timestamp of "@timestamp". The typical search is histogram on date\_time, for example, total number or in\_bytes for every minute between 9:00-10:00.

so, my question is : is there any significant difference of performance between aggregation on time\_data and @timestamp?

thanks.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [September 24, 2015, 2:00am UTC](https://discuss.elastic.co/t/number-or-date-type-for-time-serial-data/29882/2 "2015-09-24T02:00:13Z")

</div>

Doing it on `data_time` will probably give you _a lot_ of buckets as the values are more than likely unique.  
But if you add based on minute resolution timestamps you'll get a lot less.

I'd expect the latter to be more efficient.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:48pm UTC](https://discuss.elastic.co/t/number-or-date-type-for-time-serial-data/29882/3 "2017-07-05T23:48:22Z")

</div>


