# Elastic Search for storing Historical data

**URL:** <https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210>\
**Category:** Elasticsearch\
**Created:** [September 14, 2015, 1:58am UTC](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210 "2015-09-14T01:58:03Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![code\_blue](https://avatars.discourse-cdn.com/v4/letter/c/7ea924/32.png) [@code\_blue](https://discuss.elastic.co/u/code_blue)\
**Post date:** [September 14, 2015, 1:58am UTC](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210/1 "2015-09-14T01:58:03Z")

</div>

We are considering storing Historical data in ES. Is there any pattern or best practice guideline defined for this?

The way we want to do this is via a scheduled job that will pull data on a defined interval and store in Index. So one document will exist multiple times inside a mapping. We will add a timestamp to identify when the document was loaded in ES.

In such a scenario what should the \_id be defined as?

Also any suggestions regarding storing historical data in ES will be helpful.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [September 14, 2015, 2:32am UTC](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210/2 "2015-09-14T02:32:46Z")

</div>

This is pretty much what ELK is built for 🙂

So;

- Use time based indices, daily/weekly/monthly
- Specify your mappings in advance
- Look into hot+warm architecture - [https://www.elastic.co/blog/hot-warm-architecture](https://www.elastic.co/blog/hot-warm-architecture)
- Use Elasticsearch Curator [https://www.elastic.co/guide/en/elasticsearch/client/curator/current/index.html](https://www.elastic.co/guide/en/elasticsearch/client/curator/current/index.html)  
I am sure others have recommendations 🙂

But why are you worried about what `_id` would be?

---

<div class="post-metadata">

**Author:** ![code\_blue](https://avatars.discourse-cdn.com/v4/letter/c/7ea924/32.png) [@code\_blue](https://discuss.elastic.co/u/code_blue)\
**Post date:** [September 14, 2015, 7:33am UTC](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210/3 "2015-09-14T07:33:47Z")

</div>

Hey Mark

Thanks for your response.  
We want to store data in single index & run aggregations with date histogram on time stamp. Main objective is to generate analytical data using aggregations and we don't want to run aggregations over multiple indices as we are not sure how performant that will be.

Since we are looking to use a single index and same document will exist multiple times along with time stamp hence worried about \_id.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [September 14, 2015, 7:47am UTC](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210/4 "2015-09-14T07:47:06Z")

</div>

Honestly you won't notice any difference between having a single index and multiple ones that contain the same data.  
Plus it makes retention management _massively_ easier and should remove your concern about the `_id`.

However to make that even less of a problem just let ES pick the ID and the put your message ID in it's own field.

---

<div class="post-metadata">

**Author:** ![code\_blue](https://avatars.discourse-cdn.com/v4/letter/c/7ea924/32.png) [@code\_blue](https://discuss.elastic.co/u/code_blue)\
**Post date:** [September 14, 2015, 8:13am UTC](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210/5 "2015-09-14T08:13:42Z")

</div>

Creating multiple index for every time interval may lead to index maintainability overhead in our case. For e.g. if we are capturing data at end of every month then end of the year we will end up having 12 indices and every time we ran our aggregation queries we will have to add the new indices.

If i let ES pick the \_id in same index and inside same mapping wont ES replace my earlier document since it may happen that no data has changed from the earlier time stamp?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [September 14, 2015, 8:15am UTC](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210/6 "2015-09-14T08:15:02Z")

</div>

> [@code\_blue](#):
>
> Creating multiple index for every time interval may lead to index maintainability overhead in our case. For e.g. if we are capturing data at end of every month then end of the year we will end up having 12 indices and every time we ran our aggregation queries we will have to add the new indices.

So? How are you going to manage removal of old data from the index if you have a single one?

> [@code\_blue](#):
>
> If i let ES pick the \_id in same index and inside same mapping wont ES replace my earlier document since it may happen that no data has changed from the earlier time stamp?

True. How often are they likely to have the same ID though?

---

<div class="post-metadata">

**Author:** ![code\_blue](https://avatars.discourse-cdn.com/v4/letter/c/7ea924/32.png) [@code\_blue](https://discuss.elastic.co/u/code_blue)\
**Post date:** [September 14, 2015, 10:00am UTC](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210/7 "2015-09-14T10:00:17Z")

</div>

Since we are considering historical data in single index so not sure if we need to have data removal.

Every mapping inside my index corresponds to some data in my data source. As we are loading data from data source at some predefined time stamp so there will be multiple copies of same data inside my mapping.

---

<div class="post-metadata">

**Author:** ![code\_blue](https://avatars.discourse-cdn.com/v4/letter/c/7ea924/32.png) [@code\_blue](https://discuss.elastic.co/u/code_blue)\
**Post date:** [September 15, 2015, 3:19am UTC](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210/8 "2015-09-15T03:19:09Z")

</div>

Is there any convention to keep the same document multiple times within an index but with a different **\_id**

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [September 15, 2015, 4:05am UTC](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210/9 "2015-09-15T04:05:20Z")

</div>

Just assign your own ID.

---

<div class="post-metadata">

**Author:** ![vionemc](https://avatars.discourse-cdn.com/v4/letter/v/7cd45c/32.png) [@vionemc](https://discuss.elastic.co/u/vionemc)\
**Post date:** [January 20, 2016, 4:00am UTC](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210/10 "2016-01-20T04:00:13Z")

</div>

What he meant is an answer like in this article:  
[http://blog.mongodb.org/post/65517193370/schema-design-for-time-series-data-in-mongodb](http://blog.mongodb.org/post/65517193370/schema-design-for-time-series-data-in-mongodb)

Maybe what is the recommended database schema for storing historical data using ElasticSearch?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [January 20, 2016, 4:05am UTC](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210/11 "2016-01-20T04:05:28Z")

</div>

ES isn't a database 😉

It really depends what sort of historic data, is it time based, or just "old"

---

<div class="post-metadata">

**Author:** ![vionemc](https://avatars.discourse-cdn.com/v4/letter/v/7cd45c/32.png) [@vionemc](https://discuss.elastic.co/u/vionemc)\
**Post date:** [January 20, 2016, 4:16am UTC](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210/12 "2016-01-20T04:16:21Z")

</div>

I know, but since ElasticSearch doesn't support working together with MongoDB anymore, in the end we force using ElasticSearch as a database.

Storing historical data in MongoDB and ElasticSearch seems to be different. For example, if we store a string "2016-01-03 00:00:00" in MongoDB, we can process that string as a date directly. I am afraid it won't be that way in ElasticSearch. Maybe ElasticSearch will prefer timestamp, or a date object.

My need specifically is to store a history of how many twitter followers I have for every date.

This article already answer about getting the history  
[https://www.elastic.co/guide/en/elasticsearch/guide/current/\_looking\_at\_time.html](https://www.elastic.co/guide/en/elasticsearch/guide/current/_looking_at_time.html)

But that article didn't give a sample of the stored index which they process. I want to follow the sample schema and understand it before applying any history storage in my database.

---

<div class="post-metadata">

**Author:** ![vionemc](https://avatars.discourse-cdn.com/v4/letter/v/7cd45c/32.png) [@vionemc](https://discuss.elastic.co/u/vionemc)\
**Post date:** [January 20, 2016, 4:17am UTC](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210/13 "2016-01-20T04:17:09Z")

</div>

BTW, you should prepare ElasticSearch to be also a database beside a search engine. A lot of people already forcing ElasticSearch to be a database. Even including HipChat.

---

<div class="post-metadata">

**Author:** ![vionemc](https://avatars.discourse-cdn.com/v4/letter/v/7cd45c/32.png) [@vionemc](https://discuss.elastic.co/u/vionemc)\
**Post date:** [January 20, 2016, 4:18am UTC](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210/14 "2016-01-20T04:18:53Z")

</div>

My current structure in MongoDB:

`{ "_id" : "15454221-2016", #string "follower_history" : { "2016-01-01 00:00:00" : { #date, first day of the month, UTC time "values" : { "2016-01-03 00:00:00" : 1505, #key:date, first day of the month, UTC time;value: integer "2016-01-07 00:00:00" : 1508, "2016-01-08 00:00:00" : 1508 }, "num_samples" : 3, #integer "total_follower" : 4521 #integer } } }`

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [January 20, 2016, 4:20am UTC](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210/15 "2016-01-20T04:20:58Z")

</div>

A time and a date are the same thing in ES, a timestamp, and ES will detect that.

There's two ways of doing what you want;

1. Have an index per day with a single counter that you update and can simply read
2. Have an index per day and record all "new follower" events, then run an agg on it.

> [@vionemc](#):
>
> BTW, you should prepare Elasticsearch to be also a database beside a search engine. A lot of people already forcing Elasticsearch to be a database.

A lot of people use redis as a persistent store, doesn't mean that is what it is or it's right.

---

<div class="post-metadata">

**Author:** ![vionemc](https://avatars.discourse-cdn.com/v4/letter/v/7cd45c/32.png) [@vionemc](https://discuss.elastic.co/u/vionemc)\
**Post date:** [January 20, 2016, 5:12am UTC](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210/16 "2016-01-20T05:12:53Z")

</div>

> [@warkolm](#):
>
> A lot of people use redis as a persistent store, doesn't mean that is what it is or it's right.

Yup, just suggesting. I think it will be like heaven if Elasticsearch is also a database.

---

<div class="post-metadata">

**Author:** ![vionemc](https://avatars.discourse-cdn.com/v4/letter/v/7cd45c/32.png) [@vionemc](https://discuss.elastic.co/u/vionemc)\
**Post date:** [January 20, 2016, 5:14am UTC](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210/17 "2016-01-20T05:14:11Z")

</div>

> [@warkolm](#):
>
> There's two ways of doing what you want;
> 
> Have an index per day with a single counter that you update and can simply read  
> Have an index per day and record all "new follower" events, then run an agg on it.

May I ask for a sample indexed json? That I can try to aggregate. Just a simple one will suffice.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [January 20, 2016, 5:15am UTC](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210/18 "2016-01-20T05:15:17Z")

</div>

I don't have them, that was just some ideas.

---

<div class="post-metadata">

**Author:** ![vionemc](https://avatars.discourse-cdn.com/v4/letter/v/7cd45c/32.png) [@vionemc](https://discuss.elastic.co/u/vionemc)\
**Post date:** [January 20, 2016, 5:19am UTC](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210/19 "2016-01-20T05:19:46Z")

</div>

OK, thanks.

---

<div class="post-metadata">

**Author:** ![vionemc](https://avatars.discourse-cdn.com/v4/letter/v/7cd45c/32.png) [@vionemc](https://discuss.elastic.co/u/vionemc)\
**Post date:** [January 20, 2016, 5:26am UTC](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210/20 "2016-01-20T05:26:55Z")

</div>

Try this article on how to store the data:

> **[Elasticsearch as a Time Series Data Store
	  	 | Elastic](https://www.elastic.co/blog/elasticsearch-as-a-time-series-data-store)**
>
> As the project manager of stagemonitor, an open source performance monitoring tool, I've recently been looking for a database to replace the cool-but-aging Graphite Time Series DataBase (TSDB) as the ...

  
It's more on the recommended data structure.

And this article on how to view the data:  
[https://www.elastic.co/guide/en/elasticsearch/guide/current/\_looking\_at\_time.html#CO196-2](https://www.elastic.co/guide/en/elasticsearch/guide/current/_looking_at_time.html#CO196-2)  
Mainly using aggregate

[Next page](https://discuss.elastic.co/t/elastic-search-for-storing-historical-data/29210.md?page=2)
