# ES as a long-term storage system inside analytics architecture?

**URL:** https://discuss.elastic.co/t/es-as-a-long-term-storage-system-inside-analytics-architecture/84697
**Category:** Elasticsearch
**Created:** [May 5, 2017, 10:42am UTC](https://discuss.elastic.co/t/es-as-a-long-term-storage-system-inside-analytics-architecture/84697 "2017-05-05T10:42:51Z")
**Posts on this page:** 4
**Page:** 1

<div class="post-metadata">

### Author: ![ELKnewbie](https://avatars.discourse-cdn.com/v4/letter/e/5e9695/32.png) [@ELKnewbie](https://discuss.elastic.co/u/ELKnewbie)
#### Post date: [May 5, 2017, 10:42am UTC](https://discuss.elastic.co/t/es-as-a-long-term-storage-system-inside-analytics-architecture/84697/1 "2017-05-05T10:42:51Z")

</div>

Would Elasticsearch be suitable as a _long-term_ storage system (besides being a querying system) for _short_ to _mid-term_ offline batch analytics using Apache Spark ? We’re talking petabyte-scale retention over a year with terabytes of new incoming data fed into a Kafka cluster and routed to Logstash, then ES. I’m worried that ES’s overhead would make a huge difference in terms of storage space usage against alternative solutions like compressed data on HDFS/HBASE. In other terms, is there a similar consistent, automatic management system to « archive » older data on an ES cluster ?

Thanks in advance. 😇

---

<div class="post-metadata">

### Author: ![ELKnewbie](https://avatars.discourse-cdn.com/v4/letter/e/5e9695/32.png) [@ELKnewbie](https://discuss.elastic.co/u/ELKnewbie)
#### Post date: [May 7, 2017, 10:58am UTC](https://discuss.elastic.co/t/es-as-a-long-term-storage-system-inside-analytics-architecture/84697/2 "2017-05-07T10:58:54Z")

</div>

So I guess the answer is : NO?

---

<div class="post-metadata">

### Author: ![rusty](https://avatars.discourse-cdn.com/v4/letter/r/f17d59/32.png) [@rusty](https://discuss.elastic.co/u/rusty)
#### Post date: [May 10, 2017, 8:37am UTC](https://discuss.elastic.co/t/es-as-a-long-term-storage-system-inside-analytics-architecture/84697/3 "2017-05-10T08:37:36Z")

</div>

> [@How can we store large scale data with 32GB RAM / 30TB disk on machine](https://discuss.elastic.co/t/how-can-we-store-large-scale-data-with-32gb-ram-30tb-disk-on-machine/84501/3):
>
> Yeah, I had try some cases to solve my problem, include set each shards to the tens of GB in size. I found that: in my case, one shard with 70GB in size, the shard may cost about 470 MB term\_memory. So if I used all of 32GB RAM, means that I can store about 32GB/470MB=68 shards. 68 shards can only store 70GB\*68=4.7TB cry

Hello, it's really depends. You should make similar estimation to understand is it fits for your purposes (do you need to index data, do you need doc\_values, do you need replication, is best\_compression codec suitable and so on). IMHO for now ES is too memory hungry for petabyte-scale solutions.

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [June 7, 2017, 8:51am UTC](https://discuss.elastic.co/t/es-as-a-long-term-storage-system-inside-analytics-architecture/84697/4 "2017-06-07T08:51:58Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
