# Store the data more than 1TB per day

**URL:** <https://discuss.elastic.co/t/store-the-data-more-than-1tb-per-day/104511>\
**Category:** Elasticsearch\
**Created:** [October 19, 2017, 7:48am UTC](https://discuss.elastic.co/t/store-the-data-more-than-1tb-per-day/104511 "2017-10-19T07:48:23Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Sharon\_wsf](https://avatars.discourse-cdn.com/v4/letter/s/7993a0/32.png) [@Sharon\_wsf](https://discuss.elastic.co/u/Sharon_wsf)\
**Post date:** [October 19, 2017, 7:48am UTC](https://discuss.elastic.co/t/store-the-data-more-than-1tb-per-day/104511/1 "2017-10-19T07:48:23Z")

</div>

Hi Team,

Can elasticsearch store big data more than 1TB per day?  
When the data corrupt, this corrupted data have backup??

Thanks and best regards  
Sharon

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [October 19, 2017, 8:21am UTC](https://discuss.elastic.co/t/store-the-data-more-than-1tb-per-day/104511/2 "2017-10-19T08:21:53Z")

</div>

> [@Sharon\_wsf](#):
>
> Can elasticsearch store big data more than 1TB per day?

Yes. I heard about someone indexing 10m docs per second.

> [@Sharon\_wsf](#):
>
> When the data corrupt, this corrupted data have backup?

You have replicas in elasticsearch.  
You can snapshot/restore data.  
More than that you can create hourly indices for example and in the worse case you can just drop the corrupted hourly index if it’s ok for your use case.

But the team has been working very hard to reduce that risk of corruption.

---

<div class="post-metadata">

**Author:** ![Sharon\_wsf](https://avatars.discourse-cdn.com/v4/letter/s/7993a0/32.png) [@Sharon\_wsf](https://discuss.elastic.co/u/Sharon_wsf)\
**Post date:** [October 19, 2017, 8:24am UTC](https://discuss.elastic.co/t/store-the-data-more-than-1tb-per-day/104511/3 "2017-10-19T08:24:30Z")

</div>

After the data corrupted, still can snapshot/restore the data?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [October 19, 2017, 8:43am UTC](https://discuss.elastic.co/t/store-the-data-more-than-1tb-per-day/104511/4 "2017-10-19T08:43:44Z")

</div>

It depends on what you mean by corrupted I guess.  
The way elasticsearch works is by writing immutable files. If you snapshot every 10 minutes for example you will end up with valid backup.  
If for whatever reason one of the new created immutable file gets corrupted (which is unlikely going to happen) you will always being able to restore a previous point of time of your index.

But what is your fear exactly?

---

<div class="post-metadata">

**Author:** ![Sharon\_wsf](https://avatars.discourse-cdn.com/v4/letter/s/7993a0/32.png) [@Sharon\_wsf](https://discuss.elastic.co/u/Sharon_wsf)\
**Post date:** [October 19, 2017, 10:33am UTC](https://discuss.elastic.co/t/store-the-data-more-than-1tb-per-day/104511/5 "2017-10-19T10:33:52Z")

</div>

is it will dynamically and automatically be discovered? without snapshot ourself

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [October 19, 2017, 1:07pm UTC](https://discuss.elastic.co/t/store-the-data-more-than-1tb-per-day/104511/6 "2017-10-19T13:07:10Z")

</div>

I don’t understand

---

<div class="post-metadata">

**Author:** ![Sharon\_wsf](https://avatars.discourse-cdn.com/v4/letter/s/7993a0/32.png) [@Sharon\_wsf](https://discuss.elastic.co/u/Sharon_wsf)\
**Post date:** [October 20, 2017, 3:04am UTC](https://discuss.elastic.co/t/store-the-data-more-than-1tb-per-day/104511/7 "2017-10-20T03:04:46Z")

</div>

it like distribution map. Elasticsearch got this function?  
 ![Capture](https://us1.discourse-cdn.com/elastic/original/3X/7/9/79b3f5ea9a9628be4f336ac9aa3003a5a5e426f3.JPG)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 20, 2017, 6:16am UTC](https://discuss.elastic.co/t/store-the-data-more-than-1tb-per-day/104511/8 "2017-10-20T06:16:16Z")

</div>

Elastic manages this through the use of primary and replica shards, and distributes these automatically across the cluster. I would recommend that you read [this chapter from Elasticsearch: the definitive guide](https://www.elastic.co/guide/en/elasticsearch/guide/2.x/distributed-cluster.html) to get a better understanding.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 17, 2017, 6:16am UTC](https://discuss.elastic.co/t/store-the-data-more-than-1tb-per-day/104511/9 "2017-11-17T06:16:28Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
