# Data stream resilience

**URL:** <https://discuss.elastic.co/t/data-stream-resilience/320166>\
**Category:** Elasticsearch\
**Tags:** datastreams\
**Created:** [November 30, 2022, 2:44pm UTC](https://discuss.elastic.co/t/data-stream-resilience/320166 "2022-11-30T14:44:18Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![lukaszst](https://avatars.discourse-cdn.com/v4/letter/l/a9adbd/32.png) [@lukaszst](https://discuss.elastic.co/u/lukaszst)\
**Post date:** [November 30, 2022, 2:44pm UTC](https://discuss.elastic.co/t/data-stream-resilience/320166/1 "2022-11-30T14:44:18Z")

</div>

Hello

I'm migrating indices-\>datastreams in relatively large elastic cluster (15 data nodes, 80k docs per second).

For our current indices-based setup we used to monitor if currently used index get red to create new index so that no data is lost (such solution was implemented in es 2.0 days and carry over to v5 and now v8), it was very helpful during server patching, when one by one nodes were shut down for some time.

Now we're migrating to data streams (on v8 of course) and I wonder if such strategy still makes sense. I planned to change mechanism so that data stream health is monitored and if it is red rollover is forced. This mechanism seem to work on my test cluster, but when I checked underlying indices (after restarting few nodes) all of them had shards on all nodes (including ones that were temporarily down).

Are there some new mechanisms in data streams that improve resilliency and in particular handle nodes being down temporarily?

Also what are good practices you recommend for data streams to be resilient in such usecase.

Thanks,  
Lukasz

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [November 30, 2022, 10:29pm UTC](https://discuss.elastic.co/t/data-stream-resilience/320166/2 "2022-11-30T22:29:35Z")

</div>

Welcome to our community! 😃

The best approach would be to resolve why they are turning red, not forcing a rollover on red.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 1, 2022, 3:01am UTC](https://discuss.elastic.co/t/data-stream-resilience/320166/3 "2022-12-01T03:01:46Z")

</div>

How many primary and replica shards are you indexing into? Do you have enough resiliency built in? Do you have any settings that limit the number of shards for an index per node in place?

> [@lukaszst](#):
>
> Are there some new mechanisms in data streams that improve resilliency and in particular handle nodes being down temporarily?
> 
> Also what are good practices you recommend for data streams to be resilient in such usecase.

It is recommended to have a replica shard configured and ensure the nodes have enough capacity to relocate shards from a failed node if necessary.

---

<div class="post-metadata">

**Author:** ![lukaszst](https://avatars.discourse-cdn.com/v4/letter/l/a9adbd/32.png) [@lukaszst](https://discuss.elastic.co/u/lukaszst)\
**Post date:** [December 1, 2022, 8:31am UTC](https://discuss.elastic.co/t/data-stream-resilience/320166/4 "2022-12-01T08:31:01Z")

</div>

I have enough capacity and each index has enough shards to be distributed over all nodes.

Problem is that due to costs I do not have replicas. For that reason it was crucial to create new active index when one node was down so that shartds of new index shards are located on running nodes.

Do you suggest that replica is rather necessary? And second question: do you thnink having only few hours of hot index with replica and rest warm indices without would be enough?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [December 1, 2022, 8:38am UTC](https://discuss.elastic.co/t/data-stream-resilience/320166/5 "2022-12-01T08:38:37Z")

</div>

> [@lukaszst](#):
>
> Do you suggest that replica is rather necessary?

If you want resilience, yes.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 1, 2022, 8:49am UTC](https://discuss.elastic.co/t/data-stream-resilience/320166/6 "2022-12-01T08:49:16Z")

</div>

Indexing will apply backpressure but if you suffer from any type of corruption or storage loss you will lose data, so having replica in place is recommended.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [December 1, 2022, 9:58am UTC](https://discuss.elastic.co/t/data-stream-resilience/320166/7 "2022-12-01T09:58:23Z")

</div>

> [@lukaszst](#):
>
> Do you suggest that replica is rather necessary? And second question: do you thnink having only few hours of hot index with replica and rest warm indices without would be enough?

The recommended practices in this area are all covered at length in the [high availability section](https://www.elastic.co/guide/en/elasticsearch/reference/current/high-availability.html) of the manual.

For resilience you definitely need replicas on indices to which you're writing. For read-only indices you should either use replicas or [searchable snapshots](https://www.elastic.co/guide/en/elasticsearch/reference/current/searchable-snapshots.html). If you're ok with some downtime and manual recovery steps after a failure you might consider removing replicas from read-only indices after they have been snapshotted instead.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 29, 2022, 9:58am UTC](https://discuss.elastic.co/t/data-stream-resilience/320166/8 "2022-12-29T09:58:56Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
