# Index to red state cluster

**URL:** <https://discuss.elastic.co/t/index-to-red-state-cluster/335763>\
**Category:** Elasticsearch\
**Created:** [June 12, 2023, 11:36am UTC](https://discuss.elastic.co/t/index-to-red-state-cluster/335763 "2023-06-12T11:36:23Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![DJ\_Zhu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dj_zhu/32/99543_2.png) [@DJ\_Zhu](https://discuss.elastic.co/u/DJ_Zhu)\
**Post date:** [June 12, 2023, 11:36am UTC](https://discuss.elastic.co/t/index-to-red-state-cluster/335763/1 "2023-06-12T11:36:23Z")

</div>

I have a question regarding shard selection during index/bulk operations in Elasticsearch version 6.8.6.

In my cluster, I have three data nodes: A, B, and C. The shards (with no replicas) are evenly allocated across these nodes, with shards 0, 1, and 2 assigned to each node respectively.

I've noticed that when any of the data nodes crash, all the index requests are still being routed to the target shard/node, causing a timeout exception. It seems that the indexRequest route policy does not exclude the crashed node/shard from the selection process.

I would like to confirm the following:

1. Is my observation and analysis correct?
2. Why does the default route policy include the crashed node in the routing selection? Is it because Elasticsearch aims to maintain consistent route results?
3. Is there any way to route the document to an available shard?

I appreciate any insights or suggestions you can provide. Thank you.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 12, 2023, 11:40am UTC](https://discuss.elastic.co/t/index-to-red-state-cluster/335763/2 "2023-06-12T11:40:39Z")

</div>

Elasticsearch version 6.8.6 is [EOL](https://www.elastic.co/support/eol) and no longer supported. Please upgrade ASAP.

(This is an automated response from your friendly Elastic bot. Please report this post if you have any suggestions or concerns :elasticheart: )

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 12, 2023, 11:42am UTC](https://discuss.elastic.co/t/index-to-red-state-cluster/335763/3 "2023-06-12T11:42:05Z")

</div>

Elasticsearch will as far as I know not allow you to index data into an index in red state, which will be the case when you have lost a primary shard.

> [@DJ\_Zhu](#):
>
> Why does the default route policy include the crashed node in the routing selection? Is it because Elasticsearch aims to maintain consistent route results?

Index operations are assigned to a shard based on hashing of the document ID. As operations could be updates they can not necessarily be rerouted without causing problems.

> [@DJ\_Zhu](#):
>
> Is there any way to route the document to an available shard?

No, not that I am aware of. I would recommend you instead configure a replica shard.

---

<div class="post-metadata">

**Author:** ![DJ\_Zhu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dj_zhu/32/99543_2.png) [@DJ\_Zhu](https://discuss.elastic.co/u/DJ_Zhu)\
**Post date:** [June 13, 2023, 2:30am UTC](https://discuss.elastic.co/t/index-to-red-state-cluster/335763/4 "2023-06-13T02:30:35Z")

</div>

Thank you, Christian.

But I have try and confirm I do index into a red state index. (but only into online shards ofcause)

> [@Christian\_Dahlqvist](#):
>
> Elasticsearch will as far as I know not allow you to index data into an index in red state, which will be the case when you have lost a primary shard.

And the primary reason for opting to maintain zero replicas is the cost of resources. We are not concerned if some nodes experience crashes (typically, they recover quickly). During a node crash, the only operation affected is querying the offline shard, which is acceptable in our scenario. However, what troubles us is that all index requests routed to the offline shard will wait until timeout, significantly reducing the transactions per second (TPS), particularly for bulk requests. A single index timeout within the bulk request will cause the entire operation to wait until a timeout reply.

In conclusion, it appears sensible to exclusively select online shards as the index target when employing automatically generated document IDs. As you mentioned, the document ID influences shard and node selection. We can explore ways to ensure that automatically generated document IDs exclusively reference online shards.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 13, 2023, 6:32am UTC](https://discuss.elastic.co/t/index-to-red-state-cluster/335763/5 "2023-06-13T06:32:30Z")

</div>

You can override the hashing of the document ID determining the shard to write to through [document routing](https://www.elastic.co/guide/en/elasticsearch/reference/8.8/search-shard-routing.html#search-routing), which is probably easier to use rather than try to assign appropriate IDs client side. This also avoids the negative impact using custom IDs have on indexing performance.

---

<div class="post-metadata">

**Author:** ![DJ\_Zhu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dj_zhu/32/99543_2.png) [@DJ\_Zhu](https://discuss.elastic.co/u/DJ_Zhu)\
**Post date:** [June 13, 2023, 8:58am UTC](https://discuss.elastic.co/t/index-to-red-state-cluster/335763/6 "2023-06-13T08:58:19Z")

</div>

But you can not avoid to select the offline shard by change routing or hashing only.

maybe we need to do something in the server side instead of client.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 13, 2023, 11:25am UTC](https://discuss.elastic.co/t/index-to-red-state-cluster/335763/7 "2023-06-13T11:25:02Z")

</div>

> [@DJ\_Zhu](#):
>
> maybe we need to do something in the server side instead of client.

I do not think there is anything you can do. This is why adding replicas is the recommended solution.

The lack of resiliency you have experienced is an trade-off when running without replica shards enabled.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [June 13, 2023, 12:57pm UTC](https://discuss.elastic.co/t/index-to-red-state-cluster/335763/8 "2023-06-13T12:57:37Z")

</div>

> [@DJ\_Zhu](#):
>
> However, what troubles us is that all index requests routed to the offline shard will wait until timeout, significantly reducing the transactions per second (TPS), particularly for bulk requests. A single index timeout within the bulk request will cause the entire operation to wait until a timeout reply.

This timeout is under your control via the `?timeout=` query parameter to the bulk API. It defaults to `1m` but if you want a shorter (or even zero) timeout then that's fine too.

---

<div class="post-metadata">

**Author:** ![DJ\_Zhu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dj_zhu/32/99543_2.png) [@DJ\_Zhu](https://discuss.elastic.co/u/DJ_Zhu)\
**Post date:** [June 14, 2023, 1:55am UTC](https://discuss.elastic.co/t/index-to-red-state-cluster/335763/9 "2023-06-14T01:55:01Z")

</div>

Thanks David,

I have modified the default timeout to 10s which still had a considerable impact to TPS while the normal request will cost less than 1s. And the situation will become worse in bulk since any single index latency will cause the whole bulk to wait until timeout.

It seems I have to change the default pattern by rewrite the route code.

---

<div class="post-metadata">

**Author:** ![DJ\_Zhu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dj_zhu/32/99543_2.png) [@DJ\_Zhu](https://discuss.elastic.co/u/DJ_Zhu)\
**Post date:** [June 14, 2023, 2:08am UTC](https://discuss.elastic.co/t/index-to-red-state-cluster/335763/10 "2023-06-14T02:08:36Z")

</div>

Indeed, we are faced with a trade-off between the advantages of low cost and the potential risk of data loss, versus the higher cost (at least doubling) that comes with maintaining replicas. However, considering the exorbitant expense associated with doubling the cost, we must opt for the first option. In our environment, we can accept a minimal loss of data. I believe this is a reasonable and widely shared requirement for certain scenarios.

Whatever, thank you for your understanding and reply.

PS. do you guys thinks it is appropriate to propose a feature request issue regarding this matter?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [June 14, 2023, 6:59am UTC](https://discuss.elastic.co/t/index-to-red-state-cluster/335763/11 "2023-06-14T06:59:46Z")

</div>

> [@DJ\_Zhu](#):
>
> I have modified the default timeout to 10s which still had a considerable impact to TPS

Why 10s? If you want these things to fail fast, why not zero?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 12, 2023, 6:59am UTC](https://discuss.elastic.co/t/index-to-red-state-cluster/335763/12 "2023-07-12T06:59:58Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
