# A major issue with cluster state handling and persistent tasks cancellation

**URL:** <https://discuss.elastic.co/t/a-major-issue-with-cluster-state-handling-and-persistent-tasks-cancellation/386014>\
**Category:** Elasticsearch\
**Tags:** datastreams\
**Created:** [April 24, 2026, 11:06am UTC](https://discuss.elastic.co/t/a-major-issue-with-cluster-state-handling-and-persistent-tasks-cancellation/386014 "2026-04-24T11:06:50Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![sherman81](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sherman81/32/147431_2.png) [@sherman81](https://discuss.elastic.co/u/sherman81)\
**Post date:** [April 24, 2026, 11:06am UTC](https://discuss.elastic.co/t/a-major-issue-with-cluster-state-handling-and-persistent-tasks-cancellation/386014/1 "2026-04-24T11:06:50Z")

</div>

We are using ES 9.3.0.

Our cluster has many data streams and indices. The cluster state is ~350 MB (compressed on disk) under normal conditions.

I mistakenly scheduled a large number of downsampling tasks for historical data, which created ~2,200 persistent tasks.

After that, I canceled the tasks and removed the ILM policy from the data stream. However, logs show that each canceled task triggers a full cluster state update (why? I just canceled the tasks).

Each update currently takes ~20 seconds, making the cluster effectively unusable.

`[2026-04-24T13:55:53,612][WARN][o.e.g.PersistedClusterStateService] [hssf43] writing cluster state took [21473ms] which is above the warn threshold of [10s]; [wrote] global metadata, wrote [0] new mappings, removed [0] mappings and skipped [1514] unchanged mappings, wrote metadata for [0] new indices and [0] existing indices, removed metadata for [0] indices and skipped [2219] unchanged indices`

`[2026-04-24T13:56:44,213][WARN][o.e.g.PersistedClusterStateService] [hssf43] writing cluster state took [21408ms] which is above the warn threshold of [10s]; [wrote] global metadata, wrote [0] new mappings, removed [0] mappings and skipped [1514] unchanged mappings, wrote metadata for [0] new indices and [0] existing indices, removed metadata for [0] indices and skipped [2219] unchanged indices`

`[2026-04-24T13:57:35,750][WARN][o.e.g.PersistedClusterStateService] [hssf43] writing cluster state took [21693ms] which is above the warn threshold of [10s]; [wrote] global metadata, wrote [0] new mappings, removed [0] mappings and skipped [1514] unchanged mappings, wrote metadata for [0] new indices and [0] existing indices, removed metadata for [0] indices and skipped [2219] unchanged indices`

`[2026-04-24T13:58:26,707][WARN][o.e.g.PersistedClusterStateService] [hssf43] writing cluster state took [21609ms] which is above the warn threshold of [10s]; [wrote] global metadata, wrote [0] new mappings, removed [0] mappings and skipped [1514] unchanged mappings, wrote metadata for [0] new indices and [0] existing indices, removed metadata for [0] indices and skipped [2219] unchanged indices`

`[2026-04-24T13:59:16,408][WARN][o.e.g.PersistedClusterStateService] [hssf43] writing cluster state took [21609ms] which is above the warn threshold of [10s]; [wrote] global metadata, wrote [0] new mappings, removed [0] mappings and skipped [1514] unchanged mappings, wrote metadata for [0] new indices and [0] existing indices, removed metadata for [0] indices and skipped [2219] unchanged indices`

`[2026-04-24T14:00:06,732][WARN][o.e.g.PersistedClusterStateService] [hssf43] writing cluster state took [21610ms] which is above the warn threshold of [10s]; [wrote] global metadata, wrote [0] new mappings, removed [0] mappings and skipped [1514] unchanged mappings, wrote metadata for [0] new indices and [0] existing indices, removed metadata for [0] indices and skipped [2219] unchanged indices`

`[2026-04-24T14:00:57,081][WARN][o.e.g.PersistedClusterStateService] [hssf43] writing cluster state took [21810ms] which is above the warn threshold of [10s]; [wrote] global metadata, wrote [0] new mappings, removed [0] mappings and skipped [1514] unchanged mappings, wrote metadata for [0] new indices and [0] existing indices, removed metadata for [0] indices and skipped [2219] unchanged indices`

`[2026-04-24T14:01:47,891][WARN][o.e.g.PersistedClusterStateService] [hssf43] writing cluster state took [21611ms] which is above the warn threshold of [10s]; [wrote] global metadata, wrote [0] new mappings, removed [0] mappings and skipped [1514] unchanged mappings, wrote metadata for [0] new indices and [0] existing indices, removed metadata for [0] indices and skipped [2219] unchanged indices`

`[2026-04-24T14:02:38,416][WARN][o.e.g.PersistedClusterStateService] [hssf43] writing cluster state took [22010ms] which is above the warn threshold of [10s]; [wrote] global metadata, wrote [0] new mappings, removed [0] mappings and skipped [1514] unchanged mappings, wrote metadata for [0] new indices and [0] existing indices, removed metadata for [0] indices and skipped [2219] unchanged indices`

Is there a way to speed up this process or mitigate the impact?

p.s. Currently there's no write activity from users.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [April 24, 2026, 11:40am UTC](https://discuss.elastic.co/t/a-major-issue-with-cluster-state-handling-and-persistent-tasks-cancellation/386014/2 "2026-04-24T11:40:33Z")

</div>

> [@sherman81](#):
>
> However, logs show that each canceled task triggers a full cluster state update (why? I just canceled the tasks).

Persistent tasks are recorded in the cluster state, so cancelling a persistent task requires a cluster state update. These are in fact not "full" cluster state updates, they're incremental (see `skipped [1514] unchanged mappings` and `skipped [2219] unchanged indices`) but that doesn't really help you.

I don't have any other suggestions beyond waiting for these cancellations to complete. At the current rate I guess it'll be done in 12h or so, although as the number of tasks decreases they should get faster.

---

<div class="post-metadata">

**Author:** ![sherman81](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sherman81/32/147431_2.png) [@sherman81](https://discuss.elastic.co/u/sherman81)\
**Post date:** [April 24, 2026, 11:45am UTC](https://discuss.elastic.co/t/a-major-issue-with-cluster-state-handling-and-persistent-tasks-cancellation/386014/3 "2026-04-24T11:45:50Z")

</div>

Hi, David!

Thank you for answering.

If this is supposed to be a delta update, why does each cluster state update result in a new segment being written to disk that is about 90–95% of the total state size?

And one more clarification question: when all tasks have disappeared from the Tasks API and there is no ILM policy that created them, can we be sure they won’t be restored after a cluster restart?

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/d/9/d9d56bd75f585adda9c19863d35b691187270e38.png)

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [April 24, 2026, 11:52am UTC](https://discuss.elastic.co/t/a-major-issue-with-cluster-state-handling-and-persistent-tasks-cancellation/386014/4 "2026-04-24T11:52:19Z")

</div>

> [@sherman81](#):
>
> why does each cluster state update result in a new segment being written to disk that is about 90–95% of the total state size?

It's the `[wrote] global metadata` bit - this is where persistent tasks are stored, and it's not subdivided so we have to rewrite the whole thing on each change.

> [@sherman81](#):
>
> can we be sure they won’t be restored after a cluster restart?

You'd need to watch `GET _cluster/state` to see the actual state of these persistent tasks.

---

<div class="post-metadata">

**Author:** ![sherman81](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sherman81/32/147431_2.png) [@sherman81](https://discuss.elastic.co/u/sherman81)\
**Post date:** [April 24, 2026, 3:04pm UTC](https://discuss.elastic.co/t/a-major-issue-with-cluster-state-handling-and-persistent-tasks-cancellation/386014/5 "2026-04-24T15:04:23Z")

</div>

> You'd need to watch `GET _cluster/state`

Not sure I understand the situation. I canceled the tasks, and they disappeared from the `_tasks` API results, but they are still present in the cluster state.

less cluster\_state\_full.json | grep rollup-shard -c  
2234

If it matters, ILM is stopped globally.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [April 24, 2026, 3:17pm UTC](https://discuss.elastic.co/t/a-major-issue-with-cluster-state-handling-and-persistent-tasks-cancellation/386014/6 "2026-04-24T15:17:07Z")

</div>

There are (at least?) three different things called "tasks" in Elasticsearch. The ones in `GET _tasks` are things that are actively running in the system at the time. The ones in the cluster state are persistent tasks which will normally correspond to an active task but may not be assigned to a node at the time. Then there's `GET cluster/pending_tasks` which are cluster state updates. It's confusing indeed.

---

<div class="post-metadata">

**Author:** ![sherman81](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sherman81/32/147431_2.png) [@sherman81](https://discuss.elastic.co/u/sherman81)\
**Post date:** [April 24, 2026, 3:25pm UTC](https://discuss.elastic.co/t/a-major-issue-with-cluster-state-handling-and-persistent-tasks-cancellation/386014/7 "2026-04-24T15:25:29Z")

</div>

According to pending tasks API, I see my not yet canceled tasks as:

"source": " **update** project [default] task state [downsample-downsample-5m-.ds-metrics-otelcol.v1-devops-2025.10.12-000106-1-5m]",

and tasks which are already updated (after cancel) as:

"source": " **finish** project [default] persistent task [downsample-downsample-5m-.ds-metrics-otelcol.v1-devops-2025.12.22-000276-2-5m] (success)",

So, do I need to do anything additional to completely remove it from the cluster state? I’m concerned that it might be restored later.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [April 27, 2026, 8:15am UTC](https://discuss.elastic.co/t/a-major-issue-with-cluster-state-handling-and-persistent-tasks-cancellation/386014/8 "2026-04-27T08:15:53Z")

</div>

Sorry if I'm not following, but what do you mean by "restored later"? These tasks are triggered by some active process, e.g. manually or by ILM. Once the existing ones have all finished I would not expect any new ones to start.
