# Transform stability

**URL:** <https://discuss.elastic.co/t/transform-stability/274870>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-monitoring, transforms\
**Created:** [June 3, 2021, 12:43pm UTC](https://discuss.elastic.co/t/transform-stability/274870 "2021-06-03T12:43:23Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![alexs1](https://avatars.discourse-cdn.com/v4/letter/a/47e85d/32.png) [@alexs1](https://discuss.elastic.co/u/alexs1)\
**Post date:** [June 3, 2021, 12:43pm UTC](https://discuss.elastic.co/t/transform-stability/274870/1 "2021-06-03T12:43:23Z")

</div>

Hello

I've started to use transforms, and they are great- saving lots of pain of using external tools.

What concerns me, is the I had few times a load on the cluster, which caused a node to leave the cluster, and caused the transform to fail.

When I looked in the morning, it was on failed status, and when stopping/starting it, it was filling the data successfully without issues.

Then, I started thinking, what would happen if I took few days off, and didn't do that manually..  
In my case, the client would see gap in our front end charts.

Is there any way to monitor a transform, and start it automatically in such a case?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [June 7, 2021, 4:59am UTC](https://discuss.elastic.co/t/transform-stability/274870/2 "2021-06-07T04:59:20Z")

</div>

There's [cat transforms API | Elasticsearch Guide [7.13] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/cat-transforms.html), but perhaps that's something that we can also expose via Kibana(?). It might be worth creating a feature request around this 🙂

---

<div class="post-metadata">

**Author:** ![Hendrik\_Muhs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hendrik_muhs/32/25802_2.png) [@Hendrik\_Muhs](https://discuss.elastic.co/u/Hendrik_Muhs)\
**Post date:** [June 8, 2021, 8:00am UTC](https://discuss.elastic.co/t/transform-stability/274870/3 "2021-06-08T08:00:38Z")

</div>

Can you share the error you saw? The reason field in the status should tell you why it failed.

Transform is designed to be fail-safe:

- it differs between permanent and re-occurring errors. Permanent errors are e.g. configuration errors or errors in scripts, mapping issues etc., with other words everything that won't fix itself with a retry.
- if not permanent, a transform retries up to 10 times based on `frequency`. It must fail 10 times in a row to go into failed state. Every successful operation resets the counter.

You mentioned cluster load, which indicates a non permanent error. If you could share the error, I can verify transform picked the right category.

As mitigation you can increase the [number of retries](https://www.elastic.co/guide/en/elasticsearch/reference/7.13/transform-settings.html) by setting `num_transform_failure_retries` to a higher value.

In addition you can create a script that checks `_stats` (or the mentioned `cat` interface) and based on the output call `_start`. You could also call `_start` regularly and ignore the error if it's already running.

In future we plan to improve the way transform retries: Today the interval between each retry is static and based on `frequency`. We want to decouple retry from `frequency` and use exponential backoff.

Last but not least, it might be transform itself that causes your cluster issues. You might want to checkout or general guidance for optimizing transform: [Working with transforms at scale | Elasticsearch Guide [7.13] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/transform-scale.html).

---

<div class="post-metadata">

**Author:** ![alexs1](https://avatars.discourse-cdn.com/v4/letter/a/47e85d/32.png) [@alexs1](https://discuss.elastic.co/u/alexs1)\
**Post date:** [June 8, 2021, 12:18pm UTC](https://discuss.elastic.co/t/transform-stability/274870/4 "2021-06-08T12:18:43Z")

</div>

Hi @Hendrik_Muhs  
I just checked the transforms screen in Kibana, and found a transform in failed state. unfortunately this happens a lot.  
I think today I got this error when upgrading a version (Elastic cloud)

this is the message from the messages tab:

> Failed to index documents into destination index due to permanent error: [BulkIndexingException[Bulk index experienced [48] failures and at least 1 irrecoverable [pipeline with id [TRANSFORM\_NAME-pipeline] does not exist]. Other failures: ]; nested: IllegalArgumentException[pipeline with id [TRANSFORM\_NAME-pipeline] does not exist];; java.lang.IllegalArgumentException: pipeline with id [TRANSFORM\_NAME-pipeline] does not exist]
> 
> Failed to start transform. Please stop and attempt to start again. Failure: Unable to start transform [TRANSFORM\_NAME] as it is in a failed state with failure: [Failed to index documents into destination index due to permanent error: [BulkIndexingException[Bulk index experienced [48] failures and at least 1 irrecoverable [pipeline with id [TRANSFORM\_NAME-pipeline] does not exist]. Other failures: ]; nested: IllegalArgumentException[pipeline with id [TRANSFORM\_NAME-pipeline] does not exist];; java.lang.IllegalArgumentException: pipeline with id [TRANSFORM\_NAME-pipeline] does not exist]]. Use force stop and then restart the transform once error is resolved.

---

<div class="post-metadata">

**Author:** ![Hendrik\_Muhs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hendrik_muhs/32/25802_2.png) [@Hendrik\_Muhs](https://discuss.elastic.co/u/Hendrik_Muhs)\
**Post date:** [June 8, 2021, 2:00pm UTC](https://discuss.elastic.co/t/transform-stability/274870/5 "2021-06-08T14:00:15Z")

</div>

Thanks for the error message.

Can you tell me more about the failing transform? Did you created this one (a transform might also be used as part of a solution)?

As the message says, this is a _permanent error_: The configuration refers to a pipeline that does not seem to exist. Restarting the transform won't help.

I wonder, this does not fit your initial post, is this the same or another transform?

Can you check the pipeline exists by e.g. listing all your pipelines:

```auto
GET /_ingest/pipeline

```

---

<div class="post-metadata">

**Author:** ![alexs1](https://avatars.discourse-cdn.com/v4/letter/a/47e85d/32.png) [@alexs1](https://discuss.elastic.co/u/alexs1)\
**Post date:** [June 8, 2021, 2:21pm UTC](https://discuss.elastic.co/t/transform-stability/274870/6 "2021-06-08T14:21:54Z")

</div>

> [@Hendrik\_Muhs](#):
>
> `GET /_ingest/pipeline`

Hey thanks for replying so fast.  
yes the pipeline exists.  
I've seen this error before, but I suspect the error message doesnt reflect the error state of the cluster.  
Yes I did create it. I can't send the body here in the forum, but I can do that in private if you need.

I started it and it works fine now.

yes it's another transform, but similar idea: it seems that (sometimes) when ever something changes in the cluster (load\> node leave , node extending etc) the transform just fails

---

<div class="post-metadata">

**Author:** ![Hendrik\_Muhs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hendrik_muhs/32/25802_2.png) [@Hendrik\_Muhs](https://discuss.elastic.co/u/Hendrik_Muhs)\
**Post date:** [June 8, 2021, 3:44pm UTC](https://discuss.elastic.co/t/transform-stability/274870/7 "2021-06-08T15:44:02Z")

</div>

As an elastic cloud customer you can create a support request. The support engineer can request additional help from development(me) on demand, just refer to this conversation.

At least in this case, it seems to be issue with ingest pipelines being unavailable. This might not be a transform problem, but an ingest instability. Still, I like to find out what's happening. Please always check, if its this error or whether you see other failure reasons and let me know.

Unfortunately the suggested `num_transform_failure_retries` won't help, because transform classifies this as a permanent error, which never gets retried. The only workaround for now is a script that regularly checks the state.

If I remember correctly, another user uses watcher to do something like this.

---

<div class="post-metadata">

**Author:** ![alexs1](https://avatars.discourse-cdn.com/v4/letter/a/47e85d/32.png) [@alexs1](https://discuss.elastic.co/u/alexs1)\
**Post date:** [June 9, 2021, 11:00am UTC](https://discuss.elastic.co/t/transform-stability/274870/8 "2021-06-09T11:00:19Z")

</div>

thanks @Hendrik_Muhs  
Do you have an example to such a watcher script? (or anything similar)

---

<div class="post-metadata">

**Author:** ![Hendrik\_Muhs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hendrik_muhs/32/25802_2.png) [@Hendrik\_Muhs](https://discuss.elastic.co/u/Hendrik_Muhs)\
**Post date:** [June 9, 2021, 12:38pm UTC](https://discuss.elastic.co/t/transform-stability/274870/9 "2021-06-09T12:38:32Z")

</div>

I don't have an example at the moment, but I suggest to have a look at [watcher http input](https://www.elastic.co/guide/en/elasticsearch/reference/7.13/input-http.html).

---

<div class="post-metadata">

**Author:** ![alexs1](https://avatars.discourse-cdn.com/v4/letter/a/47e85d/32.png) [@alexs1](https://discuss.elastic.co/u/alexs1)\
**Post date:** [June 9, 2021, 6:01pm UTC](https://discuss.elastic.co/t/transform-stability/274870/10 "2021-06-09T18:01:23Z")

</div>

just happened again, with the same pipeline error.  
unfortunately what I had to do is stop it and then start, it again so I'm not sure it'll help to call start.  
I'll try the support but it's very limited for standard customers

---

<div class="post-metadata">

**Author:** ![Hendrik\_Muhs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hendrik_muhs/32/25802_2.png) [@Hendrik\_Muhs](https://discuss.elastic.co/u/Hendrik_Muhs)\
**Post date:** [June 10, 2021, 5:49am UTC](https://discuss.elastic.co/t/transform-stability/274870/11 "2021-06-10T05:49:40Z")

</div>

Thanks for the heads up, I will investigate it. Unfortunately a failed transform requires a _force_ stop and start. There is no _force_ start. Nevertheless, the loss of pipelines should not happen.

Was there anything of interest happening before, e.g.

- a node that dropped?
- change of master?

How many ingest nodes do you have?  
What version are you using?

---

<div class="post-metadata">

**Author:** ![alexs1](https://avatars.discourse-cdn.com/v4/letter/a/47e85d/32.png) [@alexs1](https://discuss.elastic.co/u/alexs1)\
**Post date:** [June 10, 2021, 6:38am UTC](https://discuss.elastic.co/t/transform-stability/274870/12 "2021-06-10T06:38:34Z")

</div>

Attached the cluster architecture as an image. I hope it's easy to read 🙂

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/c/2/c2499badb2ba4adce06bacbc5c85d3889d775684.png)

No changes were done. (not by us anyway)  
version 7.13.1

so to ensure stability, I need to:

1. call cat/transforms
2. if state is failed \> force stop\> start
3. else call start and ignore "already started" error?

thanks for the help

---

<div class="post-metadata">

**Author:** ![Hendrik\_Muhs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hendrik_muhs/32/25802_2.png) [@Hendrik\_Muhs](https://discuss.elastic.co/u/Hendrik_Muhs)\
**Post date:** [June 10, 2021, 7:22am UTC](https://discuss.elastic.co/t/transform-stability/274870/13 "2021-06-10T07:22:32Z")

</div>

Thanks!

> [@alexs1](#):
>
> so to ensure stability, I need to:
> 
> 1. call cat/transforms
> 2. if state is failed \> force stop\> start
> 3. else call start and ignore "already started" error?

Step 1 and 2 should be sufficient, if you already know the state, there is no need to start it again.

Instead of using `_cat` you can use [`_stats`](https://www.elastic.co/guide/en/elasticsearch/reference/current/get-transform-stats.html), it might be easier to parse json instead of text.

If I found out why your ingest pipelines drop, I let you know. It still might be good to open a support case and send diagnostics this way.

Do you also have transforms without ingest pipelines? They should not be affected, right?

---

<div class="post-metadata">

**Author:** ![alexs1](https://avatars.discourse-cdn.com/v4/letter/a/47e85d/32.png) [@alexs1](https://discuss.elastic.co/u/alexs1)\
**Post date:** [June 10, 2021, 7:43am UTC](https://discuss.elastic.co/t/transform-stability/274870/14 "2021-06-10T07:43:53Z")

</div>

good point.

> [@Hendrik\_Muhs](#):
>
> Do you also have transforms without ingest pipelines? They should not be affected, right?

correct.

I removed the pipeline from the transform and moved it to the index settings.  
I have a feeling it will solve it (at least for me 🙂 )

---

<div class="post-metadata">

**Author:** ![Hendrik\_Muhs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hendrik_muhs/32/25802_2.png) [@Hendrik\_Muhs](https://discuss.elastic.co/u/Hendrik_Muhs)\
**Post date:** [June 15, 2021, 7:47am UTC](https://discuss.elastic.co/t/transform-stability/274870/15 "2021-06-15T07:47:14Z")

</div>

Thanks for the update.

Did the error happen again after you switched to using the index template based pipeline?

(FYI I received the diagnostics from support, thanks for that.)

---

<div class="post-metadata">

**Author:** ![alexs1](https://avatars.discourse-cdn.com/v4/letter/a/47e85d/32.png) [@alexs1](https://discuss.elastic.co/u/alexs1)\
**Post date:** [June 16, 2021, 6:12pm UTC](https://discuss.elastic.co/t/transform-stability/274870/16 "2021-06-16T18:12:09Z")

</div>

Thanks.  
Correct. For now, no errors, but it didn't happen every day

---

<div class="post-metadata">

**Author:** ![alexs1](https://avatars.discourse-cdn.com/v4/letter/a/47e85d/32.png) [@alexs1](https://discuss.elastic.co/u/alexs1)\
**Post date:** [July 1, 2021, 5:40pm UTC](https://discuss.elastic.co/t/transform-stability/274870/17 "2021-07-01T17:40:26Z")

</div>

Hey @Hendrik_Muhs  
It did happen again yesterday without the pipline definition

---

<div class="post-metadata">

**Author:** ![Hendrik\_Muhs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hendrik_muhs/32/25802_2.png) [@Hendrik\_Muhs](https://discuss.elastic.co/u/Hendrik_Muhs)\
**Post date:** [July 5, 2021, 6:14am UTC](https://discuss.elastic.co/t/transform-stability/274870/18 "2021-07-05T06:14:27Z")

</div>

> [@alexs1](#):
>
> It did happen again yesterday without the pipline definition

Without the pipeline in transform, but with the one in index settings?  
Do you have the error message?

---

<div class="post-metadata">

**Author:** ![alexs1](https://avatars.discourse-cdn.com/v4/letter/a/47e85d/32.png) [@alexs1](https://discuss.elastic.co/u/alexs1)\
**Post date:** [July 5, 2021, 9:29am UTC](https://discuss.elastic.co/t/transform-stability/274870/19 "2021-07-05T09:29:23Z")

</div>

1. correct
2. Failed to start transform. Please stop and attempt to start again. Failure: Unable to start transform [xxx] as it is in a failed state with failure: [Failed to index documents into destination index due to permanent error: [BulkIndexingException[Bulk index experienced [500] failures and at least 1 irrecoverable [pipeline with id [xxx\_pipeline] does not exist]. Other failures: ]; nested: IllegalArgumentException[pipeline with id [xxx\_pipeline] does not exist];; java.lang.IllegalArgumentException: pipeline with id [xxx\_pipeline] does not exist]]. Use force stop and then restart the transform once error is resolved.

---

<div class="post-metadata">

**Author:** ![alexs1](https://avatars.discourse-cdn.com/v4/letter/a/47e85d/32.png) [@alexs1](https://discuss.elastic.co/u/alexs1)\
**Post date:** [July 5, 2021, 10:28am UTC](https://discuss.elastic.co/t/transform-stability/274870/20 "2021-07-05T10:28:40Z")

</div>

just got this again when upgrading to latest version using the elastic cloud interface

[Next page](https://discuss.elastic.co/t/transform-stability/274870.md?page=2)
