# ML job state "failed"

**URL:** <https://discuss.elastic.co/t/ml-job-state-failed/170529>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-machine-learning\
**Created:** [March 1, 2019, 3:54pm UTC](https://discuss.elastic.co/t/ml-job-state-failed/170529 "2019-03-01T15:54:39Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ant](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ant/32/33267_2.png) [@Ant](https://discuss.elastic.co/u/Ant)\
**Post date:** [March 1, 2019, 3:54pm UTC](https://discuss.elastic.co/t/ml-job-state-failed/170529/1 "2019-03-01T15:54:39Z")

</div>

I have a single node cluster used for testing which when I restart it all the ML jobs "job status" go to failed. I have tried to stop the datafeed before restarting it but this made no difference and to get the job going again I needed to clone them which with 2 or 3 isn't so bad but as I add more that will be a serious issue. My prod cluster has 3 nodes, I assume this would handle it better but if I know I'm taking the cluster down is there something I can do to allow these to stop and start in a more graceful way?

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [March 2, 2019, 12:37am UTC](https://discuss.elastic.co/t/ml-job-state-failed/170529/2 "2019-03-02T00:37:56Z")

</div>

Before a cluster restart, you could:

- [Stop](https://www.elastic.co/guide/en/elasticsearch/reference/6.6/ml-stop-datafeed.html) all running datafeeds
- [Close](https://www.elastic.co/guide/en/elasticsearch/reference/6.6/ml-close-job.html) all open jobs

If there are a large number of jobs, you could script it - for example:

```auto
#!/bin/bash
HOST='1.2.3.4'
PORT=9200
CURL_AUTH="-u elastic:changeme"

echo
echo
list=`curl $CURL_AUTH -s http://$HOST:$PORT/_xpack/ml/anomaly_detectors?pretty | awk -F" : " '/job_id/{print $2}' | sed 's/\",//g' | sed 's/\"//g'` 
while read -r JOB_ID; do
   echo
   echo "Stoping ${JOB_ID}'s datafeed..."
   curl $CURL_AUTH -s -XPOST $HOST:$PORT/_xpack/ml/datafeeds/datafeed-${JOB_ID}/_stop
   echo "Closing ${JOB_ID}... (ignore 409 error if job was already closed)" 
   curl $CURL_AUTH -s -XPOST $HOST:$PORT/_xpack/ml/anomaly_detectors/${JOB_ID}/_close
   
   echo
   echo
   echo "-------------"
   echo

done <<< "$list"

```

---

<div class="post-metadata">

**Author:** ![Ant](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ant/32/33267_2.png) [@Ant](https://discuss.elastic.co/u/Ant)\
**Post date:** [March 4, 2019, 11:27am UTC](https://discuss.elastic.co/t/ml-job-state-failed/170529/3 "2019-03-04T11:27:35Z")

</div>

@richcollier I thought closing a job was more of a finalising state so wasn't doing that first, I'd stopped them just not closed then fearing I'd not be able to re-open them. Bit green on all this ML stuff.

Thank you very much for your insights and the very useful script!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 1, 2019, 11:27am UTC](https://discuss.elastic.co/t/ml-job-state-failed/170529/4 "2019-04-01T11:27:42Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
