# Elastic cloud basic setup: Could not open job because no ML nodes with sufficient capacity were found

**URL:** <https://discuss.elastic.co/t/elastic-cloud-basic-setup-could-not-open-job-because-no-ml-nodes-with-sufficient-capacity-were-found/169866>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-machine-learning\
**Created:** [February 25, 2019, 3:35pm UTC](https://discuss.elastic.co/t/elastic-cloud-basic-setup-could-not-open-job-because-no-ml-nodes-with-sufficient-capacity-were-found/169866 "2019-02-25T15:35:50Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![dao](https://avatars.discourse-cdn.com/v4/letter/d/a6a055/32.png) [@dao](https://discuss.elastic.co/u/dao)\
**Post date:** [February 25, 2019, 3:35pm UTC](https://discuss.elastic.co/t/elastic-cloud-basic-setup-could-not-open-job-because-no-ml-nodes-with-sufficient-capacity-were-found/169866/1 "2019-02-25T15:35:51Z")

</div>

Hello,

I am running on a elastic cloud cluster, with 1Gb for the ML node (default)

I have 1 job running OK, and when I try to run a second one, I got the following.

How can I know the capacity of the ML node? when I add some jobs, how can I know the remaining capacity?

Thx

```auto
{
  "changed": false,
  "connection": "Close",
  "content": "{\"error\":{\"root_cause\":[{\"type\":\"status_exception\",\"reason\":\"Could not open job because no ML nodes with sufficient capacity were found\"}],\"type\":\"status_exception\",\"reason\":\"Could not open job because no ML nodes with sufficient capacity were found\",\"caused_by\":{\"type\":\"illegal_state_exception\",\"reason\":\"Could not open job because no suitable nodes were found, allocation explanation [Not opening job [job-rundown_time_aps5000_4] on node [instance-0000000042], because this node isn't a ml node.|Not opening job [job-rundown_time_aps5000_4] on node [tiebreaker-0000000044], because this node isn't a ml node.|Not opening job [job-rundown_time_aps5000_4] on node [{instance-0000000045}{ml.machine_memory=1073741824}{ml.max_open_jobs=20}{ml.enabled=true}], because this node has insufficient available memory. Available memory for ML [440234147], memory required by existing jobs [108105012], estimated memory required for this job [435159040]|Not opening job [job-rundown_time_aps5000_4] on node [instance-0000000043], because this node isn't a ml node.]\"}},\"status\":429}",
  "content_length": "1073",
  "content_type": "application/json; charset=UTF-8",
  "date": "Mon, 25 Feb 2019 15:22:19 GMT",
  "json": {
    "error": {
      "caused_by": {
        "reason": "Could not open job because no suitable nodes were found, allocation explanation [Not opening job [job-rundown_time_aps5000_4] on node [instance-0000000042], because this node isn't a ml node.|Not opening job [job-rundown_time_aps5000_4] on node [tiebreaker-0000000044], because this node isn't a ml node.|Not opening job [job-rundown_time_aps5000_4] on node [{instance-0000000045}{ml.machine_memory=1073741824}{ml.max_open_jobs=20}{ml.enabled=true}], because this node has insufficient available memory. Available memory for ML [440234147], memory required by existing jobs [108105012], estimated memory required for this job [435159040]|Not opening job [job-rundown_time_aps5000_4] on node [instance-0000000043], because this node isn't a ml node.]",
        "type": "illegal_state_exception"
      },
      "reason": "Could not open job because no ML nodes with sufficient capacity were found",
      "root_cause": [{
        "reason": "Could not open job because no ML nodes with sufficient capacity were found",
        "type": "status_exception"
      }],
      "type": "status_exception"
    },
    "status": 429
  },
  "msg": "Status code was 429 and not [201, 200]: HTTP Error 429: Too Many Requests",
  "redirected": false,
  "server": "fp/4xxxxx",
  "status": 429,
  "url": "https://xxxxx.eu-west-1.aws.found.io:9243/_xpack/ml/anomaly_detectors/job-rundown_time_aps5000_4/_open",
  "x_found_handling_cluster": "xxxxx",
  "x_found_handling_instance": "instance-0000000042",
  "x_found_handling_server": "xxxxx"
}

```

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [February 26, 2019, 10:22pm UTC](https://discuss.elastic.co/t/elastic-cloud-basic-setup-could-not-open-job-because-no-ml-nodes-with-sufficient-capacity-were-found/169866/2 "2019-02-26T22:22:40Z")

</div>

Currently, you have a couple of options:

1. You can try to pre-calculate the current memory usage of ML on the node and then estimate the headroom you have for a new job, or
2. Do what you did - make an attempt to open a new job and just let the node tell you that it doesn't have enough room for it.

The first obviously requires some intimate knowledge about how memory is allocated by ML on the node. Namely:

- memory is only used by jobs that are in the open state
- each job has approximately 100MB overhead
- each job has an additional "model memory" that is shown as `model_bytes` in the Counts tab of the Job Management page or via the [job stats API](https://www.elastic.co/guide/en/elasticsearch/reference/6.6/ml-get-job-stats.html). This value is approximately equal to 20kB to 30kB for every unique time series in the model (the number of splits/partitions)
- the ML processes do not have access to the entire memory space of the node, but rather approximately 30% of it (this is controlled via a [node setting](https://www.elastic.co/guide/en/elasticsearch/reference/current/ml-settings.html) called `xpack.ml.max_machine_memory_percent`. This value is a little more complicated to determine in Cloud because Cloud uses containers. I've seen this setting described as:

`xpack.ml.max_machine_memory_percent: min(max_machine_memory, round((container_size - JVM_heap_size - non_jvm_process_overhead) / container_size * 100))`

Bottom line is that it is a little complicated - and we don't yet have an easier way to figure out your future sizing needs for ML, especially in Cloud. But, keep in mind that the 1GB node on Cloud is gratis and is really meant for people to "try out" ML. It is not expected that you'd run many production-worthy, high-availability ML jobs on a single 1GB node 😉

Some more information can be found on these blogs:

> **[Sizing for Machine Learning with Elasticsearch
	  	 | Elastic](https://www.elastic.co/blog/sizing-machine-learning-with-elasticsearch)**
>
> Many organizations have started using Elastic's machine learning for their Security Analytics, Operational Analytics, and other projects that make use of anomaly detection for time series data.  ...

  

> **[Smarter Machine Learning Job Placement in Elasticsearch](https://www.elastic.co/blog/smarter-machine-learning-job-placement-in-elasticsearch)**
>
> Machine Learning in X-Pack just got a whole lot smarter about distributing your jobs

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 26, 2019, 10:34pm UTC](https://discuss.elastic.co/t/elastic-cloud-basic-setup-could-not-open-job-because-no-ml-nodes-with-sufficient-capacity-were-found/169866/3 "2019-03-26T22:34:16Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
