# Re-index using a Machine Learning with custom Trained Models

**URL:** <https://discuss.elastic.co/t/re-index-using-a-machine-learning-with-custom-trained-models/336711>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-machine-learning\
**Created:** [June 22, 2023, 4:17pm UTC](https://discuss.elastic.co/t/re-index-using-a-machine-learning-with-custom-trained-models/336711 "2023-06-22T16:17:46Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Lone\_Eagle](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lone_eagle/32/62213_2.png) [@Lone\_Eagle](https://discuss.elastic.co/u/Lone_Eagle)\
**Post date:** [June 22, 2023, 4:17pm UTC](https://discuss.elastic.co/t/re-index-using-a-machine-learning-with-custom-trained-models/336711/1 "2023-06-22T16:17:46Z")

</div>

We have currently a Production Elasticsearch using Elastic Cloud and want to experiment vectors using a Machine Learning with a custom trained model.

**Machine Learning**

- Configuration: 32 GB RAM | 16.9 vCPU
- As a pipeline defined to infer a **big** text (product description)
- Vector dimension: 768
- Trained Model:
  - Number of allocations: 18 (Could not go higher)
  - Threads per allocation: 1 (Could not go higher)

**Problem:**

- When I start a re-index async in order to fill the new embedded field dense vector, it start the process but shortly hang.

**Here what I did:**

- Reduced the batch size to 25 (default 1K) with a "requests\_per\_second=10", so a wait time between batches.

**More information:**

- Re-index like for 3K docs and then hang.
- The task does not cancel on error.
- Need to kill the task.
- Need to re-start the Machine Learning model.
- I could go bigger machine but not the smallest one and that one is already fairly expensive.

**What needed:**

- Need a way to see the Machine Learning logs and understand what going on.
- How can we re-index without overloading the machine?
- Trained Model - How much memory is taken for each:
  - Number of allocations?
  - Threads per allocation?

---

<div class="post-metadata">

**Author:** ![wei.wang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wei.wang/32/51803_2.png) [@wei.wang](https://discuss.elastic.co/u/wei.wang)\
**Post date:** [June 22, 2023, 8:45pm UTC](https://discuss.elastic.co/t/re-index-using-a-machine-learning-with-custom-trained-models/336711/2 "2023-06-22T20:45:50Z")

</div>

Before we dive into the problem, do you mind sharing a few things:

- which is you elasticsearch version?
- when it hangs, can you run this `GET _ml/trained_models/<your model>/_stats`, it should tell you the failed reason, and share here ?
- meanwhile, have you try the tip " _Set the reindex `size` option to a value smaller than the `queue_capacity` for the trained model deployment. Otherwise, requests might be rejected with a "too many requests" 429 error code._" from [this page](https://www.elastic.co/guide/en/machine-learning/master/ml-nlp-inference.html)?

---

<div class="post-metadata">

**Author:** ![Lone\_Eagle](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lone_eagle/32/62213_2.png) [@Lone\_Eagle](https://discuss.elastic.co/u/Lone_Eagle)\
**Post date:** [June 23, 2023, 1:29pm UTC](https://discuss.elastic.co/t/re-index-using-a-machine-learning-with-custom-trained-models/336711/3 "2023-06-23T13:29:35Z")

</div>

**More information:**

1. Using the latest Elasticsearch version: 8.8.1.

2. Machine Learning - Trained Model: I tried: (Look like magic number is 18)  
a. Number of allocations: 9 (Before: 18)  
b. Threads per allocation: 2 (Before: 1)  
c. Went well for a short while and then hanged:

---

<div class="post-metadata">

**Author:** ![wei.wang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wei.wang/32/51803_2.png) [@wei.wang](https://discuss.elastic.co/u/wei.wang)\
**Post date:** [June 23, 2023, 3:42pm UTC](https://discuss.elastic.co/t/re-index-using-a-machine-learning-with-custom-trained-models/336711/4 "2023-06-23T15:42:20Z")

</div>

Thanks for the info.

From the model stats api, your model stats is `"state": "started",` that tells the model is working fine. therefore, most likely you hit the queue\_capacity error.

Here are some suggestions you can try:

1. While reindex, set the a smaller batch size, which I would recommend **50**. The benefit of a small `size` value is that if you have multiple bulk upload through an ingest pipeline using the same model deployment, they all use the same queue. so for your case, the queue will be full quickly if it is using the default reindex batch size (1000)  
The reindex command should be like:

2. It looks like you only have 1 ml node. Multiple allocations are better fit to multiple ml nodes. Instead of using "9 allocations x 2 threads", please use "1 allocation x X threads". I would start with "1 allocation x 10 threads" for your case.

3. Handle [pipeline failures](https://www.elastic.co/guide/en/elasticsearch/reference/current/ingest.html#handling-pipeline-failures) in your pipeline configure. Your pipeline configure will look like:

If all the above wont help your case, as you mentioned you are running on our Elastic Cloud, feel free to create a support case, we will take a deep look at your cluster logs.

Hope it helps.

---

<div class="post-metadata">

**Author:** ![Lone\_Eagle](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lone_eagle/32/62213_2.png) [@Lone\_Eagle](https://discuss.elastic.co/u/Lone_Eagle)\
**Post date:** [June 23, 2023, 4:29pm UTC](https://discuss.elastic.co/t/re-index-using-a-machine-learning-with-custom-trained-models/336711/5 "2023-06-23T16:29:19Z")

</div>

Thanks for your help.

I am currently on Elastic Cloud. Changed the pipeline (error handling) and threading and look better.

**Machine Learning**

1. My choices for the threads are: `1 - 2 - 4 - 8` =\> So I took `8`
2. Allocations: Using `2` is working (but not higher). Should I use 1 as will not improve anything?
3. Can I downgrade the machine (`32 GB RAM | 16.9 vCPU`) or risky?

**FYI:** I created first a case with support but I got better support here! Case: [01384117](https://support.elastic.co/cases/5008X00002OJSQlQAP)

I will continue to re-index and will see.

Have a nice day! 🙂

---

<div class="post-metadata">

**Author:** ![wei.wang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wei.wang/32/51803_2.png) [@wei.wang](https://discuss.elastic.co/u/wei.wang)\
**Post date:** [June 26, 2023, 5:14pm UTC](https://discuss.elastic.co/t/re-index-using-a-machine-learning-with-custom-trained-models/336711/6 "2023-06-26T17:14:19Z")

</div>

Good day,

> 1. My choices for the threads are: `1 - 2 - 4 - 8` =\> So I took `8`
> 2. Allocations: Using `2` is working (but not higher). Should I use 1 as will not improve anything?

I just realized I had a typo in my previous reply regarding threads number, sorry.

What I meant to start with is: 1 allocation x 8 threads, then 1 allocation x 16 threads ... [this page](https://www.elastic.co/guide/en/elasticsearch/reference/current/start-trained-model-deployment.html#start-trained-model-deployment-desc) can help understand the concept of allocations & threads. Since you have one single ml node, more thread should get better performance than more allocations. Because I don't know if there are other ml features or models sharing the same ml node resource, that's why 8 threads were recommended to start with. For using more than 8 threads, you might have to use api or dev tools to send the request.

> 1. Can I downgrade the machine (`32 GB RAM | 16.9 vCPU`) or risky?

There is no risk. The trade-off will be the overall inference/reindex time. Before downgrade, Don't forget to stop your model first. After downgrade is done, you can re-start model with appropriate threads settings.

Feel free to add your questions or results to your existing support case, our support team will team up with us and provide you a best solution.

Cheers.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 24, 2023, 5:14pm UTC](https://discuss.elastic.co/t/re-index-using-a-machine-learning-with-custom-trained-models/336711/7 "2023-07-24T17:14:20Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
