# Vectorizing documents - very slow

**URL:** <https://discuss.elastic.co/t/vectorizing-documents-very-slow/327710>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-machine-learning, docker\
**Created:** [March 15, 2023, 12:42am UTC](https://discuss.elastic.co/t/vectorizing-documents-very-slow/327710 "2023-03-15T00:42:13Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Cole\_Crawford](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cole_crawford/32/118458_2.png) [@Cole\_Crawford](https://discuss.elastic.co/u/Cole_Crawford)\
**Post date:** [March 15, 2023, 12:42am UTC](https://discuss.elastic.co/t/vectorizing-documents-very-slow/327710/1 "2023-03-15T00:42:13Z")

</div>

I am trying to create an NLP pipeline using [Charangan/MedBERT · Hugging Face](https://huggingface.co/Charangan/MedBERT). Ingesting documents with this model and an ES ML pipeline is running very slowly: With a dockerized setup on my local machine with 10GB of RAM (8GB dedicated just to Elasticsearch via MEM\_LIMIT), 1.5GB swap, and 4CPUs, I'm getting roughly 1 document every 22s. We have ~1M documents, which means it will take months to index due to the vectorization. Indexing all the docs without vectorizing is reasonably fast (\< 30 minutes). Anyone have any suggestions for speeding up the vectorization process? Anything I might be missing here? I got this info log which made me think that this model isn't optimized for `task-type text_embedding`, which is what I'm trying to use.

```auto
Some weights of the model checkpoint at Charangan/MedBERT were not used when initializing BertModel: ['cls.predictions.transform.LayerNorm.bias', 'cls.predictions.bias', 'cls.seq_relationship.weight', 'cls.predictions.transform.dense.weight', 'cls.predictions.decoder.weight', 'cls.seq_relationship.bias', 'cls.predictions.transform.dense.bias', 'cls.predictions.transform.LayerNorm.weight', 'cls.predictions.decoder.bias']
- This IS expected if you are initializing BertModel from the checkpoint of a model trained on another task or with another architecture (e.g. initializing a BertForSequenceClassification model from a BertForPreTraining model).
- This IS NOT expected if you are initializing BertModel from the checkpoint of a model that you expect to be exactly identical (initializing a BertForSequenceClassification model from a BertForSequenceClassification model).

```

---

<div class="post-metadata">

**Author:** ![BenTrent](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bentrent/32/33915_2.png) [@BenTrent](https://discuss.elastic.co/u/BenTrent)\
**Post date:** [March 15, 2023, 1:39pm UTC](https://discuss.elastic.co/t/vectorizing-documents-very-slow/327710/2 "2023-03-15T13:39:36Z")

</div>

@Cole_Crawford what is your Elasticsearch version?

By default, I think model deployments only use a single thread. You must increase the number of allocations for better throughput: [Start trained model deployment API | Elasticsearch Guide [8.6] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/start-trained-model-deployment.html)

Additionally, that model is not a text\_embedding model. It is a base model that has the fill mask task. You should either optimize this model yourself text embedding or utilize a different model.

---

<div class="post-metadata">

**Author:** ![BenTrent](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bentrent/32/33915_2.png) [@BenTrent](https://discuss.elastic.co/u/BenTrent)\
**Post date:** [March 15, 2023, 1:46pm UTC](https://discuss.elastic.co/t/vectorizing-documents-very-slow/327710/3 "2023-03-15T13:46:11Z")

</div>

Also @Cole_Crawford PyTorch model inference is done off the JVM heap, so you should probably decrease the JVM size that Elasticsearch will use.

---

<div class="post-metadata">

**Author:** ![Cole\_Crawford](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cole_crawford/32/118458_2.png) [@Cole\_Crawford](https://discuss.elastic.co/u/Cole_Crawford)\
**Post date:** [March 15, 2023, 8:58pm UTC](https://discuss.elastic.co/t/vectorizing-documents-very-slow/327710/4 "2023-03-15T20:58:28Z")

</div>

Thanks @BenTrent . I decreased the JVM heap relative to the total available RAM. It looks like it's recommended to be 35-40% for ML nodes: [Sizing for Machine Learning with Elasticsearch | Elastic Blog](https://www.elastic.co/blog/sizing-machine-learning-with-elasticsearch)

I switched to [pritamdeka/S-PubMedBert-MS-MARCO · Hugging Face](https://huggingface.co/pritamdeka/S-PubMedBert-MS-MARCO) which looks to be a better fit for this application. I increased the number of allocations and threads per allocation; the way I'm understanding this is that the number of allocations \* number of threads per allocation can't exceed the total number of CPUs allocated to the container (or available on the host)? So if I have 6CPUs dedicated to Docker, and am using a couple on a web app and Kibana, then I should only allocate 4 CPUs to the ML node? As 2 allocations with 2 threads per allocation?

---

<div class="post-metadata">

**Author:** ![BenTrent](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bentrent/32/33915_2.png) [@BenTrent](https://discuss.elastic.co/u/BenTrent)\
**Post date:** [March 16, 2023, 12:07pm UTC](https://discuss.elastic.co/t/vectorizing-documents-very-slow/327710/5 "2023-03-16T12:07:36Z")

</div>

> 2 allocations with 2 threads per allocation?

Would be a good place to start. Or 4 allocations. At indexing time usually you are concerned about throughput. So more allocations is better.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 13, 2023, 12:08pm UTC](https://discuss.elastic.co/t/vectorizing-documents-very-slow/327710/6 "2023-04-13T12:08:14Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
