# Loading BERT Model

**URL:** https://discuss.elastic.co/t/loading-bert-model/327709
**Category:** Elasticsearch
**Tags:** elastic-stack-machine-learning
**Created:** [March 15, 2023, 12:39am UTC](https://discuss.elastic.co/t/loading-bert-model/327709 "2023-03-15T00:39:46Z")
**Posts on this page:** 3
**Page:** 1

<div class="post-metadata">

### Author: ![Cole\_Crawford](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cole_crawford/32/118458_2.png) [@Cole\_Crawford](https://discuss.elastic.co/u/Cole_Crawford)
#### Post date: [March 15, 2023, 12:39am UTC](https://discuss.elastic.co/t/loading-bert-model/327709/1 "2023-03-15T00:39:46Z")

</div>

I am trying to add ANN semantic search to an Elasticsearch index of scientific documents. To that end, I am trying to set up an NLP pipeline on Elasticsearch to vectorize documents on ingest. I would like to test [allenai/scibert\_scivocab\_uncased · Hugging Face](https://huggingface.co/allenai/scibert_scivocab_uncased). I am getting this error:

```auto
Traceback (most recent call last):
  File "/usr/local/bin/eland_import_hub_model", line 197, in <module>
    tm = TransformerModel(args.hub_model_id, args.task_type, args.quantize)
  File "/usr/local/lib/python3.10/site-packages/eland/ml/pytorch/transformers.py", line 551, in __init__
    self._traceable_model = self._create_traceable_model()
  File "/usr/local/lib/python3.10/site-packages/eland/ml/pytorch/transformers.py", line 661, in _create_traceable_model
    model = _DPREncoderWrapper.from_pretrained(self._model_id)
  File "/usr/local/lib/python3.10/site-packages/eland/ml/pytorch/transformers.py", line 385, in from_pretrained
    if is_compatible():
  File "/usr/local/lib/python3.10/site-packages/eland/ml/pytorch/transformers.py", line 379, in is_compatible
    has_architectures = len(config.architectures) == 1
TypeError: object of type 'NoneType' has no len()

```

I'm surprised by this as this makes it look like it's not compatible with Elasticsearch's ML capabilities even though it should fit these criteria (BERT architecture): [Compatible third party NLP models | Machine Learning in the Elastic Stack [8.6] | Elastic](https://www.elastic.co/guide/en/machine-learning/current/ml-nlp-model-ref.html). Anyone experienced something similar?

Also happy for other model recommendations - my corpus is a set of drug labels, so was thinking to use MedBERT or SciBERT.

---

<div class="post-metadata">

### Author: ![dkyle](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dkyle/32/59114_2.png) [@dkyle](https://discuss.elastic.co/u/dkyle)
#### Post date: [April 3, 2023, 1:11pm UTC](https://discuss.elastic.co/t/loading-bert-model/327709/2 "2023-04-03T13:11:18Z")

</div>

Hi @Cole_Crawford

The `eland_import_hub_model` does some work to figure out how to configure the model correctly for upload to Elasticsearch. The code path you hit is missing a None check. Once I added it in I was able to upload SciBERT with the following command.

```auto
eland_import_hub_model \
      --url http://localhost:9200/ \
      -u XXXX -p XXXX \
      --hub-model-id allenai/scibert_scivocab_uncased \
      --task-type text_embedding --insecure

```

Obviously update the --url and auth (-u, -p) params for your cluster.

The fix is in this PR [[NLP] Prevent TypeError with None check by davidkyle · Pull Request #525 · elastic/eland · GitHub](https://github.com/elastic/eland/pull/525).  
Thanks for reporting the issue.

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [May 1, 2023, 1:12pm UTC](https://discuss.elastic.co/t/loading-bert-model/327709/3 "2023-05-01T13:12:14Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
