# Vector search : when input is bigger than max\_seq

**URL:** <https://discuss.elastic.co/t/vector-search-when-input-is-bigger-than-max-seq/363753>\
**Category:** Elastic Search\
**Created:** [July 25, 2024, 3:45am UTC](https://discuss.elastic.co/t/vector-search-when-input-is-bigger-than-max-seq/363753 "2024-07-25T03:45:59Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![dan\_kim](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dan_kim/32/95741_2.png) [@dan\_kim](https://discuss.elastic.co/u/dan_kim)\
**Post date:** [July 25, 2024, 3:46am UTC](https://discuss.elastic.co/t/vector-search-when-input-is-bigger-than-max-seq/363753/1 "2024-07-25T03:46:00Z")

</div>

* * *

Hello,

I'm trying to handle long input text fields in Elasticsearch, where the token size often exceeds the maximum sequence length.

I have two questions:

1. **Applying Vector Search with Large Input Fields** :

2. **Using Non-SentenceTransformer Models in Elasticsearch Pipeline** :

Here is the command I used:

```auto
eland_import_hub_model --url https://xxxxx:9200/ --es-username xxxx --es-password xxxxx --hub-model-id monologg/kobigbird-bert-base --task-type text_embedding --insecure

```

However, the logs indicate an issue:

```auto
No sentence-transformers model found with name monologg/kobigbird-bert-base. Creating a new one with MEAN pooling.

```

Despite the log message stating that a new model is created, I cannot find any machine learning model in Elasticsearch:

```auto
GET /_ml/trained_models

```

---

<div class="post-metadata">

**Author:** ![Carlos\_D](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/carlos_d/32/126245_2.png) [@Carlos\_D](https://discuss.elastic.co/u/Carlos_D)\
**Post date:** [July 29, 2024, 9:57am UTC](https://discuss.elastic.co/t/vector-search-when-input-is-bigger-than-max-seq/363753/2 "2024-07-29T09:57:52Z")

</div>

Hey @dan_kim !

1. There are two options here:

- Truncate your text. Not ideal, but this is done by default for you on the inference processor when ingesting text.
- Use chunking, to divide your text into smaller passages that will get included as a nested field in your document. Each passage will have its embeddings calculated separately, and some overlap can be included to have semantically meaningful passages.

In order to apply chunking, you can use an external process, use a [script processor](https://www.elastic.co/search-labs/blog/chunking-via-ingest-pipelines), or use the [semantic\_text](https://www.elastic.co/search-labs/blog/semantic-search-simplified-semantic-text) field type (that is now available on serverless, and will be available on 8.15) to do automatic chunking for you.

1. Elasticsearch only supports Sentence Transformers for deploying into the Elasticsearch cluster. However, you can use the [Inference API](https://www.elastic.co/guide/en/elasticsearch/reference/current/put-inference-api.html#inference-example-hugging-face) to refer to embedding services outside of the Elasticsearch stack, using for example HuggingFace as a service provider.

Hope that helps!

---

<div class="post-metadata">

**Author:** ![dan\_kim](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dan_kim/32/95741_2.png) [@dan\_kim](https://discuss.elastic.co/u/dan_kim)\
**Post date:** [July 30, 2024, 5:41am UTC](https://discuss.elastic.co/t/vector-search-when-input-is-bigger-than-max-seq/363753/3 "2024-07-30T05:41:48Z")

</div>

Thank you for your answer,

i will try scripting soon, Thanks

---

<div class="post-metadata">

**Author:** ![dan\_kim](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dan_kim/32/95741_2.png) [@dan\_kim](https://discuss.elastic.co/u/dan_kim)\
**Post date:** [July 31, 2024, 7:25am UTC](https://discuss.elastic.co/t/vector-search-when-input-is-bigger-than-max-seq/363753/4 "2024-07-31T07:25:50Z")

</div>

@Carlos_D

I decided to change my approach and split the content into keywords at the data pipeline stage to apply it to the model.

What do you think about this approach?

1. Use the E5 model to relate both Korean and English content.
2. Use the OpenAI API to extract keywords from each sentence of the content.

In my case, it is an initial service, so there isn't much traffic or many documents in this service.

For example, for the content:

"Despite the log message stating that a new model is created, I cannot find any machine learning model in Elasticsearch,"

the OpenAI API extracted the keywords: ["log", "message", "new", "model", "machine", "Elasticsearch", "es"].

 ![스크린샷 2024-07-31 오후 4.22.48](https://us1.discourse-cdn.com/elastic/original/3X/a/0/a0a2b707c0a386d9dd286d420e8e3aa7f82534a7.jpeg)

---

<div class="post-metadata">

**Author:** ![Carlos\_D](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/carlos_d/32/126245_2.png) [@Carlos\_D](https://discuss.elastic.co/u/Carlos_D)\
**Post date:** [July 31, 2024, 4:50pm UTC](https://discuss.elastic.co/t/vector-search-when-input-is-bigger-than-max-seq/363753/5 "2024-07-31T16:50:41Z")

</div>

Hey @dan_kim :

Extracting keywords from each passage will lessen the context for the vector embeddings - keep in mind that models generate embeddings taking into account not just the words in isolation, but the overall sentence context. Generating embeddings from keywords will have much worse precision and recall on your search IMO.

You _could_ summarise your documents to a maximum number of words using OpenAI and then just generate embeddings for your document summary. But again, you will be losing context and depend on a previous summarization done by a model.

Using a chunking strategy as described above should be the way to go to maximize your search results relevancy.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 28, 2024, 4:51pm UTC](https://discuss.elastic.co/t/vector-search-when-input-is-bigger-than-max-seq/363753/6 "2024-08-28T16:51:21Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
