# Text\_embedding configured for model but rejected for query

**URL:** <https://discuss.elastic.co/t/text-embedding-configured-for-model-but-rejected-for-query/357163>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-machine-learning\
**Created:** [April 10, 2024, 5:00pm UTC](https://discuss.elastic.co/t/text-embedding-configured-for-model-but-rejected-for-query/357163 "2024-04-10T17:00:31Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![wnmills3](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wnmills3/32/133048_2.png) [@wnmills3](https://discuss.elastic.co/u/wnmills3)\
**Post date:** [April 10, 2024, 5:00pm UTC](https://discuss.elastic.co/t/text-embedding-configured-for-model-but-rejected-for-query/357163/1 "2024-04-10T17:00:31Z")

</div>

I have issued a query like below:

```auto
{
    "query": {
        "bool": {
            "should": [
                {
                    "text_embedding": {
                        "path_embedding.tokens": {
                            "model_id": "intfloat__multilingual-e5-base",
                            "model_text": "What is the purpose of the EHS Location Hierarchy?"
                        }
                    }
                },
                {
                    "text_embedding": {
                        "passage_embeddding.tokens": {
                            "model_id": "intfloat__multilingual-e5-base",
                            "model_text": "What is the purpose of the EHS Location Hierarchy?"
                        }
                    }
                }
            ]
        }
    }
}

```

and it gets rejected with this error:

```auto
{
    "error": {
        "root_cause": [
            {
                "type": "parsing_exception",
                "reason": "unknown query [text_embedding]",
                "line": 6,
                "col": 39
            }
        ],
        "type": "x_content_parse_exception",
        "reason": "[6:39] [bool] failed to parse field [should]",
        "caused_by": {
            "type": "parsing_exception",
            "reason": "unknown query [text_embedding]",
            "line": 6,
            "col": 39,
            "caused_by": {
                "type": "named_object_not_found_exception",
                "reason": "[6:39] unknown field [text_embedding]"
            }
        }
    },
    "status": 400
}

```

However, viewing the documents, they have both the passage\_embedding tokens array and path\_embedding tokens array. The model shows text\_embedding as a type:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/8/8/883cfb4c38d434ea3c82a3545c15a245ed712199.png)  
and the pipeline processors have generated these tokens using:

```auto
{
  "processors": [
    {
      "inference": {
        "model_id": "intfloat__multilingual-e5-base",
        "target_field": "passage_embedding",
        "field_map": {
          "passage": "text_field"
        },
        "inference_config": {
          "text_embedding": {
            "results_field": "tokens"
          }
        }
      }
    },
    {
      "inference": {
        "model_id": "intfloat__multilingual-e5-base",
        "target_field": "path_embedding",
        "field_map": {
          "path": "text_field"
        },
        "inference_config": {
          "text_embedding": {
            "results_field": "tokens"
          }
        }
      }
    }
  ]
}

```

Why would the query be rejected?

I was following the example [here](https://discuss.elastic.co/t/using-elser-for-multiple-fields/341347/2).

Note: if I use the incorrect text\_expansion the error returns claiming the model is configured for text\_embedding:

```auto
{
    "error": {
        "root_cause": [
            {
                "type": "status_exception",
                "reason": "Trained model [intfloat__multilingual-e5-base] is configured for task [text_embedding] but called with task [text_expansion]"
            }
        ],
        "type": "status_exception",
        "reason": "Trained model [intfloat__multilingual-e5-base] is configured for task [text_embedding] but called with task [text_expansion]",
        "caused_by": {
            "type": "status_exception",
            "reason": "Trained model [intfloat__multilingual-e5-base] is configured for task [text_embedding] but called with task [text_expansion]"
        }
    },
    "status": 403
}

```

---

<div class="post-metadata">

**Author:** ![wnmills3](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wnmills3/32/133048_2.png) [@wnmills3](https://discuss.elastic.co/u/wnmills3)\
**Post date:** [April 12, 2024, 5:32pm UTC](https://discuss.elastic.co/t/text-embedding-configured-for-model-but-rejected-for-query/357163/2 "2024-04-12T17:32:35Z")

</div>

I found lots of errors with the sequence I was building the index so the mappings were FUBAR. No need to review/answer this request. Thank you.

---

<div class="post-metadata">

**Author:** ![javigsg](https://avatars.discourse-cdn.com/v4/letter/j/c57346/32.png) [@javigsg](https://discuss.elastic.co/u/javigsg)\
**Post date:** [April 30, 2024, 11:53am UTC](https://discuss.elastic.co/t/text-embedding-configured-for-model-but-rejected-for-query/357163/3 "2024-04-30T11:53:44Z")

</div>

Hi @wnmills3 I got the same problem. I am trying to create the embeddings from a web crawl, but when I am trying to run an inference I got this error:

_**BadRequestError(400, 'parsing\_exception', 'unknown query [text\_embedding]')**_

I am seeing the embeddings in the documents tab, but I am not able to run an inference. Any advice? Thanks

---

<div class="post-metadata">

**Author:** ![wnmills3](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wnmills3/32/133048_2.png) [@wnmills3](https://discuss.elastic.co/u/wnmills3)\
**Post date:** [May 1, 2024, 1:10pm UTC](https://discuss.elastic.co/t/text-embedding-configured-for-model-but-rejected-for-query/357163/4 "2024-05-01T13:10:37Z")

</div>

I initially had separate calls to set up mappings and processors, etc. Because they were not done in the proper order the index was not created correctly. I settled on making a single call with all the details supplied at once when creating the index:

```auto
    void createIndexAndIngestData(String indexName, Path jsonlFilePath,
        String langFamily) {
        String stage = "creating index " + indexName + " for " + jsonlFilePath;
        try {
            if (_indexNames.contains(indexName) == false) {
                ESUtils.createReplaceIndex(_client, indexName, _removeIndex, langFamily);
                _indexNames.add(indexName);
                if (_thumbsucker) {
                    System.out.println("Created index: " + indexName);
                }
                ESUtils.setPipeline(_client, indexName,
                    langFamily, _logger);
            }
            stage = "Ingesting data for index " + indexName;
            System.out.println(stage);
            ingestPassagesFromFile(jsonlFilePath.toString(), indexName);
        } catch (Exception e) {
            _logger.error(stage, e);
        }
    }

```

Which makes these calls:

```auto
            IndexSettings is = getSettingsRequest(indexName, langFamily);
            TypeMapping mr = getMappingsRequest(indexName, langFamily);
            CreateIndexResponse response = client.indices()
                .create(cir -> cir.timeout(Time.of(t -> t.time("60s")))
                    .index(indexName).settings(is).mappings(mr));
            result = response.acknowledged();

```

where I read the settings from a json file:

```auto
    static public IndexSettings getSettingsRequest(String indexName,
        String langFamily) throws FileNotFoundException {
        String filename = "." + File.separator + "properties" + File.separator
            + langFamily + "_settings.json";
        // check existence
        File test = new File(filename);
        if (!test.exists()) {
            String defaultFilename = "." + File.separator + "properties"
                + File.separator + "en_settings.json";
            test = new File(defaultFilename);
            if (!test.exists()) {
                throw new FileNotFoundException("Can not find \"" + filename
                    + "\" nor default language \"" + defaultFilename + "\"");
            }
            filename = defaultFilename;
        }
        final FileReader file = new FileReader(new File(filename));

        IndexSettings req;

        req = IndexSettings.of(b -> b.withJson(file));

        return req;
    }

```

and the mappings from a json file:

```auto
    static public TypeMapping getMappingsRequest(String indexName,
        String langFamily) throws FileNotFoundException {
        String filename = "." + File.separator + "properties" + File.separator
            + langFamily + "_mappings.json";
        // check existence
        File test = new File(filename);
        if (!test.exists()) {
            String defaultFilename = "." + File.separator + "properties"
                + File.separator + "en_mappings.json";
            test = new File(defaultFilename);
            if (!test.exists()) {
                throw new FileNotFoundException("Can not find \"" + filename
                    + "\" nor default language \"" + defaultFilename + "\"");
            }
            filename = defaultFilename;
        }
        final FileReader file = new FileReader(new File(filename));

        TypeMapping req;

        req = TypeMapping.of(b -> b.withJson(file));

        return req;
    }

```

and setting the processors from a json file:

```auto
    static public void setPipeline(ElasticsearchClient client, String indexName,
        String langFamily, Logger logger) throws Exception {
        String filename = "." + File.separator + "properties" + File.separator
            + langFamily + "_processors.json";
        List<String> lines = DPUtils.loadTextFile(filename);
        StringBuffer sb = new StringBuffer();
        for (String line : lines) {
            sb.append(line + "\n");
        }
        String rdpPipeline = sb.toString();
        Request request = new Request("PUT", "/_ingest/pipeline/rdp_pipeline");
        request.setJsonEntity(rdpPipeline);
        RestClientTransport restClientTransport = (RestClientTransport) client
            ._transport();
        Response response = restClientTransport.restClient()
            .performRequest(request);
        if (response.getStatusLine().getStatusCode() == 200) {
            ObjectMapper objectMapper = new ObjectMapper();
            String acknowledged = EntityUtils.toString(response.getEntity());
            AcknowledgedResponse ak_response = objectMapper
                .readValue(acknowledged, CreatePipelineResponse.class);
            if (!ak_response.acknowledged()) {
                logger.error("Creating pipeline returned false.");
            }
        } else {
            logger.error("Could not set rdp_pipeline due to request status "
                + response.getStatusLine().getStatusCode());
        }
    }

```

---

<div class="post-metadata">

**Author:** ![dkyle](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dkyle/32/59114_2.png) [@dkyle](https://discuss.elastic.co/u/dkyle)\
**Post date:** [May 1, 2024, 2:08pm UTC](https://discuss.elastic.co/t/text-embedding-configured-for-model-but-rejected-for-query/357163/5 "2024-05-01T14:08:53Z")

</div>

`text_embedding` is an option to KNN search, see [k-nearest neighbor (kNN) search | Elasticsearch Guide [8.13] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/knn-search.html#knn-semantic-search)

The `text_expansion` query is for sparse vectors like those created by the [ELSER](https://www.elastic.co/guide/en/machine-learning/master/ml-nlp-elser.html) model. Use knn search and the text\_embedding option for dense vectors as created by the multilingual-e5-base model.

Here is a useful guide to semantic search in Elastic it covers the ELSER model and text embedding models: [Semantic search | Elasticsearch Guide [8.13] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-search.html)

---

<div class="post-metadata">

**Author:** ![javigsg](https://avatars.discourse-cdn.com/v4/letter/j/c57346/32.png) [@javigsg](https://discuss.elastic.co/u/javigsg)\
**Post date:** [May 3, 2024, 4:36pm UTC](https://discuss.elastic.co/t/text-embedding-configured-for-model-but-rejected-for-query/357163/7 "2024-05-03T16:36:08Z")

</div>

how can I solve this issue?

```auto
"[dense_vector] fields cannot be indexed if they're within [nested] mappings"

```

---

<div class="post-metadata">

**Author:** ![dkyle](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dkyle/32/59114_2.png) [@dkyle](https://discuss.elastic.co/u/dkyle)\
**Post date:** [May 5, 2024, 7:48am UTC](https://discuss.elastic.co/t/text-embedding-configured-for-model-but-rejected-for-query/357163/8 "2024-05-05T07:48:45Z")

</div>

What version of Elasticsearch are you using? Older versions do not support nested dense\_vector fields.

Here's a great blog that shows you how to set up nested dense vector field mappings and chunk large documents

> **[Chunking Large Documents via Ingest pipelines plus nested vectors equals easy...](https://www.elastic.co/search-labs/blog/chunking-via-ingest-pipelines)**
>
> In this post we'll show how to easily ingest large documents and break them up into sentences via an ingest pipeline so that they can be text embedded along with nested vector support for searching large documents semantically. Generated image of a...

---

<div class="post-metadata">

**Author:** ![javigsg](https://avatars.discourse-cdn.com/v4/letter/j/c57346/32.png) [@javigsg](https://discuss.elastic.co/u/javigsg)\
**Post date:** [May 5, 2024, 6:34pm UTC](https://discuss.elastic.co/t/text-embedding-configured-for-model-but-rejected-for-query/357163/9 "2024-05-05T18:34:10Z")

</div>

I am using 8.10, is nested dense\_vector supported? Yes that's the documentation that I am following, thanks!

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 5, 2024, 6:45pm UTC](https://discuss.elastic.co/t/text-embedding-configured-for-model-but-rejected-for-query/357163/10 "2024-05-05T18:45:50Z")

</div>

Based on the docs it seems it was added in Elasticsearch 8.11, so I would recommend upgrading to the latest version.
