# Slow handling of documents when large text in a field

**URL:** <https://discuss.elastic.co/t/slow-handling-of-documents-when-large-text-in-a-field/302679>\
**Category:** Elasticsearch\
**Created:** [April 19, 2022, 7:31am UTC](https://discuss.elastic.co/t/slow-handling-of-documents-when-large-text-in-a-field/302679 "2022-04-19T07:31:13Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![lubosvr](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lubosvr/32/104400_2.png) [@lubosvr](https://discuss.elastic.co/u/lubosvr)\
**Post date:** [April 19, 2022, 7:31am UTC](https://discuss.elastic.co/t/slow-handling-of-documents-when-large-text-in-a-field/302679/1 "2022-04-19T07:31:13Z")

</div>

Hi,  
We have problems when handling for documents with large content in a single field.

We have index with mapping like this:

```auto
"mappings" : {
  "properties" : {
    "properties" : {
      "name" : {
        "type" : "text"
      },
      "searchContentHTML" : {
        "type" : "text"
      }
    }
  }

```

The problem is that inside `searchContentHTML` can be quite a long text (MBs) that we are basically using for fulltext search only. (Hardly to be usefull for returning to clients)  
When the text is about 14MB long the `getById` query (with `_source_excludes=searchContentHTML` parameter) takes about 100ms and when the text is small it takes 30ms.

It means it is 3 times longer to simply get one field whenone filed is long..!

Are there any good practices how to handle such document with Elasticsearch?

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [April 19, 2022, 3:42pm UTC](https://discuss.elastic.co/t/slow-handling-of-documents-when-large-text-in-a-field/302679/2 "2022-04-19T15:42:48Z")

</div>

Did you look a [this](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-fields.html#search-fields-request)

You can set

` "_source": false`

and then just select the fields you want to return

`"fields": ["name"]`

```auto
GET my-index-000001/_search
{
  "query": {
....
  },
  "fields": ["name"]
  "_source": false
}

```

That may be faster....

---

<div class="post-metadata">

**Author:** ![lubosvr](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lubosvr/32/104400_2.png) [@lubosvr](https://discuss.elastic.co/u/lubosvr)\
**Post date:** [April 19, 2022, 4:02pm UTC](https://discuss.elastic.co/t/slow-handling-of-documents-when-large-text-in-a-field/302679/3 "2022-04-19T16:02:04Z")

</div>

Thanks for reply @stephenb  
I've tried

```auto
POST quickassets/_search
{
  "query": {
    "match": {
      "_id": "-YFy-n8BGruhv47CIXSL"
    }
  }, 
  "fields": ["name"],
  "_source": false
}

```

And the response took 90-110ms, so the same time ☹

Currently I see only one option. Move the `searchContentHTML` field to another index and handle fulltext search a different way - search multiple indices.

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [April 19, 2022, 4:06pm UTC](https://discuss.elastic.co/t/slow-handling-of-documents-when-large-text-in-a-field/302679/4 "2022-04-19T16:06:11Z")

</div>

Hmmm interesting I would have expected that to be much faster...

When you used `_source` ... you used exclude could you just try

` "_source": "name",`

---

<div class="post-metadata">

**Author:** ![lubosvr](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lubosvr/32/104400_2.png) [@lubosvr](https://discuss.elastic.co/u/lubosvr)\
**Post date:** [April 19, 2022, 8:50pm UTC](https://discuss.elastic.co/t/slow-handling-of-documents-when-large-text-in-a-field/302679/5 "2022-04-19T20:50:13Z")

</div>

Sure  
with

```auto
POST quickassets/_search
{
  "query": {
    "match": {
      "_id": "-YFy-n8BGruhv47CIXSL"
    }
  }, 
  "fields": ["name"],
  "_source": "name"
}

```

I'm getting the same times. It looks ES has troubles to parse such a huge documents..  
BTW here is the response I can see in DevTools in Kibana (for the long) document:

```auto
{
  "took" : 83,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 1,
      "relation" : "eq"
    },
    "max_score" : 1.0,
    "hits" : [
      {
        "_index" : "quickassets",
        "_type" : "_doc",
        "_id" : "-YFy-n8BGruhv47CIXSL",
        "_score" : 1.0,
        "_source" : {
          "name" : "long_text"
        },
        "fields" : {
          "name" : [
            "long_text"
          ]
        }
      }
    ]
  }
}

```

And for the short one

```auto
{
  "took" : 1,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 1,
      "relation" : "eq"
    },
    "max_score" : 1.0,
    "hits" : [
      {
        "_index" : "quickassets",
        "_type" : "_doc",
        "_id" : "sOjP_n8BLbh09t6Ko7bv",
        "_score" : 1.0,
        "_source" : {
          "name" : "short_text"
        },
        "fields" : {
          "name" : [
            "short_text"
          ]
        }
      }
    ]
  }
}

```

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [April 19, 2022, 8:55pm UTC](https://discuss.elastic.co/t/slow-handling-of-documents-when-large-text-in-a-field/302679/6 "2022-04-19T20:55:14Z")

</div>

> [@lubosvr](#):
>
> ```auto
> POST quickassets/_search
> {
> "query": {
> "match": {
> "_id": "-YFy-n8BGruhv47CIXSL"
> }
> }, 
> "fields": ["name"]
> }
> 
> ```

When you just did this... it was still long?  
Avoiding source all-together

---

<div class="post-metadata">

**Author:** ![lubosvr](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lubosvr/32/104400_2.png) [@lubosvr](https://discuss.elastic.co/u/lubosvr)\
**Post date:** [April 20, 2022, 8:13am UTC](https://discuss.elastic.co/t/slow-handling-of-documents-when-large-text-in-a-field/302679/7 "2022-04-20T08:13:35Z")

</div>

> [@stephenb](#):
>
> When you just did this... it was still long?  
> Avoiding source all-together

When I do this, it takes ages in Kibana, it returns the whole document including the `searchContentHTML ` field. I tried it with the short document.

```auto
{
  "took" : 1,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 1,
      "relation" : "eq"
    },
    "max_score" : 1.0,
    "hits" : [
      {
        "_index" : "quickassets",
        "_type" : "_doc",
        "_id" : "sOjP_n8BLbh09t6Ko7bv",
        "_score" : 1.0,
        "_source" : {
          "searchContentHTML" : "Deutsches Ipsum Dolor deserunt dissentias zu spät et. Tollit argumentum ius an. Kartoffelkopf lobortis elaboraret per ne, nam Schnaps probatus pertinax, impetus eripuit aliquando Guten Tag sea. Diam scripserit no vis, Hockenheim meis suscipit ea. Eam ea Freude schöner Götterfunken eleifend, ad blandit voluptatibus sed, Zauberer eius consul sanctus vix. Cu Freude schöner Götterfunken legimus veritus vim",
          "name" : "short_text"
        },
        "fields" : {
          "name" : [
            "short_text"
          ]
        }
      }
    ]
  }
}

```

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [April 20, 2022, 1:01pm UTC](https://discuss.elastic.co/t/slow-handling-of-documents-when-large-text-in-a-field/302679/8 "2022-04-20T13:01:45Z")

</div>

Apologies I left the

`"_source": false`

Out of that query... Typo

Should have been

```auto
POST quickassets/_search
{
  "query": {
    "match": {
      "_id": "-YFy-n8BGruhv47CIXSL"
    }
  }, 
  "fields": ["name"],
  "_source": false
}

```

I would think this would be the fastest option

---

<div class="post-metadata">

**Author:** ![lubosvr](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lubosvr/32/104400_2.png) [@lubosvr](https://discuss.elastic.co/u/lubosvr)\
**Post date:** [April 20, 2022, 1:38pm UTC](https://discuss.elastic.co/t/slow-handling-of-documents-when-large-text-in-a-field/302679/9 "2022-04-20T13:38:32Z")

</div>

> [@stephenb](#):
>
> "\_source": false

makes no difference

```auto
{
  "took" : 90,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 1,
      "relation" : "eq"
    },
    "max_score" : 1.0,
    "hits" : [
      {
        "_index" : "quickassets",
        "_type" : "_doc",
        "_id" : "-YFy-n8BGruhv47CIXSL",
        "_score" : 1.0,
        "fields" : {
          "name" : [
            "long_text"
          ]
        }
      }
    ]
  }
}

```

☹

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [April 20, 2022, 1:57pm UTC](https://discuss.elastic.co/t/slow-handling-of-documents-when-large-text-in-a-field/302679/10 "2022-04-20T13:57:10Z")

</div>

Have you forced merge the index?

Have you run in query profiler to see what's taking so long?

Ohh also why are you not using a term filter instead of a match?

With a filter no scoring is performed...

```auto
POST filebeat-7.15.2-2022.04.20-000142/_search
{
  "_source" : false,
  "fields": [
    "host.name"
  ], 
  "query": {
    "bool": {
      "filter": [
        {
          "term": {
            "_id": "Fh89R4ABxTyfpfWFmza8"
          }
        }
      ]
    }
  }
}

```

---

<div class="post-metadata">

**Author:** ![lubosvr](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lubosvr/32/104400_2.png) [@lubosvr](https://discuss.elastic.co/u/lubosvr)\
**Post date:** [April 20, 2022, 2:45pm UTC](https://discuss.elastic.co/t/slow-handling-of-documents-when-large-text-in-a-field/302679/11 "2022-04-20T14:45:03Z")

</div>

> [@stephenb](#):
>
> Have you forced merge the index?

Yes, with no impact on performance

> [@stephenb](#):
>
> Have you run in query profiler to see what's taking so long?

Well, in fact in production I'm using getById query, something like this: But it takes approximately the same time.  
`GET quickassets/_doc/-YFy-n8BGruhv47CIXSL?_source_excludes=searchContentHTML` or  
`GET quickassets/_doc/-YFy-n8BGruhv47CIXSL?_source_includes=name`  
It led me to the conclusion that the problem is not in the query itself, but somewhere else..  
I tried the profiler, but there is nothing interesting. Most time is spend by  
`build_scorer 23.2µs 48.5%`, which makes no sense as the query takes much longer

> [@stephenb](#):
>
> Ohh also why are you not using a term filter instead of a match?

As I wrote, I'm using get document by id..

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 18, 2022, 2:45pm UTC](https://discuss.elastic.co/t/slow-handling-of-documents-when-large-text-in-a-field/302679/12 "2022-05-18T14:45:14Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
