# Very bad performance with large text field

**URL:** <https://discuss.elastic.co/t/very-bad-performance-with-large-text-field/89924>\
**Category:** Elasticsearch\
**Created:** [June 19, 2017, 12:06pm UTC](https://discuss.elastic.co/t/very-bad-performance-with-large-text-field/89924 "2017-06-19T12:06:23Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![mos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mos/32/4898_2.png) [@mos](https://discuss.elastic.co/u/mos)\
**Post date:** [June 19, 2017, 12:06pm UTC](https://discuss.elastic.co/t/very-bad-performance-with-large-text-field/89924/1 "2017-06-19T12:06:23Z")

</div>

At one of my customer projects we work with documents containing a very large text-field (content of eBooks....).  
We saw that queries slow down more then 100 x times when such documents are queried. Even if we use the source-filter and exclude this field from the query! The only solution is to exclude the text-field from the \_source at index-time from the documents:

```
PUT my_index
{
  "mappings": {
    "_default_": {
      "_source": {
        "excludes": [
          "mylargeTextField"
        ]
      },
....

```

I found the following blog post that explains the issue in detail:

> **[Making ElasticSearch Perform Well with Large Text Fields](https://blog.ambar.cloud/making-elasticsearch-perform-well-with-large-text-fields/)**
>
> We're continuing our story about creating Ambar, and this is the second paper about ElasticSearch. The first one is Highlighting Large Documents in ElasticSearch. This paper tells the story about making ElasticSearch perform well with documents...

What I don't understand:  
Why is the search so slow even when using the source filtering in the query? There should be no need to fetch, retrieve and merge the excluded fields? I was expecting when using something like

```
GET /_search
{
    "_source": {
        "excludes": ["mylargeTextField"]
    },
    "query" : {
        "term" : { "otherField" : "something" }
    }
}

```

the large text-field shouldn't impact the performance at all and should be ignored for this specific query?

Currently we are using Elasticsearch also as a datastore. If such large documents are slow things down, Elasticsearch seems not to be a perfect datastore (in contrast to MongoDB)?

(We are using the latest ES 5.4.x)

---

<div class="post-metadata">

**Author:** ![mos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mos/32/4898_2.png) [@mos](https://discuss.elastic.co/u/mos)\
**Post date:** [June 20, 2017, 9:27am UTC](https://discuss.elastic.co/t/very-bad-performance-with-large-text-field/89924/2 "2017-06-20T09:27:39Z")

</div>

Anyone from Elastic?

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [June 20, 2017, 2:40pm UTC](https://discuss.elastic.co/t/very-bad-performance-with-large-text-field/89924/3 "2017-06-20T14:40:07Z")

</div>

There are two potential reasons:

- CPU overhead: the json parser still needs to skip over the large text field in order to exclude it from the \_source, which is linear with the size of your json doc.
- Disk overhead, since those large fields make the index larger and thus the filesystem cache can only hold a smaller ratio of the total index size.

---

<div class="post-metadata">

**Author:** ![mos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mos/32/4898_2.png) [@mos](https://discuss.elastic.co/u/mos)\
**Post date:** [June 21, 2017, 3:34pm UTC](https://discuss.elastic.co/t/very-bad-performance-with-large-text-field/89924/4 "2017-06-21T15:34:35Z")

</div>

Thanks @jpountz!  
It seems not to be the "json-parsing".  
If we use the search without any searchterms it's fast as hell (2 ms):

`GET index_with_large_text/_search`

If we use a simple searchterm, there is the performance problem (152 ms):

`GET index_with_large_text/_search?q=any_field:something`

So JSON parsing seems not be the problem. Must be something with "Disk overhead" during the search- or merge-phase , right? Do you thing increasing RAM/Heap-Size can fix the problem?

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [June 21, 2017, 4:25pm UTC](https://discuss.elastic.co/t/very-bad-performance-with-large-text-field/89924/5 "2017-06-21T16:25:38Z")

</div>

How large is your index (the size of the data dir) and how much do you give to the filesystem cache?

---

<div class="post-metadata">

**Author:** ![mos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mos/32/4898_2.png) [@mos](https://discuss.elastic.co/u/mos)\
**Post date:** [June 23, 2017, 3:44pm UTC](https://discuss.elastic.co/t/very-bad-performance-with-large-text-field/89924/6 "2017-06-23T15:44:13Z")

</div>

- The "data dir" is about 490MB.
- Disk caching is on 10G (page cache via free -mh)
- Java Heap: "heap\_max": "30.7gb" / "heap\_used": "2.1gb"

It seems to be no memory/caching issue when looking on this sizing.

We tried the query in a cluster of 4 node (same params as above) and on a single node with just one shard.  
The bad performance remains unchanged.

The content of the large-text is around 700KB for a single document.

Any suggestion what we can test or do? Is it possible that Elasticsearch is not practical usable on such large text fields in the \_source object?

---

<div class="post-metadata">

**Author:** ![mos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mos/32/4898_2.png) [@mos](https://discuss.elastic.co/u/mos)\
**Post date:** [June 26, 2017, 1:13pm UTC](https://discuss.elastic.co/t/very-bad-performance-with-large-text-field/89924/7 "2017-06-26T13:13:45Z")

</div>

@jpountz Do you have additional ideas?

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [June 27, 2017, 1:44pm UTC](https://discuss.elastic.co/t/very-bad-performance-with-large-text-field/89924/8 "2017-06-27T13:44:38Z")

</div>

> [@mos](#):
>
> GET index\_with\_large\_text/\_search?q=any\_field:something

This query executes scoring on `any_field`. Maybe you see the effect because of missing or bad stop word analyzing, or a large number of segments.

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [June 29, 2017, 10:18am UTC](https://discuss.elastic.co/t/very-bad-performance-with-large-text-field/89924/9 "2017-06-29T10:18:04Z")

</div>

Can you confirm you are not using really large `size` values? Also when you say, 100x slower, what is the order of magnitude of the response times we are talking about? Is it 100% reproducible?

---

<div class="post-metadata">

**Author:** ![mos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mos/32/4898_2.png) [@mos](https://discuss.elastic.co/u/mos)\
**Post date:** [June 29, 2017, 11:04am UTC](https://discuss.elastic.co/t/very-bad-performance-with-large-text-field/89924/10 "2017-06-29T11:04:13Z")

</div>

When indexed without storing the large text-field in \_source the "took"-time is around: 1-2ms  
With this field in \_source: 90ms -120ms (around 100x slower)  
Yes, it's always reproducible.

For our tests we are using the default size of 10 hits that should be returned.  
If using size=1 it's much faster; a size of 10000 is slowing things much more down

Currently we thinking about not storing those large texts in Elasticsearch, but using MongoDB for this. But we will loose highlighting features and some nice-to-have functions like "reindexing" and "updates" within Elasticsearch.....

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [June 29, 2017, 1:09pm UTC](https://discuss.elastic.co/t/very-bad-performance-with-large-text-field/89924/11 "2017-06-29T13:09:41Z")

</div>

It means Elasticsearch is taking about 100ms to do the source filtering for only 10 documents, which is puzzling.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 27, 2017, 1:09pm UTC](https://discuss.elastic.co/t/very-bad-performance-with-large-text-field/89924/12 "2017-07-27T13:09:47Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
