# Highlighting performance issues with stored field and fvh highlighter

**URL:** https://discuss.elastic.co/t/highlighting-performance-issues-with-stored-field-and-fvh-highlighter/353240
**Category:** Elasticsearch
**Created:** [February 14, 2024, 8:12am UTC](https://discuss.elastic.co/t/highlighting-performance-issues-with-stored-field-and-fvh-highlighter/353240 "2024-02-14T08:12:21Z")
**Posts on this page:** 1
**Showing post:** 3

<div class="post-metadata">

### Author: ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)
#### Post date: [February 14, 2024, 1:58pm UTC](https://discuss.elastic.co/t/highlighting-performance-issues-with-stored-field-and-fvh-highlighter/353240/3 "2024-02-14T13:58:42Z")

</div>

There seems to be some issue with the highlighting as you can see in this similar [post](https://discuss.elastic.co/t/highlighting-slows-down-kibana-searches-considerably/353089/2).

There are some github issues linked, mainly this one:

> <https://github.com/elastic/elasticsearch/issues/103298>
>
> \### Elasticsearch Version
> 
> 8.10.4 and 8.13.0-SNAPSHOT
> 
> \### Installed Plugins…
> 
> \_No response\_
> 
> \### Java Version
> 
> \_bundled\_
> 
> \### OS Version
> 
> ESS Cloud and MacOS
> 
> \### Problem Description
> 
> A customer has reported that highlighting (added by Kibana) is causing their queries in Kibana to extremely slowly. Profiling of the queries showed that 99% of the search time is spent in the Highlight code paths. In their case, the ConstantScoreQuery took less than 1 second, but the highlighting runs for ~850s (this is reproducible).
> 
> I was able to reproduce this to some degree locally (details below), where a the highlighting takes ~50% of the total search time.
> 
> Root cause is not known. Key points include:
> 
> \- the query being run (by Kibana) is a match\_phrase against \`match\_only\_text\`.
> 
> 
> \### Steps to Reproduce
> 
> Here is how I (partially) reproduced the issue locally.
> 
> 
> Use the following mapping to crate an index where \`message\` has \`match\_only\_text\` and 1 shard is created (I didn't test with more than 1 shard).
> 
> \<details\>
> \<summary\>\<b\>Mapping\</b\> (toggle to view)\</summary\>
> \<pre\>
> {
> "settings": {
> "index" : {
> "number\_of\_shards": 1,
> "number\_of\_replicas": 0
> }
> },
> "mappings": {
> "properties": {
> "@timestamp": {
> "type": "date"
> },
> "message": {
> "type": "match\_only\_text"
> },
> "kubernetes": {
> "properties": {
> "container": {
> "properties": {
> "name": {
> "type": "keyword",
> "ignore\_above": 1024
> }
> }
> },
> "daemonset": {
> "properties": {
> "name": {
> "type": "keyword",
> "ignore\_above": 1024
> }
> }
> },
> "deployment": {
> "properties": {
> "name": {
> "type": "keyword",
> "ignore\_above": 1024
> }
> }
> }
> }
> },
> "bytes\_sent": {
> "type": "long"
> },
> "content\_type": {
> "type": "text",
> "fields": {
> "keyword": {
> "type": "keyword",
> "ignore\_above": 256
> }
> }
> },
> "geoip\_location\_lat": {
> "type": "float"
> },
> "geoip\_location\_lon": {
> "type": "float"
> },
> "is\_https": {
> "type": "boolean"
> },
> "request": {
> "type": "text",
> "fields": {
> "keyword": {
> "type": "keyword",
> "ignore\_above": 256
> }
> }
> },
> "response": {
> "type": "long"
> },
> "runtime\_ms": {
> "type": "long"
> },
> "url": {
> "type": "keyword"
> },
> "user\_Agent": {
> "type": "text",
> "fields": {
> "keyword": {
> "type": "keyword",
> "ignore\_above": 256
> }
> }
> },
> "verb": {
> "type": "text",
> "fields": {
> "keyword": {
> "type": "keyword",
> "ignore\_above": 256
> }
> }
> }
> }
> }
> }
> \</pre\>
> \</details\>
> 
> 
> Populate this index with at least 2.5 million records in scope for the search (see my table in comments below).
> 
> \<details\>
> \<summary\>\<b\>Data generation script\</b\> (toggle to view)\</summary\>
> 
> \`\`\`
> import argparse
> import ndjson
> from faker import Faker
> from datetime import datetime, timedelta
> import random
> 
> def generate\_fake\_data(num\_lines):
> fake = Faker()
> 
> data = \[\]
> timestamp = datetime.strptime("2021-04-06T14:00:00", "%Y-%m-%dT%H:%M:%S")
> for \_ in range(num\_lines):
> timestamp\_str = timestamp.strftime("%Y-%m-%dT%H:%M:%S.000Z")
> timestamp += timedelta(seconds=5)
> 
> # Generate a message with at least 85 words
> message\_word\_count = random.randint(85, 225)
> message = fake.paragraph(nb\_sentences=1, ext\_word\_list=None)
> while len(message.split()) \< message\_word\_count:
> message += " " + fake.paragraph(nb\_sentences=1, ext\_word\_list=None)
> 
> ri = random.randint(1,4)
> if ri == 2:
> message += ". request errored 123 forty-five 12:45 foobar " + fake.word() + " " + fake.word()
> elif ri == 3:
> message = "12345 California request completed - " + message
> 
> deployment\_name = "service-integrations"
> if ri \> 2:
> deployment\_name = fake.word()
> 
> json\_data = {
> "@timestamp": timestamp\_str,
> "user\_Agent": fake.user\_agent(),
> "url": fake.uri\_path(),
> "content\_type": fake.mime\_type(),
> "is\_https": fake.boolean(),
> "response": fake.random\_int(min=100, max=599),
> "verb": fake.http\_method(),
> "geoip\_location\_lon": float(fake.longitude()),
> "geoip\_location\_lat": float(fake.latitude()),
> "bytes\_sent": fake.random\_int(min=1000, max=50000),
> "runtime\_ms": fake.random\_int(min=100, max=1000),
> "message": message,
> "kubernetes": {
> "container": {"name": fake.word()},
> "daemonset": {"name": fake.word()},
> "deployment": {"name": deployment\_name}
> }
> }
> 
> data.append(json\_data)
> 
> return data
> 
> def main():
> parser = argparse.ArgumentParser(description="Generate ndjson with Faker")
> parser.add\_argument("num\_lines", type=int, help="Number of lines to produce")
> args = parser.parse\_args()
> 
> fake\_data = generate\_fake\_data(args.num\_lines + 1)
> 
> output\_file\_name = f"sdh7687-data-{args.num\_lines}.json"
> with open(output\_file\_name, "w") as f:
> ndjson.dump(fake\_data, f, ensure\_ascii=False)
> 
> if \_\_name\_\_ == "\_\_main\_\_":
> main()
> 
> \`\`\`
> \</details\>
> 
> \<details\>
> \<summary\>\<b\>Loading script\</b\> (toggle to view)\</summary\>
> 
> \`\`\`
> #!/bin/bash
> 
> if \["$#" -ne 2 \]; then
> echo "Usage: $0 \<port\> \<input\_file\>"
> exit 1
> fi
> 
> port="$1"
> input="$2"
> output="bulk.log"
> counter=0
> max\_rows=10000
> create='{"create": {}}'
> bulk\_data=$'\\n'
> default\_port=9200
> 
> echo "Using port: $port"
> echo "Creating sdh7687 index with mappings to http://localhost:$port ..."
> /usr/bin/curl -s -XPUT "http://localhost:$port/sdh7687" -H 'Content-Type: application/json' --insecure --data-binary @sdh7687-mappings-not-match-only-text.json
> 
> echo ""
> echo "Reading sdh7687 events from $input..."
> while read -r log\_event
> do
> let "counter=counter+1"
> bulk\_data+="$create"$'\\n'"$log\_event"$'\\n'
> if \[$counter -eq $max\_rows \]
> then
> echo "Indexing $counter documents..."
> bulk\_data+=$'\\n'
> echo "$bulk\_data" | tee temp.json \> /dev/null
> /usr/bin/curl -XPOST "http://localhost:$port/sdh7687/\_bulk" -H 'Content-Type: application/x-ndjson' -# --progress-bar --insecure --data-binary @temp.json \>\> "$output"
> rm -rf temp.json
> counter=0
> bulk\_data=$'\\n'
> fi
> done \< "$input"
> 
> if \[$counter -lt $max\_rows \] && \[$counter -gt 0 \]
> then
> echo "Indexing $counter documents..."
> bulk\_data+=$'\\n'
> echo "$bulk\_data" | tee temp.json \> /dev/null
> /usr/bin/curl -XPOST "http://localhost:$port/sdh7687/\_bulk" -H 'Content-Type: application/x-ndjson' -# --progress-bar --insecure --data-binary @temp.json \>\> "$output"
> rm -rf temp.json
> fi
> 
> \`\`\`
> 
> \</details\>
> 
> Note I ran the above two scripts dozens of times to create data files with 100,000 to 200,000 entries, so that most of the data fell within a small-ish @timestamp range to index millions of documents.
> 
> 
> Run the following query (has profiling and highlighting turned on).
> 
> \<details\>
> \<summary\>\<b\>ES Query\</b\> (toggle to view)\</summary\>
> \<pre\>
> POST sdh7687/\_async\_search
> {
> "profile": true,
> "track\_total\_hits": false,
> "sort": \[
> {
> "@timestamp": {
> "order": "desc",
> "unmapped\_type": "boolean"
> }
> }
> \],
> "fields": \[
> {
> "field": "\*",
> "include\_unmapped": "true"
> },
> {
> "field": "@timestamp",
> "format": "strict\_date\_optional\_time"
> }
> \],
> "size": 10,
> "version": true,
> "script\_fields": {},
> "stored\_fields": \[
> "\*"
> \],
> "runtime\_mappings": {},
> "\_source": false,
> "query": {
> "bool": {
> "must": \[\],
> "filter": \[
> {
> "range": {
> "@timestamp": {
> "format": "strict\_date\_optional\_time",
> "gte": "2021-01-02T06:40:00.000Z",
> "lte": "2021-04-10T06:49:07.221Z"
> }
> }
> },
> {
> "match\_phrase": {
> "kubernetes.deployment.name": "service-integrations"
> }
> }
> \],
> "should": \[\],
> "must\_not": \[
> {
> "match\_phrase": {
> "message": "request errored"
> }
> },
> {
> "match\_phrase": {
> "message": "request completed"
> }
> }
> \]
> }
> },
> "highlight": {
> "pre\_tags": \[
> "@kibana-highlighted-field@"
> \],
> "post\_tags": \[
> "@/kibana-highlighted-field@"
> \],
> "fields": {
> "\*": {}
> },
> "fragment\_size": 2147483647
> }
> }
> \</pre\>
> \</details\>
> 
> 
> 
> \### Logs (if relevant)
> 
> \_No response\_

Your issue may be related to this.

---

_[View the full topic](https://discuss.elastic.co/t/highlighting-performance-issues-with-stored-field-and-fvh-highlighter/353240)._
