# Update performance - very low indexing rate

**URL:** <https://discuss.elastic.co/t/update-performance-very-low-indexing-rate/103183>\
**Category:** Elasticsearch\
**Created:** [October 9, 2017, 7:07am UTC](https://discuss.elastic.co/t/update-performance-very-low-indexing-rate/103183 "2017-10-09T07:07:52Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![matw](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/matw/32/13913_2.png) [@matw](https://discuss.elastic.co/u/matw)\
**Post date:** [October 9, 2017, 7:07am UTC](https://discuss.elastic.co/t/update-performance-very-low-indexing-rate/103183/1 "2017-10-09T07:07:52Z")

</div>

We're using a redis -\> logstash -\> elasticsearch pipeline

Our test system is a single instance installation with 8CPU, 12 GB RAM, running on VMware  
Currently we can’t separate the components, so no clustering is possible (will change in the future)

Here’s the Logstash redis input

```
redis { 
    data_type => "list“  
    host => "${REDIS_HOST:127.0.0.1}“  
    key => "import“  
    password => „xxx“  
    threads => "2"  
    codec => json { charset => "ASCII“ }
}

```

The messages are saved in 2 indexes (2 Outputs in logstash)

**single-index** collects all messages  
**summary-index** collects special messages, a groovy script is used to create a summary record, that is updated frequently by id (5-x) times

Here's the Logstash output for the single-index

```
 elasticsearch {
      index => "single-index"
      hosts => ["127.0.0.1"]
    }

```

Here's the Logstash output for the summary-index

```
elasticsearch {
            action => "update"
            document_id => "%{uniqueId}"
            index => "summary-index"
            script => "summarize"
            script_lang => "groovy"
            script_type => "file"
            scripted_upsert => true
            retry_on_conflict => 5
            hosts => ["127.0.0.1"]
          }

```

Originally only a part of the messages were sent to the **summary-index**

For this scenario, the indexing rate was ok (max about 9000/s )

Now we’ve got data, that is stored in both indexes, and while it’s clear that the performance  
can’t be like in the mixed scenario, we didn’t expect the performance numbers we’ve got

Sending **1000 msg/s** to redis (20 msg per id -\> 50 summary records)

Result:  
Only **1600/s** indexing rate (2000 would be enough to keep pace). The odd thing is, that the system has a CPU Usage of 50%, Load average of 4, so there seems to be headroom for a higher rate

By deactivating the **single-message** pipeline:  
**700/s**

By deactivating the groovy script (single-message-pipeline still deactivated):  
**1000/s**

So the question is, how could we improve this performance?  
It’s clear that the upserts are the bottleneck.

thx

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [October 9, 2017, 7:28am UTC](https://discuss.elastic.co/t/update-performance-very-low-indexing-rate/103183/2 "2017-10-09T07:28:13Z")

</div>

What does the script do?

---

<div class="post-metadata">

**Author:** ![matw](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/matw/32/13913_2.png) [@matw](https://discuss.elastic.co/u/matw)\
**Post date:** [October 9, 2017, 7:34am UTC](https://discuss.elastic.co/t/update-performance-very-low-indexing-rate/103183/3 "2017-10-09T07:34:55Z")

</div>

it extracts values of the message sent to a summary records (firstMessage, lastMessage, computed state, List of IPs, etc.)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 9, 2017, 7:40am UTC](https://discuss.elastic.co/t/update-performance-very-low-indexing-rate/103183/4 "2017-10-09T07:40:30Z")

</div>

Which version of Elasticsearch are you using?

---

<div class="post-metadata">

**Author:** ![matw](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/matw/32/13913_2.png) [@matw](https://discuss.elastic.co/u/matw)\
**Post date:** [October 9, 2017, 7:41am UTC](https://discuss.elastic.co/t/update-performance-very-low-indexing-rate/103183/5 "2017-10-09T07:41:48Z")

</div>

sry, forgot to mention, logstash + elasticsearch 5.6.1

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 9, 2017, 7:54am UTC](https://discuss.elastic.co/t/update-performance-very-low-indexing-rate/103183/6 "2017-10-09T07:54:23Z")

</div>

There was a change that affected certain types of update scenarios for ES 5.X as outlined in the [release notes](https://www.elastic.co/guide/en/elasticsearch/reference/5.5/breaking_50_document_api_changes.html#_get_api). This was also discussed in [this thread](https://discuss.elastic.co/t/bulk-updates-are-extremely-slow-after-upgrading-to-5-2/79450/26). Does this match how you are doing updates?

---

<div class="post-metadata">

**Author:** ![matw](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/matw/32/13913_2.png) [@matw](https://discuss.elastic.co/u/matw)\
**Post date:** [October 9, 2017, 2:59pm UTC](https://discuss.elastic.co/t/update-performance-very-low-indexing-rate/103183/7 "2017-10-09T14:59:15Z")

</div>

ok, yes, thanks, so the best way is to solve very frequent updates at application level, right? Didn't find a fitting solution using logstash for this use case (anybody knows how?), working on a POC by using another service.

One message of the thread you posted

> As of 5.0.0 the get API will issue a refresh if the requested document has been changed since the last refresh but the change hasn’t been refreshed yet. This will also make all other changes visible immediately. This can have an impact on performance if the same document is updated very frequently using a read modify update pattern since it might create many small segments. This behavior can be disabled by passing realtime=false to the get request.

realtime=false

could that be an option to speed up the current solution (Until i developed a new one), can i add this the the logstash output?

thank you very much

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 6, 2017, 2:59pm UTC](https://discuss.elastic.co/t/update-performance-very-low-indexing-rate/103183/8 "2017-11-06T14:59:37Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
