# Elasticsearch indexing rate from Logstash

**URL:** <https://discuss.elastic.co/t/elasticsearch-indexing-rate-from-logstash/134083>\
**Category:** Elasticsearch\
**Created:** [May 31, 2018, 4:01pm UTC](https://discuss.elastic.co/t/elasticsearch-indexing-rate-from-logstash/134083 "2018-05-31T16:01:20Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![SebScoFr](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sebscofr/32/31623_2.png) [@SebScoFr](https://discuss.elastic.co/u/SebScoFr)\
**Post date:** [May 31, 2018, 4:01pm UTC](https://discuss.elastic.co/t/elasticsearch-indexing-rate-from-logstash/134083/1 "2018-05-31T16:01:20Z")

</div>

So let me try to provide as much context as possible.

We are using ECE. Both ES and Logstash version 6.2.4.  
We have a cluster of **10 nodes, 16GB memory per node**.  
Everyday we have a process responsible for indexing a bunch of documents (~5M) into ES by sending them to logstash (we use pubsub input / es output).  
Our Logstash instance is configured with **16GB of RAM, with -Xms8g -Xmx8g**

logstash.yml has the following:  
**pipeline.workers: 40**  
**pipeline.batch.size: 5000**

Finally the logstash pipeline looks like this:

```
input {
  google_pubsub {
    project_id => "my-project-id"
    topic => "logstash"
    subscription => "logstash"
    max_messages => 5000
  }
}
filter {
  json {
    source => "message"
  }
  mutate {
    remove_field => ["message", "@version", "@timestamp"]
  }
}
output {
  elasticsearch {
    hosts => ["my-es-host"]
    user => "elastic"
    password => "password"
    index => "my-index"
    document_id => "%{[@metadata][uuid]}"
    action => "update"
    scripted_upsert => true
    script => "import"
    script_type => "indexed"
    script_lang => ""
  }
}

```

So basically we are seeing a fairly low indexing rate. Check out the graph below. On the left side is the rate at which messages are posted to logstash, on the right side are messages remaining to be processed by logstash over time.

![58](https://us1.discourse-cdn.com/elastic/original/3X/b/1/b1887bfe391442aaa1315d741d0040a80cece1fb.png)

According to the following monitoring graph, the indexing rate seems to plateau at around 1.5k/s, with a few spikes here and there but they never last.

 ![24](https://us1.discourse-cdn.com/elastic/original/3X/e/2/e266ec7e75f0a89e0ba118a19dd5b44de87f74ce.png)

Finally during this import process, CPU and memory usage are nowhere near capacity.

 ![50](https://us1.discourse-cdn.com/elastic/original/3X/e/e/ee0524512d0570eeb4214fd14f5b52c159711db1.png)

Which makes me feel like the ingestion rate could potentially be much faster. I'm trying to figure out what the bottleneck is. Any thoughts?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 31, 2018, 4:21pm UTC](https://discuss.elastic.co/t/elasticsearch-indexing-rate-from-logstash/134083/2 "2018-05-31T16:21:33Z")

</div>

Wondering what you can see in the Logstash Monitoring part...  
You should be able to see how LS itself is performing IMO.

I'm moving your question to #logstash

---

<div class="post-metadata">

**Author:** ![SebScoFr](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sebscofr/32/31623_2.png) [@SebScoFr](https://discuss.elastic.co/u/SebScoFr)\
**Post date:** [June 1, 2018, 3:19pm UTC](https://discuss.elastic.co/t/elasticsearch-indexing-rate-from-logstash/134083/3 "2018-06-01T15:19:02Z")

</div>

Hi @dadoonet, thanks for the reply.

Here is a comparison of Logstash metrics and Elasticsearch indexing rate when the process is running

 ![57](https://us1.discourse-cdn.com/elastic/original/3X/8/c/8c0b5d187f605b86dbde56b2c7b3017bd8cd1f3f.png)

![06](https://us1.discourse-cdn.com/elastic/original/3X/2/b/2b2cb6bd28a393ea4049cfb44d4429ae9b584799.png)

Those graphs match perfectly which leads me to believe that Logstash is indeed the bottleneck here. Any idea why the **Events Received Rate** would be so low?  
Logstash CPU usage is very low (5%) and heap is also fine.

---

<div class="post-metadata">

**Author:** ![SebScoFr](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sebscofr/32/31623_2.png) [@SebScoFr](https://discuss.elastic.co/u/SebScoFr)\
**Post date:** [June 3, 2018, 6:56pm UTC](https://discuss.elastic.co/t/elasticsearch-indexing-rate-from-logstash/134083/4 "2018-06-03T18:56:36Z")

</div>

Hi @dadoonet, I think you can move that back to the ES section.

I was wrong in thinking that Logstash was the issue. I configured my pipeline to send the output to stdout (for testing purpose) and now the LS monitoring looks like this.

 ![00](https://us1.discourse-cdn.com/elastic/original/3X/b/a/baa30182689bf4291b977153ea5135d783073981.png)

Receive Rate is now 5 times higher which makes it obvious that ES is actually my bottleneck here.

---

<div class="post-metadata">

**Author:** ![SebScoFr](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sebscofr/32/31623_2.png) [@SebScoFr](https://discuss.elastic.co/u/SebScoFr)\
**Post date:** [June 4, 2018, 3:15pm UTC](https://discuss.elastic.co/t/elasticsearch-indexing-rate-from-logstash/134083/5 "2018-06-04T15:15:31Z")

</div>

After further investigation, I really suspect it must have to do with the fact that I'm using a **scripted\_upsert**

```
output {
  elasticsearch {
    hosts => ["my-es-host"]
    user => "elastic"
    password => "password"
    index => "my-index"
    document_id => "%{[@metadata][uuid]}"
    action => "update"
    scripted_upsert => true
    script => "import"
    script_type => "indexed"
    script_lang => ""
  }
}

```

I did some tests with a normal index (not scripted upsert) instead and it was much faster.  
Is there a limit on how many scripts ES can execute on a given period?

The script itself looks like this, nothing too funky

```
POST _scripts/import
{
  "script": {
    "lang": "painless",
    "source": "if (ctx._source.created == null) { ctx._source.created = params.event.get('updated'); }"
  }
}

```

The reason I'm using a script is to be able to keep track of the created timestamp while doing an upsert.

---

<div class="post-metadata">

**Author:** ![SebScoFr](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sebscofr/32/31623_2.png) [@SebScoFr](https://discuss.elastic.co/u/SebScoFr)\
**Post date:** [June 9, 2018, 1:03pm UTC](https://discuss.elastic.co/t/elasticsearch-indexing-rate-from-logstash/134083/6 "2018-06-09T13:03:31Z")

</div>

Would anyone be able to confirm that a **scripted\_upsert** is indeed much slower than an index or even a regular upsert?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 7, 2018, 1:03pm UTC](https://discuss.elastic.co/t/elasticsearch-indexing-rate-from-logstash/134083/7 "2018-07-07T13:03:33Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
