# Max Indexing rate per single Elasticsearch server

**URL:** <https://discuss.elastic.co/t/max-indexing-rate-per-single-elasticsearch-server/91340>\
**Category:** Elasticsearch\
**Created:** [June 29, 2017, 7:49pm UTC](https://discuss.elastic.co/t/max-indexing-rate-per-single-elasticsearch-server/91340 "2017-06-29T19:49:31Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![jakesjohn](https://avatars.discourse-cdn.com/v4/letter/j/90ced4/32.png) [@jakesjohn](https://discuss.elastic.co/u/jakesjohn)\
**Post date:** [June 29, 2017, 7:49pm UTC](https://discuss.elastic.co/t/max-indexing-rate-per-single-elasticsearch-server/91340/1 "2017-06-29T19:49:31Z")

</div>

I have a single ES node(5.4.3) for doing benchmarks. It is a 64 GB machine with having 32 G as Heap size. Others are default settings with mlocking kept as true

I am using the following settings for the benchmark tests  
No of written indices 10  
total number of documents 10 -  
concurrent clients 10  
No of-shards 1  
Number-of-replicas 0  
Bulk-size of 5000 with each document with max-fields of 10 and max-size-per-field as 50

I am getting just 30000 requests per second(around 7MB/sec). Indexing rate remains more or less same if I change above test settings.  
I know that they are default settings. But, how do I know that I have reached max out of a single node and need to scale? How do I find some theoretical max or can someone share their maximum indexing rate that was achieved?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 29, 2017, 7:57pm UTC](https://discuss.elastic.co/t/max-indexing-rate-per-single-elasticsearch-server/91340/2 "2017-06-29T19:57:42Z")

</div>

Do you have monitoring installed? What does CPU and disk I/O looking like during indexing? How large are your documents? What type of data do they contain? What do your mappings look like? How are you loading the data?

---

<div class="post-metadata">

**Author:** ![jakesjohn](https://avatars.discourse-cdn.com/v4/letter/j/90ced4/32.png) [@jakesjohn](https://discuss.elastic.co/u/jakesjohn)\
**Post date:** [June 29, 2017, 9:06pm UTC](https://discuss.elastic.co/t/max-indexing-rate-per-single-elasticsearch-server/91340/3 "2017-06-29T21:06:09Z")

</div>

What does CPU and disk I/O looking like during indexing?  
CPU and IO are around 20-30% . Below is the iostat output

```
avg-cpu: %user %nice %system %iowait %steal %idle
          14.30 0.00 2.03 0.84 0.00 82.83

Device: rrqm/s wrqm/s r/s w/s rkB/s wkB/s avgrq-sz avgqu-sz await r_await w_await svctm %util
sda 0.00 665.00 0.00 291.00 0.00 11808.00 81.15 1.01 3.46 0.00 3.46 1.03 30.00

```

How large are your documents?  
There are only 10 unique documents in whole benchmark. Each document contains maximum of 10 fields with key of 10 characters and value of 50 characters. This means that a document has max of 600 characters. And documents are sent to ES in a bulk of 5000 documents continuously(each batch contains 5000 docs).

What type of data do they contain?  
Each character is a alphabet

What do your mappings look like?  
It is dynamic mapping. I haven't created any static mappings.

How are you loading the data?  
I am using elasticsearch python module Elasticsearch().bulk() to push bulk requests

Here is the output after 30 second run

> ```
> {
> 
> ```
> 
> "\_nodes" : {  
> "total" : 1,  
> "successful" : 1,  
> "failed" : 0  
> },  
> "cluster\_name" : "cluster",  
> "timestamp" : 1498770136805,  
> "status" : "green",  
> "indices" : {  
> "count" : 10,  
> "shards" : {  
> "total" : 10,  
> "primaries" : 10,  
> "replication" : 0.0,  
> "index" : {  
> "shards" : {  
> "min" : 1,  
> "max" : 1,  
> "avg" : 1.0  
> },  
> "primaries" : {  
> "min" : 1,  
> "max" : 1,  
> "avg" : 1.0  
> },  
> "replication" : {  
> "min" : 0.0,  
> "max" : 0.0,  
> "avg" : 0.0  
> }  
> }  
> },  
> "docs" : {  
> "count" : 905000,  
> "deleted" : 0  
> },  
> "store" : {  
> "size\_in\_bytes" : 100841187,  
> "throttle\_time\_in\_millis" : 0  
> },  
> "fielddata" : {  
> "memory\_size\_in\_bytes" : 0,  
> "evictions" : 0  
> },  
> "query\_cache" : {  
> "memory\_size\_in\_bytes" : 0,  
> "total\_count" : 0,  
> "hit\_count" : 0,  
> "miss\_count" : 0,  
> "cache\_size" : 0,  
> "cache\_count" : 0,  
> "evictions" : 0  
> },  
> "completion" : {  
> "size\_in\_bytes" : 0  
> },  
> "segments" : {  
> "count" : 55,  
> "memory\_in\_bytes" : 1652155,  
> "terms\_memory\_in\_bytes" : 1430943,  
> "stored\_fields\_memory\_in\_bytes" : 53352,  
> "term\_vectors\_memory\_in\_bytes" : 0,  
> "norms\_memory\_in\_bytes" : 154880,  
> "points\_memory\_in\_bytes" : 0,  
> "doc\_values\_memory\_in\_bytes" : 12980,  
> "index\_writer\_memory\_in\_bytes" : 0,  
> "version\_map\_memory\_in\_bytes" : 0,  
> "fixed\_bit\_set\_memory\_in\_bytes" : 0,  
> "max\_unsafe\_auto\_id\_timestamp" : -1,  
> "file\_sizes" : { }  
> }  
> },

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 30, 2017, 5:32am UTC](https://discuss.elastic.co/t/max-indexing-rate-per-single-elasticsearch-server/91340/4 "2017-06-30T05:32:10Z")

</div>

In benchmarks I have done I am often able to saturate a node with around the same number of connections you are using, so I suspect it may be something with how you run the benchmark.

Are you allowing Elasticsearch to assign a document id or are you updating the same documents over and over? If you are in effect updating, how does throughput differ if you do not specify document ids in the bulk requests?

In order to get a baseline, I would recommend running a few of the default benchmarks available with [rally](https://github.com/elastic/rally).

---

<div class="post-metadata">

**Author:** ![jakesjohn](https://avatars.discourse-cdn.com/v4/letter/j/90ced4/32.png) [@jakesjohn](https://discuss.elastic.co/u/jakesjohn)\
**Post date:** [June 30, 2017, 8:12pm UTC](https://discuss.elastic.co/t/max-indexing-rate-per-single-elasticsearch-server/91340/5 "2017-06-30T20:12:58Z")

</div>

@Christian_Dahlqvist Thanks for the reply. I am allowing elasticsearch to automatically assign ids. I am using mostly default settings. What custom settings do you use? It would be helpful

Can you also help me in finding the interesting parameters to be tuned for each of the following system resources

1. Node CPU is underutilized
2. Node disk io is underutilized
3. Node memory is underutilized

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 1, 2017, 6:27am UTC](https://discuss.elastic.co/t/max-indexing-rate-per-single-elasticsearch-server/91340/6 "2017-07-01T06:27:38Z")

</div>

I would recommend that you run Rally to get a comparison. A good track might be the default `logging` track. I have used Rally to saturate nodes and we know how it works and can therefore compare your results to benchmarks that we have run.

Is Elasticsearch installed on a bare-metal server in your environment or is it a VM?

---

<div class="post-metadata">

**Author:** ![jakesjohn](https://avatars.discourse-cdn.com/v4/letter/j/90ced4/32.png) [@jakesjohn](https://discuss.elastic.co/u/jakesjohn)\
**Post date:** [July 1, 2017, 8:07am UTC](https://discuss.elastic.co/t/max-indexing-rate-per-single-elasticsearch-server/91340/7 "2017-07-01T08:07:20Z")

</div>

Currently, It is a bare metal installation. Any particular reason? I will try rally and update. Thanks for the suggestion

---

<div class="post-metadata">

**Author:** ![jakesjohn](https://avatars.discourse-cdn.com/v4/letter/j/90ced4/32.png) [@jakesjohn](https://discuss.elastic.co/u/jakesjohn)\
**Post date:** [July 1, 2017, 6:16pm UTC](https://discuss.elastic.co/t/max-indexing-rate-per-single-elasticsearch-server/91340/8 "2017-07-01T18:16:42Z")

</div>

@Christian_Dahlqvist Rally is showing throughput of 120k average for logging track. Where can I get the rally benchmark configurations(indices, shards etc) and how data is bulk sent to ES from rally benchmark? I will try to replicate the same.

One other thing I noticed was CPU usage. Rally results showed 500 % while iostat showed 50% of idle time.

```
| Lap | Metric | Operation | Value | Unit |
|------:|--------------------------------:|-------------:|----------:|-------:|
| All | Indexing time | | 389.108 | min |
| All | Merge time | | 121.269 | min |
| All | Refresh time | | 15.4617 | min |
| All | Flush time | | 7.4838 | min |
| All | Merge throttle time | | 49.7276 | min |
| All | Median CPU usage | | 500.4 | % |
| All | Total Young Gen GC | | 233.184 | s |
| All | Total Old Gen GC | | 14.96 | s |
| All | Index size | | 19.2082 | GB |
| All | Totally written | | 182.407 | GB |
| All | Heap used for segments | | 71.0596 | MB |
| All | Heap used for doc values | | 0.134235 | MB |
| All | Heap used for terms | | 58.6108 | MB |
| All | Heap used for norms | | 0.0319214 | MB |
| All | Heap used for points | | 4.62425 | MB |
| All | Heap used for stored fields | | 7.65836 | MB |
| All | Segment count | | 523 | |
| All | Min Throughput | index-append | 111915 | docs/s |
| All | Median Throughput | index-append | 116063 | docs/s |
| All | Max Throughput | index-append | 125471 | docs/s |
| All | 50th percentile latency | index-append | 286.41 | ms |
| All | 90th percentile latency | index-append | 568.223 | ms |
| All | 99th percentile latency | index-append | 1352.39 | ms |
| All | 99.9th percentile latency | index-append | 2526.76 | ms |
| All | 99.99th percentile latency | index-append | 3299.07 | ms |
| All | 100th percentile latency | index-append | 3329.02 | ms |
| All | 50th percentile service time | index-append | 286.41 | ms |
| All | 90th percentile service time | index-append | 568.223 | ms |
| All | 99th percentile service time | index-append | 1352.39 | ms |
| All | 99.9th percentile service time | index-append | 2526.76 | ms |
| All | 99.99th percentile service time | index-append | 3299.07 | ms |
| All | 100th percentile service time | index-append | 3329.02 | ms |

```

Thanks for your help.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 1, 2017, 6:36pm UTC](https://discuss.elastic.co/t/max-indexing-rate-per-single-elasticsearch-server/91340/9 "2017-07-01T18:36:40Z")

</div>

If you have a very powerful server you may need to tweak the default settings in order to fully saturate the node, e.g. by increasing the level of concurrency. The operations and settings are defined in the [rally-tracks repository](https://github.com/elastic/rally-tracks).

---

<div class="post-metadata">

**Author:** ![jakesjohn](https://avatars.discourse-cdn.com/v4/letter/j/90ced4/32.png) [@jakesjohn](https://discuss.elastic.co/u/jakesjohn)\
**Post date:** [July 1, 2017, 8:35pm UTC](https://discuss.elastic.co/t/max-indexing-rate-per-single-elasticsearch-server/91340/10 "2017-07-01T20:35:17Z")

</div>

@Christian_Dahlqvist Thanks for your help. I am constantly seeing that CPUs are less than 50 % utilized. Which settings should i look for in order to change the **concurrency**? Do you mean threadpool settings?

I see that default threadpool type for bulk and index is fixed and min/max is very high(32). I increased queue size of index and bulk to 100000. Any other settings that can effectively use more CPUs for faster ingest?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 2, 2017, 6:05am UTC](https://discuss.elastic.co/t/max-indexing-rate-per-single-elasticsearch-server/91340/11 "2017-07-02T06:05:32Z")

</div>

I was referring to the number of concurrent connections that rally used, which is specified in the track. I see no evidence that you need to change anything in the Elasticsearch config at this point.

Go to the host where Rally is running and go to `~/.rally/benchmarks/tracks/default/logging/challenges` and edit the `default.json` file. Here you can increase [the number of clients Rally will use for indexing](https://github.com/elastic/rally-tracks/blob/master/logging/challenges/default.json#L12). I would recommend to start by doubling it and see what difference that makes.

When comparing the indexing rates achieved in this Rally benchmark to what you get with your data, note that the difference in documents size will have a significant impact. This set of tests should however at least show you how to go about saturating your server and it should not be too hard to create a custom track for Rally that uses your data.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 30, 2017, 6:06am UTC](https://discuss.elastic.co/t/max-indexing-rate-per-single-elasticsearch-server/91340/12 "2017-07-30T06:06:09Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
