# How does indexing performance vary over increase in number of nodes?

**URL:** <https://discuss.elastic.co/t/how-does-indexing-performance-vary-over-increase-in-number-of-nodes/56246>\
**Category:** Elasticsearch\
**Created:** [July 24, 2016, 4:02pm UTC](https://discuss.elastic.co/t/how-does-indexing-performance-vary-over-increase-in-number-of-nodes/56246 "2016-07-24T16:02:35Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![sw.jung](https://avatars.discourse-cdn.com/v4/letter/s/67e7ee/32.png) [@sw.jung](https://discuss.elastic.co/u/sw.jung)\
**Post date:** [July 24, 2016, 4:02pm UTC](https://discuss.elastic.co/t/how-does-indexing-performance-vary-over-increase-in-number-of-nodes/56246/1 "2016-07-24T16:02:35Z")

</div>

How does indexing performance vary over increase in number of nodes?

I plan to add more nodes for performance issue. I have read from a guide that it is effective for searching, but I wasn't sure with indexing. Does it become faster?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 24, 2016, 4:36pm UTC](https://discuss.elastic.co/t/how-does-indexing-performance-vary-over-increase-in-number-of-nodes/56246/2 "2016-07-24T16:36:16Z")

</div>

It depends on many factors.  
Like if you have only one shard, increasing the number of nodes won't change indexing throughput.

May be you could describe a bit what is your current issue?  
And what is your platform? In details.

---

<div class="post-metadata">

**Author:** ![sw.jung](https://avatars.discourse-cdn.com/v4/letter/s/67e7ee/32.png) [@sw.jung](https://discuss.elastic.co/u/sw.jung)\
**Post date:** [July 24, 2016, 11:03pm UTC](https://discuss.elastic.co/t/how-does-indexing-performance-vary-over-increase-in-number-of-nodes/56246/3 "2016-07-24T23:03:39Z")

</div>

Thanks for your reply 🙂

I set up 6 or more data node instances with these settings.

number\_of\_shard=18  
Number\_of\_replica=0  
indices.throttle.max\_bytes\_per\_second=512mb

How about this?

In my case shown, indexing performance are not increase.

---

<div class="post-metadata">

**Author:** ![DiscussBuster](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/discussbuster/32/11042_2.png) [@DiscussBuster](https://discuss.elastic.co/u/DiscussBuster)\
**Post date:** [July 25, 2016, 12:06am UTC](https://discuss.elastic.co/t/how-does-indexing-performance-vary-over-increase-in-number-of-nodes/56246/4 "2016-07-25T00:06:03Z")

</div>

> I set up 6 or more data node instances with these settings.

6 more than what?

I think these details will be instructive to someone trying to understand your system:

- version of ES
- how many client/data nodes before and after
- how many CPU cores per node
- how many indices
- how many shards per index
- what kind of storage on data nodes
- how much RAM per node, how much heap
- do you have paging/swapping disabled
- what is your indexing client, and how is it configured (posting all to one node, or all nodes, or ?)

> In my case shown, indexing performance are not increase.

And what is that performance? How are you measuring it? Are you sure the data pipeline ahead of it isn't bottlenecked?

---

<div class="post-metadata">

**Author:** ![sw.jung](https://avatars.discourse-cdn.com/v4/letter/s/67e7ee/32.png) [@sw.jung](https://discuss.elastic.co/u/sw.jung)\
**Post date:** [July 25, 2016, 12:59am UTC](https://discuss.elastic.co/t/how-does-indexing-performance-vary-over-increase-in-number-of-nodes/56246/5 "2016-07-25T00:59:55Z")

</div>

Sorry, I did not explain in detail ☹

- version of ES  
: 2.3.4
- how many client/data nodes before and after  
: start with 1 client node, 1 data node. and sequentially increase number of data node to 6.
- how many CPU cores per node  
: 32 core with hyper threading
- how many indices  
: for 1 index.
- how many shards per index  
: 18 shards
- what kind of storage on data nodes  
: SAS HDD
- how much RAM per node, how much heap  
: 64gb RAM per nodes and 32gb heap size.
- do you have paging/swapping disabled  
: disabled
- what is your indexing client, and how is it configured (posting all to one node, or all nodes, or ?)  
: single logstash -\> 1 client node on same host -\> other data nodes on other hosts.

Thanks. 😃

---

<div class="post-metadata">

**Author:** ![DiscussBuster](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/discussbuster/32/11042_2.png) [@DiscussBuster](https://discuss.elastic.co/u/DiscussBuster)\
**Post date:** [July 25, 2016, 1:07am UTC](https://discuss.elastic.co/t/how-does-indexing-performance-vary-over-increase-in-number-of-nodes/56246/6 "2016-07-25T01:07:02Z")

</div>

> [@sw.jung](#):
>
> : SAS HDD

This could be a problem...

> [@sw.jung](#):
>
> : 64gb RAM per nodes and 32gb heap size.

This isn't _the_ problem, but you should probably reduce that to, say, 30GB, to make sure you are getting compressed ordinary object pointers.

> [@sw.jung](#):
>
> : single logstash -\> 1 client node on same host -\> other data nodes on other hosts.

Are you getting "backpressure" in Logstash? That is, are you getting bulk rejections from Elasticsearch because it's at its indexing capacity?

How are you invoking Logstash? (what are you providing for `-w`?)  
What does your Elasticsearch output configuration look like?

Again:

> And what is that performance? How are you measuring it? Are you sure the data pipeline ahead of it isn't bottlenecked?

---

<div class="post-metadata">

**Author:** ![sw.jung](https://avatars.discourse-cdn.com/v4/letter/s/67e7ee/32.png) [@sw.jung](https://discuss.elastic.co/u/sw.jung)\
**Post date:** [July 25, 2016, 5:32am UTC](https://discuss.elastic.co/t/how-does-indexing-performance-vary-over-increase-in-number-of-nodes/56246/7 "2016-07-25T05:32:13Z")

</div>

> [@DiscussBuster](#):
>
> Are you getting "backpressure" in Logstash? That is, are you getting bulk rejections from Elasticsearch because it's at its indexing capacity?
> 
> How are you invoking Logstash? (what are you providing for -w?)What does your Elasticsearch output configuration look like?

No, I have already set logstash batch size, workers, and elasticsearch output flush size enough .

> [@DiscussBuster](#):
>
> And what is that performance? How are you measuring it? Are you sure the data pipeline ahead of it isn't bottlenecked?

It is indexing speed was not fast enough than I expected. It does not seem to changed by number of nodes. Elasticsearch seems to be a bottleneck.

Thanks.

---

<div class="post-metadata">

**Author:** ![german23](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/german23/32/11052_2.png) [@german23](https://discuss.elastic.co/u/german23)\
**Post date:** [July 25, 2016, 5:46am UTC](https://discuss.elastic.co/t/how-does-indexing-performance-vary-over-increase-in-number-of-nodes/56246/8 "2016-07-25T05:46:20Z")

</div>

My suggestion:

Install plugins like  
bigdesk:

> **[abrahamduran/bigdesk](https://github.com/abrahamduran/bigdesk)**
>
> bigdesk - Live charts and statistics for Elasticsearch cluster.

or  
kopf

> **[lmenezes/elasticsearch-kopf](https://github.com/lmenezes/elasticsearch-kopf)**
>
> elasticsearch-kopf - web admin interface for elasticsearch

if u have not already.

Look in kopf in the first place if the index get spread evenly across all of the nodes - if not try to adjust allocation settings.

In Bigdesk u should look for the bulk thread graphs. See if the bulk-queue is getting used and if the number of bulk-threads match your cpu-cores(or even more) - u can also check there if u get heavy GC-Collection pauses.

The last thing i would recommend is enable the DEBUG-Logging in ES and check the log if merging is falling behind cause of the HDDs.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 25, 2016, 6:07am UTC](https://discuss.elastic.co/t/how-does-indexing-performance-vary-over-increase-in-number-of-nodes/56246/9 "2016-07-25T06:07:13Z")

</div>

> [@sw.jung](#):
>
> It is indexing speed was not fast enough than I expected. It does not seem to changed by number of nodes. Elasticsearch seems to be a bottleneck.

How have you determined that Elasticsearch is the bottleneck? Have you replaced the Elasticsearch output in Logstash with e.g. a stdout output with the dots codec, to verify that Logstash with your current configuration is able to achieve higher throughput than Elasticsearch can handle?

---

<div class="post-metadata">

**Author:** ![DiscussBuster](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/discussbuster/32/11042_2.png) [@DiscussBuster](https://discuss.elastic.co/u/DiscussBuster)\
**Post date:** [July 25, 2016, 12:28pm UTC](https://discuss.elastic.co/t/how-does-indexing-performance-vary-over-increase-in-number-of-nodes/56246/10 "2016-07-25T12:28:50Z")

</div>

> [@sw.jung](#):
>
> It is indexing speed was not fast enough than I expected. It does not seem to changed by number of nodes. Elasticsearch seems to be a bottleneck.

Yes, you made this hypothesis clear in your introduction. The questions I've asked are an effort to discover how you arrived at that conclusion, further test your hypothesis, and explore other (frankly, more likely) reasons for the behavior you are witnessing.

> [@sw.jung](#):
>
> It is indexing speed was not fast enough than I expected. It does not seem to changed by number of nodes. Elasticsearch seems to be a bottleneck.

How do you use _that_ evidence to come to _that_ conclusion? It's precisely the reverse of how one normally goes about finding a bottleneck. You've got a process with several serial components, A -\> B -\> C, and the process fully processes N events/second. You decide you want the process to complete 4\*N events/second, so you multiply the number of C resources by 4. In response, you still get N events/second processed. The normal suspicion is then that A or B are only capable of N events/second, not that C is broken. Instead, you are guessing that C is broken or misconfigured. What's being requested is evidence for that supposition, and, in parallel, additional information about part B, a thus far (from just the evidence volunteered) more likely bottleneck.

Let's try again.

What is your indexing speed? How are you measuring it? Are you getting backpressure (bulk rejections) from Elasticsearch? What are you using for `-w` for Logstash? What does your output configuration look like?

As mentioned above, try using a "local" output (file or stdout) to test your Logstash throughput. Also, consider using the metrics filter to help measure that. [Try these tips](https://www.elastic.co/blog/logstash-configuration-tuning) to optimize and test your Logstash configuration.

If you already have additional information that leads you to the conclusion that Elasticsearch is actually the bottleneck, please share it.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:33pm UTC](https://discuss.elastic.co/t/how-does-indexing-performance-vary-over-increase-in-number-of-nodes/56246/11 "2017-07-05T22:33:10Z")

</div>


