# Spikes in Indexing Rate - Methods to Increase Indexing Rate?

**URL:** <https://discuss.elastic.co/t/spikes-in-indexing-rate-methods-to-increase-indexing-rate/68740>\
**Category:** Elasticsearch\
**Created:** [December 12, 2016, 4:58pm UTC](https://discuss.elastic.co/t/spikes-in-indexing-rate-methods-to-increase-indexing-rate/68740 "2016-12-12T16:58:09Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![seth.yes](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/seth.yes/32/11788_2.png) [@seth.yes](https://discuss.elastic.co/u/seth.yes)\
**Post date:** [December 12, 2016, 4:58pm UTC](https://discuss.elastic.co/t/spikes-in-indexing-rate-methods-to-increase-indexing-rate/68740/1 "2016-12-12T16:58:09Z")

</div>

I have six ES nodes (ES 2.4.1); two client nodes, two master nodes, and two data nodes.

I have setup Logstash (2.2.4) to push data to the two client nodes. I have also set up Logstash to push the data to both the two client nodes and the two master nodes. Either way, I'm seeing a dramatic spike in Indexing rates, every minute the indexing rate will jump from 0 events/sec to 5000 events/sec, leaving the average around roughly 2500 events/sec.

The indexing rate is shown via Marvel as this;

 ![](https://us1.discourse-cdn.com/elastic/original/2X/2/298e0d4c41b2e64ac5105e02cf709542487fecb9.png)  
The arrow is where I modified the config from LS pushing to Master/Client nodes to LS just pushing to Client nodes.

This appears to be some sort of batch processing. I have a couple questions associated with this setup;

1. Is it best practice to push Logstash output (from 4 LS instances) to the two client nodes, the two master nodes, or all four? I've read that you want your data nodes to not participate in directly receiving LS data.

2. Is there a way to smooth out this indexing rate? My primary goal here is to increase throughput, and it appears that there are downtimes where I could be processing additional data.  
-- This is an issue because the log shippers (Filebeat 5.1.1) are sending data at a faster rate than I'm currently ingesting, I know this because everything is timestamped -- there's a two or three day delay between the time the log was created and the time it's ingested into ELK.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [December 13, 2016, 5:19am UTC](https://discuss.elastic.co/t/spikes-in-indexing-rate-methods-to-increase-indexing-rate/68740/2 "2016-12-13T05:19:04Z")

</div>

Send requests to the data or client nodes, not to the data nodes.

> [@seth.yes](#):
>
> there's a two or three day delay between the time the log was created and the time it's ingested into ELK

That seems irregular.  
What's the load on ES like?

> [@seth.yes](#):
>
> two master nodes

That's not ideal as there is no majority for a quorum.

---

<div class="post-metadata">

**Author:** ![seth.yes](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/seth.yes/32/11788_2.png) [@seth.yes](https://discuss.elastic.co/u/seth.yes)\
**Post date:** [December 13, 2016, 3:20pm UTC](https://discuss.elastic.co/t/spikes-in-indexing-rate-methods-to-increase-indexing-rate/68740/3 "2016-12-13T15:20:19Z")

</div>

> [@warkolm](#):
>
> Send requests to the data or client nodes, not to the data nodes.

Do you mean 'send requests to the master or client nodes'? I'm curious if there's a logical preference between the two..  
I modified my config from sending LS data to both the two masters and the two client nodes to just sending the data to the two client nodes -- really saw no difference in the indexing rate.

Also, I've read that having too many client nodes can be detrimental to performance as any data allocation process by one client node must be confirmed with the other before data is committed. I'm considering removing one of the client nodes and just directing all LS data to the one client node.

> [@warkolm](#):
>
> What's the load on ES like?

The load doesn't seem too large, the master & data nodes all have 24 CPUs, the VMs each have 8.

 ![](https://us1.discourse-cdn.com/elastic/original/2X/4/484ee54bb2eb64d8fc852e31096702814d1aa9f2.png)

@warkolm while I have your attention, does the numbers above look okay? Any tweaking you'd do in order to get more performance out of my cluster?

> [@warkolm](#):
>
> [Two Master nodes] is not ideal as there is no majority for a quorum.

I agree and you've told me that before but I had no additional physical machines to allocate. At the beginning of the year I'm getting new hardware and will be utilizing three master nodes as per Elastic' recommendation.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [December 13, 2016, 8:17pm UTC](https://discuss.elastic.co/t/spikes-in-indexing-rate-methods-to-increase-indexing-rate/68740/4 "2016-12-13T20:17:54Z")

</div>

> [@seth.yes](#):
>
> Do you mean 'send requests to the master or client nodes'? I'm curious if there's a logical preference between the two..

Sorry, I meant "send requests to the data and client nodes, not the master nodes".

> [@seth.yes](#):
>
> Also, I've read that having too many client nodes can be detrimental to performance as any data allocation process by one client node must be confirmed with the other before data is committed.

Where did you read this?

> [@seth.yes](#):
>
> The load doesn't seem too large, the master & data nodes all have 24 CPUs, the VMs each have 8.

Yeah that looks ok.

---

<div class="post-metadata">

**Author:** ![seth.yes](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/seth.yes/32/11788_2.png) [@seth.yes](https://discuss.elastic.co/u/seth.yes)\
**Post date:** [December 13, 2016, 10:47pm UTC](https://discuss.elastic.co/t/spikes-in-indexing-rate-methods-to-increase-indexing-rate/68740/5 "2016-12-13T22:47:43Z")

</div>

> [@warkolm](#):
>
> Sorry, I meant "send requests to the data and client nodes, not the master nodes".

Thanks, I'll send LS data to the client & data nodes to see if that increases throughput -- it held pretty steady when I moved to just client nodes receiving the LS data, as compared to both Client & Master nodes.

> [@warkolm](#):
>
> Where did you read this?

> **[Node | Elasticsearch Guide \[2.4\] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/2.4/modules-node.html#client-node)**

 ![](https://us1.discourse-cdn.com/elastic/original/2X/c/cab148a112527ae80bb7e106aa7c01e105adb453.png)

Is there any reason for the spiking in the throughput as seen in my first post? Is it batching up data or something? I ask because this is still occurring and I'm unsure whether it's the desired behavior or not..

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [December 14, 2016, 2:16am UTC](https://discuss.elastic.co/t/spikes-in-indexing-rate-methods-to-increase-indexing-rate/68740/6 "2016-12-14T02:16:48Z")

</div>

Ahh ok, so that's not going to really impact data allocation, more cluster state changes.

Regarding the sawtooth, it could be a data sampling problem, it could be LS batching things up. It might be worth moving this to the X-Pack category to see what the team there thinks.

---

<div class="post-metadata">

**Author:** ![seth.yes](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/seth.yes/32/11788_2.png) [@seth.yes](https://discuss.elastic.co/u/seth.yes)\
**Post date:** [December 15, 2016, 6:35pm UTC](https://discuss.elastic.co/t/spikes-in-indexing-rate-methods-to-increase-indexing-rate/68740/7 "2016-12-15T18:35:16Z")

</div>

I'll do that, thanks for your help!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 12, 2017, 6:35pm UTC](https://discuss.elastic.co/t/spikes-in-indexing-rate-methods-to-increase-indexing-rate/68740/8 "2017-01-12T18:35:52Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
