# Optimizing Elastic-Agent Performance

**URL:** <https://discuss.elastic.co/t/optimizing-elastic-agent-performance/359322>\
**Category:** Elastic Search\
**Created:** [May 12, 2024, 1:16pm UTC](https://discuss.elastic.co/t/optimizing-elastic-agent-performance/359322 "2024-05-12T13:16:54Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![wangsubo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wangsubo/32/133865_2.png) [@wangsubo](https://discuss.elastic.co/u/wangsubo)\
**Post date:** [May 12, 2024, 1:16pm UTC](https://discuss.elastic.co/t/optimizing-elastic-agent-performance/359322/1 "2024-05-12T13:16:54Z")

</div>

I am currently evaluating the use of ELK to collect Checkpoint data. On a VM, I have a QRadar Event Collector and an Elastic-Agent, both configured with 8 cores, 8GB RAM, and HDD.

Checkpoint logs are being sent at a rate of 5000 EPS, split between the QRadar Event Collector and the Elastic-Agent. I have observed that the QRadar Event Collector consistently receives about 5% more events than the Elastic-Agent during the same time frame.

Here is my current configuration for the elastic-agent.yml:

bulk\_max\_size: 10000  
worker: 6  
queue.mem.events: 12800  
queue.mem.flush.min\_events: 10000  
queue.mem.flush.timeout: 50ms  
compression\_level: 1  
connection\_idle\_timeout: 15s

How can I optimize the settings so that the event volume discrepancy between Elastic-Agent and QRadar Event Collector is within 1%？

Thanks

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [May 12, 2024, 1:30pm UTC](https://discuss.elastic.co/t/optimizing-elastic-agent-performance/359322/2 "2024-05-12T13:30:54Z")

</div>

Hello,

You need to also provide context about your Elasticsearch cluster, like how many nodes do you have, what are their specs, also, what is the output configuration of your Elastic Agent? How many nodes do you have configured in the output?

Also, QRadar and Elasticsearch are different tools that work in different ways, not sure if it makes any sense comparing the two of them here.

---

<div class="post-metadata">

**Author:** ![wangsubo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wangsubo/32/133865_2.png) [@wangsubo](https://discuss.elastic.co/u/wangsubo)\
**Post date:** [May 12, 2024, 1:50pm UTC](https://discuss.elastic.co/t/optimizing-elastic-agent-performance/359322/3 "2024-05-12T13:50:30Z")

</div>

Initially, the Elastic-Agent was configured with only one output node, but I found the EPS to be too low, which I suspected might be related to Elasticsearch performance.

Later, I increased this to three nodes, and there was an improvement in the event capture rate, but there was still about a 5% discrepancy in event volume compared to the QRadar Event Collector.

The Elasticsearch cluster consists of three nodes, each with a 6-core CPU, 32 GB RAM, and a 1 TB HDD.

The reason for comparing QRadar and Elasticsearch is to evaluate the possibility of replacing QRadar with Elasticsearch. Without comparing event volumes, it's difficult to determine whether a replacement is feasible.

Thanks

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [May 12, 2024, 2:20pm UTC](https://discuss.elastic.co/t/optimizing-elastic-agent-performance/359322/4 "2024-05-12T14:20:40Z")

</div>

> [@wangsubo](#):
>
> The Elasticsearch cluster consists of three nodes, each with a 6-core CPU, 32 GB RAM, and a 1 TB HDD.

HDD is pretty bad for performance in Elasticsearch, it can impact in the indexing rate, the recommendation is to use SSD.

The performance of the Elastic Agent is also influenced by the performance of the destination cluster, if your cluster cannot write to the disk fast enough, it will tell the Elastic Agent to backoff a little.

Have you checked this [documentation](https://www.elastic.co/guide/en/elasticsearch/reference/current/tune-for-indexing-speed.html)? There are probably some things that you can try to improve it, even using HDD.

If you haven't read it yet, I strongly recommend that you do.

Per default the Elastic Agent will use 1 primary shard and 1 replica and a refresh interval of 1s, you may need to create a custom template to change the number of primary shards to 3 (the number of nodes you have) and increase the refresh interval for something like 10s, 15s.

I do not know QRadar, does it index every field? And does it have some processing to parse the data? Elasticsearch will do both things.

> [@wangsubo](#):
>
> I increased this to three nodes, and there was an improvement in the event capture rate, but there was still about a 5% discrepancy in event volume compared to the QRadar Event Collector.

How did you arrive to this metric? Can you provide more context and some evidences?

---

<div class="post-metadata">

**Author:** ![wangsubo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wangsubo/32/133865_2.png) [@wangsubo](https://discuss.elastic.co/u/wangsubo)\
**Post date:** [May 12, 2024, 3:07pm UTC](https://discuss.elastic.co/t/optimizing-elastic-agent-performance/359322/5 "2024-05-12T15:07:05Z")

</div>

I haven't read the document you provided, but I have attempted some benchmark tests. In the CheckPoint logs, there are Rule Names.

I sent the top ten Rule Names by "event count" in CheckPoint (e.g., Rule Name = allow trust to untrust), starting with the one with the fewest events, to both QRadar Event Collector and Elastic-Agent.

Using QRadar Search and Kibana data view to compare the event counts of the same time range from CheckPoint, I found that for Rule Names with low EPS (events per second), the event counts between QRadar Event Collector and Elastic-Agent are exactly the same. However, when the Rule Name EPS reaches 3000-4000, Elastic-Agent begins to show about a 5% discrepancy compared to QRadar.

QRadar requires regular expressions to parse the data.

I also suspect that the bottleneck might be in the HDD, and I plan to try switching to SSD.

Additionally, under default settings, Elastic Agent uses 1 primary shard and 1 replica with a refresh interval of 1 second. Why should increasing the number of primary shards to 3 involve increasing the refresh interval to about 10 or 15 seconds?

Could you provide me with reference materials on how to increase the refresh interval?

Thanks

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 9, 2024, 3:07pm UTC](https://discuss.elastic.co/t/optimizing-elastic-agent-performance/359322/6 "2024-06-09T15:07:20Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
