# Recommended Elasticsearch Node requirements for monitoring 2000 nodes

**URL:** <https://discuss.elastic.co/t/recommended-elasticsearch-node-requirements-for-monitoring-2000-nodes/224057>\
**Category:** Metrics\
**Created:** [March 18, 2020, 9:25am UTC](https://discuss.elastic.co/t/recommended-elasticsearch-node-requirements-for-monitoring-2000-nodes/224057 "2020-03-18T09:25:32Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Bingu\_Shim](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bingu_shim/32/57949_2.png) [@Bingu\_Shim](https://discuss.elastic.co/u/Bingu_Shim)\
**Post date:** [March 18, 2020, 9:25am UTC](https://discuss.elastic.co/t/recommended-elasticsearch-node-requirements-for-monitoring-2000-nodes/224057/1 "2020-03-18T09:25:32Z")

</div>

Hello, I'm starting to research Metricbeat + Kibana Inventory to monitor system resources.  
The total number machine that we are planning to monitor is about 2,000.

Currently I was able to deploy metricbeat to 79 hosts and explore the metric on Kibana Inventory whithout any problem.

Next step will be expanding it to more nodes, and I want to prepare enough Elasticsearch Node for it.

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/8/e/8ed1011325bc2b2ed4698684f61eed491ea8ffc9.png)

Currently I've tested with on this environment.

- 5 Data Nodes (16G/16Core) + 3 Master Node(16G/16Core)

metricbeat.yml

```auto
logging.level: info

output.elasticsearch:
  ...
  worker: 2
  bulk_max_size : 1024

queue:
  mem :
    events: 4096
    flush.min_events: 2048
max_procs : 1

setup.ilm:
  enabled : auto

setup.dashboards.enabled: false
setup.template.settings:
  index:
    codec: best_compression
    number_of_shards: 5
    number_of_replicas: 1
    refresh_interval: 10s

setup.kibana:
  ...

#------ metric beat specific configuration

metricbeat.max_start_delay: 10s
metricbeat.modules:
  - module: system
    metricsets:
      - cpu # CPU usage
      - load # CPU load averages
      - memory # Memory usage
      - network # Network IO
      - uptime # System Uptime
      - fsstat # File system summary metrics
      - diskio # Disk IO
      - process_summary # process 요약

    enabled: true
    period: 10s
    processes: ['.*']

    # Configure the metric types that are included by these metricsets.
    cpu.metrics: ["percentages", "normalized_percentages"] # The other available options are normalized_percentages and ticks.
    core.metrics: ["percentages"] # The other available option is ticks.
processors:
- add_host_metadata:

```

---

<div class="post-metadata">

**Author:** ![Kerry](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kerry/32/40330_2.png) [@Kerry](https://discuss.elastic.co/u/Kerry)\
**Post date:** [March 19, 2020, 3:56pm UTC](https://discuss.elastic.co/t/recommended-elasticsearch-node-requirements-for-monitoring-2000-nodes/224057/2 "2020-03-19T15:56:38Z")

</div>

Hi @Bingu_Shim, I've reached out internally regarding your question, and hope to have some information for you soon.

---

<div class="post-metadata">

**Author:** ![Kerry](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kerry/32/40330_2.png) [@Kerry](https://discuss.elastic.co/u/Kerry)\
**Post date:** [March 20, 2020, 10:50am UTC](https://discuss.elastic.co/t/recommended-elasticsearch-node-requirements-for-monitoring-2000-nodes/224057/3 "2020-03-20T10:50:56Z")

</div>

Hi, after speaking with a colleague, unfortunately it's hard for us to answer this as it's often a case of "it depends".

Our reccomendation would be to experiment with as much (realistic) data as possible, and see if there's a bottleneck. Serving 2,000 nodes in the UI shouldn't be a problem, but your cluster might not be able to handle the write load. If this is the case, you may need to add more shards to optimise for a write heavy environment.

---

<div class="post-metadata">

**Author:** ![Bingu\_Shim](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bingu_shim/32/57949_2.png) [@Bingu\_Shim](https://discuss.elastic.co/u/Bingu_Shim)\
**Post date:** [March 21, 2020, 2:56am UTC](https://discuss.elastic.co/t/recommended-elasticsearch-node-requirements-for-monitoring-2000-nodes/224057/4 "2020-03-21T02:56:15Z")

</div>

Hello @Kerry

Thank you for your response.  
With my configurations above, the number of documents written per minutes are as follow.

| event.dataset | docs per minutes |
| --- | --- |
| system.diskio | 6 |
| system.fsstat | 6 |
| system.load | 6 |
| system.uptime | 6 |
| system.cpu | 6 |
| system.memory | 6 |
| system.process.summary | 6 |
| system.network | 96 |
| total | 138 |

So, I can get required write performance as follow.

- 2.3 tps per 1 node
- 4,600 tps for 2,000 nodes.

This write load should be handled with our elasticsearch cluster. (I've done over 30K write performance test on our environment. The use cases was different though)

What I want to know is Metric UI sides latency.  
Since I got UI latency problem with using Elastic APM ([THIS ISSUE](https://discuss.elastic.co/t/kibana-apm-app-pages-are-too-slow/220570), [THIS ISSUE](https://discuss.elastic.co/t/apm-ui-kibana-internal-server-error/223601/21)), and found out the team just started [architectural improvement](https://github.com/elastic/apm-server/issues/3485) to solve latency problem.

So we just want to make it sure that scalability of Metric UI, before going further.

As you mentioned as follow, there won't be scalability problem with UI side.

> Serving 2,000 nodes in the UI shouldn't be a problem

We will trying to apply Metricbeat more machines.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 18, 2020, 2:56am UTC](https://discuss.elastic.co/t/recommended-elasticsearch-node-requirements-for-monitoring-2000-nodes/224057/5 "2020-04-18T02:56:33Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
