# Monitor cluster with elastic agent

**URL:** <https://discuss.elastic.co/t/monitor-cluster-with-elastic-agent/342413>\
**Category:** Elastic Agent\
**Tags:** elastic-stack-monitoring, metricbeat\
**Created:** [September 6, 2023, 9:23am UTC](https://discuss.elastic.co/t/monitor-cluster-with-elastic-agent/342413 "2023-09-06T09:23:27Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![lduvnjak](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lduvnjak/32/77724_2.png) [@lduvnjak](https://discuss.elastic.co/u/lduvnjak)\
**Post date:** [September 6, 2023, 9:23am UTC](https://discuss.elastic.co/t/monitor-cluster-with-elastic-agent/342413/1 "2023-09-06T09:23:27Z")

</div>

Hey Everyone,

We're trying to move from the legacy exporters over to Elastic Agent. Our data pipeline is `Elastic Agent > Logstash > Kafka > Elastic`. I have a couple of questions and would appreciate any and all knowledge anyone is wishing to share.

- Will monitoring work if the underlying data stream names are changed
- If it is a big cluster with dedicated master, how, warm, cold and coordianting nodes should I use the `scope` option for monitoring
- If the scope option is selected, how do you make the agent collecting the metrics HA (highly available)

To briefly explain the first question, we are using [Kafka's Elasticsearch connectors](https://www.confluent.io/blog/elastic-data-streams-support-with-confluents-elasticsearch-connector/#step-2) which require a `type` and `dataset name` to be specified.  
As such the datastream name in ES would be `metrics-something-kibana.stack_monitoring.stats-prod` for example.

Regarding the second and third - What I concluded from the documentation is that you should use the `scope` option and add a LB url that balances across nodes which are not master-eligible (in our case those would be coordinating-only nodes). But how do you make it so if the elastic-agent collecting the metrics goes offline the metrics still keep coming.  
If you were to have two agents, both collecting the same data from the same cluster, would they duplicate the metric data or does every pipeline have a static way to generate the doc `_id` field?

Sorry for the long post, here's a cookie 🍪 for those of you who made it, and thanks for any help in advance!

Cheers,  
Luka

---

<div class="post-metadata">

**Author:** ![miltonhultgren](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/miltonhultgren/32/100401_2.png) [@miltonhultgren](https://discuss.elastic.co/u/miltonhultgren)\
**Post date:** [September 6, 2023, 11:21am UTC](https://discuss.elastic.co/t/monitor-cluster-with-elastic-agent/342413/2 "2023-09-06T11:21:44Z")

</div>

Hi Luka!

I'll try to answer the parts I can.

For the data stream name, the Stack Monitoring UI looks at the following pattern for data collected with Elastic Agent (or rather, data streams matching the naming conventions): `metrics-elasticsearch.stack_monitoring.DATASET-*` (where DATASET is for example `cluster_stats`.  
As long as your data streams match this pattern the UI should work. From your example, the `something` part would break it. I'm afraid this isn't something you can configure.

The size of your cluster isn't really what dedicates the value of the `scope` config, but rather how you want to deploy your collection. ~~If you use `cluster` you only need to deploy one agent which should target the master node to fetch all the needed data sets. But you might want to avoid this in a large cluster since that will cause more load on your master. So in that case using `node` might be better but then you need to deploy one agent per node (at least this is my understanding).~~  
(see discussion below)

How you achieve HA on the agent likely depends on how you deploy your agent and stack, this is somewhat outside of the realm of the Elastic stack. If it's running as a side car in a Kubernetes pod for example then we would expect Kubernetes to take responsibility for that. In a raw deployment, I don't know. Who watches the watcher so to say. That's why we have alerting rules about monitoring data missing, so that's one option.

Having two agents collecting the same metrics would lead to issues since they have no way to coordinate around duplications.

---

<div class="post-metadata">

**Author:** ![lduvnjak](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lduvnjak/32/77724_2.png) [@lduvnjak](https://discuss.elastic.co/u/lduvnjak)\
**Post date:** [September 6, 2023, 11:42am UTC](https://discuss.elastic.co/t/monitor-cluster-with-elastic-agent/342413/3 "2023-09-06T11:42:15Z")

</div>

Hi @miltonhultgren, thank you for the prompt response!

> For the data stream name, the Stack Monitoring UI looks at the following pattern for data collected with Elastic Agent (or rather, data streams matching the naming conventions): `metrics-elasticsearch.stack_monitoring.DATASET-*` (where DATASET is for example `cluster_stats` .  
> As long as your data streams match this pattern the UI should work. From your example, the `something` part would break it. I'm afraid this isn't something you can configure.

Unfortunately you are correct about the `something` part breaking it. After disabling the legacy collection via:

```auto
PUT _cluster/settings
{
  "persistent": {
    "xpack.monitoring.collection.enabled": false
  }
}

```

The Stack Monitoring data completely stopped updating. Currently I'm monitoring one hot node, one Kibana instance and one Logstash instance. All the monitoring data is being ingested properly into ES (afaik), since the data streams are consantly getting new documents.

Is there any potential on "fixing" (_not litreally since it ain't broke_) the stack monitoring in the future to work like a dashboard? In the sense that it would look for `logs-*/metrics-*`, and then filter further using the `datastream.dataset` field, if that's even possible? Considering all the prebuilt dashboards work in this manner it might be a good idea to streamline that as well.

Guess I'll just jump over Kafka with metrics and only write logs, don't see another solution atm.

> The size of your cluster isn't really what dedicates the value of the `scope` config, but rather how you want to deploy your collection. If you use `cluster` you only need to deploy one agent which should target the master node to fetch all the needed data sets. But you might want to avoid this in a large cluster since that will cause more load on your master. So in that case using `node` might be better but then you need to deploy one agent per node (at least this is my understanding).

That was my original understanding as well, but after reading this from the [docs](https://www.elastic.co/guide/en/elasticsearch/reference/current/configuring-elastic-agent.html#_add_elasticsearch_monitoring_data) I'm a bit confused:

_Elastic Agent will collect most of the metrics from the elected master of the cluster, so you must scale up all your master-eligible nodes to account for this extra load. Do not use this `node` if you have dedicated master nodes._

What I'm getting from this is that regardless of if you choose to monitor the cluster via the master, or a single node, it will still go to the master to collect cluster related information (_this is an educated guess_). Meaning that in my case with 100+ nodes it would probably outright crash it.

> Having two agents collecting the same metrics would lead to issues since they have no way to coordinate around duplications.

Yeah, I figured as much after looking at the ingest pipelines. Thank you nevertheless, this is extremely valuable info!

---

<div class="post-metadata">

**Author:** ![miltonhultgren](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/miltonhultgren/32/100401_2.png) [@miltonhultgren](https://discuss.elastic.co/u/miltonhultgren)\
**Post date:** [September 6, 2023, 11:56am UTC](https://discuss.elastic.co/t/monitor-cluster-with-elastic-agent/342413/4 "2023-09-06T11:56:03Z")

</div>

I agree that it would be a **good** idea to align the Stack Monitoring UI with how dashboards work, that's however not on our roadmap I'm afraid.

One option might be to use an ingest pipeline and the newly added [rerouting processor](https://www.elastic.co/guide/en/elasticsearch/reference/current/reroute-processor.html) to direct the documents to the "right" data stream.

For the `scope` setting, my experience is that there are some limits to cluster size that we can monitor regardless of which mode you use simply because the collecting depends on getting the cluster state which is owned by the master eligible nodes and this is usually the dataset which becomes largest as the cluster grows huge.  
So there isn't that much we can do to work around that beyond scaling up the master nodes, changing the polling frequency of the collection, increasing the timeout settings so the collection waits for the master node to compute and transfer the cluster state. Or, change the topology of the cluster (which usually isn't an option).

@DavidTurner Is my understanding correct here?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [September 6, 2023, 12:19pm UTC](https://discuss.elastic.co/t/monitor-cluster-with-elastic-agent/342413/5 "2023-09-06T12:19:31Z")

</div>

> [@miltonhultgren](#):
>
> Is my understanding correct here?

Not really, most of the APIs that Metricbeat hits will do all the hard work on the node handling the HTTP request and not the master. Using `scope: cluster` with a single Beat connected to a node other than the elected master is a very good idea and will _significantly_ reduce the load on the elected master.

---

<div class="post-metadata">

**Author:** ![miltonhultgren](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/miltonhultgren/32/100401_2.png) [@miltonhultgren](https://discuss.elastic.co/u/miltonhultgren)\
**Post date:** [September 6, 2023, 12:24pm UTC](https://discuss.elastic.co/t/monitor-cluster-with-elastic-agent/342413/6 "2023-09-06T12:24:16Z")

</div>

Great, thanks for clarifying!

Then, for a larger cluster, it is best to use `scope: cluster` and load balance across all the non-master nodes?

For my own understanding, does that include resolving the cluster state as well? Meaning the non-master node can resolve that without involving the master node? Or is it more that **most** of the other work is not being done by the master node which leaves resources left on the master node to resolve the cluster state request?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [September 6, 2023, 12:43pm UTC](https://discuss.elastic.co/t/monitor-cluster-with-elastic-agent/342413/7 "2023-09-06T12:43:18Z")

</div>

> [@miltonhultgren](#):
>
> Then, for a larger cluster, it is best to use `scope: cluster` and load balance across all the non-master nodes?

Right, that'd work well.

> [@miltonhultgren](#):
>
> For my own understanding, does that include resolving the cluster state as well?

Not sure what you mean by "resolving" here. Metricbeat doesn't request the full cluster state AFAICT, only bits of it. It'd be even better if those requests used the `?local` query parameter to keep that work completely off the elected master, but they're not really the expensive requests anyway. It's other things like shard-level stats that tend to cause the bigger problems, and they don't hit the master at all.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [September 6, 2023, 12:48pm UTC](https://discuss.elastic.co/t/monitor-cluster-with-elastic-agent/342413/8 "2023-09-06T12:48:04Z")

</div>

> [@miltonhultgren](#):
>
> How you achieve HA on the agent likely depends on how you deploy your agent and stack, this is somewhat outside of the realm of the Elastic stack. If it's running as a side car in a Kubernetes pod for example then we would expect Kubernetes to take responsibility for that. In a raw deployment, I don't know. Who watches the watcher so to say. That's why we have alerting rules about monitoring data missing, so that's one option.

FWIW this concern is independent of the choice of `scope: node` or `scope: cluster`, you have a single point of failure either way. If using `scope: node` then things will stop working if the Metricbeat instance targetting the elected master fails.

---

<div class="post-metadata">

**Author:** ![lduvnjak](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lduvnjak/32/77724_2.png) [@lduvnjak](https://discuss.elastic.co/u/lduvnjak)\
**Post date:** [September 6, 2023, 1:07pm UTC](https://discuss.elastic.co/t/monitor-cluster-with-elastic-agent/342413/9 "2023-09-06T13:07:17Z")

</div>

Thanks for all the info guys!

Just to clarify, the `scope: node` option will make it so it only gathers the metrics from said node, while `scope: cluster` will make it so the node which gets the requests coordinates it further to get the statistics from every node. Did I get that right?

> FWIW this concern is independent of the choice of `scope: node` or `scope: cluster` , you have a single point of failure either way. If using `scope: node` then things will stop working if the Metricbeat instance targetting the elected master fails.

Yeah that I get. In which case I'd much rather prefer to go with `scope: node` so that even if it fails it only fails for that one node. Every node will have a local elastic agent since they need to collect the logs either way, and the installation + enrollment is automated so not a big deal.

My main concern is that if we pick `scope: node` and have a local agent on each one targeting that node's respective API that the consequent requests which hit the master don't overload it (_which shouldn't happen from what I've gathered in this thread_), since the requests that do make it to the master node are not expensive.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [September 6, 2023, 1:24pm UTC](https://discuss.elastic.co/t/monitor-cluster-with-elastic-agent/342413/10 "2023-09-06T13:24:25Z")

</div>

> [@lduvnjak](#):
>
> Just to clarify, the `scope: node` option will make it so it only gathers the metrics from said node,

That's not quite right. `scope: node` means it collects a few node-specific metrics from each node, but most of the metrics are cluster-wide things rather than node-specific ones, and `scope: node` means that those cluster-wide metrics are all only collected directly from the elected master. That imposes an enormous amount of extra load on the elected master in a big cluster, and means that if this Beat fails you get very few metrics.

You want `scope: cluster`.

---

<div class="post-metadata">

**Author:** ![lduvnjak](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lduvnjak/32/77724_2.png) [@lduvnjak](https://discuss.elastic.co/u/lduvnjak)\
**Post date:** [September 6, 2023, 1:30pm UTC](https://discuss.elastic.co/t/monitor-cluster-with-elastic-agent/342413/11 "2023-09-06T13:30:48Z")

</div>

Alright, that makes sense. Just figured that out due to this log:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/b/7/b7253bd59292519a458a24628360a31408ae9d42.png)

Thank you once again for all the help!

TLDR for those who don't want to read everything:

- The monitoring data streams it writes to **HAS** to match the standard schema `<type>-<dataset>-<namespace>`, only `namespace` can be changed
- If you have a large cluster with dedicated master-eligible nodes use `scope: cluster` which points to a LB that balances across master-ineligible nodes
- The agent is a SPOF (Single Point of Failure), unless on a platform that provides HA natively, like k8s

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [September 6, 2023, 1:58pm UTC](https://discuss.elastic.co/t/monitor-cluster-with-elastic-agent/342413/12 "2023-09-06T13:58:01Z")

</div>

Just a comment, unless something has changed, using the scope cluster will show the same transport address for every elasticsearch node.

So for every node the ip address you will see in Kibana monitoring will be the one for the node where the requests are being made.

We used scope cluster in the past, but this was confusing and we want back to using scope node and have one metricbeat per node, which was not an issue because we use it to get system metrics as well.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [September 6, 2023, 2:17pm UTC](https://discuss.elastic.co/t/monitor-cluster-with-elastic-agent/342413/13 "2023-09-06T14:17:17Z")

</div>

> [@leandrojmp](#):
>
> using the scope cluster will show the same transport address for every elasticsearch node.

That sounds like a pretty bad bug, @miltonhultgren are you aware of this?

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [September 6, 2023, 2:31pm UTC](https://discuss.elastic.co/t/monitor-cluster-with-elastic-agent/342413/14 "2023-09-06T14:31:14Z")

</div>

> [@DavidTurner](#):
>
> That sounds like a pretty bad bug

I reported it on a case to support on October last year, on one of the answers they mentioned that an internal enhancemente was created, this one: [https://github.com/elastic/enhancements/issues/17248](https://github.com/elastic/enhancements/issues/17248)

I was on 8.4 at the time, not sure if this still persists because I'm using one metricbeat per node.

---

<div class="post-metadata">

**Author:** ![lduvnjak](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lduvnjak/32/77724_2.png) [@lduvnjak](https://discuss.elastic.co/u/lduvnjak)\
**Post date:** [September 6, 2023, 2:34pm UTC](https://discuss.elastic.co/t/monitor-cluster-with-elastic-agent/342413/15 "2023-09-06T14:34:05Z")

</div>

I'm on 8.8.1 atm. Can't really say if it's the case in Stack Monitoring, but it's not the case in the documents.  
I switched from `node` to `cluster` atm just to see what it looks like, and the nodes' correct transport address is showing for each one.

---

<div class="post-metadata">

**Author:** ![miltonhultgren](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/miltonhultgren/32/100401_2.png) [@miltonhultgren](https://discuss.elastic.co/u/miltonhultgren)\
**Post date:** [September 6, 2023, 2:48pm UTC](https://discuss.elastic.co/t/monitor-cluster-with-elastic-agent/342413/16 "2023-09-06T14:48:20Z")

</div>

I was not aware of this, I'll verify if this happens in either case (in the docs or the UI) and follow up with a fix if needed. @leandrojmp thanks for brining that to our attention!

---

<div class="post-metadata">

**Author:** ![miltonhultgren](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/miltonhultgren/32/100401_2.png) [@miltonhultgren](https://discuss.elastic.co/u/miltonhultgren)\
**Post date:** [September 13, 2023, 6:39pm UTC](https://discuss.elastic.co/t/monitor-cluster-with-elastic-agent/342413/17 "2023-09-13T18:39:17Z")

</div>

@leandrojmp Just following up with that indeed the bug still exists, I've opened a PR to address it [[elasticsearch] Always report transport address in node\_stats by miltonhultgren · Pull Request #36582 · elastic/beats · GitHub](https://github.com/elastic/beats/pull/36582)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 11, 2023, 6:39pm UTC](https://discuss.elastic.co/t/monitor-cluster-with-elastic-agent/342413/18 "2023-10-11T18:39:19Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
