# Stack Monitoring - Elasticsearch - Elastic

**URL:** <https://discuss.elastic.co/t/stack-monitoring-elasticsearch-elastic/312485>\
**Category:** Kibana\
**Tags:** elastic-stack-monitoring\
**Created:** [August 19, 2022, 4:19pm UTC](https://discuss.elastic.co/t/stack-monitoring-elasticsearch-elastic/312485 "2022-08-19T16:19:36Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![alexus](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alexus/32/12696_2.png) [@alexus](https://discuss.elastic.co/u/alexus)\
**Post date:** [August 19, 2022, 4:19pm UTC](https://discuss.elastic.co/t/stack-monitoring-elasticsearch-elastic/312485/1 "2022-08-19T16:19:36Z")

</div>

Hello World!

Via `Kibana` -\> `Stack Monitoring`

 ![Screen Shot 2022-08-19 at 12.14.40 PM](https://us1.discourse-cdn.com/elastic/original/3X/5/b/5bbcbaf942706f4d6a8bc035b413e480ead02209.png)

Any ideas what causing so many spaces in between, why the line isn't solid?

Thanks in advance!

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [August 19, 2022, 6:27pm UTC](https://discuss.elastic.co/t/stack-monitoring-elasticsearch-elastic/312485/2 "2022-08-19T18:27:43Z")

</div>

Are you using the legacy self-monitoring or metricbeat to collect the metrics?

This normally happens when the service (elasticsearch, logstash etc) is under heavy load and doesn't answer the metric requests.

---

<div class="post-metadata">

**Author:** ![rugenl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rugenl/32/12887_2.png) [@rugenl](https://discuss.elastic.co/u/rugenl)\
**Post date:** [August 19, 2022, 7:28pm UTC](https://discuss.elastic.co/t/stack-monitoring-elasticsearch-elastic/312485/3 "2022-08-19T19:28:48Z")

</div>

Zoom out to at least 30 minutes. What versions?

---

<div class="post-metadata">

**Author:** ![alexus](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alexus/32/12696_2.png) [@alexus](https://discuss.elastic.co/u/alexus)\
**Post date:** [August 20, 2022, 4:10am UTC](https://discuss.elastic.co/t/stack-monitoring-elasticsearch-elastic/312485/4 "2022-08-20T04:10:32Z")

</div>

thank you for looking into my topic

- regards to the metrics: I do believe the cluster is still using legacy self-monitoring at the moment..
- regards to the load: I don't think the elasticsearch cluster is under heavy load.

how do I check the cluster's load anyways?

---

<div class="post-metadata">

**Author:** ![alexus](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alexus/32/12696_2.png) [@alexus](https://discuss.elastic.co/u/alexus)\
**Post date:** [August 20, 2022, 4:13am UTC](https://discuss.elastic.co/t/stack-monitoring-elasticsearch-elastic/312485/5 "2022-08-20T04:13:15Z")

</div>

An interesting observation is that if I zoom out to 30mins instead of the default 15mins, there is no break in the lines at all, yet when I zoom back in, there are a lot of blank spaces in that line... it's almost like it's not rendering it properly w/ 15 mins for whatever reason, but the data _is_ there...

the version is: _7.17.5_

---

<div class="post-metadata">

**Author:** ![matschaffer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/matschaffer/32/95396_2.png) [@matschaffer](https://discuss.elastic.co/u/matschaffer)\
**Post date:** [August 26, 2022, 2:26am UTC](https://discuss.elastic.co/t/stack-monitoring-elasticsearch-elastic/312485/6 "2022-08-26T02:26:54Z")

</div>

Usually when I see this it means the ES master is having trouble keeping up. Doubly so when using internal-monitoring since it needs to do all the work of gathering and shipping data to the monitoring cluster.

Things I'd watch for:

- GC warnings in the ES logs
- long running \_cat/tasks
- master node hot threads & circuit breakers
- Large number of fields relative to master memory

That "many fields" has been a hard one for me since it sometimes won't manifest via any log.

It's just that when fields get added the master needs to do _that_ work in addition to everything else, so everything just gets just that little bit slower (monitoring reporting, indexing new docs, etc).

Hope that helps!

---

<div class="post-metadata">

**Author:** ![alexus](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alexus/32/12696_2.png) [@alexus](https://discuss.elastic.co/u/alexus)\
**Post date:** [August 26, 2022, 6:05pm UTC](https://discuss.elastic.co/t/stack-monitoring-elasticsearch-elastic/312485/7 "2022-08-26T18:05:13Z")

</div>

- GC warnings in the elasticsearch log:

> {"type": "server", "timestamp": "2022-08-26T15:25:56,579Z", "level": "WARN", "component": "o.e.i.SystemIndexManager", "cluster.name": "elastic", "node.name": "elastic-es-master-2", "message": "Missing \_meta field in mapping [doc] of index [.watches], assuming mappings update required", "cluster.uuid": "90tky-UNTRCfA42cAFK-aw", "node.id": "HqIuhom5RDOxyN7I8\_rfwg" }

- long running \_cat/tasks

```auto
data_frame/transforms[c] sXu62qwyQvK-sty1vjXWIQ:257560 cluster:117 persistent 1661283558215 19:39:18 2.9d 10.202.77.2 elastic-es-data-9
geoip-downloader[c] snl_FAzBQcaQ43HEDNBCEA:308961 cluster:121 persistent 1661284170370 19:49:30 2.9d 10.202.78.130 elastic-es-data-8
data_frame/transforms[c] snl_FAzBQcaQ43HEDNBCEA:674625 cluster:123 persistent 1661286361074 20:26:01 2.8d 10.202.78.130 elastic-es-data-8

```

- master node hot threads & circuit breakers

i'm not sure how to check for that, if you could advise I would appreciate that 😉  
I did check `/_cat/thread_pool` and most I see has all three 0s.

- Large number of fields relative to master memory

also not sure how to check for that either(

I did check indexing via monitoring Kibana app and it seems low (the highest index is only ~50/s)

Please advise) and "thank you" in advance!

---

<div class="post-metadata">

**Author:** ![matschaffer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/matschaffer/32/95396_2.png) [@matschaffer](https://discuss.elastic.co/u/matschaffer)\
**Post date:** [August 29, 2022, 1:15am UTC](https://discuss.elastic.co/t/stack-monitoring-elasticsearch-elastic/312485/8 "2022-08-29T01:15:09Z")

</div>

That warning might be fine, though hopefully it's only happening once.

GC warnings will look more like this:

```auto
[gc][359] overhead, spent [7.5s] collecting in the last [8.3s]

```

Also those tasks are expected to be long running. What you'll want to watch for are long running write or search tasks.

Hot threads API doc is at [Nodes hot threads API | Elasticsearch Guide [8.4] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/cluster-nodes-hot-threads.html), good that thread pool is low. Rejections are the thing to watch for there. See [cat thread pool API | Elasticsearch Guide [8.4] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/cat-thread-pool.html#cat-thread-pool-api-ex-headings) for info.

For field count I usually check the field caps API. [Field capabilities API | Elasticsearch Guide [8.4] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-field-caps.html)

```auto
curl -s -u USER:PASS https://ES/_field_caps'?fields=*' | jq '.fields | keys | length'

```

As that number grows, the master can slow down especially if a lot of fields are getting added frequently. One time I saw this happen because http request header names were being added as fields.

I've never seen very specific measurement/guidance but generally if I see more than 10k per 1gb of master heap, I might look for ways to reduce the field count.

---

<div class="post-metadata">

**Author:** ![matschaffer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/matschaffer/32/95396_2.png) [@matschaffer](https://discuss.elastic.co/u/matschaffer)\
**Post date:** [August 29, 2022, 1:24am UTC](https://discuss.elastic.co/t/stack-monitoring-elasticsearch-elastic/312485/9 "2022-08-29T01:24:03Z")

</div>

If you're able, try giving the master node more heap for starters.

If that helps make the graphs consistent then you know where to focus attention (alleviating master pressure).

---

<div class="post-metadata">

**Author:** ![alexus](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alexus/32/12696_2.png) [@alexus](https://discuss.elastic.co/u/alexus)\
**Post date:** [August 29, 2022, 10:25am UTC](https://discuss.elastic.co/t/stack-monitoring-elasticsearch-elastic/312485/10 "2022-08-29T10:25:04Z")

</div>

> GC warnings will look more like this:

I checked master's "Elasticsearch" log and did not find warnings like the one you listed, the only warning I did find is the one that I listed previously and I also did research for it and it's safe to ignore.

> cluster-nodes-hot-threads:

```auto
   100.3% [cpu=100.3%, other=0.0%] (501.2ms out of 500ms) cpu usage by thread 'elasticsearch[elastic-es-data-11][[.monitoring-es-7-2022.08.29][0]: Lucene Merge Thread #2343]'

```

&

```auto
   100.4% [cpu=100.4%, other=0.0%] (501.8ms out of 500ms) cpu usage by thread 'elasticsearch[elastic-es-data-1][system_read][T#2]'

```

&

```auto
   100.0% [cpu=27.6%, other=72.4%] (500ms out of 500ms) cpu usage by thread 'elasticsearch[elastic-es-coor-0][transport_worker][T#4]'

```

&

```auto
   100.3% [cpu=100.3%, other=0.0%] (501.6ms out of 500ms) cpu usage by thread 'elasticsearch[elastic-es-data-9][write][T#9]'

```

> threat\_pool

low (almost all 0s)

> field counts:

```auto
% curl -s -k -u elastic:$es_password https://localhost:9200/_field_caps'?fields=*' | jq '.fields | keys | length'
12688
%

```

after killing `auditbeat-*`, `filebeat-*`, `metricbeat-*` and `logstash-*`, now it's:

```auto
% curl -s -k -u elastic:$es_password https://localhost:9200/_field_caps'?fields=*' | jq '.fields | keys | length' 
4806
% 

```

> heap

master node has heap of 14GB and using under 9GB..

---

<div class="post-metadata">

**Author:** ![matschaffer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/matschaffer/32/95396_2.png) [@matschaffer](https://discuss.elastic.co/u/matschaffer)\
**Post date:** [August 31, 2022, 1:10am UTC](https://discuss.elastic.co/t/stack-monitoring-elasticsearch-elastic/312485/11 "2022-08-31T01:10:43Z")

</div>

100% CPU in merge for .monitoring-es-7-2022.08.29 is curious. I wonder if the nodes handling that index are struggling rather than the master. It might help to try adding shards to that index.

The quickest way I know to do that is to update the template and delete the index, then when new documents show up it'll have more shards.

Note that it's a bit destructive, so you may want to make sure you have backups.

I don't have an internal-collection cluster handy right now to provide exact steps, but hopefully the above can help you find the underlying issue.

---

<div class="post-metadata">

**Author:** ![alexus](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alexus/32/12696_2.png) [@alexus](https://discuss.elastic.co/u/alexus)\
**Post date:** [August 31, 2022, 3:26pm UTC](https://discuss.elastic.co/t/stack-monitoring-elasticsearch-elastic/312485/12 "2022-08-31T15:26:04Z")

</div>

there are several data nodes and each of the data node is running on GCP: n1-standard-16 (vCPU:16 RAM:60GB), heap size is 31gb on each of the data nodes.

the busiest node is using during the peaks about 4 vCPU (out of 16vCPU)

I checked and previous day for `.monitoring-es-7-YYYY.MM.DD` and it is about 10gb in size and about 10M of docs and it has 1 primary and 1 replica shard.

---

<div class="post-metadata">

**Author:** ![matschaffer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/matschaffer/32/95396_2.png) [@matschaffer](https://discuss.elastic.co/u/matschaffer)\
**Post date:** [September 5, 2022, 5:26am UTC](https://discuss.elastic.co/t/stack-monitoring-elasticsearch-elastic/312485/13 "2022-09-05T05:26:58Z")

</div>

Yeah, that all sounds healthy. If you only see it at 15min zoom, I guess it's possible we have some rendering bug.

If you're able, have a look after updating to a more recent version and if it persists, we can move this to a github issue for further investigation.

---

<div class="post-metadata">

**Author:** ![matschaffer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/matschaffer/32/95396_2.png) [@matschaffer](https://discuss.elastic.co/u/matschaffer)\
**Post date:** [September 5, 2022, 5:44am UTC](https://discuss.elastic.co/t/stack-monitoring-elasticsearch-elastic/312485/14 "2022-09-05T05:44:48Z")

</div>

I had 7.17 booted today so I enabled internal collection and 15 min interval seems okay so far.

 ![Screen Shot 2022-09-05 at 14.43.53](https://us1.discourse-cdn.com/elastic/original/3X/1/8/18d6620b329a2ed730e7fba26256e314fdcfcddc.jpeg)

---

<div class="post-metadata">

**Author:** ![alexus](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alexus/32/12696_2.png) [@alexus](https://discuss.elastic.co/u/alexus)\
**Post date:** [September 5, 2022, 7:53am UTC](https://discuss.elastic.co/t/stack-monitoring-elasticsearch-elastic/312485/15 "2022-09-05T07:53:25Z")

</div>

I doubt this is a bug, rather some misconfiguration within the cluster (it's like looking for the needle in a haystack 🙃 )

> **[Advanced configuration | Elasticsearch Guide \[7.17\] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/7.17/advanced-configuration.html#set-jvm-heap-size)**

```auto
% curl --silent --insecure https://elastic:$es_password@localhost:9200/_nodes/_all/jvm | jq | grep using_compressed_ordinary_object_pointers | uniq
        "using_compressed_ordinary_object_pointers": "true",
% 

```

> **[Disable swapping | Elasticsearch Guide \[7.17\] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/7.17/setup-configuration-memory.html#swappiness)**

> **[Disable swapping | Elasticsearch Guide \[7.17\] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/7.17/setup-configuration-memory.html#bootstrap-memory_lock)**

> **[Virtual memory | Elasticsearch Guide \[7.17\] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/7.17/vm-max-map-count.html)**

```auto
root@elastic-es-data-1:/usr/share/elasticsearch# sysctl vm.max_map_count
vm.max_map_count = 65530
root@elastic-es-data-1:/usr/share/elasticsearch# 

```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 3, 2022, 7:53am UTC](https://discuss.elastic.co/t/stack-monitoring-elasticsearch-elastic/312485/16 "2022-10-03T07:53:44Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
