# Set up watcher for alerting high CPU usage by some process

**URL:** <https://discuss.elastic.co/t/set-up-watcher-for-alerting-high-cpu-usage-by-some-process/139353>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-alerting\
**Created:** [July 10, 2018, 1:29pm UTC](https://discuss.elastic.co/t/set-up-watcher-for-alerting-high-cpu-usage-by-some-process/139353 "2018-07-10T13:29:34Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Oleksandr\_Novozhylov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/oleksandr_novozhylov/32/33191_2.png) [@Oleksandr\_Novozhylov](https://discuss.elastic.co/u/Oleksandr_Novozhylov)\
**Post date:** [July 10, 2018, 1:29pm UTC](https://discuss.elastic.co/t/set-up-watcher-for-alerting-high-cpu-usage-by-some-process/139353/1 "2018-07-10T13:29:34Z")

</div>

Hello!

I'm trying to create a Watcher Alert that will be triggered when some process on a node uses over 0.95% of CPU for the last one hour.

Here is an example of my config:

```
{
  "trigger": {
    "schedule": {
      "interval": "10m"
    }
  },
  "input": {
    "search": {
      "request": {
        "search_type": "query_then_fetch",
        "indices": [
          "metricbeat*"
        ],
        "types": [],
        "body": {
          "size": 0,
          "query": {
            "bool": {
              "must": [
                {
                  "range": {
                    "system.process.cpu.total.norm.pct": {
                      "gte": 0.95
                    }
                  }
                },
                {
                  "range": {
                    "system.process.cpu.start_time": {
                      "gte": "now-1h"
                    }
                  }
                },
                {
                  "match": {
                    "environment": "test"
                  }
                }
              ]
            }
          }
        }
      }
    }
  },
  "condition": {
    "compare": {
      "ctx.payload.hits.total": {
        "gt": 0
      }
    }
  },
  "actions": {
    "send-to-slack": {
      "throttle_period_in_millis": 1800000,
      "webhook": {
        "scheme": "https",
        "host": "hooks.slack.com",
        "port": 443,
        "method": "post",
        "path": "{{ctx.metadata.onovozhylov-test}}",
        "params": {},
        "headers": {
          "Content-Type": "application/json"
        },
        "body": "{ \"text\": \" ==========\nTest parameters:\n\tthrottle_period_in_millis: 60000\n\tInterval: 1m\n\tcpu.total.norm.pct: 0.5\n\tcpu.start_time: now-1m\n\nThe watcher:*{{ctx.watch_id}}* in env:*{{ctx.metadata.env}}* found that the process *{{ctx.system.process.name}}* has been utilizing CPU over 95% for the past 1 hr on node:\n{{#ctx.payload.nodes}}\t{{.}}\n\n{{/ctx.payload.nodes}}\n\nThe runbook entry is here: *{{ctx.metadata.runbook}}* \"}"
      }
    }
  },
  "metadata": {
    "onovozhylov-test": "/services/T0U0CFMT4/BBK1A2AAH/MlHAF2QuPjGZV95dvO11111111",
    "env": "{{ grains.get('environment') }}",
    "runbook": "http://mytest.com"
  }
}

```

This Watcher doesn't work when I set the metric `system.process.cpu.start_time`. Perhaps this metric is not a correct one...

And another issue is that I don't know how to add the `system.process.name` into a message body.

Thanks in advance for any help!

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [July 13, 2018, 9:30am UTC](https://discuss.elastic.co/t/set-up-watcher-for-alerting-high-cpu-usage-by-some-process/139353/2 "2018-07-13T09:30:35Z")

</div>

can you elaborate what does not work? What do you mean with 'set the metric'? What do you want to do with the start time? Should it be part of the query?

In order to access the process name of the first hit, you can access the hits array from the response like `ctx.payload.hits.hits[0]._source.system.process.name`. You probably want to add an aggregation on your query to collect all the process names instead of going through the hits though.

Also, there is a dedicated `slack` action that you could use instead.

Hope this helps!

--alex

---

<div class="post-metadata">

**Author:** ![Oleksandr\_Novozhylov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/oleksandr_novozhylov/32/33191_2.png) [@Oleksandr\_Novozhylov](https://discuss.elastic.co/u/Oleksandr_Novozhylov)\
**Post date:** [July 15, 2018, 7:46pm UTC](https://discuss.elastic.co/t/set-up-watcher-for-alerting-high-cpu-usage-by-some-process/139353/3 "2018-07-15T19:46:45Z")

</div>

Thank you for your answer.

I used `system.process.cpu.start_time` in the query to alert about a process that used over 0.95% of CPU for a particular period of time (e.g. "gte": "now-1h"). However, it didn't work for this purpose because no alerts were sent. So I'm not sure whether this field can be used for such a case.

My issue is that I can't find either a CPU-specific field or some other field to track a process that uses over 0.95% of CPU for a particular period of time.

> [@spinscale](#):
>
> In order to access the process name of the first hit, you can access the hits array from the response like `ctx.payload.hits.hits[0]._source.system.process.name` . You probably want to add an aggregation on your query to collect all the process names instead of going through the hits though.

Thanks, I'll give a try to it!

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [July 17, 2018, 7:05am UTC](https://discuss.elastic.co/t/set-up-watcher-for-alerting-high-cpu-usage-by-some-process/139353/4 "2018-07-17T07:05:26Z")

</div>

Hey,

in order to pinpoint the problem of 'does not work', can we step away from the watch for a second and ensure the query is working as expected?

Can you share your full query and the response? Finding out why there are no responses/data being returned is the first step here I think.

--Alex

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 14, 2018, 7:05am UTC](https://discuss.elastic.co/t/set-up-watcher-for-alerting-high-cpu-usage-by-some-process/139353/5 "2018-08-14T07:05:37Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
