# Question on how to create a simple ML job

**URL:** <https://discuss.elastic.co/t/question-on-how-to-create-a-simple-ml-job/142155>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-machine-learning\
**Created:** [July 30, 2018, 11:08am UTC](https://discuss.elastic.co/t/question-on-how-to-create-a-simple-ml-job/142155 "2018-07-30T11:08:20Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![GioXmen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gioxmen/32/33916_2.png) [@GioXmen](https://discuss.elastic.co/u/GioXmen)\
**Post date:** [July 30, 2018, 11:08am UTC](https://discuss.elastic.co/t/question-on-how-to-create-a-simple-ml-job/142155/1 "2018-07-30T11:08:20Z")

</div>

Hello!

I am new to machine learning in elastic products, it was released here recently and we are on version 5.4 .

I want to create a machine learning job, but I am not sure on how to do it.

Sample data (obfuscated):

Let's say we have data coming in (filebeat--\>logstash--\>elastic search), field number one is - person\_id ( with this field we can identify the unique person), every person either eats an apple or an orange (field - "food")(apple means good and orange means bad). I need to see if there are irregularities between the two (either check/look at both, or just count "orange", though eating too many apples may also prove to be important to know/bad) .

I want to create a machine learning job that would find anomalies **in the top 10 most frequent/common** persons eating apples or oranges.

(The top 10 is for eg.: Jimmy ate an apple 10 times (he has 10 data points) and Alex has 9 data points and the others have 3 - 5 points, so Jimmy and Alex is now Top 2).

The issue: I can only see count "events" and "offset" in "fields" and not any specific fields, I can select them in "Key fields" but that doesn't seem to affect/do anything (or can I do that after creating the job?).

I can choose to split data (this separates the persons individually, which is good since I want to track their activity individually and not as a group, but there is more than ten of them (probably more, since I can only see ten in the preview)

Do I create 10 different jobs and define a specific "person\_id" or can it be done in a single job? Can I seperate the persons and have machine learning look at them either having field "food = apple" or field "food = orange".

An ideal way would be to have an event count for the 10 persons, and set the "food" field as the influencer. One main issue is how do I split data, while only keeping the top 10 persons?

Right now, I am going through the advanced editor, since the basic one "seems" to not be applicable to this scenario... perhaps I may be wrong though.

Related a bit: [Related topic I found](https://discuss.elastic.co/t/how-create-a-machine-learning-job-depends-a-particular-value-in-a-field/142185)

Thank you!

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [July 30, 2018, 8:07pm UTC](https://discuss.elastic.co/t/question-on-how-to-create-a-simple-ml-job/142155/2 "2018-07-30T20:07:45Z")

</div>

May I suggest first watching these getting started videos:

8 Min Tutorial #1 - How to create a single metric job:

8 Min Tutorial #2 - How to create a multi-metric job:

8 Min Tutorial #3 - Detect outliers in a population:

And also working through the sample tutorial would be beneficial:

[https://www.elastic.co/guide/en/x-pack/current/ml-getting-started.html](https://www.elastic.co/guide/en/x-pack/current/ml-getting-started.html)

---

<div class="post-metadata">

**Author:** ![GioXmen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gioxmen/32/33916_2.png) [@GioXmen](https://discuss.elastic.co/u/GioXmen)\
**Post date:** [July 31, 2018, 7:22am UTC](https://discuss.elastic.co/t/question-on-how-to-create-a-simple-ml-job/142155/3 "2018-07-31T07:22:23Z")

</div>

Thanks!  
This answers most of my questions! One left though - how can I "split" data? Can I define for eg.: look for anomalous activity in person\_id: "john" in his field "food"? ("food=apple, orange ...).

End result:

Machine learning job that looks for anomalous activity in food consumption:  
10 specific "top 10" people"

John - apple,orange, apple, orange ...  
Jimmy orange,orange ,orange, **apple**....  
Bob apple, apple, apple, apple ...  
....

Now I know how to get it to look at the food consumption, but I can only get it to look at it as one whole group and not specific people.

---

<div class="post-metadata">

**Author:** ![dkyle](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dkyle/32/59114_2.png) [@dkyle](https://discuss.elastic.co/u/dkyle)\
**Post date:** [July 31, 2018, 9:41am UTC](https://discuss.elastic.co/t/question-on-how-to-create-a-simple-ml-job/142155/4 "2018-07-31T09:41:20Z")

</div>

If you want to limit the data to specific people use a terms query in the datafeed

```auto
 {
  "query": {
    "terms": {
      "person_id": ["John", "Jimmy", "Bob"]
    }
  }
}

```

Create an advance job then enter the query in the datafeed section ![df_query](https://us1.discourse-cdn.com/elastic/original/3X/0/c/0ce242b3873c6acf05605c7a30fc0a64a7df50e1.jpeg)

---

<div class="post-metadata">

**Author:** ![GioXmen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gioxmen/32/33916_2.png) [@GioXmen](https://discuss.elastic.co/u/GioXmen)\
**Post date:** [July 31, 2018, 10:56am UTC](https://discuss.elastic.co/t/question-on-how-to-create-a-simple-ml-job/142155/5 "2018-07-31T10:56:14Z")

</div>

Perfect! That was what I was looking for!

The last thing that I am not sure of is, what analysis type to use? Do I just use count (high count) to look for anomalously high occurrences of let's say "apple" in field "food" (There can be three of them: "apple", "orange", "pear"). But I want to look for anomalies in all three occurrences, which would require creating three different jobs for a single person (with query filtering of apple, orange and pear).

"person\_id": ["John", "Jimmy", "Bob"...]  
"food": ["apple", "orange", "pear"]

Job1: high\_count for "apple" in person\_id: "John"  
Job2: high\_count for "orange" in person\_id: "John"  
Job1: high\_count for "pear" in person\_id: "John"

Is there any other way of doing this? Since I realize that if I want to monitor specific "people" I will have to create separate jobs for them, but do I have to create separate jobs for every variable I want to monitor too?

Perhaps distinct, high\_distinct or varp ? Goal: find anomalies in the number of occurrences of apple, orange and pear.

Data:  
Monday 1PM "John" "Apple"  
Monday 1PM "John" "Pear"  
Monday 2PM "John" "Apple"  
Monday 2PM "John" "Pear"  
Monday 3PM "John" "Orange"  
Monday 3PM "John" "Pear"  
Monday 4PM "John" "Orange"  
**Monday 4PM "John" "Apple"** I would consider this anomalous - higher than usual occurrence of "Orange

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [July 31, 2018, 1:27pm UTC](https://discuss.elastic.co/t/question-on-how-to-create-a-simple-ml-job/142155/6 "2018-07-31T13:27:36Z")

</div>

If you want to "split" the data/analysis on a field, then set the `partition_field_name` to `person_id` (or in a multi-metric job, make that the Split field). In this way, you don't need separate jobs per `person_id` - it can be done in one job.

What Dave suggests with a query filter is only relevant if you ONLY want to analyze "John", "Jimmy" and "Bob", but want to exclude others in the data. If you don't do this filter, then you'll have separate baselines/analyses for every `person_id`

In terms of the function to use, yes, you can certainly use `high_count` to detect more-than-usual occurrences of documents for a particular `person_id`. However, the `count` functions DO NOT COUNT FIELDS - it counts documents.

So, in your example, you want to count the occurrent of `food` per `person_id` then you need to do a double-split, which is only available via the Advanced Job Wizard.

To accomplish a double-split, set `partition_field_name=person_id` and `by_field_name=food`

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/b/c/bcdd9f3c6c4bbf68e411d205c98c7fa72be87473.png)

---

<div class="post-metadata">

**Author:** ![GioXmen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gioxmen/32/33916_2.png) [@GioXmen](https://discuss.elastic.co/u/GioXmen)\
**Post date:** [July 31, 2018, 1:48pm UTC](https://discuss.elastic.co/t/question-on-how-to-create-a-simple-ml-job/142155/7 "2018-07-31T13:48:07Z")

</div>

Wow! Great!

> [@richcollier](#):
>
> What Dave suggests with a query filter is only relevant if you ONLY want to analyze "John", "Jimmy" and "Bob", but want to exclude others in the data. If you don't do this filter, then you'll have separate baselines/analyses for every `person_id`

Yes! I do want to **do specific analysis** since I have many many person's, **but I only care about the most active ones** , since if something "wrong" happens to them, I have to look at them quickly and solve it, hence the "top 10" thing. People who only come up rarely are not a "high" priority 🙂

> [@richcollier](#):
>
> So, in your example, you want to count the occurrent of `food` per `person_id` then you need to do a double-split, which is only available via the Advanced Job Wizard.
> 
> To accomplish a double-split, set `partition_field_name=person_id` and `by_field_name=food`

Okay! That looks like it makes sense!, now I will set that detector and **who should I set as the influencer?** should I set **person\_id** as the influencer? or **food**? or perhaps **both**? 😅 Though I believe it won't let me set "food" as the influencer 🙂

What has happened:

Goal: analyse specific 10 people and see their eating habits (apple, orange, pear), find anomalies. IF an anomaly is found I have to know which person is having it.

Execution: Will limit input with a query to those specific ten people (or make separate jobs for each person to know which person is having issues?) and then follow above given example.

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [July 31, 2018, 2:16pm UTC](https://discuss.elastic.co/t/question-on-how-to-create-a-simple-ml-job/142155/8 "2018-07-31T14:16:39Z")

</div>

You can put both `person_id` and `food` as an influencer. As a general rule, any field that you're splitting on is a good candidate to be an influencer.

---

<div class="post-metadata">

**Author:** ![GioXmen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gioxmen/32/33916_2.png) [@GioXmen](https://discuss.elastic.co/u/GioXmen)\
**Post date:** [August 1, 2018, 7:29am UTC](https://discuss.elastic.co/t/question-on-how-to-create-a-simple-ml-job/142155/9 "2018-08-01T07:29:19Z")

</div>

While trying to save my new configuration I get:

Save failed: [illegal\_argument\_exception] Can't merge a non object mapping [food] with an object mapping [food]

Solution: [https://www.elastic.co/guide/en/x-pack/current/ml-mappingclash.html](https://www.elastic.co/guide/en/x-pack/current/ml-mappingclash.html)

When using this solution, I get:

[no\_shard\_available\_action\_exception] No shard available for [get [][data\_counts][]: routing [null]]

Upgrading my version to above 5.4, will report back if issue is solved.

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [August 1, 2018, 12:00pm UTC](https://discuss.elastic.co/t/question-on-how-to-create-a-simple-ml-job/142155/10 "2018-08-01T12:00:25Z")

</div>

Using a dedicated results index for the ML job avoids this problem, but I also do heavily suggest that you get off of v5.4 as that was the "beta" version of ML. Can you upgrade to at least v6.1?

---

<div class="post-metadata">

**Author:** ![GioXmen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gioxmen/32/33916_2.png) [@GioXmen](https://discuss.elastic.co/u/GioXmen)\
**Post date:** [August 1, 2018, 12:22pm UTC](https://discuss.elastic.co/t/question-on-how-to-create-a-simple-ml-job/142155/11 "2018-08-01T12:22:12Z")

</div>

Sadly only to 5.6 for now, update still in progress, using a dedicated index caused an issue earlier, though I believe it was 5.4 related OR because we ran out of available memory. Will report once upgrade is finished

---

<div class="post-metadata">

**Author:** ![GioXmen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gioxmen/32/33916_2.png) [@GioXmen](https://discuss.elastic.co/u/GioXmen)\
**Post date:** [August 2, 2018, 12:18pm UTC](https://discuss.elastic.co/t/question-on-how-to-create-a-simple-ml-job/142155/12 "2018-08-02T12:18:55Z")

</div>

The update to 5.6 has fixed all of our issues, no more bugs and UI weirdness in general. Data analysis seems to be working perfectly!

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [October 29, 2018, 10:04pm UTC](https://discuss.elastic.co/t/question-on-how-to-create-a-simple-ml-job/142155/13 "2018-10-29T22:04:20Z")

</div>


