# Create jobs with field combinations

**URL:** <https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-machine-learning\
**Created:** [June 18, 2019, 1:50pm UTC](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268 "2019-06-18T13:50:46Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Bharath\_Kumar\_R](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bharath_kumar_r/32/45474_2.png) [@Bharath\_Kumar\_R](https://discuss.elastic.co/u/Bharath_Kumar_R)\
**Post date:** [June 18, 2019, 1:50pm UTC](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268/1 "2019-06-18T13:50:46Z")

</div>

Hey, I have 2 fields: field1 and field2 in my data. Right now, Im filtering the data for some combinations of field1 and field2 and creating the jobs for those saved searches. What modifications/configurations are to be made such that my job automatically filters data for every combination of field1 and field2 and creates models for every such combination? Is it possible with multi metric job(or any other way) or must be implemented via any language client?

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [June 18, 2019, 5:55pm UTC](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268/2 "2019-06-18T17:55:27Z")

</div>

One possibility would be to dynamically create a `script_field` that is the concatenation of field1 and field2:

```auto
PUT _xpack/ml/anomaly_detectors/my_job
{
    "analysis_config": {
        "bucket_span": "1h",
        "detectors": [{
            "detector_description": "count per method_status",
            "function": "count",
            "partition_field_name": "method_status"
        }],
        "influencers": ["method", "status"]
    },
    "data_description": {
        "time_field": "@timestamp"
    }
}

```

```auto
PUT _xpack/ml/datafeeds/datafeed-my_job/
{
  "job_id": "my_job",
  "indices": [
    "gallery-*"
  ],
      "query": {
        "match_all": {
        }
      },
      "script_fields": {
        "method_status": {
          "script": {
            "source": "doc['method'].value + '_' + doc['status'].value",
            "lang": "painless"
          },
          "ignore_failure": false
        }
      }

}

```

```auto
GET _xpack/ml/datafeeds/datafeed-my_job/_preview/

```

```auto
...
 {
    "@timestamp" : 1483244920000,
    "method" : "POST",
    "method_status" : "POST_200",
    "status" : "200"
  },
  {
    "@timestamp" : 1483244949000,
    "method" : "GET",
    "method_status" : "GET_200",
    "status" : "200"
  },
  {
    "@timestamp" : 1483245000000,
    "method" : "GET",
    "method_status" : "GET_200",
    "status" : "200"
  },
...

```

Result:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/9/a/9a6554c14ef75ef8442e6832620be9eef02c091c.png)

---

<div class="post-metadata">

**Author:** ![Bharath\_Kumar\_R](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bharath_kumar_r/32/45474_2.png) [@Bharath\_Kumar\_R](https://discuss.elastic.co/u/Bharath_Kumar_R)\
**Post date:** [June 19, 2019, 10:10am UTC](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268/3 "2019-06-19T10:10:12Z")

</div>

I tried to create a job with the foll. request:

PUT /\_ml/anomaly\_detectors/my\_job  
{  
"analysis\_config": {  
"bucket\_span": "15m",  
"detectors": [{  
"detector\_description": "count per method\_status",  
"function": "low\_count"  
}],  
"influencers": ["SHIPPERID", "CARRIERID"]  
},  
"data\_description": {  
"time\_field": "EVENTTIME"  
}  
}

But I got an error as follows:

{  
"error": {  
"root\_cause": [  
{  
"type": "status\_exception",  
"reason": "This job would cause a mapping clash with existing field [CARRIERID] - avoid the clash by assigning a dedicated results index"  
}  
],  
"type": "status\_exception",  
"reason": "This job would cause a mapping clash with existing field [CARRIERID] - avoid the clash by assigning a dedicated results index",  
"caused\_by": {  
"type": "illegal\_argument\_exception",  
"reason": "Can't merge a non object mapping [CARRIERID] with an object mapping [CARRIERID]"  
}  
},  
"status": 400  
}

Can you explain what is the cause for the error and how to resolve it?

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [June 19, 2019, 10:28am UTC](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268/4 "2019-06-19T10:28:38Z")

</div>

The destination index for the results of your jobs is a shared index called `.ml-anomalies-shared`. There apparently is already a field within that index (from some other job you've run apparently) with the name `CARRIERID`. This field has a mapping (assignment to a data type) that is different than the `CARRIERID` field mapping type from your new job. A single index cannot have two fields with the same name with different mapping types.

To avoid this, add the following to make a dedicated new results index just for that job:

```auto
  "results_index_name": "mynewresultsindexname"

```

for example:

```auto
PUT _xpack/ml/anomaly_detectors/my_job
{
    "analysis_config": {
        "bucket_span": "1h",
        "detectors": [{
            "detector_description": "count per method_status",
            "function": "count",
            "partition_field_name": "method_status"
        }],
        "influencers": ["method", "status"]
    },
    "data_description": {
        "time_field": "@timestamp"
    },
    "results_index_name": "mynewresultsindexname"
}

```

---

<div class="post-metadata">

**Author:** ![Bharath\_Kumar\_R](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bharath_kumar_r/32/45474_2.png) [@Bharath\_Kumar\_R](https://discuss.elastic.co/u/Bharath_Kumar_R)\
**Post date:** [June 19, 2019, 11:49am UTC](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268/5 "2019-06-19T11:49:24Z")

</div>

PUT /_ml/datafeeds/datafeed-my\_job/  
{  
"job\_id": "my\_job",  
"indices": [  
"ab-\*"  
],  
"query": {  
"match\_all": {  
}  
},  
"script\_fields": {  
"method\_status": {  
"script": {  
"source": "doc['SHIPPERID'].value+ '_' + doc['CARRIERID'].value",  
"lang": "painless"  
},  
"ignore\_failure": false  
}  
}

}

In the above request, does Elasticsearch filters the documents for every combination of the fields given?  
For eg, if SHIPPERID="abcd" and CARRIERID="efgh", then does it automatically filter the documents with the given field values?  
Does it create separate model for every combination of the given fields? I need separate models to be created for every collection of documents filtered with a combination of the given fields?

As of right now, I'm directly adding some combinations of the given fields as filters and creating individual single metric jobs for each of the saved searches

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [June 19, 2019, 12:11pm UTC](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268/6 "2019-06-19T12:11:29Z")

</div>

The query in the datafeed does not filter - it simply creates a new field that is the concatenation of two other fields in the documents.

method: GET  
status:200  
method\_status: GET\_200

It is the ML job configuration, specifically the:

```auto
            "partition_field_name": "method_status"

```

that creates an independent baseline analysis for every instance of `method_status`, thus, every observed combination of those two fields

---

<div class="post-metadata">

**Author:** ![Bharath\_Kumar\_R](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bharath_kumar_r/32/45474_2.png) [@Bharath\_Kumar\_R](https://discuss.elastic.co/u/Bharath_Kumar_R)\
**Post date:** [June 19, 2019, 12:36pm UTC](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268/7 "2019-06-19T12:36:13Z")

</div>

PUT /_ml/datafeeds/datafeed-my\_job/  
{  
"job\_id": "my\_job",  
"indices": [  
"ab-\*"  
],  
"query": {  
"match\_all": {  
}  
},  
"script\_fields": {  
"method\_status": {  
"script": {  
"source": "doc['SHIPPERID'].value+ '_' + doc['CARRIERID'].value",  
"lang": "painless"  
},  
"ignore\_failure": false  
}  
}

}

I get the foll error:  
"error": {  
"root\_cause": [  
{  
"type": "status\_exception",  
"reason": "[datafeed-my\_job] cannot retrieve field [SHIPPERID\_CARRIERID] because it has no mappings"  
}  
],

Is the error produced because the index does not have the field SHIPPERID\_CARRIERID or any other reason?

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [June 19, 2019, 12:51pm UTC](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268/8 "2019-06-19T12:51:24Z")

</div>

it is because you called your scripted field `method_status`, not `SHIPPERID_CARRIERID`

This is where you did that:

"script\_fields": {  
**"method\_status": {**  
"script": {

You can see what your datafeed returns by:

```auto
GET _xpack/ml/datafeeds/datafeed-my_job/_preview/

```

---

<div class="post-metadata">

**Author:** ![Bharath\_Kumar\_R](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bharath_kumar_r/32/45474_2.png) [@Bharath\_Kumar\_R](https://discuss.elastic.co/u/Bharath_Kumar_R)\
**Post date:** [June 19, 2019, 1:19pm UTC](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268/9 "2019-06-19T13:19:07Z")

</div>

Can you suggest me a blog which explains creation of jobs like these via requests on console?

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [June 19, 2019, 1:34pm UTC](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268/10 "2019-06-19T13:34:35Z")

</div>

There is no specific blog on this, but the [online API docs](https://www.elastic.co/guide/en/elasticsearch/reference/7.1/ml-apis.html) show everything

---

<div class="post-metadata">

**Author:** ![Bharath\_Kumar\_R](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bharath_kumar_r/32/45474_2.png) [@Bharath\_Kumar\_R](https://discuss.elastic.co/u/Bharath_Kumar_R)\
**Post date:** [June 19, 2019, 1:41pm UTC](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268/11 "2019-06-19T13:41:07Z")

</div>

Hey, as you suggested, I created the job and started the Datafeeds for the job. Under the Single Metric Viewer, I couldn't view my job. Is it because open jobs can't be viewed under Single Metric Viewer?

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [June 19, 2019, 1:54pm UTC](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268/12 "2019-06-19T13:54:49Z")

</div>

No, jobs can be viewed in Single Metric Viewer as long as they have results.

Check the Job Management page for your newly created job and check its status there. I'm guessing that if you tried to start the job from the API, you may have started the datafeed but neglected to "open" the job first.

By the way, even if you set the config of the job/datafeed with the API, you can still use the Job Management UI to start/stop the job.

---

<div class="post-metadata">

**Author:** ![Bharath\_Kumar\_R](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bharath_kumar_r/32/45474_2.png) [@Bharath\_Kumar\_R](https://discuss.elastic.co/u/Bharath_Kumar_R)\
**Post date:** [June 19, 2019, 2:00pm UTC](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268/13 "2019-06-19T14:00:58Z")

</div>

No, I first opened the job with the request:  
POST \_ml/anomaly\_detectors/my\_job/\_open

and started Datafeeds with request:  
POST \_ml/datafeeds/datafeed-my\_job/\_start

This is how my job appears in Job management page and as you can see Single Metric Viewer has been disabled

 ![22%20PM](https://us1.discourse-cdn.com/elastic/original/3X/b/4/b4c97d00735b5f493c184ad69274eaf15b8945ec.png)

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [June 19, 2019, 2:36pm UTC](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268/14 "2019-06-19T14:36:13Z")

</div>

Ah yes - I forgot. This is because of the datafeed creates a `script_field`, making the Single Metric Viewer incapable of reconstructing the query to paint the time series.

We'll support that in v7.2: [https://github.com/elastic/kibana/pull/34079](https://github.com/elastic/kibana/pull/34079)

---

<div class="post-metadata">

**Author:** ![Bharath\_Kumar\_R](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bharath_kumar_r/32/45474_2.png) [@Bharath\_Kumar\_R](https://discuss.elastic.co/u/Bharath_Kumar_R)\
**Post date:** [June 21, 2019, 1:47pm UTC](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268/15 "2019-06-21T13:47:52Z")

</div>

Hey, I tried to fetch the anomaly results of a specific value from field SHIPPERID\_CARRIERID. But I'm still getting all the results. What needs to be corrected in the following query:

```auto
GET .ml-anomalies-.write-my_job_low_sum/_search
{
    "size": 10000,
    "query": {
            "bool": {
              "should": [
                {
                  "match": {
                    "SHIPPERID_CARRIERID": "abcd"
                  }
                }
              ], 
              "filter": [
                  { "term" : { "result_type" : "record"}},
                  { "range" : { "record_score" : { "gte": "75" } } },
                  { "range" : { "multi_bucket_impact" : { "lt": "-4" } } }
                  ]
            }
    }
}

```

How do I get only the results from the job which satisfies "SHIPPERID\_CARRIERID": "abcd"

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [June 21, 2019, 2:38pm UTC](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268/16 "2019-06-21T14:38:34Z")

</div>

```auto
GET .ml-anomalies-my_job_low_sum/_search
{
    "size": 10000,
    "query": {
            "bool": {
              "filter": [
                  { "term" : { "result_type" : "record"}},
                  { "term" : { "partition_field_value" : "abcd"}},
                  { "range" : { "record_score" : { "gte": "75" } } },
                  { "range" : { "multi_bucket_impact" : { "lt": "-4" } } }
                  ]
            }
    }
}

```

---

<div class="post-metadata">

**Author:** ![Bharath\_Kumar\_R](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bharath_kumar_r/32/45474_2.png) [@Bharath\_Kumar\_R](https://discuss.elastic.co/u/Bharath_Kumar_R)\
**Post date:** [June 26, 2019, 12:05pm UTC](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268/17 "2019-06-26T12:05:13Z")

</div>

Is it possible to create a scripted job for some particular values of a field alongside the fields SHIPPERID\_CARRIERID? Like example, I have a value field3= "xyz" and I wanted to create job with the field combinations of SHIPPERID, CARRIERID and field3="xyz"?

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [June 26, 2019, 12:48pm UTC](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268/18 "2019-06-26T12:48:45Z")

</div>

Not sure I fully understand. You want to continue to use your scripted field of `SHIPPERID_CARRIERID` but only analyze this for values of field3="xyz"?

If so, then in your datafeed, you'd need to replace the `match_all` part with a query that limits to only that field value, i.e something like.

```auto
            "bool": {
              "filter": [         
                   { "term" : { "field3" : "xyz"}}
               ]
           }

```

---

<div class="post-metadata">

**Author:** ![Bharath\_Kumar\_R](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bharath_kumar_r/32/45474_2.png) [@Bharath\_Kumar\_R](https://discuss.elastic.co/u/Bharath_Kumar_R)\
**Post date:** [June 26, 2019, 2:40pm UTC](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268/19 "2019-06-26T14:40:24Z")

</div>

Hey, I have a doubt regarding anomalies. Will the anomalies' score reduce or change as more and more data is input to the job. Right now, I'm getting over 2000 anomalies for the scripted job I ran for data with different combinations of 2 fields. Will the number of anomalies detected change over time?

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [June 27, 2019, 10:31am UTC](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268/20 "2019-06-27T10:31:04Z")

</div>

In general, yes - with more data, the more mature the modeling of that data becomes, and the more accurate the anomaly detection results get.

Plus, keep in mind that not all anomalies are created equal - use the scoring ranges in order to rank the anomalies by severity and obtain the appropriate amount of anomalies you'd like to deal with.

[Next page](https://discuss.elastic.co/t/create-jobs-with-field-combinations/186268.md?page=2)
