# ML: difference between partition\_field\_name and by\_field\_name?

**URL:** <https://discuss.elastic.co/t/ml-difference-between-partition-field-name-and-by-field-name/280013>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-machine-learning\
**Created:** [July 29, 2021, 8:13pm UTC](https://discuss.elastic.co/t/ml-difference-between-partition-field-name-and-by-field-name/280013 "2021-07-29T20:13:08Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![CameronCenic](https://avatars.discourse-cdn.com/v4/letter/c/df788c/32.png) [@CameronCenic](https://discuss.elastic.co/u/CameronCenic)\
**Post date:** [July 29, 2021, 8:13pm UTC](https://discuss.elastic.co/t/ml-difference-between-partition-field-name-and-by-field-name/280013/1 "2021-07-29T20:13:08Z")

</div>

I do not really understand the difference between these two settings. They seem to perform the same function from my perspective.

---

<div class="post-metadata">

**Author:** ![BenTrent](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bentrent/32/33915_2.png) [@BenTrent](https://discuss.elastic.co/u/BenTrent)\
**Post date:** [July 29, 2021, 8:28pm UTC](https://discuss.elastic.co/t/ml-difference-between-partition-field-name-and-by-field-name/280013/2 "2021-07-29T20:28:38Z")

</div>

This other discuss thread may shed some light: [ML Kibana: difference between by\_field\_name and partition\_field\_name - #4 by richcollier](https://discuss.elastic.co/t/ml-kibana-difference-between-by-field-name-and-partition-field-name/193385/4)

---

<div class="post-metadata">

**Author:** ![CameronCenic](https://avatars.discourse-cdn.com/v4/letter/c/df788c/32.png) [@CameronCenic](https://discuss.elastic.co/u/CameronCenic)\
**Post date:** [July 29, 2021, 8:45pm UTC](https://discuss.elastic.co/t/ml-difference-between-partition-field-name-and-by-field-name/280013/3 "2021-07-29T20:45:58Z")

</div>

Alright thanks for the link. As I understand it, the partition\_field\_name is going to be a harder split in the model, then? So if I want the anomaly scores to be solely based on data matching the split field, I should use partition\_field\_name. And I should only use by\_field\_name if I want a softer split that is going to let data from the whole population affect anomaly scores.

---

<div class="post-metadata">

**Author:** ![richcollier](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/richcollier/32/115035_2.png) [@richcollier](https://discuss.elastic.co/u/richcollier)\
**Post date:** [July 30, 2021, 2:01pm UTC](https://discuss.elastic.co/t/ml-difference-between-partition-field-name-and-by-field-name/280013/4 "2021-07-30T14:01:24Z")

</div>

Yes, that's pretty much it. Think of using `partition_field_name` as practically the equivalent of N number of single metric jobs, one for every value of `partition_field_name` (with a cardinality of N). The scoring for anomalies in a partition ([since version 6.5](https://www.elastic.co/blog/changes-to-elastic-machine-learning-anomaly-scoring-in-6-5)) is very independent of anomalies in other partitions.

So, utilize `partition_field_name` for logical splits that should be more independent from each other.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 27, 2021, 2:01pm UTC](https://discuss.elastic.co/t/ml-difference-between-partition-field-name-and-by-field-name/280013/5 "2021-08-27T14:01:29Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
