# Two Transform jobs overwrite result doc ID of each other (duplicate \_id)

**URL:** <https://discuss.elastic.co/t/two-transform-jobs-overwrite-result-doc-id-of-each-other-duplicate-id/303064>\
**Category:** Elasticsearch\
**Tags:** transforms\
**Created:** [April 22, 2022, 10:30pm UTC](https://discuss.elastic.co/t/two-transform-jobs-overwrite-result-doc-id-of-each-other-duplicate-id/303064 "2022-04-22T22:30:21Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![goldsky](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/goldsky/32/40178_2.png) [@goldsky](https://discuss.elastic.co/u/goldsky)\
**Post date:** [April 22, 2022, 10:30pm UTC](https://discuss.elastic.co/t/two-transform-jobs-overwrite-result-doc-id-of-each-other-duplicate-id/303064/1 "2022-04-22T22:30:21Z")

</div>

Hello,

during my journey with migration of existing, "home-made" transformations to Transform API (continuous mode), I faced one minor issue - ability to add static fields into result documents, on the beginning I tried with runtime\_mappings but it was not so effective as then I basically bypass static field in source instead of result object, then I implement such functionality using ingest pipelines, but my main issue is that 2 from 4 transform job overwrite results of each other by \_id

Goal of those jobs is to calculate aggregated data for same period (month), due to the case that different query/grouping conditions apply I can't make it with single Transform job so I split them on 4 jobs, logically they could be explained like:

1. a \> b
2. a \> c
3. b \> c
4. a \> c

where each symbol are index, each has or different query or grouping and for sure each job produce different result doc with different result fields names, but somehow last two even with different sources, conditions and result fields overwrite each other by result \_id, how I can ensure uniqueness of result doc \_id in this scenario?

I test this behaviour for several times, if I start all of them by order in list results of job 3 disappear (overwritten) by job 4, if I recreate transform job 3 and trigger it - same doc \_id that were holding data from 4 replaced with data from 3.

Job 3:

```auto
PUT _transform/foo-3
{
  "source": {
    "index": [
      "foo-hourly"
    ]
  },
  "pivot": {
    "group_by": {
      "@timestamp": {
        "date_histogram": {
          "field": "@timestamp",
          "calendar_interval": "1M"
        }
      },
    },
    "aggregations": {
      "count.percentiles": {
        "percentiles": {
          "field": "count",
          "percents": [
            50
          ]
        }
      },
      "count.max": {
        "max": {
          "field": "count"
        }
      }
    }
  },
  "frequency": "1h",
  "dest": {
    "index": "foo-monthly",
    "pipeline": "foo"
  },
  "sync": {
    "time": {
      "field": "@timestamp",
      "delay": "1d"
    }
  },
  "settings": {
    "max_page_search_size": 500
  }
}

```

Job 4:

```auto
PUT _transform/foo-4
{
  "source": {
    "index": [
      "foo"
    ],
    "query": {
      "bool": {
        "should": [
          {
              "exists": {
                "field": "boo"
              }
          }
        ],
        "minimum_should_match": 1
      }
    }
  },
  "pivot": {
    "group_by": {
      "@timestamp": {
        "date_histogram": {
          "field": "@timestamp",
          "calendar_interval": "1M"
        }
      },
      "somekey": {
        "terms": {
          "field": "somekey"
        }
      }
    },
    "aggregations": {
      "all.bucket": {
        "filter": {
          "query_string": {
            "query": "*",
            "analyze_wildcard": true
          }
        },
        "aggs": {
          "count": {
            "value_count": {
              "field": "someresult"
            }
          }
        }
      },
      "success.bucket": {
        "filter": {
          "query_string": {
            "query": "NOT someresult:2 AND NOT someresult:3 AND NOT someresult:4",
            "analyze_wildcard": true
          }
        },
        "aggs": {
          "count": {
            "value_count": {
              "field": "someresult"
            }
          }
        }
      },
      "success.rate": {
        "bucket_script": {
          "buckets_path": {
            "total_count": "all.bucket>count",
            "success_count": "success.bucket>count"
          },
          "script": "params.success_count / params.total_count * 100"
        }
      }
    }
  },
  "frequency": "1h",
  "dest": {
    "index": "foo-monthly",
    "pipeline": "foo"
  },
  "sync": {
    "time": {
      "field": "@timestamp",
      "delay": "1d"
    }
  },
  "settings": {
    "max_page_search_size": 500
  }
}

```

how I can be sure that transform jobs will not overwrite each other results?

for now I'm using "workaround" and adding and bypassing transform job id to resolve such \_id collision, like:

```auto
PUT _transform/foo-3
{
  "source": {
    "index": [
      "foo"
    ],
    "runtime_mappings": {
      "transform.id": {
        "type": "keyword",
        "script": {
          "source": "emit('foo-3')"
        }
      }
    }
  },
  "pivot": {
    "group_by": {
      "@timestamp": {
        ...
      },
      "transform.id": {
        "terms": {
          "field": "transform.id"
        }
      },
      ...
   ...
}

```

Thanks

---

<div class="post-metadata">

**Author:** ![Hendrik\_Muhs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hendrik_muhs/32/25802_2.png) [@Hendrik\_Muhs](https://discuss.elastic.co/u/Hendrik_Muhs)\
**Post date:** [April 25, 2022, 6:14am UTC](https://discuss.elastic.co/t/two-transform-jobs-overwrite-result-doc-id-of-each-other-duplicate-id/303064/2 "2022-04-25T06:14:10Z")

</div>

What you see is by design, we do not recommend to let multiple transforms write into the same destination index. Transform calculates the _doc\_id_ from the _group\_by_ bucket values, if 2 transforms produce the same bucket values, the same _doc\_id_ gets produced.

I understand your use case and your workaround looks good to me. As an alternative I can list 2 more options:

- _recommended_: write into separate indices and query using a pattern, e.g. `dest-*`. The overhead of heaving 4 instead of 1 destination index is negligible.
- calculate and overwrite the _doc\_id_ yourself by using a [fingerprint processor](https://www.elastic.co/guide/en/elasticsearch/reference/current/fingerprint-processor.html) using a ingest pipeline. Similar to transform use the bucket values of the _group\_by_ fields, not the produced values from the aggregation part. The _doc\_id_ must be created in a way that is deterministic and repeatable, so documents can get overridden. By using a different `salt` value for each transform, the individual transform results won't override each other.

In future we consider making the doc id generation in transform configurable, e.g. providing a `salt` value similar to the fingerprint processor.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 23, 2022, 6:15am UTC](https://discuss.elastic.co/t/two-transform-jobs-overwrite-result-doc-id-of-each-other-duplicate-id/303064/3 "2022-05-23T06:15:04Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
