# Ingest pipeline: \_id generation

**URL:** <https://discuss.elastic.co/t/ingest-pipeline-id-generation/149189>\
**Category:** Elasticsearch\
**Created:** [September 19, 2018, 8:11pm UTC](https://discuss.elastic.co/t/ingest-pipeline-id-generation/149189 "2018-09-19T20:11:04Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![jetnet](https://avatars.discourse-cdn.com/v4/letter/j/a87d85/32.png) [@jetnet](https://discuss.elastic.co/u/jetnet)\
**Post date:** [September 19, 2018, 8:11pm UTC](https://discuss.elastic.co/t/ingest-pipeline-id-generation/149189/1 "2018-09-19T20:11:04Z")

</div>

is it safe to generate doc IDs in a pipeline as following:

```
  {
    "script": {
      "ignore_failure": false,
      "lang": "painless",
      "source": "ctx._id = 'prfx-'.concat(ctx.message.hashCode().toString())"
    }
  }
```

---

<div class="post-metadata">

**Author:** ![jetnet](https://avatars.discourse-cdn.com/v4/letter/j/a87d85/32.png) [@jetnet](https://discuss.elastic.co/u/jetnet)\
**Post date:** [September 19, 2018, 8:46pm UTC](https://discuss.elastic.co/t/ingest-pipeline-id-generation/149189/2 "2018-09-19T20:46:09Z")

</div>

or is it better to use UUID:

```
  {
    "script": {
      "ignore_failure": false,
      "lang": "painless",
      "source": """
      	 char[] buffer = ctx.message.toCharArray();
      	 byte[] b = new byte[buffer.length];
      	 for (int i = 0; i < b.length; i++) {
      	  b[i] = (byte) buffer[i];
      	 }
          ctx._id= UUID.nameUUIDFromBytes(b).toString();
      """
    }
  }
```

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [September 20, 2018, 11:57am UTC](https://discuss.elastic.co/t/ingest-pipeline-id-generation/149189/3 "2018-09-20T11:57:32Z")

</div>

both approaches can lead to duplicate IDs and overwrite data. is this something you are ok with? Could you use the automate ID generation within Elasticsearch?

---

<div class="post-metadata">

**Author:** ![jetnet](https://avatars.discourse-cdn.com/v4/letter/j/a87d85/32.png) [@jetnet](https://discuss.elastic.co/u/jetnet)\
**Post date:** [September 20, 2018, 1:45pm UTC](https://discuss.elastic.co/t/ingest-pipeline-id-generation/149189/4 "2018-09-20T13:45:05Z")

</div>

I need to re-index **some pieces** of data quite often, that's why I'd like to compute a hash based on the whole message in order to avoid duplicates.  
Of course, I don't want to overwrite data by creating non-unique IDs too.  
Do you know, if there are other hash-methods available in [painless API](https://www.elastic.co/guide/en/elasticsearch/painless/current/painless-api-reference.html)?

---

<div class="post-metadata">

**Author:** ![jetnet](https://avatars.discourse-cdn.com/v4/letter/j/a87d85/32.png) [@jetnet](https://discuss.elastic.co/u/jetnet)\
**Post date:** [September 21, 2018, 9:28am UTC](https://discuss.elastic.co/t/ingest-pipeline-id-generation/149189/5 "2018-09-21T09:28:33Z")

</div>

I ended up with a "paranoid" ID generator - `_id = MD5(text) + "." + MD5(base64(text))`:

Stored script `generate_id`:

```
POST _scripts/generate_id
{
  "script": {
	"lang": "painless",
	"source": """
		// function getUUID
		String getUUID(def str) {
		  def res = null;
		  if (str != null) {
			 char[] buffer = str.toCharArray();
			 byte[] b = new byte[buffer.length];
			 for (int i = 0; i < b.length; i++) {
			  b[i] = (byte) buffer[i];
			 }
			 res = UUID.nameUUIDFromBytes(b).toString();
		  } else {
			  // randomUUID does not work
			  //res = UUID.randomUUID().toString();
			  res = Math.random().toString();
		  }
		  return res;
		}
		
		//
		// Main
		// params.field - the field name.
		// Note: "doted" names (like "content.message") will not work.
		//
		if (ctx.containsKey(params.field) && !ctx[params.field].isEmpty()) {
		  ctx._id = getUUID(ctx[params.field]) + "." + getUUID(ctx[params.field].encodeBase64());
		} else {
		  ctx._id = params.field + "_empty_" + getUUID(null);
		}
	"""
  }
}

```

how to call it in a pipeline:

```
"processors": [
...
{
    "script": {
      "ignore_failure": false,
        "id": "generate_id",
        "params": {
             "field": "message"
        }
    }
},
...
```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 19, 2018, 9:28am UTC](https://discuss.elastic.co/t/ingest-pipeline-id-generation/149189/6 "2018-10-19T09:28:33Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
