# Handling unique field (Other than the ID)

**URL:** <https://discuss.elastic.co/t/handling-unique-field-other-than-the-id/96034>\
**Category:** Elasticsearch\
**Created:** [August 7, 2017, 2:33am UTC](https://discuss.elastic.co/t/handling-unique-field-other-than-the-id/96034 "2017-08-07T02:33:45Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![mattkallo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mattkallo/32/109746_2.png) [@mattkallo](https://discuss.elastic.co/u/mattkallo)\
**Post date:** [August 7, 2017, 2:33am UTC](https://discuss.elastic.co/t/handling-unique-field-other-than-the-id/96034/1 "2017-08-07T02:33:46Z")

</div>

Hi All

Need help/suggestions on how to handle this

From what I understand reading documentation and previous forum postings, ES supports only the ID field as an unique value field. Pls correct me if I am wrong.

We are trying to index web pages and want to avoid duplicate indexing of URLs. It can be done by making the URL, the ID of the document, however because of some constraints with the usage of ID in rest of our application, we cannot go with URL as the ID. We need an UUID as the ID of the document.

We are considering below options - which one would be better? or are they all bad ideas and there is another way?

1. Keep URL as a separate field and then decide to insert/update by doing a lookup on the URL each time a new document is indexed
2. Maintain another index with just URL and the UUID mapping. Do a lookup in this to find the UUID for incoming URL. (each time a new document comes in)
3. Have a batch job that looks for duplicates (using aggregation). But in this case we will have duplicate documents till the batch job kicks in

We will be indexing 2-3 million document every 24 hrs and the index will have around 100 million documents in total.

Thnx

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [August 7, 2017, 11:22am UTC](https://discuss.elastic.co/t/handling-unique-field-other-than-the-id/96034/2 "2017-08-07T11:22:07Z")

</div>

If you have a good hashing algorithm (fast, few collisions as possible) on the client side, you could hash the URL and use that as the key of the document and then you would not need to have any lookup mechanism on the Elasticsearch side. Isnt that an option?

---

<div class="post-metadata">

**Author:** ![mattkallo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mattkallo/32/109746_2.png) [@mattkallo](https://discuss.elastic.co/u/mattkallo)\
**Post date:** [August 9, 2017, 4:35am UTC](https://discuss.elastic.co/t/handling-unique-field-other-than-the-id/96034/3 "2017-08-09T04:35:14Z")

</div>

@spinscale Thanks for the suggestion - will explore this option as well.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 6, 2017, 4:43am UTC](https://discuss.elastic.co/t/handling-unique-field-other-than-the-id/96034/4 "2017-09-06T04:43:19Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
