# Exactly-once guarantee for Spark Structured Streaming

**URL:** <https://discuss.elastic.co/t/exactly-once-guarantee-for-spark-structured-streaming/200746>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [September 23, 2019, 7:49pm UTC](https://discuss.elastic.co/t/exactly-once-guarantee-for-spark-structured-streaming/200746 "2019-09-23T19:49:37Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![danielyahn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/danielyahn/32/41575_2.png) [@danielyahn](https://discuss.elastic.co/u/danielyahn)\
**Post date:** [September 23, 2019, 7:49pm UTC](https://discuss.elastic.co/t/exactly-once-guarantee-for-spark-structured-streaming/200746/1 "2019-09-23T19:49:37Z")

</div>

The [following blog](https://www.elastic.co/blog/structured-streaming-elasticsearch-for-hadoop-6-0) that @james.baiera wrote says,

> Users of ES-Hadoop can specify one of their fields to be used as the document’s ID, and Elasticsearch manages ID based writes consistently. We can’t know ahead of time what your documents’ IDs are, so it’s up to each user to ensure their streaming data contains an ID of some sort.

Are there some best practices on how to generate the ID for time series data and to optimize for ingest? I'm starting by introducing UUID v4, but seeing that there are better implementation. Such as [one](http://blog.mikemccandless.com/2014/05/choosing-fast-unique-identifier-uuid.html) and [two](https://github.com/elastic/elasticsearch/issues/33049#issuecomment-415296769).

I already understand that using [auto-generated ID](https://www.elastic.co/guide/en/elasticsearch/reference/master/tune-for-indexing-speed.html#_use_auto_generated_ids) will skip duplicate check, thus saving lookup cost. Here I'm looking for an ID generation scheme for exactly-once guarantee and ingest performance.

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [September 23, 2019, 8:26pm UTC](https://discuss.elastic.co/t/exactly-once-guarantee-for-spark-structured-streaming/200746/2 "2019-09-23T20:26:36Z")

</div>

@danielyahn your linked options pretty much hit the nail on the head as far as I can see. UUID v4 is a great way to ensure uniqueness of the ID without too much thinking about the implementation. Since it's so easy, I would suggest giving it a shot and seeing the performance implications before spending too much time on other ID strategies.

I've seen a few ID creation techniques across multiple projects. Making a compound ID out of other field values on a document in my experience provides a decent mix of ease of design and performance, though ensuring the uniqueness can be an issue if you don't know your data inside and out.

---

<div class="post-metadata">

**Author:** ![danielyahn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/danielyahn/32/41575_2.png) [@danielyahn](https://discuss.elastic.co/u/danielyahn)\
**Post date:** [September 23, 2019, 9:07pm UTC](https://discuss.elastic.co/t/exactly-once-guarantee-for-spark-structured-streaming/200746/3 "2019-09-23T21:07:49Z")

</div>

Thanks @james.baiera for a quick reply.

I'm seeing 20% drop in performance with the UUID V4, which seems to be aligned with the benchmark shown [here](https://github.com/elastic/elasticsearch/issues/33049) (25%).

Could you elaborate on ID creation techniques you've seen? I'm interested in learning what was done to improve the performance.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 21, 2019, 9:07pm UTC](https://discuss.elastic.co/t/exactly-once-guarantee-for-spark-structured-streaming/200746/4 "2019-10-21T21:07:51Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
