# Duplicate in Dataset while reading from elasticsearch index with SPARK

**URL:** <https://discuss.elastic.co/t/duplicate-in-dataset-while-reading-from-elasticsearch-index-with-spark/176379>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [April 11, 2019, 10:32am UTC](https://discuss.elastic.co/t/duplicate-in-dataset-while-reading-from-elasticsearch-index-with-spark/176379 "2019-04-11T10:32:56Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![pverma0312](https://avatars.discourse-cdn.com/v4/letter/p/45deac/32.png) [@pverma0312](https://discuss.elastic.co/u/pverma0312)\
**Post date:** [April 11, 2019, 10:32am UTC](https://discuss.elastic.co/t/duplicate-in-dataset-while-reading-from-elasticsearch-index-with-spark/176379/1 "2019-04-11T10:32:57Z")

</div>

Hi,

I am trying to read an unstructured index from es and loading it to RDD of SPARK, sometimes in RDD the complete data is getting duplicated, did somebody faced the similar issue, any recommendation.

Thanks,  
Prashant Verma

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 9, 2019, 10:33am UTC](https://discuss.elastic.co/t/duplicate-in-dataset-while-reading-from-elasticsearch-index-with-spark/176379/2 "2019-05-09T10:33:03Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
