# Duplicates result with elasticsearch hadoop spark

**URL:** <https://discuss.elastic.co/t/duplicates-result-with-elasticsearch-hadoop-spark/83794>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [April 27, 2017, 5:28am UTC](https://discuss.elastic.co/t/duplicates-result-with-elasticsearch-hadoop-spark/83794 "2017-04-27T05:28:23Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![suanmeiguo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/suanmeiguo/32/11758_2.png) [@suanmeiguo](https://discuss.elastic.co/u/suanmeiguo)\
**Post date:** [April 27, 2017, 5:28am UTC](https://discuss.elastic.co/t/duplicates-result-with-elasticsearch-hadoop-spark/83794/1 "2017-04-27T05:28:23Z")

</div>

Elasticsearch cluster: 5.2.2  
es-hadoop: 5.2.2 java

```auto
<dependency>
      <groupId>org.elasticsearch</groupId>
      <artifactId>elasticsearch-spark-20_2.11</artifactId>
      <version>5.2.2</version>
    </dependency>
    <dependency>
      <groupId>org.elasticsearch</groupId>
      <artifactId>elasticsearch</artifactId>
      <version>5.2.2</version>
    </dependency>
    <dependency>
      <groupId>org.elasticsearch.client</groupId>
      <artifactId>transport</artifactId>
      <version>5.2.2</version>
    </dependency>

```

I have a query built by elasticsearch dsl (using `QueryBuilder`), if I print that query and run directly in elasticsearch dev tool, it shows me 200M result. But if I run it through es-hadoop spark and save them on s3, I got duplicated records. One thing worth notice is: there are same number of records exported to s3 (which means some duplicates replaced some other records, I verified this).

Here are my code snippet, there's no transformation I just dump the records to s3 directly:

```auto
JavaPairRDD<String, Map<String, Object>> esRDD = JavaEsSpark.esRDD(sc, "index_name/type_name", "query_str");
esRDD.saveAsTextFile("a_path_to_s3");

```

In the result data on s3, I find duplicated rows, indicated by same \_id, which I set explicitly in elasticsearch (so not using auto generated \_id).

Here is my query\_str

```auto
{
  "query": {
    "bool": {
      "filter": [
        {
          "range": {
            "integer_field_1": {
              "from": 15,
              "to": null,
              "include_lower": true,
              "include_upper": true,
              "boost": 1.0
            }
          }
        }
      ],
      "should": [
        {
          "range": {
            "date_field_1": {
              "from": "now-100d/d",
              "to": null,
              "include_lower": false,
              "include_upper": true,
              "boost": 1.0
            }
          }
        },
        {
          "bool": {
            "must": [
              {
                "range": {
                  "date_field_1": {
                    "from": "now-300d/d",
                    "to": null,
                    "include_lower": false,
                    "include_upper": true,
                    "boost": 1.0
                  }
                }
              },
              {
                "range": {
                  "integer_field_2": {
                    "from": 500,
                    "to": null,
                    "include_lower": false,
                    "include_upper": true,
                    "boost": 1.0
                  }
                }
              }
            ],
            "disable_coord": false,
            "adjust_pure_negative": true,
            "boost": 1.0
          }
        },
        {
          "range": {
            "integer_field_2": {
              "from": 2000,
              "to": null,
              "include_lower": false,
              "include_upper": true,
              "boost": 1.0
            }
          }
        }
      ],
      "disable_coord": false,
      "adjust_pure_negative": true,
      "minimum_should_match": "1",
      "boost": 1.0
    }
  },
  "_source": {
    "includes": [
      "some_field_1",
      "some_field_2"
    ],
    "excludes": []
  }
}

```

Any helps are appreciated!

---

<div class="post-metadata">

**Author:** ![suanmeiguo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/suanmeiguo/32/11758_2.png) [@suanmeiguo](https://discuss.elastic.co/u/suanmeiguo)\
**Post date:** [April 27, 2017, 4:29pm UTC](https://discuss.elastic.co/t/duplicates-result-with-elasticsearch-hadoop-spark/83794/2 "2017-04-27T16:29:41Z")

</div>

I've tried with `Dataset<Row>` and I got duplicate as well. Here's my Dataset code:

```auto
Dataset<Row> df = sqlContext.read().format("es").load("index_name/type_name");
df.write().mode(SaveMode.Overwrite).csv("a_path_to_s3");

```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 25, 2017, 4:31pm UTC](https://discuss.elastic.co/t/duplicates-result-with-elasticsearch-hadoop-spark/83794/3 "2017-05-25T16:31:41Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
