# Elasticsearch Reindex API - reindex only missing docs

**URL:** <https://discuss.elastic.co/t/elasticsearch-reindex-api-reindex-only-missing-docs/83154>\
**Category:** Elasticsearch\
**Created:** [April 21, 2017, 6:52am UTC](https://discuss.elastic.co/t/elasticsearch-reindex-api-reindex-only-missing-docs/83154 "2017-04-21T06:52:48Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Dom\_Sie](https://avatars.discourse-cdn.com/v4/letter/d/9e8a1a/32.png) [@Dom\_Sie](https://discuss.elastic.co/u/Dom_Sie)\
**Post date:** [April 21, 2017, 6:52am UTC](https://discuss.elastic.co/t/elasticsearch-reindex-api-reindex-only-missing-docs/83154/1 "2017-04-21T06:52:48Z")

</div>

Hi together,

we try to reindex big Indices ( about 10 million docs per index ) with the curl command:

```
curl -XPOST 'http://localhost:9200/_reindex?slices=5&refresh' -d '{
  "conflicts": "proceed",
  "source": {
    "index": "'.$index.'",
    "size": 10000
  },
  "dest": {
    "index": "'.$index.$version_string.'",
    "op_type": "create"
  }
}'

```

The reindex process has done a good and completely job in most indices.  
Two indices, however, the Reindex breaks off again and again.  
In an affected index are still 500 docs to reindexing and I am trying again and again to reindex the missing docs. Unfortunately unsuccessful.

How can i reindex only the missing docs between two indices or how must I have to modify my Reindex command to the effect that the process goes completely through the reindex?

Sometimes the Reindex process throws "SearchContextMissingExceptions" - if it can not resolve an Scroll-ID ; or sometimes a data-store leaves temporarily the cluster and there comes a "node\_not\_connected\_exception".

**This was my last unsuccessful try:**

count v1: 8039457  
count v2: 8038957

```
+++ COUNT IS DIFFERENT ;; Start reindexing 2016_10 for the 64 time

  % Total % Received % Xferd Average Speed Time Time Time Current
                                 Dload Upload Total Spent Left Speed
100 276 0 152 0 124 0 0 --:--:-- 1:10:11 --:--:-- 6

+++RESPONSE: {"took":4211061,"timed_out":false,"total":8039457,"updated":0,"created":0,"batches":804,"version_conflicts":8035457,"noops":0,"retries":0,"failures":[]}

```

* * *

**For more background-informations the response of the current reindex-tasks for this two indices:**

```
{
  "nodes": {
    "LG1ycx-6STKYenLnqSMZIg": {
      "name": "client_xx",
      "transport_address": "x.x.x.x:9300",
      "host": "x.x.x.x",
      "ip": "x.x.x.x:9300",
      "attributes": {
        "rack": "xxx",
        "rack_id": "xxx",
        "data": "false",
        "master": "false"
      },
      "tasks": {
        "LG1ycx-6STKYenLnqSMZIg:5559629": {
          "node": "LG1ycx-6STKYenLnqSMZIg",
          "id": 5559629,
          "type": "transport",
          "action": "indices:data/write/reindex",
          "status": {
            "total": 5150349,
            "updated": 0,
            "created": 0,
            "deleted": 0,
            "batches": 231,
            "version_conflicts": 2310000,
            "noops": 0,
            "retries": 0
          },
          "description": "",
          "start_time_in_millis": 1492754781226,
          "running_time_in_nanos": 1222719041139
        },
        "LG1ycx-6STKYenLnqSMZIg:5551067": {
          "node": "LG1ycx-6STKYenLnqSMZIg",
          "id": 5551067,
          "type": "transport",
          "action": "indices:data/write/reindex",
          "status": {
            "total": 8039457,
            "updated": 0,
            "created": 0,
            "deleted": 0,
            "batches": 706,
            "version_conflicts": 7060000,
            "noops": 0,
            "retries": 0
          },
          "description": "",
          "start_time_in_millis": 1492752458445,
          "running_time_in_nanos": 3545465607841
        }
      }
    }
  }
}
```

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [April 21, 2017, 1:47pm UTC](https://discuss.elastic.co/t/elasticsearch-reindex-api-reindex-only-missing-docs/83154/2 "2017-04-21T13:47:04Z")

</div>

> [@Dom\_Sie](#):
>
> Sometimes the Reindex process throws "SearchContextMissingExceptions" - if it can not resolve an Scroll-ID ; or sometimes a data-store leaves temporarily the cluster and there comes a "node\_not\_connected\_exception".

There isn't really anything reindex can do about this. The scrolls aren't resumable on another node. It is probably worth figuring out why this happens in your cluster and fixing it. But you should be able to work around it by chunking the reindex processes by filtering on some field in your documents. Like time or some keyword field or something.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 19, 2017, 1:48pm UTC](https://discuss.elastic.co/t/elasticsearch-reindex-api-reindex-only-missing-docs/83154/3 "2017-05-19T13:48:27Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
