# REINDEX API - Node choice when using TASK API

**URL:** <https://discuss.elastic.co/t/reindex-api-node-choice-when-using-task-api/89595>\
**Category:** Elasticsearch\
**Created:** [June 15, 2017, 4:38pm UTC](https://discuss.elastic.co/t/reindex-api-node-choice-when-using-task-api/89595 "2017-06-15T16:38:41Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![timost](https://avatars.discourse-cdn.com/v4/letter/t/e95f7d/32.png) [@timost](https://discuss.elastic.co/u/timost)\
**Post date:** [June 15, 2017, 4:38pm UTC](https://discuss.elastic.co/t/reindex-api-node-choice-when-using-task-api/89595/1 "2017-06-15T16:38:41Z")

</div>

I'm running an ES 5.4.1 cluster with only two nodes. One uses HDDs (2TB) the other uses SSDs (500GB), both can be masters and data nodes. The HDD node holds roughly 75% of the data and the remaining is on the SSD node.

I'm in the process of reindexing the data on the cluster using the REINDEX API and the TASK API with slices as described in the [documentation](https://www.elastic.co/guide/en/elasticsearch/reference/5.4/docs-reindex.html#docs-reindex-task-api). All the reindexing tasks seem to run on the HDD node which I think make them slower than if they would be ran on the SSD node.

My questions are the following:

1. How does Elasticsearch choose the node on which to run the reindex task ?  
My guess would be that it runs on the node where the shard is but I could not find anything regarding this in the doc.

2. Is there a way to force the task to be run on a given node of the cluster ?

Thank you for your help !

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [June 20, 2017, 7:43am UTC](https://discuss.elastic.co/t/reindex-api-node-choice-when-using-task-api/89595/2 "2017-06-20T07:43:19Z")

</div>

It'll reindex on whichever nodes it needs to based on the the allocation of the shards, as you thought.

If you want to force it use allocation awareness/filtering to put the index where you want it.

---

<div class="post-metadata">

**Author:** ![timost](https://avatars.discourse-cdn.com/v4/letter/t/e95f7d/32.png) [@timost](https://discuss.elastic.co/u/timost)\
**Post date:** [June 20, 2017, 7:58am UTC](https://discuss.elastic.co/t/reindex-api-node-choice-when-using-task-api/89595/3 "2017-06-20T07:58:35Z")

</div>

Thank you for your answer @warkolm , I'll try this !

---

<div class="post-metadata">

**Author:** ![timost](https://avatars.discourse-cdn.com/v4/letter/t/e95f7d/32.png) [@timost](https://discuss.elastic.co/u/timost)\
**Post date:** [June 23, 2017, 7:14am UTC](https://discuss.elastic.co/t/reindex-api-node-choice-when-using-task-api/89595/4 "2017-06-23T07:14:56Z")

</div>

In case anyone is interested, I've been able to try this procedure. Reindexing on the SSD node was roughly 5 times faster than on the HDD node ( 1500-2000 Doc/s vs 10 000 Doc/s) see the screenshot below.

 ![](https://us1.discourse-cdn.com/elastic/original/3X/e/d/ed2714ba74cd75c8672b73a27566464843a46b0a.png)

The things I had to do in order to make this happen are:

- Reroute source index shards to the SSD node using the [cluster reroute api](https://www.elastic.co/guide/en/elasticsearch/reference/current/cluster-reroute.html) or simply the [shard allocation filtering](https://www.elastic.co/guide/en/elasticsearch/reference/current/shard-allocation-filtering.html) mecanism. The reroute api is useful to understand why a shard isn't relocated.
- Check indices settings do not conflict with the reroute order. Especially check that replica shards aren't on the target node
- Make sure target indices are also on the SSD node (there is a bottleneck on both read and write operations on disks)
- Increase `cluster.routing.allocation.node_concurrent_recoveries` to make rerouting shards faster

One thing that amazed me is that you can relocate the shards _while_ the reindexing process on those shards is running.

In my particular case, I had limited storage available on the SSD node so my workflow was:

1. Move all the shards to be reindexed to the SSD node.
2. Move or create the target indices on the SSD node.
3. Start the reindexing process.
4. Delete the source index.
5. Move the target indices back to the HDD node if needed to save space.
6. Repeat for next index.

Even with the overhead of relocating the shards (~5 minutes per index), it was still _much_ faster to do this than to reindex on the HDD node.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [June 23, 2017, 7:31am UTC](https://discuss.elastic.co/t/reindex-api-node-choice-when-using-task-api/89595/5 "2017-06-23T07:31:00Z")

</div>

Great to hear, thanks for sharing such a detailed outcome too!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 21, 2017, 7:31am UTC](https://discuss.elastic.co/t/reindex-api-node-choice-when-using-task-api/89595/6 "2017-07-21T07:31:05Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
