# Reindex throttled after a few hours

**URL:** <https://discuss.elastic.co/t/reindex-throttled-after-a-few-hours/355059>\
**Category:** Elasticsearch\
**Tags:** reindex\
**Created:** [March 8, 2024, 6:49pm UTC](https://discuss.elastic.co/t/reindex-throttled-after-a-few-hours/355059 "2024-03-08T18:49:53Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![elijah\_voigt](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/elijah_voigt/32/132487_2.png) [@elijah\_voigt](https://discuss.elastic.co/u/elijah_voigt)\
**Post date:** [March 8, 2024, 6:49pm UTC](https://discuss.elastic.co/t/reindex-throttled-after-a-few-hours/355059/1 "2024-03-08T18:49:53Z")

</div>

I am on my 3rd attempt at re-indexing an index. The re-index task reliably demonstrates the following failure mode:

- The re-index task starts with high throughput, processing ~12,000 documents per second.
- After ~12-24 hours the task suddenly slows down by an order of magnitude processing ~1,500 documents per second.
- The task _may_ increase throughput after a cooldown of ~24 hours for a period of ~3-6 hours, but I have only seen this once so it may be a fluke.

This has been difficult to troubleshoot because the issue is **not** that the re-index is _slow_, it's that it is _fast_ and then _slows down_ without any intervention on my part.

# Information

## Indexes

### Source Index

Some information about the source index:

- 2.4TB of storage
- 9.1 Billion Documents
- 6 Shards
- 1 Replica

This is _not_ an ideal shard-size. One goal of this re-index is increasing the shard count.

### Destination Index

Here are some settings on the destination index:

- `routing.allocation.include._tier_preference: "data_content"`
- `refesh_interval: -1`
- `number_of_shards: 60`
- ` translog.durability: "async"`
- `number_of_replcias: 0`

Nothing is writing to the destination index outside of the re-index task, so it has `refresh_interval` disabled and replicas set to 0.

## Cluster

I am running a cluster in ElasticCloud with 19 nodes total. Here is the cluster information:

- 9 "Hot" Data Tier; `aws.data.highio.i3`: 58GB RAM, 1.69TB Disk
- 3 "Warm" Data Tier; `aws.data.highstorage.d3`: 8GB RAM, 1.48TB RAM
- 3 Coordinating; `aws.coordinating.m5d`
- 3 Master; `aws.master.r5d`
- 1 Kibana; `aws.kibana.r5d`

My understanding with ElasticCloud is that I **cannot** _increase_ the _size_ of my "hot" or "warm" instances any further, I can only _add more_ instances.

Before I started re-indexing I had 6 "Hot" instances but added 3 more for additional storage space.

## Re-Index Task

The goal of the re-index is to switch from the now deprecated `dateOptionalTime` date format to `date_optional_time`. This is a blocker for upgrading our cluster from Elasticsearch 7.17 to 8.x.

The re-index task is running on a `coordinating` instance.

Here is the status of the latest, running, re-index task:

```json
    {
          // ...
          "action" : "indices:data/write/reindex",
          "status" : {
            "total" : 9123281565,
            "updated" : 1800710708,
            "created" : 1384292,
            "deleted" : 0,
            "batches" : 1802096,
            "version_conflicts" : 0,
            "noops" : 0,
            "retries" : {
              "bulk" : 0,
              "search" : 0
            },
            "throttled_millis" : 1491643,
            "requests_per_second" : -1.0,
            "throttled_until_millis" : 0
          },
          "description" : "reindex from [source-index-name] to [destination-index-name][_doc]",
          "start_time_in_millis" : 1709747775686,
          "running_time_in_nanos" : 173147997512004,
          "cancellable" : true,
          "cancelled" : false,
          "headers" : { }
        }
      }
    }

```

> I temporarily set `requests_per_second` to a non-`-1` value which is why `throttled_millis` is non-zero.

- The re-index task started at Wednesday at 9:30AM and was reliably processing at a rate of ~12,000 documents/second.
- At 6AM Thursday morning the throughput suddenly dropped to ~1,500 documents/second. No changes to the cluster or task occurred around this time.

Before the re-index I did the following tuning:

- Destination Index: Set `replicas` to `0`
- Destination Index: Set `refresh_interval` to `-1`

I have since done the following tuning:

- Task: Set `requests_per_second` to `-1`
  - This was the default, but I set it to a non-zero value temporarily for troubleshooting.

- Destination Index: Set `translog.durability` to `async`
  - Done just a few hours ago. Increased throughput from ~1500 doc/sec to ~1600 doc/sec.

The task is currently processing ~1,600 documents/second which is still an order of magnitude slower than the initial re-index speed.

The re-index running at the current slow rate will take ~60 days to complete.  
Were it to run at the initial speed it would take ~6 days.

My question for y'all is this: **Why did the re-index slow down and what can I do to speed it up again?**

I contacted Elastic Support about this but have not gotten a straight answer to the above question and I have implemented all suggestions they had.

- Disable replicas on the destination index.
- Set `refresh_interval` to `30s` (or `-1`).
- Set `requests_per_second` to `-1`.
- Set `translog.durability` to `async` (suggestion from [this blog post](https://developers.soundcloud.com/blog/how-to-reindex-1-billion-documents-in-1-hour-at-soundcloud/)).

I'm happy to provide any more information y'all need from me to help troubleshoot. Thank you for your time.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [March 8, 2024, 7:19pm UTC](https://discuss.elastic.co/t/reindex-throttled-after-a-few-hours/355059/2 "2024-03-08T19:19:40Z")

</div>

> [@elijah\_voigt](#):
>
> routing.allocation.include.\_tier\_preference: "data\_content"

Ideally both the shards you are reading from as well as the ones you are writing to should be located on the hot nodes, so you may want to change the tier preference (unless `data_content` in reality is equivalent to `data_hot` in your cluster). I would recommend checking how the shards are distributed across the cluster using the cat shards API.

Are you specifying [slicing](https://www.elastic.co/guide/en/elasticsearch/reference/8.11/docs-reindex.html#docs-reindex-slice) when you start the reindexing task so you get some parallelism?

---

<div class="post-metadata">

**Author:** ![elijah\_voigt](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/elijah_voigt/32/132487_2.png) [@elijah\_voigt](https://discuss.elastic.co/u/elijah_voigt)\
**Post date:** [March 8, 2024, 8:46pm UTC](https://discuss.elastic.co/t/reindex-throttled-after-a-few-hours/355059/3 "2024-03-08T20:46:23Z")

</div>

> unless `data_content` in reality is equivalent to `data_hot` in your cluster

Yes, `data_content` and `data_hot` are equivalent. Someone before my time added that label.

Thanks for the tip! I will look into Slicing and follow-up.

EDIT: Do you know why the task's throughput went down without intervention? That is still confusing for me...

EDIT 2: I am starting a re-index with `slices=auto` which has resulted in 6 slices. Will report back with throughput over the weekend.

---

<div class="post-metadata">

**Author:** ![elijah\_voigt](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/elijah_voigt/32/132487_2.png) [@elijah\_voigt](https://discuss.elastic.co/u/elijah_voigt)\
**Post date:** [March 11, 2024, 4:45pm UTC](https://discuss.elastic.co/t/reindex-throttled-after-a-few-hours/355059/4 "2024-03-11T16:45:42Z")

</div>

Great news! Enabling `slices=auto` was the silver bullet I was looking for. It increased throughput of the re-index job by spinning up 6 sub-tasks and the whole job completed over the weekend.

This topic can be considered "Solved" as far as I am concerned.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 8, 2024, 4:46pm UTC](https://discuss.elastic.co/t/reindex-throttled-after-a-few-hours/355059/5 "2024-04-08T16:46:14Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
