# Transaction Log when making new replica

**URL:** <https://discuss.elastic.co/t/transaction-log-when-making-new-replica/37871>\
**Category:** Elasticsearch\
**Created:** [December 23, 2015, 4:20pm UTC](https://discuss.elastic.co/t/transaction-log-when-making-new-replica/37871 "2015-12-23T16:20:04Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![haldrich](https://avatars.discourse-cdn.com/v4/letter/h/b19c9b/32.png) [@haldrich](https://discuss.elastic.co/u/haldrich)\
**Post date:** [December 23, 2015, 4:20pm UTC](https://discuss.elastic.co/t/transaction-log-when-making-new-replica/37871/1 "2015-12-23T16:20:04Z")

</div>

Our system requires that we reindex potentially 1m+ records several times per day.  
To improve indexing time, we set the replicas to 0 while indexing... then, after the indexing is done, we call a flush and then set the replicas to 1.

When the new replica is created, it causes Marvel to show the index in stage "TRANSLOG" from one node to another. This lasts for some time, often over an hour.. causing the cluster to be yellow during the process.

I'm confused as to why the replica needs to play back the transaction log to be generated... the index hasn't changed when the replica is made. (it hasn't even been made an active index in our production system at the time the replica is added).

Is this the normal process for creating a replica? If so, is there some more efficient process that can be used?

Our cluster is basic: 4 proc VMs w/ 12GB RAM (6 allocated to JVM) and SAN storage (so it isn't as fast as SSDs or local 15k storage ~100MB/s throughput ). However, we can't do much about the setup we have.

Thanks!  
Heath

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [December 23, 2015, 8:55pm UTC](https://discuss.elastic.co/t/transaction-log-when-making-new-replica/37871/2 "2015-12-23T20:55:14Z")

</div>

> [@haldrich](#):
>
> When the new replica is created, it causes Marvel to show the index in stage "TRANSLOG" from one node to another. This lasts for some time, often over an hour.. causing the cluster to be yellow during the process.

That's not the reason why it's yellow, it's yellow because there are replica shards in a non-STARTED state.

> [@haldrich](#):
>
> I'm confused as to why the replica needs to play back the transaction log to be generated... the index hasn't changed when the replica is made. (it hasn't even been made an active index in our production system at the time the replica is added).
> 
> Is this the normal process for creating a replica? If so, is there some more efficient process that can be used?

When we index a document it gets sent to the primary shard, then the entire document is then sent to the replica shard and reindexed from scratch. We do not simply send the indexed outcome from the primary to the replica.  
When you add a new replica, we create that using the translog, as it holds the complete action that was done on the document.

---

<div class="post-metadata">

**Author:** ![haldrich](https://avatars.discourse-cdn.com/v4/letter/h/b19c9b/32.png) [@haldrich](https://discuss.elastic.co/u/haldrich)\
**Post date:** [December 25, 2015, 12:25am UTC](https://discuss.elastic.co/t/transaction-log-when-making-new-replica/37871/3 "2015-12-25T00:25:38Z")

</div>

Hi Mark

Thanks for the info.

So to confirm, what you're saying is that this is a normal process to be expected when adding a replica to an existing index, and there isn't any way to make it more efficient?

It is confusing that we have seen a better result by simply indexing with a refresh of -1 but a replica count of 1.

This adds about 10% to the overall indexing time but avoids the translog state and the cluster more quickly goes green, completing the replication.

I don't understand why the translog state of adding a replica takes longer than the initial indexing.  
We noted about 1mm records indexed in 18 minutes (bulk) but the similar translog process takes almost 2 hours.

Thanks again for the clarifications.  
Heath

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [December 25, 2015, 2:00am UTC](https://discuss.elastic.co/t/transaction-log-when-making-new-replica/37871/4 "2015-12-25T02:00:17Z")

</div>

Unless you are still writing to the index after you make the replica I  
wouldn't expect much translog.

So it sounds like a bug. I'd file it on github.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:28pm UTC](https://discuss.elastic.co/t/transaction-log-when-making-new-replica/37871/5 "2017-07-05T23:28:39Z")

</div>


