# Shard copying performance

**URL:** https://discuss.elastic.co/t/shard-copying-performance/17262
**Category:** Elasticsearch
**Created:** [April 29, 2014, 1:50pm UTC](https://discuss.elastic.co/t/shard-copying-performance/17262 "2014-04-29T13:50:05Z")
**Posts on this page:** 5
**Page:** 1

<div class="post-metadata">

### Author: ![Michael\_Salmon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/michael_salmon/32/5330_2.png) [@Michael\_Salmon](https://discuss.elastic.co/u/Michael_Salmon)
#### Post date: [April 29, 2014, 1:50pm UTC](https://discuss.elastic.co/t/shard-copying-performance/17262/1 "2014-04-29T13:50:05Z")

</div>

I am having trouble replicating a shard and I cannot see any possible  
reason for it. After 15 minutes I get a timeout in phase 2.

The shard isn't that large about 60,000K, 5GB and 22 segments and the  
translog directories are empty.  
The computers in question are lightly loaded as is the network between them.  
Copying all the files in the shard from all 4 disks between the two  
computers with rsync takes about 40 seconds.  
I can't run checkIndex on the source machine as it can't handle shards that  
are spread over multiple disks but it runs quite happily on the files I  
copied with rsync although it took a bit over 12 minutes to run the check.  
I have ES 1.1.0 installed.  
I changed some settings but none of them seem to make much difference:

"transient": {  
"logger": {  
"level": "TRACE"  
},  
"indices": {  
"store": {  
"throttle": {  
"type": "none"  
}  
},  
"recovery": {  
"translog\_size": "256MB",  
"concurrent\_streams": "16",  
"translog\_ops": "10000",  
"max\_bytes\_per\_sec": "250MB"  
}  
}  
}

Does anyone have any tips on how I should proceed?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/a85c76cb-72d5-45c4-82cf-d8c8867a2151%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/a85c76cb-72d5-45c4-82cf-d8c8867a2151%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

### Author: ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)
#### Post date: [May 5, 2014, 10:25am UTC](https://discuss.elastic.co/t/shard-copying-performance/17262/2 "2014-05-05T10:25:50Z")

</div>

Hey,

you could change your default loglevel to find out, if those settings are  
actually applied (either DEBUG or TRACE). Depending on the elasticsearch  
version you are using, you might want to try with a lower-cased setting of  
max\_bytes\_per\_sec and set it to "250mb". Also, can you show the exception  
which contains the "timeout in phase 2"?

--Alex

On Tue, Apr 29, 2014 at 3:50 PM, Michael Salmon [michael.salmon@inovia.nu](mailto:michael.salmon@inovia.nu)wrote:

> I am having trouble replicating a shard and I cannot see any possible  
> reason for it. After 15 minutes I get a timeout in phase 2.
> 
> The shard isn't that large about 60,000K, 5GB and 22 segments and the  
> translog directories are empty.  
> The computers in question are lightly loaded as is the network between  
> them.  
> Copying all the files in the shard from all 4 disks between the two  
> computers with rsync takes about 40 seconds.  
> I can't run checkIndex on the source machine as it can't handle shards  
> that are spread over multiple disks but it runs quite happily on the files  
> I copied with rsync although it took a bit over 12 minutes to run the check.  
> I have ES 1.1.0 installed.  
> I changed some settings but none of them seem to make much difference:
> 
> "transient": {  
> "logger": {  
> "level": "TRACE"  
> },  
> "indices": {  
> "store": {  
> "throttle": {  
> "type": "none"  
> }  
> },  
> "recovery": {  
> "translog\_size": "256MB",  
> "concurrent\_streams": "16",  
> "translog\_ops": "10000",  
> "max\_bytes\_per\_sec": "250MB"  
> }  
> }  
> }
> 
> Does anyone have any tips on how I should proceed?
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/a85c76cb-72d5-45c4-82cf-d8c8867a2151%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/a85c76cb-72d5-45c4-82cf-d8c8867a2151%40googlegroups.com)[https://groups.google.com/d/msgid/elasticsearch/a85c76cb-72d5-45c4-82cf-d8c8867a2151%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/a85c76cb-72d5-45c4-82cf-d8c8867a2151%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAGCwEM9gmWVVV8y8FtqC4ESVkjPoc4Giqp4feX2x4znEBDaYyg%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAGCwEM9gmWVVV8y8FtqC4ESVkjPoc4Giqp4feX2x4znEBDaYyg%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

### Author: ![Michael\_Salmon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/michael_salmon/32/5330_2.png) [@Michael\_Salmon](https://discuss.elastic.co/u/Michael_Salmon)
#### Post date: [May 5, 2014, 11:55am UTC](https://discuss.elastic.co/t/shard-copying-performance/17262/3 "2014-05-05T11:55:48Z")

</div>

This is the exception that I posted earlier:

[2014-04-28 13:40:15,039][WARN][cluster.action.shard] [eis05]  
[ds\_clearcase-vob-heat-analyzer][2] sending failed shard for  
[ds\_clearcase-vob-heat-analyzer][2], node[QyeTlW2YQbG27zrsdjBBGA], [R],  
s[INITIALIZING], indexUUID [ms7jQeuMQduNIHCmjxsKjQ], reason [Failed to  
start shard, message  
[RecoveryFailedException[[ds\_clearcase-vob-heat-analyzer][2]: Recovery  
failed from [eis09][p8-\_fzHeTR22pSlsBsYm8A][eis09.rnditlab.ericsson.se][inet[/137.58.184.239:9300]]{datacenter=PoCC}  
into [eis05][QyeTlW2YQbG27zrsdjBBGA][eis05.rnditlab.ericsson.se][inet[  
eis05.rnditlab.ericsson.se/137.58.184.235:9300]]{datacenter=PoCC}[http://eis05.rnditlab.ericsson.se/137.58.184.235:9300]]{datacenter=PoCC}](http://eis05.rnditlab.ericsson.se/137.58.184.235:9300%5D%5D%7Bdatacenter=PoCC%7D)];  
nested:  
RemoteTransportException[[eis09][inet[/137.58.184.239:9300]][index/shard/recovery/startRecovery]];  
nested: RecoveryEngineException[[ds\_clearcase-vob-heat-analyzer][2]  
Phase[2] Execution failed]; nested:  
ReceiveTimeoutTransportException[[eis05][inet[/137.58.184.235:9300]][index/shard/recovery/prepareTranslog]  
request\_id [6809886] timed out after [900000ms]]; ]]  
[2014-04-28 14:00:11,614][WARN][indices.cluster] [eis05]  
[ds\_clearcase-vob-heat-analyzer][0] failed to start shard  
org.elasticsearch.indices.recovery.RecoveryFailedException:  
[ds\_clearcase-vob-heat-analyzer][0]: Recovery failed from  
[eis07][Q8ZWgDIXRGiUej1oMoH8Jg][eis07.rnditlab.ericsson.se][inet[/137.58.184.237:9300]]{datacenter=PoCC}  
into [eis05][QyeTlW2YQbG27zrsdjBBGA][eis05.rnditlab.ericsson.se][inet[  
eis05.rnditlab.ericsson.se/137.58.184.235:9300]]{datacenter=PoCC}[http://eis05.rnditlab.ericsson.se/137.58.184.235:9300]]{datacenter=PoCC}](http://eis05.rnditlab.ericsson.se/137.58.184.235:9300%5D%5D%7Bdatacenter=PoCC%7D)  
at  
org.elasticsearch.indices.recovery.RecoveryTarget.doRecovery(RecoveryTarget.java:307)  
at  
org.elasticsearch.indices.recovery.RecoveryTarget.access$300(RecoveryTarget.java:65)  
at  
org.elasticsearch.indices.recovery.RecoveryTarget$3.run(RecoveryTarget.java:184)  
at  
java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)  
at  
java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)  
at java.lang.Thread.run(Thread.java:744)  
Caused by: org.elasticsearch.transport.RemoteTransportException:  
[eis07][inet[/137.58.184.237:9300]][index/shard/recovery/startRecovery]  
Caused by: org.elasticsearch.index.engine.RecoveryEngineException:  
[ds\_clearcase-vob-heat-analyzer][0] Phase[2] Execution failed  
at  
org.elasticsearch.index.engine.internal.InternalEngine.recover(InternalEngine.java:1098)  
at  
org.elasticsearch.index.shard.service.InternalIndexShard.recover(InternalIndexShard.java:627)  
at  
org.elasticsearch.indices.recovery.RecoverySource.recover(RecoverySource.java:117)  
at  
org.elasticsearch.indices.recovery.RecoverySource.access$1600(RecoverySource.java:61)  
at  
org.elasticsearch.indices.recovery.RecoverySource$StartRecoveryTransportRequestHandler.messageReceived(RecoverySource.java:337)  
at  
org.elasticsearch.indices.recovery.RecoverySource$StartRecoveryTransportRequestHandler.messageReceived(RecoverySource.java:323)  
at  
org.elasticsearch.transport.netty.MessageChannelHandler$RequestHandler.run(MessageChannelHandler.java:270)  
at  
java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)  
at  
java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)  
at java.lang.Thread.run(Thread.java:744)  
Caused by: org.elasticsearch.transport.ReceiveTimeoutTransportException:  
[eis05][inet[/137.58.184.235:9300]][index/shard/recovery/prepareTranslog]  
request\_id [154592652] timed out after [900000ms]  
at  
org.elasticsearch.transport.TransportService$TimeoutHandler.run(TransportService.java:356)  
... 3 more

I checked that the max\_bytes\_per\_sec changed in the log, it accepts both MB  
and mb.

I am also changing my log level to trace but restarting the servers takes a  
long while.

On Monday, 5 May 2014 12:25:50 UTC+2, Alexander Reelsen wrote:

> Hey,
> 
> you could change your default loglevel to find out, if those settings are  
> actually applied (either DEBUG or TRACE). Depending on the elasticsearch  
> version you are using, you might want to try with a lower-cased setting of  
> max\_bytes\_per\_sec and set it to "250mb". Also, can you show the exception  
> which contains the "timeout in phase 2"?
> 
> --Alex
> 
> On Tue, Apr 29, 2014 at 3:50 PM, Michael Salmon \<michael...@inovia.nu\<javascript:\>
> 
> > wrote:
> 
> > I am having trouble replicating a shard and I cannot see any possible  
> > reason for it. After 15 minutes I get a timeout in phase 2.
> > 
> > The shard isn't that large about 60,000K, 5GB and 22 segments and the  
> > translog directories are empty.  
> > The computers in question are lightly loaded as is the network between  
> > them.  
> > Copying all the files in the shard from all 4 disks between the two  
> > computers with rsync takes about 40 seconds.  
> > I can't run checkIndex on the source machine as it can't handle shards  
> > that are spread over multiple disks but it runs quite happily on the files  
> > I copied with rsync although it took a bit over 12 minutes to run the check.  
> > I have ES 1.1.0 installed.  
> > I changed some settings but none of them seem to make much difference:
> > 
> > "transient": {  
> > "logger": {  
> > "level": "TRACE"  
> > },  
> > "indices": {  
> > "store": {  
> > "throttle": {  
> > "type": "none"  
> > }  
> > },  
> > "recovery": {  
> > "translog\_size": "256MB",  
> > "concurrent\_streams": "16",  
> > "translog\_ops": "10000",  
> > "max\_bytes\_per\_sec": "250MB"  
> > }  
> > }  
> > }
> > 
> > Does anyone have any tips on how I should proceed?
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/a85c76cb-72d5-45c4-82cf-d8c8867a2151%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/a85c76cb-72d5-45c4-82cf-d8c8867a2151%40googlegroups.com)[https://groups.google.com/d/msgid/elasticsearch/a85c76cb-72d5-45c4-82cf-d8c8867a2151%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/a85c76cb-72d5-45c4-82cf-d8c8867a2151%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > .  
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/c04725d9-ef92-4c67-ac33-cb8fd96def06%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/c04725d9-ef92-4c67-ac33-cb8fd96def06%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

### Author: ![Michael\_Salmon](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/michael_salmon/32/5330_2.png) [@Michael\_Salmon](https://discuss.elastic.co/u/Michael_Salmon)
#### Post date: [March 17, 2015, 12:50pm UTC](https://discuss.elastic.co/t/shard-copying-performance/17262/4 "2015-03-17T12:50:56Z")

</div>

We recently removed index.shard.check\_on\_startup:fix from our settings and  
haven't had this problem since. The guide says "Should shard consistency be  
checked upon opening" but it appears to also affect replication. I'm not  
going to say that that is wrong although it isn't what I want but I think  
that the guide should be more explicit as to when the checking is done.

On Tuesday, 29 April 2014 15:50:05 UTC+2, Michael Salmon wrote:

> I am having trouble replicating a shard and I cannot see any possible  
> reason for it. After 15 minutes I get a timeout in phase 2.
> 
> The shard isn't that large about 60,000K, 5GB and 22 segments and the  
> translog directories are empty.  
> The computers in question are lightly loaded as is the network between  
> them.  
> Copying all the files in the shard from all 4 disks between the two  
> computers with rsync takes about 40 seconds.  
> I can't run checkIndex on the source machine as it can't handle shards  
> that are spread over multiple disks but it runs quite happily on the files  
> I copied with rsync although it took a bit over 12 minutes to run the check.  
> I have ES 1.1.0 installed.  
> I changed some settings but none of them seem to make much difference:
> 
> "transient": {  
> "logger": {  
> "level": "TRACE"  
> },  
> "indices": {  
> "store": {  
> "throttle": {  
> "type": "none"  
> }  
> },  
> "recovery": {  
> "translog\_size": "256MB",  
> "concurrent\_streams": "16",  
> "translog\_ops": "10000",  
> "max\_bytes\_per\_sec": "250MB"  
> }  
> }  
> }
> 
> Does anyone have any tips on how I should proceed?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/bde6fb91-7b3c-42d7-8e31-7fdb7bd5555b%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/bde6fb91-7b3c-42d7-8e31-7fdb7bd5555b%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 6, 2017, 12:26am UTC](https://discuss.elastic.co/t/shard-copying-performance/17262/5 "2017-07-06T00:26:17Z")

</div>


