# Failed to indices:data/write/bulk\[s\] on replica because of Netty4TcpChannel / CompositeBytesReference more than 2GB

**URL:** https://discuss.elastic.co/t/failed-to-indices-data-write-bulk-s-on-replica-because-of-netty4tcpchannel-compositebytesreference-more-than-2gb/349797
**Category:** Elasticsearch
**Created:** [December 21, 2023, 1:20pm UTC](https://discuss.elastic.co/t/failed-to-indices-data-write-bulk-s-on-replica-because-of-netty4tcpchannel-compositebytesreference-more-than-2gb/349797 "2023-12-21T13:20:38Z")
**Posts on this page:** 7
**Page:** 1

<div class="post-metadata">

### Author: ![Martin\_Berlin](https://avatars.discourse-cdn.com/v4/letter/m/958977/32.png) [@Martin\_Berlin](https://discuss.elastic.co/u/Martin_Berlin)
#### Post date: [December 21, 2023, 1:20pm UTC](https://discuss.elastic.co/t/failed-to-indices-data-write-bulk-s-on-replica-because-of-netty4tcpchannel-compositebytesreference-more-than-2gb/349797/1 "2023-12-21T13:20:38Z")

</div>

While Indexing to our Cluster sometimes this error occures turning the cluster in red & yellow state:

One node is trying to "**perform indices:data/write/bulk[s] on replica**" on another node but fails because of "exception caught on transport layer [ **Netty4TcpChannel**...closing connection java.lang.IllegalArgumentException: CompositeBytesReference **cannot hold more than 2GB**"

We did not set any specific "network.tcp.send\_buffer\_size" or "network.tcp.receive\_buffer\_size" and the system values are:  
sysctl -a | grep rmem  
net.core.rmem\_default = 212992  
net.core.rmem\_max = 212992  
net.ipv4.tcp\_rmem = 4096 131072 6291456  
net.ipv4.udp\_rmem\_min = 4096

Any ideas? Google searches for this problem resulted only in Source code results...

Full Error node writing/sending:

```auto
[WARN][o.e.t.OutboundHandler] [cluster_II_node_6] sending transport message [Request{indices:data/write/bulk[s][r]}{6210854}{false}{true}{false}] of size [669561038] on [Netty4TcpChannel{localAddress=/10.10.1.35:40702, remoteAddress=10.10.1.31/10.10.1.31:9300, profile=default}] took [9023ms] which is above the warn threshold of [5000ms] with success [true]
[2023-12-21T12:47:12,800][WARN][o.e.t.OutboundHandler] [cluster_II_node_6] sending transport message [Request{indices:data/write/bulk[s][r]}{6210881}{false}{true}{false}] of size [690408] on [Netty4TcpChannel{localAddress=/10.10.1.35:40702, remoteAddress=10.10.1.31/10.10.1.31:9300, profile=default}] took [9014ms] which is above the warn threshold of [5000ms] with success [true]
[2023-12-21T12:47:12,801][WARN][o.e.t.OutboundHandler] [cluster_II_node_6] sending transport message [Request{indices:data/write/bulk[s][r]}{6210885}{false}{true}{false}] of size [2319] on [Netty4TcpChannel{localAddress=/10.10.1.35:40702, remoteAddress=10.10.1.31/10.10.1.31:9300, profile=default}] took [9014ms] which is above the warn threshold of [5000ms] with success [true]
[2023-12-21T12:47:12,801][WARN][o.e.t.OutboundHandler] [cluster_II_node_6] sending transport message [Request{indices:data/write/bulk[s][r]}{6210888}{false}{true}{false}] of size [343027] on [Netty4TcpChannel{localAddress=/10.10.1.35:40702, remoteAddress=10.10.1.31/10.10.1.31:9300, profile=default}] took [9011ms] which is above the warn threshold of [5000ms] with success [true]
[2023-12-21T12:47:12,801][WARN][o.e.t.OutboundHandler] [cluster_II_node_6] sending transport message [Request{indices:data/write/bulk[s][r]}{6210892}{false}{true}{false}] of size [2248] on [Netty4TcpChannel{localAddress=/10.10.1.35:40702, remoteAddress=10.10.1.31/10.10.1.31:9300, profile=default}] took [9010ms] which is above the warn threshold of [5000ms] with success [true]
[2023-12-21T12:47:12,801][WARN][o.e.t.OutboundHandler] [cluster_II_node_6] sending transport message [Request{indices:data/write/bulk[s][r]}{6210896}{false}{true}{false}] of size [2416] on [Netty4TcpChannel{localAddress=/10.10.1.35:40702, remoteAddress=10.10.1.31/10.10.1.31:9300, profile=default}] took [9004ms] which is above the warn threshold of [5000ms] with success [true]
[2023-12-21T12:47:12,801][WARN][o.e.t.OutboundHandler] [cluster_II_node_6] sending transport message [Request{indices:data/write/bulk[s][r]}{6210899}{false}{true}{false}] of size [2776] on [Netty4TcpChannel{localAddress=/10.10.1.35:40702, remoteAddress=10.10.1.31/10.10.1.31:9300, profile=default}] took [8754ms] which is above the warn threshold of [5000ms] with success [true]
[2023-12-21T12:47:12,801][WARN][o.e.t.OutboundHandler] [cluster_II_node_6] sending transport message [Request{indices:data/write/bulk[s][r]}{6210906}{false}{true}{false}] of size [3078] on [Netty4TcpChannel{localAddress=/10.10.1.35:40702, remoteAddress=10.10.1.31/10.10.1.31:9300, profile=default}] took [8200ms] which is above the warn threshold of [5000ms] with success [true]
[2023-12-21T12:47:12,801][WARN][o.e.t.OutboundHandler] [cluster_II_node_6] sending transport message [Request{indices:data/write/bulk[s][r]}{6210909}{false}{true}{false}] of size [3647] on [Netty4TcpChannel{localAddress=/10.10.1.35:40702, remoteAddress=10.10.1.31/10.10.1.31:9300, profile=default}] took [8200ms] which is above the warn threshold of [5000ms] with success [true]
[2023-12-21T12:47:12,809][INFO][o.e.t.ClusterConnectionManager] [cluster_II_node_6] transport connection to [{cluster_II_node_2}{L6USlqkNSAKjUKJ_iNC3yA}{Sd-dtRctQBWKfn2OdHvadQ}{cluster_II_node_2}{10.10.1.31}{10.10.1.31:9300}{d}{8.11.1}{7000099-8500003}] closed by remote
[2023-12-21T12:47:12,810][WARN][o.e.t.OutboundHandler] [cluster_II_node_6] sending transport message [Request{indices:data/write/bulk[s][r]}{6210912}{false}{true}{false}] of size [19389305] on [Netty4TcpChannel{localAddress=/10.10.1.35:40702, remoteAddress=10.10.1.31/10.10.1.31:9300, profile=default}] took [8037ms] which is above the warn threshold of [5000ms] with success [false]
[2023-12-21T12:47:18,705][WARN][o.e.a.b.TransportShardBulkAction] [cluster_II_node_6] [[cluster_node_2023_12_20_16_46_55][44]] failed to perform indices:data/write/bulk[s] on replica [cluster_node_2023_12_20_16_46_55][44], node[L6USlqkNSAKjUKJ_iNC3yA], [R], s[STARTED], a[id=a5hzS6-1TmGG4qHt3_GQxw], failed_attempts[0]
org.elasticsearch.transport.NodeNotConnectedException: [cluster_II_node_2][10.10.1.31:9300] Node not connected
	at org.elasticsearch.transport.ClusterConnectionManager.getConnection(ClusterConnectionManager.java:283) ~[elasticsearch-8.11.1.jar:?]
	at org.elasticsearch.transport.TransportService.getConnection(TransportService.java:869) ~[elasticsearch-8.11.1.jar:?]
	at org.elasticsearch.transport.TransportService.getConnectionOrFail(TransportService.java:764) ~[elasticsearch-8.11.1.jar:?]
	at org.elasticsearch.transport.TransportService.sendRequest(TransportService.java:750) ~[elasticsearch-8.11.1.jar:?]
	at org.elasticsearch.action.support.replication.TransportReplicationAction$ReplicasProxy.performOn(TransportReplicationAction.java:1272) ~[elasticsearch-8.11.1.jar:?]
	at org.elasticsearch.action.support.replication.ReplicationOperation$3.tryAction(ReplicationOperation.java:303) ~[elasticsearch-8.11.1.jar:?]
	at org.elasticsearch.action.support.RetryableAction$1.doRun(RetryableAction.java:111) ~[elasticsearch-8.11.1.jar:?]
	at org.elasticsearch.common.util.concurrent.ThreadContext$ContextPreservingAbstractRunnable.doRun(ThreadContext.java:983) ~[elasticsearch-8.11.1.jar:?]
	at org.elasticsearch.common.util.concurrent.AbstractRunnable.run(AbstractRunnable.java:26) ~[elasticsearch-8.11.1.jar:?]
	at org.elasticsearch.threadpool.ThreadPool$1.run(ThreadPool.java:481) ~[elasticsearch-8.11.1.jar:?]
	at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:572) ~[?:?]
	at java.util.concurrent.FutureTask.run(FutureTask.java:317) ~[?:?]
	at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:304) ~[?:?]
	at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1144) ~[?:?]
	at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:642) ~[?:?]
	at java.lang.Thread.run(Thread.java:1583) ~[?:?]
	Suppressed: org.elasticsearch.transport.NodeDisconnectedException: [cluster_II_node_2][10.10.1.31:9300][indices:data/write/bulk[s][r]] disconnected
	Suppressed: org.elasticsearch.transport.NodeNotConnectedException: [cluster_II_node_2][10.10.1.31:9300] Node not connected
		at org.elasticsearch.transport.ClusterConnectionManager.getConnection(ClusterConnectionManager.java:283) ~[elasticsearch-8.11.1.jar:?]
		at org.elasticsearch.transport.TransportService.getConnection(TransportService.java:869) ~[elasticsearch-8.11.1.jar:?]
		at org.elasticsearch.transport.TransportService.getConnectionOrFail(TransportService.java:764) ~[elasticsearch-8.11.1.jar:?]
		at org.elasticsearch.transport.TransportService.sendRequest(TransportService.java:750) ~[elasticsearch-8.11.1.jar:?]
		at org.elasticsearch.action.support.replication.TransportReplicationAction$ReplicasProxy.performOn(TransportReplicationAction.java:1272) ~[elasticsearch-8.11.1.jar:?]
		at org.elasticsearch.action.support.replication.ReplicationOperation$3.tryAction(ReplicationOperation.java:303) ~[elasticsearch-8.11.1.jar:?]
		at org.elasticsearch.action.support.RetryableAction$1.doRun(RetryableAction.java:111) ~[elasticsearch-8.11.1.jar:?]
		at org.elasticsearch.common.util.concurrent.ThreadContext$ContextPreservingAbstractRunnable.doRun(ThreadContext.java:983) ~[elasticsearch-8.11.1.jar:?]
		at org.elasticsearch.common.util.concurrent.AbstractRunnable.run(AbstractRunnable.java:26) ~[elasticsearch-8.11.1.jar:?]
		at org.elasticsearch.threadpool.ThreadPool$1.run(ThreadPool.java:481) ~[elasticsearch-8.11.1.jar:?]
		at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:572) ~[?:?]
		at java.util.concurrent.FutureTask.run(FutureTask.java:317) ~[?:?]
		at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:304) ~[?:?]
		at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1144) ~[?:?]
		at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:642) ~[?:?]
		at java.lang.Thread.run(Thread.java:1583) ~[?:?]
	Suppressed: org.elasticsearch.transport.NodeNotConnectedException: [cluster_II_node_2][10.10.1.31:9300] Node not connected
		at org.elasticsearch.transport.ClusterConnectionManager.getConnection(ClusterConnectionManager.java:283) ~[elasticsearch-8.11.1.jar:?]
		at org.elasticsearch.transport.TransportService.getConnection(TransportService.java:869) ~[elasticsearch-8.11.1.jar:?]
		at org.elasticsearch.transport.TransportService.getConnectionOrFail(TransportService.java:764) ~[elasticsearch-8.11.1.jar:?]
		at org.elasticsearch.transport.TransportService.sendRequest(TransportService.java:750) ~[elasticsearch-8.11.1.jar:?]
		at org.elasticsearch.action.support.replication.TransportReplicationAction$ReplicasProxy.performOn(TransportReplicationAction.java:1272) ~[elasticsearch-8.11.1.jar:?]
		at org.elasticsearch.action.support.replication.ReplicationOperation$3.tryAction(ReplicationOperation.java:303) ~[elasticsearch-8.11.1.jar:?]
		at org.elasticsearch.action.support.RetryableAction$1.doRun(RetryableAction.java:111) ~[elasticsearch-8.11.1.jar:?]
		at org.elasticsearch.common.util.concurrent.ThreadContext$ContextPreservingAbstractRunnable.doRun(ThreadContext.java:983) ~[elasticsearch-8.11.1.jar:?]
		at org.elasticsearch.common.util.concurrent.AbstractRunnable.run(AbstractRunnable.java:26) ~[elasticsearch-8.11.1.jar:?]
		at org.elasticsearch.threadpool.ThreadPool$1.run(ThreadPool.java:481) ~[elasticsearch-8.11.1.jar:?]
		at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:572) ~[?:?]
		at java.util.concurrent.FutureTask.run(FutureTask.java:317) ~[?:?]
		at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:304) ~[?:?]
		at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1144) ~[?:?]
		at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:642) ~[?:?]
		at java.lang.Thread.run(Thread.java:1583) ~[?:?]

```

Full Error Node receiving:

```auto
[2023-12-21T12:47:12,808][WARN][o.e.t.TcpTransport] [cluster_II_node_2] exception caught on transport layer [Netty4TcpChannel{localAddress=/10.10.1.31:9300, remoteAddress=/10.10.1.35:40702, profile=default}], closing connection
java.lang.IllegalArgumentException: CompositeBytesReference cannot hold more than 2GB
	at org.elasticsearch.common.bytes.CompositeBytesReference.ofMultiple(CompositeBytesReference.java:59) ~[elasticsearch-8.11.1.jar:?]
	at org.elasticsearch.common.bytes.CompositeBytesReference.of(CompositeBytesReference.java:40) ~[elasticsearch-8.11.1.jar:?]
	at org.elasticsearch.transport.InboundAggregator.finishAggregation(InboundAggregator.java:104) ~[elasticsearch-8.11.1.jar:?]
	at org.elasticsearch.transport.InboundPipeline.forwardFragments(InboundPipeline.java:121) ~[elasticsearch-8.11.1.jar:?]
	at org.elasticsearch.transport.InboundPipeline.doHandleBytes(InboundPipeline.java:96) ~[elasticsearch-8.11.1.jar:?]
	at org.elasticsearch.transport.InboundPipeline.handleBytes(InboundPipeline.java:61) ~[elasticsearch-8.11.1.jar:?]
	at org.elasticsearch.transport.netty4.Netty4MessageInboundHandler.channelRead(Netty4MessageInboundHandler.java:48) ~[?:?]
	at io.netty.channel.AbstractChannelHandlerContext.invokeChannelRead(AbstractChannelHandlerContext.java:444) ~[?:?]
	at io.netty.channel.AbstractChannelHandlerContext.invokeChannelRead(AbstractChannelHandlerContext.java:420) ~[?:?]
	at io.netty.channel.AbstractChannelHandlerContext.fireChannelRead(AbstractChannelHandlerContext.java:412) ~[?:?]
	at io.netty.handler.codec.MessageToMessageDecoder.channelRead(MessageToMessageDecoder.java:103) ~[?:?]
	at io.netty.channel.AbstractChannelHandlerContext.invokeChannelRead(AbstractChannelHandlerContext.java:444) ~[?:?]
	at io.netty.channel.AbstractChannelHandlerContext.invokeChannelRead(AbstractChannelHandlerContext.java:420) ~[?:?]
	at io.netty.channel.AbstractChannelHandlerContext.fireChannelRead(AbstractChannelHandlerContext.java:412) ~[?:?]
	at io.netty.channel.DefaultChannelPipeline$HeadContext.channelRead(DefaultChannelPipeline.java:1410) ~[?:?]
	at io.netty.channel.AbstractChannelHandlerContext.invokeChannelRead(AbstractChannelHandlerContext.java:440) ~[?:?]
	at io.netty.channel.AbstractChannelHandlerContext.invokeChannelRead(AbstractChannelHandlerContext.java:420) ~[?:?]
	at io.netty.channel.DefaultChannelPipeline.fireChannelRead(DefaultChannelPipeline.java:919) ~[?:?]
	at io.netty.channel.nio.AbstractNioByteChannel$NioByteUnsafe.read(AbstractNioByteChannel.java:166) ~[?:?]
	at io.netty.channel.nio.NioEventLoop.processSelectedKey(NioEventLoop.java:788) ~[?:?]
	at io.netty.channel.nio.NioEventLoop.processSelectedKeysPlain(NioEventLoop.java:689) ~[?:?]
	at io.netty.channel.nio.NioEventLoop.processSelectedKeys(NioEventLoop.java:652) ~[?:?]
	at io.netty.channel.nio.NioEventLoop.run(NioEventLoop.java:562) ~[?:?]
	at io.netty.util.concurrent.SingleThreadEventExecutor$4.run(SingleThreadEventExecutor.java:997) ~[?:?]
	at io.netty.util.internal.ThreadExecutorMap$2.run(ThreadExecutorMap.java:74) ~[?:?]
	at java.lang.Thread.run(Thread.java:1583) ~[?:?]

```

---

<div class="post-metadata">

### Author: ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)
#### Post date: [December 21, 2023, 5:41pm UTC](https://discuss.elastic.co/t/failed-to-indices-data-write-bulk-s-on-replica-because-of-netty4tcpchannel-compositebytesreference-more-than-2gb/349797/2 "2023-12-21T17:41:15Z")

</div>

Looks like another instance of [Transport messages exceeding 2GiB are not handled gracefully · Issue #94137 · elastic/elasticsearch · GitHub](https://github.com/elastic/elasticsearch/issues/94137) - the workaround is to send smaller bulk requests and/or avoid ingest pipelines and other scripts which might blow up the sizes of the documents to be replicated.

---

<div class="post-metadata">

### Author: ![Martin\_Berlin](https://avatars.discourse-cdn.com/v4/letter/m/958977/32.png) [@Martin\_Berlin](https://discuss.elastic.co/u/Martin_Berlin)
#### Post date: [December 22, 2023, 9:49am UTC](https://discuss.elastic.co/t/failed-to-indices-data-write-bulk-s-on-replica-because-of-netty4tcpchannel-compositebytesreference-more-than-2gb/349797/3 "2023-12-22T09:49:58Z")

</div>

Thanks @DavidTurner  
our bulk sizes are less than 100mb and we don't use ingest pipelines.  
But we use script updates to fill nested documents in a document.  
Unfortunately we need those nested documents and the Bug you referenced doesn't seem to be fixed. We are using Version 8.11.1

---

<div class="post-metadata">

### Author: ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)
#### Post date: [December 22, 2023, 11:09am UTC](https://discuss.elastic.co/t/failed-to-indices-data-write-bulk-s-on-replica-because-of-netty4tcpchannel-compositebytesreference-more-than-2gb/349797/4 "2023-12-22T11:09:19Z")

</div>

> [@Martin\_Berlin](#):
>
> our bulk sizes are less than 100mb

That's still pretty large, especially if you're running scripts that blow up the document size massively. Make them smaller.

---

<div class="post-metadata">

### Author: ![Martin\_Berlin](https://avatars.discourse-cdn.com/v4/letter/m/958977/32.png) [@Martin\_Berlin](https://discuss.elastic.co/u/Martin_Berlin)
#### Post date: [December 22, 2023, 1:46pm UTC](https://discuss.elastic.co/t/failed-to-indices-data-write-bulk-s-on-replica-because-of-netty4tcpchannel-compositebytesreference-more-than-2gb/349797/5 "2023-12-22T13:46:16Z")

</div>

OK - we were using the defaults of the python client:

```auto
elasticsearch-py 7.x
elasticsearch.helpers.streaming_bulk(client, actions, chunk_size=500, max_chunk_bytes=104857600, raise_on_error=True, expand_action_callback=<function expand_action>, raise_on_exception=True, max_retries=0, initial_backoff=2, max_backoff=600, yield_ok=True, ignore_status=(), *args, **kwargs)

```

= _chunk\_size=500_ , _max\_chunk\_bytes=104857600_

what would be an appropriate size?

---

<div class="post-metadata">

### Author: ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)
#### Post date: [December 22, 2023, 3:11pm UTC](https://discuss.elastic.co/t/failed-to-indices-data-write-bulk-s-on-replica-because-of-netty4tcpchannel-compositebytesreference-more-than-2gb/349797/6 "2023-12-22T15:11:59Z")

</div>

Impossible to say without knowing how much your scripted updates expand the documents.

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [January 19, 2024, 3:12pm UTC](https://discuss.elastic.co/t/failed-to-indices-data-write-bulk-s-on-replica-because-of-netty4tcpchannel-compositebytesreference-more-than-2gb/349797/7 "2024-01-19T15:12:52Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
