# Too many files open, recovering cluster

**URL:** <https://discuss.elastic.co/t/too-many-files-open-recovering-cluster/35048>\
**Category:** Elasticsearch\
**Created:** [November 19, 2015, 12:24pm UTC](https://discuss.elastic.co/t/too-many-files-open-recovering-cluster/35048 "2015-11-19T12:24:30Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![Felipe\_Santos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/felipe_santos/32/3511_2.png) [@Felipe\_Santos](https://discuss.elastic.co/u/Felipe_Santos)\
**Post date:** [November 19, 2015, 12:24pm UTC](https://discuss.elastic.co/t/too-many-files-open-recovering-cluster/35048/1 "2015-11-19T12:24:31Z")

</div>

3 x master node 8GB 2vCPU  
3X data note 30GB 8vCPU

I am recovering a cluster and I am getting this error

`2015-11-18 02:45:20,374][WARN][action.bulk] [Bruiser] failed to perform indices:data/write/bulk[s] on remote replica [Douglas Birely][qTqgsv_STVG3je5Fn7tEeg][zupme-1b-elasticsearch003.aws.zup.com.br][inet[/***]][events-vivo-2-20151118][1] org.elasticsearch.transport.RemoteTransportException: [Douglas Birely][inet[/*****]][indices:data/write/bulk[s][r]] Caused by: org.elasticsearch.index.engine.CreateFailedEngineException: [events-vivo-2-20151118][1] Create failed for [events#AVEY6UeqCpZT8FxkSXyC] at org.elasticsearch.index.engine.InternalEngine.create(InternalEngine.java:264) at org.elasticsearch.index.shard.IndexShard.create(IndexShard.java:483) at org.elasticsearch.action.bulk.TransportShardBulkAction.shardOperationOnReplica(TransportShardBulkAction.java:569) at org.elasticsearch.action.support.replication.TransportShardReplicationOperationAction$ReplicaOperationTransportHandler.messageReceived(TransportShardReplicationOperationAction.java:250) at org.elasticsearch.action.support.replication.TransportShardReplicationOperationAction$ReplicaOperationTransportHandler.messageReceived(TransportShardReplicationOperationAction.java:229) at org.elasticsearch.transport.netty.MessageChannelHandler$RequestHandler.doRun(MessageChannelHandler.java:279) at org.elasticsearch.common.util.concurrent.AbstractRunnable.run(AbstractRunnable.java:36) at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) at java.lang.Thread.run(Thread.java:745) Caused by: java.io.FileNotFoundException: /data/elasticsearch/zupme/nodes/0/indices/events-vivo-2-20151118/1/index/_gb.fdt (Too many open files) at java.io.FileOutputStream.open0(Native Method) at java.io.FileOutputStream.open(FileOutputStream.java:270) at java.io.FileOutputStream.<init>(FileOutputStream.java:213) at java.io.FileOutputStream.<init>(FileOutputStream.java:162) .....`

File descriptor are set to high values.

`{ "cluster_name" : "zupme", "nodes" : { "BnoAILz0Q3KQjSWtE2KNKw" : { "name" : "Sebastian Shaw", "transport_address" : "inet[*****]", "host" : "*****", "ip" : "****", "version" : "1.7.2", "build" : "e43676b", "http_address" : "inet[****]", "attributes" : { "master" : "false" }, "process" : { "refresh_interval_in_millis" : 1000, "id" : 18701, "max_file_descriptors" : 131072, "mlockall" : true } }, "o__dvgL7QfyIM-jRwlJlHg" : { "name" : "Milan", "transport_address" : "inet[****]", "host" : "*****", "ip" : "*****", "version" : "1.7.2", "build" : "e43676b", "http_address" : "inet[******]", "attributes" : { "data" : "false", "master" : "true" }, "process" : { "refresh_interval_in_millis" : 1000, "id" : 16397, "max_file_descriptors" : 65536, "mlockall" : true } }, "C11qTS23R5aX2t6TTSCGSA" : { "name" : "Seeker", "transport_address" : "inet[*****]", "host" : "*****", "ip" : "*****", "version" : "1.7.2", "build" : "e43676b", "http_address" : "inet[****]", "attributes" : { "master" : "false" }, "process" : { "refresh_interval_in_millis" : 1000, "id" : 7885, "max_file_descriptors" : 131072, "mlockall" : true } }, "o_5zCidLQpiXsjJGsBPbXw" : { "name" : "Phantom Eagle", "transport_address" : "inet[*****]", "host" : "*****", "ip" : "*****", "version" : "1.7.2", "build" : "e43676b", "http_address" : "inet[****]", "attributes" : { "data" : "false", "master" : "true" }, "process" : { "refresh_interval_in_millis" : 1000, "id" : 16777, "max_file_descriptors" : 65536, "mlockall" : true } }, "QbbbxsqmTlWpSRtWzdIhgg" : { "name" : "Lilith, the Daughter of Dracula ", "transport_address" : "inet[****]", "host" : "****", "ip" : "****", "version" : "1.7.2", "build" : "e43676b", "http_address" : "inet[****]", "attributes" : { "master" : "false" }, "process" : { "refresh_interval_in_millis" : 1000, "id" : 25335, "max_file_descriptors" : 131072, "mlockall" : true } } } }`

If I have too many indices, but these indices is not beeing use now, neither for search nor index, It remains file descriptors open?

---

<div class="post-metadata">

**Author:** ![Felipe\_Santos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/felipe_santos/32/3511_2.png) [@Felipe\_Santos](https://discuss.elastic.co/u/Felipe_Santos)\
**Post date:** [November 19, 2015, 1:39pm UTC](https://discuss.elastic.co/t/too-many-files-open-recovering-cluster/35048/2 "2015-11-19T13:39:59Z")

</div>

All the machines are mostly idle

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [November 19, 2015, 2:03pm UTC](https://discuss.elastic.co/t/too-many-files-open-recovering-cluster/35048/3 "2015-11-19T14:03:50Z")

</div>

> [@Felipe\_Santos](#):
>
> If I have too many indices, but these indices is not beeing use now, neither for search nor index, It remains file descriptors open?

Yes. Indices need to be closed in order to use fewer file descriptors.

---

<div class="post-metadata">

**Author:** ![Felipe\_Santos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/felipe_santos/32/3511_2.png) [@Felipe\_Santos](https://discuss.elastic.co/u/Felipe_Santos)\
**Post date:** [November 19, 2015, 2:10pm UTC](https://discuss.elastic.co/t/too-many-files-open-recovering-cluster/35048/4 "2015-11-19T14:10:25Z")

</div>

> [@jpountz](#):
>
> .

Could you explain why It remain open if it is not beeing used?

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [November 19, 2015, 2:16pm UTC](https://discuss.elastic.co/t/too-many-files-open-recovering-cluster/35048/5 "2015-11-19T14:16:50Z")

</div>

This is how databases work in general. Opening files is a costly operation, so elasticsearch opens files when opening the index and then keeps them open.

---

<div class="post-metadata">

**Author:** ![Felipe\_Santos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/felipe_santos/32/3511_2.png) [@Felipe\_Santos](https://discuss.elastic.co/u/Felipe_Santos)\
**Post date:** [November 19, 2015, 2:53pm UTC](https://discuss.elastic.co/t/too-many-files-open-recovering-cluster/35048/6 "2015-11-19T14:53:06Z")

</div>

I am recovering the cluster,

{  
"cluster\_name" : "zupme",  
"status" : "red",  
"timed\_out" : false,  
"number\_of\_nodes" : 6,  
"number\_of\_data\_nodes" : 3,  
"active\_primary\_shards" : 15125,  
"active\_shards" : 24647,  
"relocating\_shards" : 0,  
"initializing\_shards" : 6,  
"unassigned\_shards" : 5599,  
"delayed\_unassigned\_shards" : 0,  
"number\_of\_pending\_tasks" : 6845,  
"number\_of\_in\_flight\_fetch" : 0  
}

And when unassigned\_shards achieve \< 700 it raise too many open files and stop on this numeber(cpu is mostly idle), the only way to solve this is to close indices? The master has one CPU at 100% but others CPUs are idle, and recovery is too slow

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [November 19, 2015, 3:00pm UTC](https://discuss.elastic.co/t/too-many-files-open-recovering-cluster/35048/7 "2015-11-19T15:00:16Z")

</div>

This is a lot of shards for only 3 data nodes, you should try to have fewer indices and/or fewer shards per index. I'm afraid open files are just the first thing that breaks, but even if you were able to fix this issue eg. by letting the OS allocate more open files, something else would break.

---

<div class="post-metadata">

**Author:** ![Felipe\_Santos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/felipe_santos/32/3511_2.png) [@Felipe\_Santos](https://discuss.elastic.co/u/Felipe_Santos)\
**Post date:** [November 19, 2015, 3:01pm UTC](https://discuss.elastic.co/t/too-many-files-open-recovering-cluster/35048/8 "2015-11-19T15:01:36Z")

</div>

There is some document that has a `formula` to get number of shards? I don't think we can have fewer indices

---

<div class="post-metadata">

**Author:** ![Felipe\_Santos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/felipe_santos/32/3511_2.png) [@Felipe\_Santos](https://discuss.elastic.co/u/Felipe_Santos)\
**Post date:** [November 19, 2015, 3:06pm UTC](https://discuss.elastic.co/t/too-many-files-open-recovering-cluster/35048/9 "2015-11-19T15:06:41Z")

</div>

And why reallocating unassigned\_shards are too slow 1 shard per second, and none errors on logs?

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [November 19, 2015, 3:07pm UTC](https://discuss.elastic.co/t/too-many-files-open-recovering-cluster/35048/10 "2015-11-19T15:07:41Z")

</div>

There isn't really a formula, this would depend on the hardware, mappings, etc. but hundreds of shards per node is already a lot.

Why can't you have fewer indices? Sometimes you can share data eg. for several users in the same index. See eg. [https://vimeo.com/44716955](https://vimeo.com/44716955) from 13'45

---

<div class="post-metadata">

**Author:** ![Felipe\_Santos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/felipe_santos/32/3511_2.png) [@Felipe\_Santos](https://discuss.elastic.co/u/Felipe_Santos)\
**Post date:** [November 19, 2015, 3:10pm UTC](https://discuss.elastic.co/t/too-many-files-open-recovering-cluster/35048/11 "2015-11-19T15:10:26Z")

</div>

Because an user could have millions of event per day, so the search and index will be slow. I will take a look the video..

Thanks a lot

---

<div class="post-metadata">

**Author:** ![Felipe\_Santos](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/felipe_santos/32/3511_2.png) [@Felipe\_Santos](https://discuss.elastic.co/u/Felipe_Santos)\
**Post date:** [November 19, 2015, 3:50pm UTC](https://discuss.elastic.co/t/too-many-files-open-recovering-cluster/35048/12 "2015-11-19T15:50:22Z")

</div>

Any tips why cluster recovery is too slow? And CPUs are mostly idle.. Its because the number of shards?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:37pm UTC](https://discuss.elastic.co/t/too-many-files-open-recovering-cluster/35048/13 "2017-07-05T23:37:10Z")

</div>


