# How do I find the reason for a failed data node? (Elasticsearch 6.5)

**URL:** <https://discuss.elastic.co/t/how-do-i-find-the-reason-for-a-failed-data-node-elasticsearch-6-5/185057>\
**Category:** Elasticsearch\
**Created:** [June 10, 2019, 9:14pm UTC](https://discuss.elastic.co/t/how-do-i-find-the-reason-for-a-failed-data-node-elasticsearch-6-5/185057 "2019-06-10T21:14:39Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Eric\_Paul](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/eric_paul/32/47825_2.png) [@Eric\_Paul](https://discuss.elastic.co/u/Eric_Paul)\
**Post date:** [June 10, 2019, 9:14pm UTC](https://discuss.elastic.co/t/how-do-i-find-the-reason-for-a-failed-data-node-elasticsearch-6-5/185057/1 "2019-06-10T21:14:39Z")

</div>

We have been running elasticsearch for a few years now. We ran a 2 node simple cluster (version 1.7). This cluster supported some internal utilities so it was relatively low use. In the last 4 years that cluster has never crashed, been restarted or even hiccuped.

We decided to set up a more production focused cluster. I did a lot of research and this is that I came up with for the new cluster:

```auto
2 Client Nodes (a.k.a. Coordinating nodes) [4 core, 8GB memory, 300GB HD, Virtual]
3 Master Nodes[4 core, 8GB memory, 300GB HD, Virtual] 
3 Data Nodes[48 core, 64GB memory, 3TB HD (Raid 0), Physical] 

```

This cluster is running ES 6.5.4 CENTOS 7 (I am planning to upgrade to 7.1 soon). All of nodes are operating on essentially vanilla configurations. We only have about 5 million documents and less than 60GB of data total for the cluster. The configuration looks something like this:

```auto
# Example Master Config
cluster.name: MYCLUSTER
node.name: MASTER01
node.master: true
node.data: false
node.ingest: false
cluster.remote.connect: false
path.repo: /repo/nfs/path

# Example Data Config
cluster.name: MYCLUSTER
node.name: DATA01
node.master: false
node.data: true
node.ingest: false
cluster.remote.connect: false
path.repo: /repo/nfs/path

# Example Client Config
cluster.name: MYCLUSTER
node.name: CLIENT01
node.master: false
node.data: false
node.ingest: false
cluster.remote.connect: false

# All have
http.port: MY_ES_PORT
discovery.zen.ping.unicast.hosts: MY_LIST_OF_SERVERS(8)
discovery.zen.minimum_master_nodes: 2

```

In jvm.options the heap space is at the default 1G for all nodes except the data nodes which are at 26G.

The problem is that my data nodes keep crashing. 3 times in the last 3 days one of my 3 data nodes has crashed. Bringing it back online and ridding myself of corrupted pieces of indexes has been many hours of work and learning. I can't figure out what is making them crash. I see errors in the log that refer to the "Failed Node" and "CorruptIndexException" but I have no idea what caused the actual node to fail. I have examined log files of all of the severs and while they all show errors none seem to have anything that helps me pinpoint the cause of the failure. 2 of the 3 data nodes have failed.

The interwebs seem to suggest that the most common reason for this is hardware. Unfortunately, I don't see any evidence that the hardware is the problem.

Can anyone offer any advice on how I can figure out why the data nodes are crashing?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [June 10, 2019, 9:59pm UTC](https://discuss.elastic.co/t/how-do-i-find-the-reason-for-a-failed-data-node-elasticsearch-6-5/185057/2 "2019-06-10T21:59:38Z")

</div>

Welcome! 🙂

> [@Eric\_Paul](#):
>
> discovery.zen.ping.unicast.hosts: MY\_LIST\_OF\_SERVERS(8)

Just set that to be your master nodes, it's much easier to maintain.

> [@Eric\_Paul](#):
>
> In jvm.options the heap space is at the default 1G for all nodes except the data nodes which are at 26G.

I would suggest you increase that. On the master+client nodes you can go to 3GB easily enough. For the data nodes you want to be [just under 32GB](https://www.elastic.co/guide/en/elasticsearch/reference/7.1/heap-size.html).

> [@Eric\_Paul](#):
>
> Can anyone offer any advice on how I can figure out why the data nodes are crashing?

Posting the logs, or using gist/pastebin/etc and linking, would be really helpful.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 8, 2019, 9:59pm UTC](https://discuss.elastic.co/t/how-do-i-find-the-reason-for-a-failed-data-node-elasticsearch-6-5/185057/3 "2019-07-08T21:59:53Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
