# Elasticsearch cause linux kernel crash

**URL:** https://discuss.elastic.co/t/elasticsearch-cause-linux-kernel-crash/187461
**Category:** Elasticsearch
**Created:** [June 26, 2019, 2:41am UTC](https://discuss.elastic.co/t/elasticsearch-cause-linux-kernel-crash/187461 "2019-06-26T02:41:14Z")
**Posts on this page:** 10
**Page:** 1

<div class="post-metadata">

### Author: ![john\_am](https://avatars.discourse-cdn.com/v4/letter/j/5f9b8f/32.png) [@john\_am](https://discuss.elastic.co/u/john_am)
#### Post date: [June 26, 2019, 2:41am UTC](https://discuss.elastic.co/t/elasticsearch-cause-linux-kernel-crash/187461/1 "2019-06-26T02:41:15Z")

</div>

Hi,guys, thanks for your reading my poor english.  
I got a situation that elasticsearch cause my linux kernel crash.However, what information i got from logs(linux messsage log) is Garbled. What's worse i found nothing from my elasticsearch's logs (info level).The crash will happen at least once time each day.

linux message log:  
(plenty of that info ,i ask for red hed help, they just say io is too heavey. And the cloud serivices say their software is ok.)  
[12655.166750] sd 14:65535:11:0: [sdi] CDB: Write(10) 2a 00 40 87 42 d8 00 01 70 00  
[12655.188749] SD100EP: [ERR][epfront\_scmd\_printk][1096]: scsi\_cmnd retrying: serial\_number[4192555] retries[1], allowed[5]  
[12655.188756] sd 14:65535:11:0: [sdi] CDB: Write(10) 2a 00 40 87 45 78 00 00 50 00  
[12655.211419] SD100EP: [ERR][epfront\_scmd\_printk][1096]: scsi\_cmnd retrying: serial\_number[4192614] retries[1], allowed[5]  
[12655.211424] sd 14:65535:11:0: [sdi] CDB: Write(10) 2a 00 40 87 47 38 00 00 78 00  
[12655.224413] SD100EP: [ERR][epfront\_scmd\_printk][1096]: scsi\_cmnd retrying: serial\_number[4192637] retries[1], allowed[5]  
[12655.224418] sd 14:65535:11:0: [sdi] CDB: Write(10) 2a 00 40 87 47 a8 00 02 70 00

Here is my linux env:  
Linux version 3.10.0-514.el7.x86\_64 ([builder@kbuilder.dev.centos.org](mailto:builder@kbuilder.dev.centos.org)) (gcc version 4.8.5 20150623 (Red Hat 4.8.5-11)

12655.400653] SD100EP: [ERR][epfront\_io\_send][2198]: alloc\_iod failed, no memory  
[12655.402837] SD100EP: [ERR][epfront\_io\_send][2198]: alloc\_iod failed, no memory

after that ,linux kernel dead.

jvm version:  
1.8.0\_91-b14  
30 g for each instance，each computer run four instance.

hardware:  
15_56core 256G ram 2_2t ssd 2\*6t sas

logstash conf:  
http.port: 9600-9700  
pipeline.workers: 80  
pipeline.batch.size: 1500  
#pipeline.batch.delay: 100  
config.reload.automatic: true  
config.reload.interval: 3s

elasticsearch conf:  
xpack.security.enabled: true  
xpack.security.transport.ssl.enabled: true  
xpack.security.transport.ssl.verification\_mode: certificate  
xpack.security.transport.ssl.keystore.path: certs/elastic-certificates.p12  
xpack.security.transport.ssl.truststore.path: certs/elastic-certificates.p12  
node.max\_local\_storage\_nodes: 10  
thread\_pool.get.queue\_size: 10000  
thread\_pool.write.queue\_size: 10000  
thread\_pool.analyze.queue\_size: 1000  
thread\_pool.search.queue\_size: 10000  
thread\_pool.listener.queue\_size: 10000  
bootstrap.system\_call\_filter: false  
node.attr.disk\_type: ssd  
node.attr.zone: data-3-90  
cluster.routing.allocation.awareness.attributes: zone  
cluster.routing.allocation.awareness.force.zone.values: data-2-52,data-3-90,data-3-222,data-3-19,data-2-126,data-2-89,data-2-153，data-2-211,data-2-4,data-2-146,data-3-220,data-2-38  
#transport.compress: true

I check my jvm gc suck like is ok. But es takes my memory nerlly 99% of my service(es jvm+buffer). The crash will happen at least once time each day.

---

<div class="post-metadata">

### Author: ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)
#### Post date: [June 26, 2019, 2:59am UTC](https://discuss.elastic.co/t/elasticsearch-cause-linux-kernel-crash/187461/2 "2019-06-26T02:59:38Z")

</div>

Which version of Elasticsearch are you running? Is there anything in the Elasticsearch logs?

> [@john\_am](#):
>
> jvm version:  
> 1.8.0\_91-b14

This seems quite old and not supported according to the [support matrix](https://www.elastic.co/support/matrix#matrix_jvm).

---

<div class="post-metadata">

### Author: ![john\_am](https://avatars.discourse-cdn.com/v4/letter/j/5f9b8f/32.png) [@john\_am](https://discuss.elastic.co/u/john_am)
#### Post date: [June 26, 2019, 5:26am UTC](https://discuss.elastic.co/t/elasticsearch-cause-linux-kernel-crash/187461/3 "2019-06-26T05:26:21Z")

</div>

Thanks for your time. My es version is 7.0.1, and my es log will show as follow(the crash service is data-3-90-1, and this is 3-90 log, what i get sometime) :

[2019-06-26T12:02:42,393][WARN][o.e.c.a.s.ShardStateAction] [data-3-90-1] node closed while execution action [internal:cluster/shard/failure] for shard entry [shard id [[cloud-edrive-download][2]], allocation id [BR53DpZ1QSy9lFxqzP9-nw], primary term [10], message [failed to perform indices:data/write/bulk[s] on replica [cloud-edrive-download][2], node[mIwqy245SwieG4xwck9k3A], [R], recovery\_source[peer recovery], s[INITIALIZING], a[id=BR53DpZ1QSy9lFxqzP9-nw], unassigned\_info[[reason=NODE\_LEFT], at[2019-06-26T03:39:47.526Z], delayed=false, details[node\_left [mIwqy245SwieG4xwck9k3A]], allocation\_status[no\_attempt]]], failure [NodeDisconnectedException[[data-3-19-1][10.10.3.19:9302][indices:data/write/bulk[s][r]] disconnected]], markAsStale [true]]

[2019-06-26T11:21:32,154][WARN][o.e.c.NodeConnectionsService] [data-3-90-1] failed to connect to node {data-2-89-2}{f0RQuUDoQKG3AG0KFuZ9jQ}{pohYLyp0RFGZd9PM0YVGCw}{10.10.2.89}{10.10.2.89:9300}{ml.machine\_memory=269889724416, ml.max\_open\_jobs=20, xpack.installed=true, zone=data-2-89, disk\_type=sas} (tried [1] times)

what puzzle me is sometime before crash i just got a log few hours ago(gc log，gc is ok，we can see monitoring gc young count is 1/per minute, old is 30 minute. )

I can give you the total log by email .  
Really appreciate for your great help,by john.

---

<div class="post-metadata">

### Author: ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)
#### Post date: [June 26, 2019, 5:43am UTC](https://discuss.elastic.co/t/elasticsearch-cause-linux-kernel-crash/187461/4 "2019-06-26T05:43:24Z")

</div>

What is the full output of the [cluster stats API](https://www.elastic.co/guide/en/elasticsearch/reference/current/cluster-stats.html)?

> [@john\_am](#):
>
> thread\_pool.get.queue\_size: 10000  
> thread\_pool.write.queue\_size: 10000  
> thread\_pool.analyze.queue\_size: 1000  
> thread\_pool.search.queue\_size: 10000  
> thread\_pool.listener.queue\_size: 10000

Why have you bumped these up so far?

This may not be directly related to the issue, but as it sems I/O or scsi related I am not able to help there...

---

<div class="post-metadata">

### Author: ![john\_am](https://avatars.discourse-cdn.com/v4/letter/j/5f9b8f/32.png) [@john\_am](https://discuss.elastic.co/u/john_am)
#### Post date: [June 26, 2019, 6:20am UTC](https://discuss.elastic.co/t/elasticsearch-cause-linux-kernel-crash/187461/5 "2019-06-26T06:20:22Z")

</div>

I/O problem i get from red hat and my cloud service reply is : My service is suffer from system loding. Which cause i/o is terrible(because the disk is not local disk in cloud service, it need few system source to output to my disk. This situation casue only if system loding over 60 by top showing. ).

---

<div class="post-metadata">

### Author: ![john\_am](https://avatars.discourse-cdn.com/v4/letter/j/5f9b8f/32.png) [@john\_am](https://discuss.elastic.co/u/john_am)
#### Post date: [June 26, 2019, 6:23am UTC](https://discuss.elastic.co/t/elasticsearch-cause-linux-kernel-crash/187461/6 "2019-06-26T06:23:36Z")

</div>

stats api show are as follow(To protect Linux we use cpulimit in each instance just from today. ):  
{  
"\_nodes" : {  
"total" : 53,  
"successful" : 53,  
"failed" : 0  
},  
"cluster\_name" : "elastic",  
"cluster\_uuid" : "I1z8p5Q-SPe-Y4OFw8AXTQ",  
"timestamp" : 1561529621384,  
"status" : "red",  
"indices" : {  
"count" : 1450,  
"shards" : {  
"total" : 5204,  
"primaries" : 4075,  
"replication" : 0.27705521472392636,  
"index" : {  
"shards" : {  
"min" : 1,  
"max" : 10,  
"avg" : 3.5889655172413795  
},  
"primaries" : {  
"min" : 1,  
"max" : 5,  
"avg" : 2.810344827586207  
},  
"replication" : {  
"min" : 0.0,  
"max" : 1.0,  
"avg" : 0.27089655172413796  
}  
}  
},  
"docs" : {  
"count" : 53501026655,  
"deleted" : 5322372  
},  
"store" : {  
"size" : "29tb",  
"size\_in\_bytes" : 31969924566435  
},  
"fielddata" : {  
"memory\_size" : "226.6kb",  
"memory\_size\_in\_bytes" : 232072,  
"evictions" : 0  
},  
"query\_cache" : {  
"memory\_size" : "1.4gb",  
"memory\_size\_in\_bytes" : 1576283504,  
"total\_count" : 138023014,  
"hit\_count" : 20357130,  
"miss\_count" : 117665884,  
"cache\_size" : 29687,  
"cache\_count" : 75719,  
"evictions" : 46032  
},  
"completion" : {  
"size" : "0b",  
"size\_in\_bytes" : 0  
},  
"segments" : {  
"count" : 125847,  
"memory" : "68gb",  
"memory\_in\_bytes" : 73067214664,  
"terms\_memory" : "56.2gb",  
"terms\_memory\_in\_bytes" : 60417899019,  
"stored\_fields\_memory" : "10.5gb",  
"stored\_fields\_memory\_in\_bytes" : 11354919656,  
"term\_vectors\_memory" : "0b",  
"term\_vectors\_memory\_in\_bytes" : 0,  
"norms\_memory" : "151.9mb",  
"norms\_memory\_in\_bytes" : 159356800,  
"points\_memory" : "1gb",  
"points\_memory\_in\_bytes" : 1103197673,  
"doc\_values\_memory" : "30.3mb",  
"doc\_values\_memory\_in\_bytes" : 31841516,  
"index\_writer\_memory" : "1.4gb",  
"index\_writer\_memory\_in\_bytes" : 1551804982,  
"version\_map\_memory" : "1.8mb",  
"version\_map\_memory\_in\_bytes" : 1973242,  
"fixed\_bit\_set" : "23.2mb",  
"fixed\_bit\_set\_memory\_in\_bytes" : 24373568,  
"max\_unsafe\_auto\_id\_timestamp" : 1561527126190,  
"file\_sizes" : { }  
}  
},  
"nodes" : {  
"count" : {  
"total" : 53,  
"data" : 48,  
"coordinating\_only" : 0,  
"master" : 5,  
"ingest" : 52  
},  
"versions" : [  
"7.0.1"  
],  
"os" : {  
"available\_processors" : 2952,  
"allocated\_processors" : 2952,  
"names" : [  
{  
"name" : "Linux",  
"count" : 53  
}  
],  
"pretty\_names" : [  
{  
"pretty\_name" : "CentOS Linux 7 (Core)",  
"count" : 53  
}  
],  
"mem" : {  
"total" : "12.8tb",  
"total\_in\_bytes" : 14168874549248,  
"free" : "143gb",  
"free\_in\_bytes" : 153556291584,  
"used" : "12.7tb",  
"used\_in\_bytes" : 14015318257664,  
"free\_percent" : 1,  
"used\_percent" : 99  
}  
},  
"process" : {  
"cpu" : {  
"percent" : 111  
},  
"open\_file\_descriptors" : {  
"min" : 2074,  
"max" : 9432,  
"avg" : 4674  
}  
},  
"jvm" : {  
"max\_uptime" : "23.6d",  
"max\_uptime\_in\_millis" : 2044555574,  
"versions" : [  
{  
"version" : "1.8.0\_91",  
"vm\_name" : "Java HotSpot(TM) 64-Bit Server VM",  
"vm\_version" : "25.91-b14",  
"vm\_vendor" : "Oracle Corporation",  
"bundled\_jdk" : true,  
"using\_bundled\_jdk" : false,  
"count" : 48  
},  
{  
"version" : "12.0.1",  
"vm\_name" : "OpenJDK 64-Bit Server VM",  
"vm\_version" : "12.0.1+12",  
"vm\_vendor" : "Oracle Corporation",  
"bundled\_jdk" : true,  
"using\_bundled\_jdk" : true,  
"count" : 5  
}  
],  
"mem" : {  
"heap\_used" : "349.2gb",  
"heap\_used\_in\_bytes" : 374988337728,  
"heap\_max" : "1.4tb",  
"heap\_max\_in\_bytes" : 1636081139712  
},  
"threads" : 17227  
},  
"fs" : {  
"total" : "64.4tb",  
"total\_in\_bytes" : 70832357376000,  
"free" : "54.5tb",  
"free\_in\_bytes" : 59953917485056,  
"available" : "54.5tb",  
"available\_in\_bytes" : 59953917485056  
},  
"plugins" : ,  
"network\_types" : {  
"transport\_types" : {  
"security4" : 53  
},  
"http\_types" : {  
"security4" : 53  
}  
},  
"discovery\_types" : {  
"zen" : 53  
}  
}  
}

---

<div class="post-metadata">

### Author: ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)
#### Post date: [June 26, 2019, 7:13am UTC](https://discuss.elastic.co/t/elasticsearch-cause-linux-kernel-crash/187461/7 "2019-06-26T07:13:39Z")

</div>

How many indices and shards are you actively indexing into? How large are your documents? What bulk size are you using?

I can also see that 5 of your nodes are using a bundled newer JVM. All nodes should run the same version so I would recommend switching to the more up to date and supported one that is bundled with the distribution.

based on the stats it looks like you have reasonably large documents and that you may be updating them as well as indexing. Is this correct? If so, can you describe the load and use case in more detail?

---

<div class="post-metadata">

### Author: ![john\_am](https://avatars.discourse-cdn.com/v4/letter/j/5f9b8f/32.png) [@john\_am](https://discuss.elastic.co/u/john_am)
#### Post date: [June 26, 2019, 7:58am UTC](https://discuss.elastic.co/t/elasticsearch-cause-linux-kernel-crash/187461/8 "2019-06-26T07:58:01Z")

</div>

Your are right .The lagest index is nearly 1 tb(contain ont Replica) with 5 parimary shard . In most case we control the index with 5 shard , and the size limit 50g(each shard without Replica.). We are trying to using rollover to control the shard size.

---

<div class="post-metadata">

### Author: ![john\_am](https://avatars.discourse-cdn.com/v4/letter/j/5f9b8f/32.png) [@john\_am](https://discuss.elastic.co/u/john_am)
#### Post date: [June 26, 2019, 8:09am UTC](https://discuss.elastic.co/t/elasticsearch-cause-linux-kernel-crash/187461/9 "2019-06-26T08:09:18Z")

</div>

Thanks for your help. Firstly , i will fix the java version.  
We input our log by logstash to es , which daily insert nerly 1 tb( original data)（in the coming day ,we will input 3 tb per/day），here is our template setting:  
"settings" : {  
"index" : {  
"lifecycle" : {  
"name" : "cloud-edrive-download",  
"rollover\_alias" : "cloud-edrive-download"  
},  
"routing" : {  
"allocation" : {  
"require" : {  
"disk\_type" : "ssd"  
},  
"total\_shards\_per\_node" : "2"  
}  
},  
"refresh\_interval" : "10s",  
"number\_of\_shards" : "5",  
"number\_of\_replicas" : "1"

What's worse we just have 12 \* 2t ssd. So we just can using 6 instance as the hot node(each node with 2 \* 2 tb ssd).

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 24, 2019, 8:09am UTC](https://discuss.elastic.co/t/elasticsearch-cause-linux-kernel-crash/187461/10 "2019-07-24T08:09:18Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
