I'm experiencing a recurring kernel crash on a Debian 13 VM running Kibana 9.5.4 (I have had this issue for a while now and have upgraded Kibana and the Linux kernel multiple times). The crash occurs in the Linux kernel while a Kibana V8 worker is exiting, and leaves the VM partially unresponsive.
Environment
- Debian 13
- Linux kernel:
6.12.107+deb13-amd64(6.12.107-1) - Kibana:
9.5.4 - Node.js:
24.19.0(Kibana's bundled Node runtime) - V8:
13.6.233.17-node.51 - Elasticsearch:
9.5.4 - VM: 4 vCPUs, 32 GiB RAM
- Proxmox/QEMU virtual machine
virtio-scsistorage- x86-64
Kibana is running with its bundled Node executable:
/usr/share/kibana/node/default/bin/node
process.versions reports:
node: '24.19.0'
v8: '13.6.233.17-node.51'
uv: '1.52.1'
openssl: '3.5.7'
Failure
The first significant kernel error is:
Oct 01 03:32:03 elkstack kernel: BUG: kernel NULL pointer dereference, address: 0000000000000650
Oct 01 03:32:03 elkstack kernel: #PF: supervisor read access in kernel mode
Oct 01 03:32:03 elkstack kernel: #PF: error_code(0x0000) - not-present page
Oct 01 03:32:03 elkstack kernel: Oops: Oops: 0000 [#1] PREEMPT SMP NOPTI
The process was:
CPU: 2 UID: 102 PID: 1747 Comm: V8Worker
The fault occurred here:
RIP: __lruvec_stat_mod_folio+0x60/0xd0
The relevant call trace was:
__lruvec_stat_mod_folio
folio_remove_rmap_ptes
unmap_page_range
unmap_vmas
exit_mmap
__mmput
do_exit
do_group_exit
get_signal
arch_do_signal_or_restart
irqentry_exit_to_user_mode
The kernel then reported:
note: V8Worker[1747] exited with irqs disabled
note: V8Worker[1747] exited with preempt_count 1
Fixing recursive fault but reboot is needed!
Immediately after there was an RCU warning:
Voluntary context switch within RCU read-side critical section!
WARNING: CPU: 2 PID: 1747 at kernel/rcu/tree_plugin.h:331
and approximately 20 seconds later:
rcu: INFO: rcu_preempt detected stalls on CPUs/tasks:
rcu: Tasks blocked on level-0 rcu-node (CPUs 0-3): P1747/2:b..l
The RCU stall then repeated approximately every 63 seconds with the same task/generation and t increases by ~15,755 jiffies each time. This has persisted continuously since October 1st.
Resulting behaviour
The VM remained reachable by ICMP/ping, but SSH stopped responding and the console became unusable. The VM remained in this state until it was rebooted.
This appears consistent with the kernel Oops occurring first and the subsequent RCU stalls being a consequence of the kernel being left in an unrecoverable state.
Resource usage
This does not appear to be a conventional CPU or memory exhaustion issue.
At the time of investigation, the Proxmox host had approximately:
- 128 GiB RAM
- 105 GiB free
- ~85% CPU idle
- negligible I/O wait
- ~8 GiB swap configured, almost entirely free
The VM itself has 32 GiB RAM and 4 vCPUs.
The Proxmox host also showed no obvious CPU, memory, or I/O pressure associated with the VM.