# "stacktrace": \["org.apache.lucene.index.CorruptIndexException: compound sub-files must have a valid codec header and footer: file is too small (0 bytes) (resource=BufferedChecksumIndexInput)

**URL:** <https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323>\
**Category:** Elasticsearch\
**Created:** [July 13, 2023, 11:40am UTC](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323 "2023-07-13T11:40:00Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Aravindh\_M](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aravindh_m/32/121876_2.png) [@Aravindh\_M](https://discuss.elastic.co/u/Aravindh_M)\
**Post date:** [July 13, 2023, 11:40am UTC](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323/1 "2023-07-13T11:40:00Z")

</div>

Recently, we have been encountering the "CorruptIndexException" frequently, accompanied by the following stacktrace: "**org.apache.lucene.index.CorruptIndexException: compound sub-files must have a valid codec header and footer: file is too small (0 bytes) (resource=BufferedChecksumIndexInput(NIOFSIndexInput(path="/usr/share/elasticsearch/data/nodes/0/indices/UAS5VDw1Sv6xrrrvoN39Bw/1/index/\_52.kdm")))**".

Upon checking the "\_52.kdm" file, we found that it actually contains 143 bytes. Has anyone else encountered a similar issue?

We are currently using Elasticsearch version 7.17.5 and spring-data-elasticsearch version 4.4.2.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [July 13, 2023, 1:17pm UTC](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323/2 "2023-07-13T13:17:09Z")

</div>

Almost certainly that means your storage doesn't work correctly under concurrent access. Are you using local disks or something network-attached?

See [these docs](https://www.elastic.co/guide/en/elasticsearch/reference/current/corruption-troubleshooting.html) for more information.

---

<div class="post-metadata">

**Author:** ![Aravindh\_M](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aravindh_m/32/121876_2.png) [@Aravindh\_M](https://discuss.elastic.co/u/Aravindh_M)\
**Post date:** [July 13, 2023, 1:35pm UTC](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323/3 "2023-07-13T13:35:56Z")

</div>

Thanks for your response @DavidTurner. We are using VM local disk[Gluster file system].

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 13, 2023, 1:41pm UTC](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323/4 "2023-07-13T13:41:24Z")

</div>

Have a look at the folowing, potentially related issues:

> [@ElastiSearch Supports GlusterFS and Rook Storage System?](https://discuss.elastic.co/t/elastisearch-supports-glusterfs-and-rook-storage-system/193319):
>
> Hi, Does ElastiSearch Supports GlusterFS and Rook Storage System? Regards, Varun S

> [@Elk on Docker Swarm and glusterFS crash](https://discuss.elastic.co/t/elk-on-docker-swarm-and-glusterfs-crash/127940):
>
> Hi, I'm trying to deploy an ELK stack on docker swarm. If I bind the elastic dir data to a Docker volume there is no problem. The problems comes as soon as I try to bind the elstastic data dir to a glusterFS volume. I use glusterFS to synchronise the data between all the swarm nodes in the cluster. I deploy ELK using the following code: elasticsearch: image: docker.elastic.co/elasticsearch/elasticsearch:6.2.3 # container\_name: elasticsearch environment: - "http.host=0.0.…

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [July 13, 2023, 1:48pm UTC](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323/5 "2023-07-13T13:48:28Z")

</div>

Yeah GlusterFS isn't at all a local disk and the error you're seeing indicates it does not behave like a local disk accurately enough for Elasticsearch. See [these docs](https://www.elastic.co/guide/en/elasticsearch/reference/current/modules-node.html#data-path) for more information:

> Elasticsearch requires the filesystem to act as if it were backed by a local disk, but this means that it will work correctly on properly-configured remote block devices (e.g. a SAN) and remote filesystems (e.g. NFS) as long as the remote storage behaves no differently from local storage.

---

<div class="post-metadata">

**Author:** ![Aravindh\_M](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aravindh_m/32/121876_2.png) [@Aravindh\_M](https://discuss.elastic.co/u/Aravindh_M)\
**Post date:** [July 13, 2023, 1:58pm UTC](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323/6 "2023-07-13T13:58:47Z")

</div>

Thanks for your response @Christian_Dahlqvist. I'll have a look.

---

<div class="post-metadata">

**Author:** ![Aravindh\_M](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aravindh_m/32/121876_2.png) [@Aravindh\_M](https://discuss.elastic.co/u/Aravindh_M)\
**Post date:** [July 13, 2023, 2:02pm UTC](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323/7 "2023-07-13T14:02:38Z")

</div>

Thanks @DavidTurner. We will go through the docs and will review our setup.

---

<div class="post-metadata">

**Author:** ![Aravindh\_M](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aravindh_m/32/121876_2.png) [@Aravindh\_M](https://discuss.elastic.co/u/Aravindh_M)\
**Post date:** [July 14, 2023, 11:40am UTC](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323/8 "2023-07-14T11:40:48Z")

</div>

@DavidTurner,  
We created a disk on SAN storage and set up a new VM with the disk. Docker and Docker Compose have been installed on the VM, and the same VM functions as a worker node within a swarm cluster.

On the VM, we created a directory named "/data" and mounted it as a GlusterFS volume to ensure persistent data across the worker nodes.

I would like to clarify that the Gluster setup is not based on a shared volume or disk. Each worker node has its own local disk, and the GlusterFS directory is mounted on each node to achieve data persistence.

Could you please confirm if this setup appears to be fine or if it might be the cause of the issue?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [July 14, 2023, 12:16pm UTC](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323/9 "2023-07-14T12:16:45Z")

</div>

I can confirm that there is definitely something in your storage setup that is _not_ fine (i.e. doesn't behave like a local disk as ES requires). Since the problem is outside of ES, I can't really help you pin it down further. But I am suspicious of GlusterFS because it has had problems like this before, and GlusterFS 7 in particular has been [EOL and unmaintained for years](https://www.gluster.org/release-schedule/).

As per the docs I linked above:

> To narrow down the source of the corruptions, systematically change components in your cluster’s environment until the corruptions stop.

In particular try using a more common filesystem instead of GlusterFS and see if the problems go away.

---

<div class="post-metadata">

**Author:** ![JeyakumarKarunanithi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jeyakumarkarunanithi/32/123825_2.png) [@JeyakumarKarunanithi](https://discuss.elastic.co/u/JeyakumarKarunanithi)\
**Post date:** [July 25, 2023, 6:41am UTC](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323/10 "2023-07-25T06:41:06Z")

</div>

Hi @DavidTurner ,  
We are planning to move On-Prem servers to Azure cloud and planning to have the Elasticsearch data in Azure files as persistent storage.

Do we have any known issues to have the Elasticsearch data on Azure files storage?

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [July 25, 2023, 6:59am UTC](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323/11 "2023-07-25T06:59:37Z")

</div>

I know of no issues with Azure persistent storage, but I am also not familiar with its various configuration options and also think we don't run many (any?) tests with it. If you encounter problems, you'll need to contact the Azure folks for help.

---

<div class="post-metadata">

**Author:** ![JeyakumarKarunanithi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jeyakumarkarunanithi/32/123825_2.png) [@JeyakumarKarunanithi](https://discuss.elastic.co/u/JeyakumarKarunanithi)\
**Post date:** [July 25, 2023, 7:19am UTC](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323/12 "2023-07-25T07:19:59Z")

</div>

Sure thank you, what are your recommendations as Shared storage options for running Elasticsearch with Docker swarm setup.?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 25, 2023, 7:34am UTC](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323/13 "2023-07-25T07:34:13Z")

</div>

It looks like [Azure File is distributed storage accessed via SMB or NFS](https://azure.microsoft.com/en-gb/products/storage/files/). This type of storage can often result in very poor performance and may not necessarily behave like local storage like David described is required. I therefore would not rule out that you may experience similar corruption issues with this but am also not aware of any reported issues.

I would recommend using premium or standard storage.

---

<div class="post-metadata">

**Author:** ![JeyakumarKarunanithi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jeyakumarkarunanithi/32/123825_2.png) [@JeyakumarKarunanithi](https://discuss.elastic.co/u/JeyakumarKarunanithi)\
**Post date:** [July 25, 2023, 7:43am UTC](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323/14 "2023-07-25T07:43:01Z")

</div>

Thanks @Christian_Dahlqvist , even though we goes with Standard/Premium disk, I cant move the container across the other worker node in the cluster since I do not the have the shared mount path and I will lost my data when the container moves to another node, So Im looking for a solution that I can run Elasticsearch container on any of the available node without losing the data/corrupting the data.

Looking for some suggestions on my need. Kindly help...

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [July 25, 2023, 7:45am UTC](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323/15 "2023-07-25T07:45:54Z")

</div>

> [@JeyakumarKarunanithi](#):
>
> options for running Elasticsearch with Docker swarm setup.?

I think the problem here is likely Docker Swarm - AIUI it doesn't really work very well with stateful applications like Elasticsearch.

---

<div class="post-metadata">

**Author:** ![JeyakumarKarunanithi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jeyakumarkarunanithi/32/123825_2.png) [@JeyakumarKarunanithi](https://discuss.elastic.co/u/JeyakumarKarunanithi)\
**Post date:** [July 25, 2023, 7:47am UTC](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323/16 "2023-07-25T07:47:14Z")

</div>

How about Kubernetes/Minikube..?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 25, 2023, 7:47am UTC](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323/17 "2023-07-25T07:47:20Z")

</div>

I have no experience using Docker Swarm so can not comment on that.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 25, 2023, 7:48am UTC](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323/18 "2023-07-25T07:48:31Z")

</div>

There are lots of users running Elasticsearch successfully on kubernetes. It does support persistent volumes so does not require shared storage the same way.

---

<div class="post-metadata">

**Author:** ![JeyakumarKarunanithi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jeyakumarkarunanithi/32/123825_2.png) [@JeyakumarKarunanithi](https://discuss.elastic.co/u/JeyakumarKarunanithi)\
**Post date:** [July 25, 2023, 7:51am UTC](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323/19 "2023-07-25T07:51:55Z")

</div>

Oh okay, when it comes to Kubernetes persistent volumes, there might be need to share the volume with other nodes to keep the pod highly available, also when we go with more than 1 replicas on multiple nodes, we need to keep the volume as shared between the nodes. then only the same data can be available across all the replicas.  
What type of volumes can be used for this requirement?

---

<div class="post-metadata">

**Author:** ![JeyakumarKarunanithi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jeyakumarkarunanithi/32/123825_2.png) [@JeyakumarKarunanithi](https://discuss.elastic.co/u/JeyakumarKarunanithi)\
**Post date:** [July 28, 2023, 7:14am UTC](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323/20 "2023-07-28T07:14:16Z")

</div>

@Christian_Dahlqvist @DavidTurner Looking for your kind suggestions on this

[Next page](https://discuss.elastic.co/t/stacktrace-org-apache-lucene-index-corruptindexexception-compound-sub-files-must-have-a-valid-codec-header-and-footer-file-is-too-small-0-bytes-resource-bufferedchecksumindexinput/338323.md?page=2)
