# Why use checksum to verify segment files？

**URL:** <https://discuss.elastic.co/t/why-use-checksum-to-verify-segment-files/208636>\
**Category:** Elasticsearch\
**Created:** [November 20, 2019, 7:16am UTC](https://discuss.elastic.co/t/why-use-checksum-to-verify-segment-files/208636 "2019-11-20T07:16:34Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![cigarl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cigarl/32/56745_2.png) [@cigarl](https://discuss.elastic.co/u/cigarl)\
**Post date:** [November 20, 2019, 7:16am UTC](https://discuss.elastic.co/t/why-use-checksum-to-verify-segment-files/208636/1 "2019-11-20T07:16:34Z")

</div>

When I restart the cluster,I found that many copies are copied from 0，the reason is that the checksum of the segment files is inconsistent,and then I wrote a demo by using the Lucene interface，the result is the same operation, and the checksum is inconsistent.  
So why use checksum to verify the consistency of Lucene files？This may cause all segment files to be copied each time when the replica is recovering？

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [November 20, 2019, 8:03am UTC](https://discuss.elastic.co/t/why-use-checksum-to-verify-segment-files/208636/2 "2019-11-20T08:03:24Z")

</div>

> [@cigarl](#):
>
> So why use checksum to verify the consistency of Lucene files?

Because it is vital that the data read from disk is the data that was previously written. A checksum mismatch indicates that some of the data has been changed since it was written, which means it cannot be trusted. The usual reason for this is faulty storage hardware.

---

<div class="post-metadata">

**Author:** ![cigarl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cigarl/32/56745_2.png) [@cigarl](https://discuss.elastic.co/u/cigarl)\
**Post date:** [November 20, 2019, 8:21am UTC](https://discuss.elastic.co/t/why-use-checksum-to-verify-segment-files/208636/4 "2019-11-20T08:21:00Z")

</div>

But I found that consistent operations can also produce different checksums. Is it too easy to use cheksum alone?  
In addition, sync\_id cannot guarantee that the primary shard and copies are consistent when the cluster is restarted.Because after the sync\_id of the primary shard is updated, the replicas may not be restored and the sync\_id has not been updated synchronously.

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [November 20, 2019, 9:30am UTC](https://discuss.elastic.co/t/why-use-checksum-to-verify-segment-files/208636/6 "2019-11-20T09:30:33Z")

</div>

Sorry, I did not really understand your question because it seems to be confusing a number of distinct concepts. Checksums are used to verify segment files (and other things) locally to protect against corruption. But you seem to be asking about peer recovery too. Checksums are verified during peer recovery (also to protect against corruption) but this is unrelated to the sync id used to detect whether any recovery is needed.

On closer reading I think you are actually asking about replica allocation after a restart. We recently merged a big improvement to how replicas are allocated after a restart in [#46959](https://github.com/elastic/elasticsearch/pull/46959) to compare the contents of shard copies rather than using the checksums of the underlying segment files. Is this what you are asking about?

---

<div class="post-metadata">

**Author:** ![cigarl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cigarl/32/56745_2.png) [@cigarl](https://discuss.elastic.co/u/cigarl)\
**Post date:** [November 20, 2019, 9:55am UTC](https://discuss.elastic.co/t/why-use-checksum-to-verify-segment-files/208636/8 "2019-11-20T09:55:09Z")

</div>

Thank you for your reply！Actually，I have two questions：

1. After restart, when allocating replica shards, the checksum will be checked to ensure whether the current node is the most matching node.I found that the replica shard was assigned to a new node because the checksum verification was inconsistent（The verification shows that the checksums of the segment files that with same operations are inconsistent,Is it appropriate to use checksum in this case？）.
2. The _phase1_ of _PEER RECOVERY_ ，will check whether the checksum is consistent again.If I want to skip _phase1_ ,then I have to make sure that the _sync\_id_ is consistent（As I said above, it's hard to ensure that the _sync\_id_ are consistent）.

Thanks again！

---

<div class="post-metadata">

**Author:** ![cigarl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cigarl/32/56745_2.png) [@cigarl](https://discuss.elastic.co/u/cigarl)\
**Post date:** [November 20, 2019, 10:06am UTC](https://discuss.elastic.co/t/why-use-checksum-to-verify-segment-files/208636/9 "2019-11-20T10:06:48Z")

</div>

In addition, I have seen the optimization content of the new version, but the version I am using has not modified this part of the logic. Therefore, I want to know what the purpose of the previous version by using _checksum_ to find the best macthing node

> [@cigarl](#):
>
> The verification shows that the checksums of the segment files that with same operations are inconsistent,Is it appropriate to use checksum in this case？

Thank you

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [November 20, 2019, 10:25am UTC](https://discuss.elastic.co/t/why-use-checksum-to-verify-segment-files/208636/10 "2019-11-20T10:25:04Z")

</div>

> [@cigarl](#):
>
> I have two questions：

Sorry, it is unclear what you're asking. You have (fairly accurately) described how things work in older versions, but there aren't really any questions to answer there.

> [@cigarl](#):
>
> what the purpose of the previous version by using _checksum_ to find the best macthing node

Looking for identical segments is easy and fast and reliable and often quite accurate too: it's often highly likely that older indices will have identical segments since they will have been subject to a file-based recovery at some point in the past. In contrast, it took many many years of effort to build the groundwork needed to implement [#46959](https://github.com/elastic/elasticsearch/pull/46959).

---

<div class="post-metadata">

**Author:** ![cigarl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cigarl/32/56745_2.png) [@cigarl](https://discuss.elastic.co/u/cigarl)\
**Post date:** [November 20, 2019, 11:47am UTC](https://discuss.elastic.co/t/why-use-checksum-to-verify-segment-files/208636/12 "2019-11-20T11:47:40Z")

</div>

In other words, _checksum_ does not guarantee the consistency of segment files.Therefore, is it too strict to use _checksum_ to judge whether the segment files are consistent in the older verison ?

For example，the content of my segment files are consistent，but I still can't guarantee that their _checksums_ are consistent.

Because I did the **same operation** on the **same index** by using _The Interface Of Lucene_ , then compared the _\_0.cfs_ files and find that the _cheksum_ of the two files ( _\_0.cfs_ ) **is not the same.**

Maybe this is what must be done in the process of version evolution.

Thank you for your patience！

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 4, 2022, 7:34am UTC](https://discuss.elastic.co/t/why-use-checksum-to-verify-segment-files/208636/13 "2022-11-04T07:34:05Z")

</div>


