# Indexes seem corrupted

**URL:** <https://discuss.elastic.co/t/indexes-seem-corrupted/3581>\
**Category:** Elasticsearch\
**Created:** [November 17, 2010, 5:33pm UTC](https://discuss.elastic.co/t/indexes-seem-corrupted/3581 "2010-11-17T17:33:47Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![John\_Chang](https://avatars.discourse-cdn.com/v4/letter/j/ecc23a/32.png) [@John\_Chang](https://discuss.elastic.co/u/John_Chang)\
**Post date:** [November 17, 2010, 5:33pm UTC](https://discuss.elastic.co/t/indexes-seem-corrupted/3581/1 "2010-11-17T17:33:47Z")

</div>

We are worried are indexes are corrupted for a number of reasons. We are looking through the logs to see what might have happened, but are still without a grasp on it. Any advice on understanding, trouble-shooting, and preventing what we are seeing would be greatly appreciated. Thanks.

1. We keep 4 document types; they all used to have desired mappings, now 2 of the 4 seem to be missing the mappings. Our system maps all 4 types at once and we are confident those mappings used to be there for all types.

2. We lost a lot of documents; we do a count, and there are fraction remaining of what used to be there.

3. We are getting an error we've never seen before (see below). The document type in question here does still seem to have the correct mappings.

[Failed to execute main query]]; nested: CompileException[[Error: Invalid shift value in prefixCoded string (is encoded value really an INT?)]\n[Near : {... Unknown ....}]\n ^\n[Line: 1, Column: 0]]; nested: NumberFormatException[Invalid shift value in prefixCoded string (is encoded value really an INT?)]; }{[6fe786b4-de13-451c-8296-7803b8bbe1d8][index0][2]: RemoteTransportException[[Angel][inet[/10.198.109.171:9300]][search/phase/query]]; nested: QueryPhaseExecutionException[[index0][2]: query[custom score (+userId:4c6b25774f8bd5147ab46cf4 +(body:"john smith" subject:"john smith" to:"john smith" from:"john smith" cc:\john smith"),function=org.elasticsearch.index.query.xcontent.CustomScoreQueryParser$ScriptScoreFunction@7daf32e3)],from[0],size[100]: Query Failed

---

<div class="post-metadata">

**Author:** ![John\_Chang](https://avatars.discourse-cdn.com/v4/letter/j/ecc23a/32.png) [@John\_Chang](https://discuss.elastic.co/u/John_Chang)\
**Post date:** [November 17, 2010, 5:43pm UTC](https://discuss.elastic.co/t/indexes-seem-corrupted/3581/2 "2010-11-17T17:43:59Z")

</div>

I should add that the index was created on Elastic Search 0.11 and we upgraded to 0.12.1, without reindexing (which we understood to be not necessary as we are not doing geo searches). We tested it after the upgrade and it seemed fine then; not sure when it went off the rails.

Not expecting this has to do with the upgrade, but just wanted to call it out just in case it was useful info.

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [November 17, 2010, 6:01pm UTC](https://discuss.elastic.co/t/indexes-seem-corrupted/3581/3 "2010-11-17T18:01:13Z")

</div>

Hi John

On Wed, 2010-11-17 at 09:43 -0800, John Chang wrote:

> I should add that the index was created on Elastic Search 0.11 and we  
> upgraded to 0.12.1, without reindexing (which we understood to be not  
> necessary as we are not doing geo searches). We tested it after the upgrade  
> and it seemed fine then; not sure when it went off the rails.
> 
> Not expecting this has to do with the upgrade, but just wanted to call it  
> out just in case it was useful info.

This does sound like your indexed have been corrupted somewhere along  
the way. You may have been hit by this bug:

> <https://github.com/elastic/elasticsearch/issues/466>
>
> The problem stems from the reusing of existing index files when doing recovery. …Checksums should be added, but, in order to have it performant, they are computed on write.
> 
> This menas that existing indices will still work, but might suffer from it. New internal index files will get checksummed, and eventually the index will be fully checksummed, though, it is recommended to reindex the data.

Although I'm not sure if that would result in you losing mappings.

Would be worth gist'ing your logs: [https://gist.github.com/](https://gist.github.com/)

clint

---

<div class="post-metadata">

**Author:** ![John\_Chang](https://avatars.discourse-cdn.com/v4/letter/j/ecc23a/32.png) [@John\_Chang](https://discuss.elastic.co/u/John_Chang)\
**Post date:** [November 17, 2010, 8:23pm UTC](https://discuss.elastic.co/t/indexes-seem-corrupted/3581/4 "2010-11-17T20:23:52Z")

</div>

Here is a gist of the elastic search logs. However, I don't know if they will useful; they just log some activity about 2 hours before I started seeing the problems noted above in my application logs, and they seem pretty tame:

> <https://gist.github.com/anonymous/703964>

Here is some more info from my application log. It is basically more of what I put in the original post:

> <https://gist.github.com/anonymous/704009>

I don't know if this is useful, but I can't think of anything more to post. Let me know if there's something else that I'm missing.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [November 17, 2010, 8:57pm UTC](https://discuss.elastic.co/t/indexes-seem-corrupted/3581/5 "2010-11-17T20:57:37Z")

</div>

It might relate to the possible corruption that might happen that was fixed  
in master (upcoming 0.13). I also fixed a possible race condition between  
the recovery of an index and the creation of its mappings and an index  
operation getting in between the two (the new full cluster and index level  
blocks). It sounds like you might have hit both of them... . I assume you  
use local gateway?

-shay.banon

On Wed, Nov 17, 2010 at 10:23 PM, John Chang [jchangkihtest2@gmail.com](mailto:jchangkihtest2@gmail.com)wrote:

> Here is a gist of the Elasticsearch logs. However, I don't know if they  
> will useful; they just log some activity about 2 hours before I started  
> seeing the problems noted above in my application logs, and they seem  
> pretty  
> tame:  
> [Elastic Search Nodes Logs · GitHub](https://gist.github.com/703964)
> 
> Here is some more info from my application log. It is basically more of  
> what I put in the original post:  
> [Logs from my search app with a data only node · GitHub](https://gist.github.com/704009)
> 
> I don't know if this is useful, but I can't think of anything more to post.  
> Let me know if there's something else that I'm missing.
> 
> --  
> View this message in context:  
> [http://elasticsearch-users.115913.n3.nabble.com/Indexes-seem-corrupted-tp1918553p1919499.html](http://elasticsearch-users.115913.n3.nabble.com/Indexes-seem-corrupted-tp1918553p1919499.html)  
> Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).

---

<div class="post-metadata">

**Author:** ![John\_Chang](https://avatars.discourse-cdn.com/v4/letter/j/ecc23a/32.png) [@John\_Chang](https://discuss.elastic.co/u/John_Chang)\
**Post date:** [November 17, 2010, 10:29pm UTC](https://discuss.elastic.co/t/indexes-seem-corrupted/3581/6 "2010-11-17T22:29:15Z")

</div>

I think that's the problem. Yes, we are using local search. Also, what you (kimchy) write makes sense, as the Elastic Search data node logs here [https://gist.github.com/703964](https://gist.github.com/703964) show initialization at times that correspond perfectly to when the searches started going bad in our application log (which uses the no-data nodes).

The only thing I wonder is...why did the Elastic Search data nodes decide to reinitialize at that time; we did restart the data node cluster, but that was over 2 hours before this initialization in those logs. What kicks off the initialization other than a service restart?

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [November 17, 2010, 10:35pm UTC](https://discuss.elastic.co/t/indexes-seem-corrupted/3581/7 "2010-11-17T22:35:34Z")

</div>

It seems like the network connection got completely broken between the nodes  
(you see the transport disconnect reason for nodes being identified as  
failed).

You can try and set: discovery.zen.fd.connect\_on\_network\_disconnect to true,  
which in such event will try and connect again to the node in question to  
make sure it can't be connected.

-shay.banon

On Thu, Nov 18, 2010 at 12:29 AM, John Chang [jchangkihtest2@gmail.com](mailto:jchangkihtest2@gmail.com)wrote:

> I think that's the problem. Yes, we are using local search. Also, what  
> you  
> (kimchy) write makes sense, as the Elastic Search data node logs here  
> [Elastic Search Nodes Logs · GitHub](https://gist.github.com/703964) show initialization at times that  
> correspond  
> perfectly to when the searches started going bad in our application log  
> (which uses the no-data nodes).
> 
> ## The only thing I wonder is...why did the Elastic Search data nodes decide to reinitialize at that time; we did restart the data node cluster, but that was over 2 hours before this initialization in those logs. What kicks off the initialization other than a service restart?
> 
> View this message in context:  
> [http://elasticsearch-users.115913.n3.nabble.com/Indexes-seem-corrupted-tp1918553p1920227.html](http://elasticsearch-users.115913.n3.nabble.com/Indexes-seem-corrupted-tp1918553p1920227.html)  
> Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:16am UTC](https://discuss.elastic.co/t/indexes-seem-corrupted/3581/8 "2017-07-06T04:16:21Z")

</div>


