# Timeouts and index corruptions

**URL:** <https://discuss.elastic.co/t/timeouts-and-index-corruptions/22433>\
**Category:** Elasticsearch\
**Created:** [February 27, 2015, 2:18pm UTC](https://discuss.elastic.co/t/timeouts-and-index-corruptions/22433 "2015-02-27T14:18:29Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Gabriel\_Flory](https://avatars.discourse-cdn.com/v4/letter/g/8491ac/32.png) [@Gabriel\_Flory](https://discuss.elastic.co/u/Gabriel_Flory)\
**Post date:** [February 27, 2015, 2:18pm UTC](https://discuss.elastic.co/t/timeouts-and-index-corruptions/22433/1 "2015-02-27T14:18:29Z")

</div>

Hi,

We have quite a few problems with our ES cluster here after upgrading from  
1.2.1 to to 1.4.1 .  
The upgrade was done quite brutally, as the new version of es was installed  
and the cluster restarted by a scm.

First of all, after upgrading and restarting ES, multiple shards started  
giving us errors regarding checksums and preexisting corrupted indexes.

The lucene checkIndex tool claimed that said shards contained corrupted  
data, and recreated segments (with a bit of document loss, but that's all  
normal).  
The data that was lost was too important to us though and we decided to  
restore the index from a snapshot, and it worked (yay!).

All went well but a few hours later, when ES started getting a little bit  
more load, it started timing out on some queries.  
The timeout behavior seems random (when it comes to query types).  
Then the garbage collector started warning us on all indices :  
[2015-02-27 10:25:09,814][WARN][monitor.jvm] [node1]  
[gc][old][83449][548] duration [27s], collections [2]/[13.3s], total  
[27s]/[1.2h], memory [4.9gb]-\>[4.8gb]/[31.8gb], all\_pools {[young]  
[85.2mb]-\>[27.9mb]/[865.3mb]}{[survivor] [0b]-\>[0b]/[108.1mb]}{[old]  
[4.8gb]-\>[4.6gb]/[30.9gb]}  
I don't know exactly how to read these but it seems like something isn't  
right, right ?

So after loading the snapshot and writing the delta back into the index we  
performed another snapshot.  
That snapshot is actually written (although dreadfully slow to create)  
claimed "OK" by ES but when we try to restore an index from it, we get  
corrupted shards again. (duh !)

Sadly it appears that our java version doesn't fit the recommandations  
(7#51)  
Is there anything that we would need to do to upgrade the cluster  
"properly" ?

Also, the snapshots now take FOREVER to perform, is it normal ?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/539e8fe7-d879-4514-ade1-8f4e43d06369%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/539e8fe7-d879-4514-ade1-8f4e43d06369%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Gabriel\_Flory](https://avatars.discourse-cdn.com/v4/letter/g/8491ac/32.png) [@Gabriel\_Flory](https://discuss.elastic.co/u/Gabriel_Flory)\
**Post date:** [February 28, 2015, 1:15am UTC](https://discuss.elastic.co/t/timeouts-and-index-corruptions/22433/2 "2015-02-28T01:15:42Z")

</div>

Maybe this can help,

Since we restarted in the new version, with the same configured heap the  
memory consumption changed A LOT  
See memory graphs here

> **[WeTransfer - Send Large Files & Share Photos Online - Up to 2GB Free](https://wetransfer.com)**
>
> WeTransfer is the simplest way to send your files around the world

The node stats shows very little heap consumed (like 15%) but a lot of the  
other metrics seem off :  
"mem" : {  
"heap\_used\_in\_bytes" : 4934366152,  
"heap\_used\_percent" : 14,  
"heap\_committed\_in\_bytes" : 34246361088,  
"heap\_max\_in\_bytes" : 34246361088,  
"non\_heap\_used\_in\_bytes" : 108007424,  
"non\_heap\_committed\_in\_bytes" : 119799808,  
"pools" : {  
"young" : {  
"used\_in\_bytes" : 500486864,  
"max\_in\_bytes" : 907345920,  
"peak\_used\_in\_bytes" : 907345920,  
"peak\_max\_in\_bytes" : 907345920  
},  
"survivor" : {  
"used\_in\_bytes" : 63272264,  
"max\_in\_bytes" : 113377280,  
"peak\_used\_in\_bytes" : 113377280,  
"peak\_max\_in\_bytes" : 113377280  
},  
"old" : {  
"used\_in\_bytes" : 4370607024,  
"max\_in\_bytes" : 33225637888,  
"peak\_used\_in\_bytes" : 4443735936,  
"peak\_max\_in\_bytes" : 33225637888  
}  
}  
},

Is it helping ?

Le vendredi 27 février 2015 15:18:29 UTC+1, Gabriel Flory a écrit :

> Hi,
> 
> We have quite a few problems with our ES cluster here after upgrading from  
> 1.2.1 to to 1.4.1 .  
> The upgrade was done quite brutally, as the new version of es was  
> installed and the cluster restarted by a scm.
> 
> First of all, after upgrading and restarting ES, multiple shards started  
> giving us errors regarding checksums and preexisting corrupted indexes.
> 
> The lucene checkIndex tool claimed that said shards contained corrupted  
> data, and recreated segments (with a bit of document loss, but that's all  
> normal).  
> The data that was lost was too important to us though and we decided to  
> restore the index from a snapshot, and it worked (yay!).
> 
> All went well but a few hours later, when ES started getting a little bit  
> more load, it started timing out on some queries.  
> The timeout behavior seems random (when it comes to query types).  
> Then the garbage collector started warning us on all indices :  
> [2015-02-27 10:25:09,814][WARN][monitor.jvm] [node1]  
> [gc][old][83449][548] duration [27s], collections [2]/[13.3s], total  
> [27s]/[1.2h], memory [4.9gb]-\>[4.8gb]/[31.8gb], all\_pools {[young]  
> [85.2mb]-\>[27.9mb]/[865.3mb]}{[survivor] [0b]-\>[0b]/[108.1mb]}{[old]  
> [4.8gb]-\>[4.6gb]/[30.9gb]}  
> I don't know exactly how to read these but it seems like something isn't  
> right, right ?
> 
> So after loading the snapshot and writing the delta back into the index we  
> performed another snapshot.  
> That snapshot is actually written (although dreadfully slow to create)  
> claimed "OK" by ES but when we try to restore an index from it, we get  
> corrupted shards again. (duh !)
> 
> Sadly it appears that our java version doesn't fit the recommandations  
> (7#51)  
> Is there anything that we would need to do to upgrade the cluster  
> "properly" ?
> 
> Also, the snapshots now take FOREVER to perform, is it normal ?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/b72c996b-89f0-46a4-ba29-71ad549d5075%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/b72c996b-89f0-46a4-ba29-71ad549d5075%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:29am UTC](https://discuss.elastic.co/t/timeouts-and-index-corruptions/22433/3 "2017-07-06T00:29:19Z")

</div>


