# Rebuilding corrupted index metadata

**URL:** <https://discuss.elastic.co/t/rebuilding-corrupted-index-metadata/5489>\
**Category:** Elasticsearch\
**Created:** [September 30, 2011, 5:01pm UTC](https://discuss.elastic.co/t/rebuilding-corrupted-index-metadata/5489 "2011-09-30T17:01:29Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Curtis\_Caravone](https://avatars.discourse-cdn.com/v4/letter/c/e47c2d/32.png) [@Curtis\_Caravone](https://discuss.elastic.co/u/Curtis_Caravone)\
**Post date:** [September 30, 2011, 5:01pm UTC](https://discuss.elastic.co/t/rebuilding-corrupted-index-metadata/5489/1 "2011-09-30T17:01:29Z")

</div>

Hey all,

I thought I would share an experience I had recently in case it helps  
anybody having a similar problem:

We are in the process up upgrading from es 0.16.2 to the latest 0.17.7. We  
did this by shutting down the cluster  
and starting up the 0.17.7 instances pointing to our existing data  
directories (we use local gateway).

We're not sure how it happened, but when the new cluster came up, it never  
left red state and it showed no indices existing.  
Switching back to 0.16.2, even restoring the data directory from a backup  
didn't help.

So, here's what we did:

1. Delete the data directories completely (we had a backup elsewhere)
2. Start up a clean es 0.17.7 and wait for green (no indices to wait for,  
of course)
3. Issue create index / put mapping commands to recreate all our index  
definitions and mappings (same # of shards, same mappings, etc. as before)
4. Shutdown the cluster and copy the data index directories only (not the  
\_state directories) over from the backup
5. Startup the cluster -- all indices came up green and had all our data!
  - Note that es seems to delete any index directories that don't match  
up with existing indices, so make sure you hang on the the backup until you  
are sure

On the next environment we tried this on, we first flushed the transaction  
logs before shutting down the cluster and upgrading. Everything went  
smoothly.  
I don't know if flushing had anything to do with it, or if the first problem  
was kind of a freak occurrence, but I thought I would mention it.

Curtis

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [October 2, 2011, 12:59pm UTC](https://discuss.elastic.co/t/rebuilding-corrupted-index-metadata/5489/2 "2011-10-02T12:59:34Z")

</div>

this method will only work if it ends up with the same shard distribution  
across the cluster. did any upgrade you tried from 0.16.2 to 0.17.7 caused  
missing data?

On Fri, Sep 30, 2011 at 8:01 PM, Curtis Caravone [caravone@gmail.com](mailto:caravone@gmail.com) wrote:

> Hey all,
> 
> I thought I would share an experience I had recently in case it helps  
> anybody having a similar problem:
> 
> We are in the process up upgrading from es 0.16.2 to the latest 0.17.7. We  
> did this by shutting down the cluster  
> and starting up the 0.17.7 instances pointing to our existing data  
> directories (we use local gateway).
> 
> We're not sure how it happened, but when the new cluster came up, it never  
> left red state and it showed no indices existing.  
> Switching back to 0.16.2, even restoring the data directory from a backup  
> didn't help.
> 
> So, here's what we did:
> 
> 1. Delete the data directories completely (we had a backup elsewhere)
> 2. Start up a clean es 0.17.7 and wait for green (no indices to wait for,  
> of course)
> 3. Issue create index / put mapping commands to recreate all our index  
> definitions and mappings (same # of shards, same mappings, etc. as before)
> 4. Shutdown the cluster and copy the data index directories only (not the  
> \_state directories) over from the backup
> 5. Startup the cluster -- all indices came up green and had all our data!
> - Note that es seems to delete any index directories that don't match  
> up with existing indices, so make sure you hang on the the backup until you  
> are sure
> 
> On the next environment we tried this on, we first flushed the transaction  
> logs before shutting down the cluster and upgrading. Everything went  
> smoothly.  
> I don't know if flushing had anything to do with it, or if the first  
> problem was kind of a freak occurrence, but I thought I would mention it.
> 
> Curtis

---

<div class="post-metadata">

**Author:** ![Curtis\_Caravone](https://avatars.discourse-cdn.com/v4/letter/c/e47c2d/32.png) [@Curtis\_Caravone](https://discuss.elastic.co/u/Curtis_Caravone)\
**Post date:** [October 2, 2011, 4:25pm UTC](https://discuss.elastic.co/t/rebuilding-corrupted-index-metadata/5489/3 "2011-10-02T16:25:12Z")

</div>

That's a good point. In our case, we are starting with three nodes and two  
replicas, so every node has all the shards.

To answer your question:

The first one we tried was a single-node dev instance. It failed and we had  
to rebuild the metatdata.

The second one we tried was also single-node, and it succeeded without a  
problem.

The third one we tried was the three-node production instance. It had the  
same failure as the first dev instance.

In the failure cases, there didn't seem to be anything in the logs, even at  
trace level, except a "starting" message.

Curtis

On Sun, Oct 2, 2011 at 5:59 AM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:

> this method will only work if it ends up with the same shard distribution  
> across the cluster. did any upgrade you tried from 0.16.2 to 0.17.7 caused  
> missing data?
> 
> On Fri, Sep 30, 2011 at 8:01 PM, Curtis Caravone [caravone@gmail.com](mailto:caravone@gmail.com)wrote:
> 
> > Hey all,
> > 
> > I thought I would share an experience I had recently in case it helps  
> > anybody having a similar problem:
> > 
> > We are in the process up upgrading from es 0.16.2 to the latest 0.17.7.  
> > We did this by shutting down the cluster  
> > and starting up the 0.17.7 instances pointing to our existing data  
> > directories (we use local gateway).
> > 
> > We're not sure how it happened, but when the new cluster came up, it never  
> > left red state and it showed no indices existing.  
> > Switching back to 0.16.2, even restoring the data directory from a backup  
> > didn't help.
> > 
> > So, here's what we did:
> > 
> > 1. Delete the data directories completely (we had a backup elsewhere)
> > 2. Start up a clean es 0.17.7 and wait for green (no indices to wait for,  
> > of course)
> > 3. Issue create index / put mapping commands to recreate all our index  
> > definitions and mappings (same # of shards, same mappings, etc. as before)
> > 4. Shutdown the cluster and copy the data index directories only (not the  
> > \_state directories) over from the backup
> > 5. Startup the cluster -- all indices came up green and had all our data!
> > - Note that es seems to delete any index directories that don't match  
> > up with existing indices, so make sure you hang on the the backup until you  
> > are sure
> > 
> > On the next environment we tried this on, we first flushed the transaction  
> > logs before shutting down the cluster and upgrading. Everything went  
> > smoothly.  
> > I don't know if flushing had anything to do with it, or if the first  
> > problem was kind of a freak occurrence, but I thought I would mention it.
> > 
> > Curtis

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [October 2, 2011, 10:14pm UTC](https://discuss.elastic.co/t/rebuilding-corrupted-index-metadata/5489/4 "2011-10-02T22:14:32Z")

</div>

Strange regarding the failure... . Can you recreate it in some way. I will  
try myself to run an upgrade from 0.16.2 to 0.17.7 in different scenarios,  
would help if you can try and pin point the steps taken to recreate it.

On Sun, Oct 2, 2011 at 6:25 PM, Curtis Caravone [caravone@gmail.com](mailto:caravone@gmail.com) wrote:

> That's a good point. In our case, we are starting with three nodes and two  
> replicas, so every node has all the shards.
> 
> To answer your question:
> 
> The first one we tried was a single-node dev instance. It failed and we  
> had to rebuild the metatdata.
> 
> The second one we tried was also single-node, and it succeeded without a  
> problem.
> 
> The third one we tried was the three-node production instance. It had the  
> same failure as the first dev instance.
> 
> In the failure cases, there didn't seem to be anything in the logs, even at  
> trace level, except a "starting" message.
> 
> Curtis
> 
> On Sun, Oct 2, 2011 at 5:59 AM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> 
> > this method will only work if it ends up with the same shard distribution  
> > across the cluster. did any upgrade you tried from 0.16.2 to 0.17.7 caused  
> > missing data?
> > 
> > On Fri, Sep 30, 2011 at 8:01 PM, Curtis Caravone [caravone@gmail.com](mailto:caravone@gmail.com)wrote:
> > 
> > > Hey all,
> > > 
> > > I thought I would share an experience I had recently in case it helps  
> > > anybody having a similar problem:
> > > 
> > > We are in the process up upgrading from es 0.16.2 to the latest 0.17.7.  
> > > We did this by shutting down the cluster  
> > > and starting up the 0.17.7 instances pointing to our existing data  
> > > directories (we use local gateway).
> > > 
> > > We're not sure how it happened, but when the new cluster came up, it  
> > > never left red state and it showed no indices existing.  
> > > Switching back to 0.16.2, even restoring the data directory from a backup  
> > > didn't help.
> > > 
> > > So, here's what we did:
> > > 
> > > 1. Delete the data directories completely (we had a backup elsewhere)
> > > 2. Start up a clean es 0.17.7 and wait for green (no indices to wait  
> > > for, of course)
> > > 3. Issue create index / put mapping commands to recreate all our index  
> > > definitions and mappings (same # of shards, same mappings, etc. as before)
> > > 4. Shutdown the cluster and copy the data index directories only (not  
> > > the \_state directories) over from the backup
> > > 5. Startup the cluster -- all indices came up green and had all our  
> > > data!
> > > - Note that es seems to delete any index directories that don't  
> > > match up with existing indices, so make sure you hang on the the backup  
> > > until you are sure
> > > 
> > > On the next environment we tried this on, we first flushed the  
> > > transaction logs before shutting down the cluster and upgrading. Everything  
> > > went smoothly.  
> > > I don't know if flushing had anything to do with it, or if the first  
> > > problem was kind of a freak occurrence, but I thought I would mention it.
> > > 
> > > Curtis

---

<div class="post-metadata">

**Author:** ![Curtis\_Caravone](https://avatars.discourse-cdn.com/v4/letter/c/e47c2d/32.png) [@Curtis\_Caravone](https://discuss.elastic.co/u/Curtis_Caravone)\
**Post date:** [October 2, 2011, 10:16pm UTC](https://discuss.elastic.co/t/rebuilding-corrupted-index-metadata/5489/5 "2011-10-02T22:16:36Z")

</div>

Ok, I'll see what I can do in recreating it.

Curtis

On Sun, Oct 2, 2011 at 3:14 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:

> Strange regarding the failure... . Can you recreate it in some way. I will  
> try myself to run an upgrade from 0.16.2 to 0.17.7 in different scenarios,  
> would help if you can try and pin point the steps taken to recreate it.
> 
> On Sun, Oct 2, 2011 at 6:25 PM, Curtis Caravone [caravone@gmail.com](mailto:caravone@gmail.com)wrote:
> 
> > That's a good point. In our case, we are starting with three nodes and  
> > two replicas, so every node has all the shards.
> > 
> > To answer your question:
> > 
> > The first one we tried was a single-node dev instance. It failed and we  
> > had to rebuild the metatdata.
> > 
> > The second one we tried was also single-node, and it succeeded without a  
> > problem.
> > 
> > The third one we tried was the three-node production instance. It had the  
> > same failure as the first dev instance.
> > 
> > In the failure cases, there didn't seem to be anything in the logs, even  
> > at trace level, except a "starting" message.
> > 
> > Curtis
> > 
> > On Sun, Oct 2, 2011 at 5:59 AM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> > 
> > > this method will only work if it ends up with the same shard distribution  
> > > across the cluster. did any upgrade you tried from 0.16.2 to 0.17.7 caused  
> > > missing data?
> > > 
> > > On Fri, Sep 30, 2011 at 8:01 PM, Curtis Caravone [caravone@gmail.com](mailto:caravone@gmail.com)wrote:
> > > 
> > > > Hey all,
> > > > 
> > > > I thought I would share an experience I had recently in case it helps  
> > > > anybody having a similar problem:
> > > > 
> > > > We are in the process up upgrading from es 0.16.2 to the latest 0.17.7.  
> > > > We did this by shutting down the cluster  
> > > > and starting up the 0.17.7 instances pointing to our existing data  
> > > > directories (we use local gateway).
> > > > 
> > > > We're not sure how it happened, but when the new cluster came up, it  
> > > > never left red state and it showed no indices existing.  
> > > > Switching back to 0.16.2, even restoring the data directory from a  
> > > > backup didn't help.
> > > > 
> > > > So, here's what we did:
> > > > 
> > > > 1. Delete the data directories completely (we had a backup elsewhere)
> > > > 2. Start up a clean es 0.17.7 and wait for green (no indices to wait  
> > > > for, of course)
> > > > 3. Issue create index / put mapping commands to recreate all our index  
> > > > definitions and mappings (same # of shards, same mappings, etc. as before)
> > > > 4. Shutdown the cluster and copy the data index directories only (not  
> > > > the \_state directories) over from the backup
> > > > 5. Startup the cluster -- all indices came up green and had all our  
> > > > data!
> > > > - Note that es seems to delete any index directories that don't  
> > > > match up with existing indices, so make sure you hang on the the backup  
> > > > until you are sure
> > > > 
> > > > On the next environment we tried this on, we first flushed the  
> > > > transaction logs before shutting down the cluster and upgrading. Everything  
> > > > went smoothly.  
> > > > I don't know if flushing had anything to do with it, or if the first  
> > > > problem was kind of a freak occurrence, but I thought I would mention it.
> > > > 
> > > > Curtis

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [October 3, 2011, 8:47am UTC](https://discuss.elastic.co/t/rebuilding-corrupted-index-metadata/5489/6 "2011-10-03T08:47:52Z")

</div>

On Mon, 2011-10-03 at 00:14 +0200, Shay Banon wrote:

> Strange regarding the failure... . Can you recreate it in some way. I  
> will try myself to run an upgrade from 0.16.2 to 0.17.7 in different  
> scenarios, would help if you can try and pin point the steps taken to  
> recreate it.

Might this not be the bug that was incorrectly adding delete-by-query to  
the translogs?

clint

> On Sun, Oct 2, 2011 at 6:25 PM, Curtis Caravone [caravone@gmail.com](mailto:caravone@gmail.com)  
> wrote:  
> That's a good point. In our case, we are starting with three  
> nodes and two replicas, so every node has all the shards.
> 
> ```
> To answer your question:
>     
>     
> The first one we tried was a single-node dev instance. It
> failed and we had to rebuild the metatdata.
>     
>     
> The second one we tried was also single-node, and it succeeded
> without a problem.
>     
>     
> The third one we tried was the three-node production
> instance. It had the same failure as the first dev instance.
>     
>     
> In the failure cases, there didn't seem to be anything in the
> logs, even at trace level, except a "starting" message.
>     
>     
> Curtis
>     
>     
>     
>     
> On Sun, Oct 2, 2011 at 5:59 AM, Shay Banon <kimchy@gmail.com>
> wrote:
> this method will only work if it ends up with the same
> shard distribution across the cluster. did any upgrade
> you tried from 0.16.2 to 0.17.7 caused missing data?
>             
>             
>             
> On Fri, Sep 30, 2011 at 8:01 PM, Curtis Caravone
> <caravone@gmail.com> wrote:
> Hey all,
>                     
>                     
> I thought I would share an experience I had
> recently in case it helps anybody having a
> similar problem:
>                     
>                     
> We are in the process up upgrading from es
> 0.16.2 to the latest 0.17.7. We did this by
> shutting down the cluster
> and starting up the 0.17.7 instances pointing
> to our existing data directories (we use local
> gateway).
>                     
>                     
> We're not sure how it happened, but when the
> new cluster came up, it never left red state
> and it showed no indices existing.
> Switching back to 0.16.2, even restoring the
> data directory from a backup didn't help.
>                     
>                     
> So, here's what we did:
>                     
>                     
> 1) Delete the data directories completely (we
> had a backup elsewhere)
> 2) Start up a clean es 0.17.7 and wait for
> green (no indices to wait for, of course)
> 3) Issue create index / put mapping commands
> to recreate all our index definitions and
> mappings (same # of shards, same mappings,
> etc. as before)
> 4) Shutdown the cluster and copy the data
> index directories only (not the _state
> directories) over from the backup
> 5) Startup the cluster -- all indices came up
> green and had all our data!
> * Note that es seems to delete any index
> directories that don't match up with existing
> indices, so make sure you hang on the the
> backup until you are sure
>                     
>                     
> On the next environment we tried this on, we
> first flushed the transaction logs before
> shutting down the cluster and upgrading.
> Everything went smoothly.
> I don't know if flushing had anything to do
> with it, or if the first problem was kind of a
> freak occurrence, but I thought I would
> mention it.
>                     
>                     
> Curtis
> 
> ```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:53am UTC](https://discuss.elastic.co/t/rebuilding-corrupted-index-metadata/5489/7 "2017-07-06T03:53:01Z")

</div>


