# Elasticsearch failed to start shard and deletes all index files

**URL:** <https://discuss.elastic.co/t/elasticsearch-failed-to-start-shard-and-deletes-all-index-files/9751>\
**Category:** Elasticsearch\
**Created:** [November 17, 2012, 9:03pm UTC](https://discuss.elastic.co/t/elasticsearch-failed-to-start-shard-and-deletes-all-index-files/9751 "2012-11-17T21:03:00Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Pinar\_Yanardag](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/pinar_yanardag/32/2611_2.png) [@Pinar\_Yanardag](https://discuss.elastic.co/u/Pinar_Yanardag)\
**Post date:** [November 17, 2012, 9:03pm UTC](https://discuss.elastic.co/t/elasticsearch-failed-to-start-shard-and-deletes-all-index-files/9751/1 "2012-11-17T21:03:00Z")

</div>

Hi,

I am trying to index Wikipedia dump with a Python script I wrote. I simply  
read the Wikipedia dump and for each article, I send a curl request to  
index id, title and text properties of the page. Here is the mapping I use:

{  
"article" : {  
"properties" : {  
"text" : {  
"type" : "string"  
},  
"title" : {  
"type" : "string"  
},  
"wid" : {  
"type" : "date",  
"format" : "dateOptionalTime"  
}  
}  
}

After my script finishes its work, I am left with a 55GB data index  
(extracted wikipedia dump is 39GB). However, after I start Elasticsearch  
again, it deletes all the indexed files and give the following error:  
"failed to start  
shard org.elasticsearch.indices.recovery.RecoveryFailedException:" (more  
log: [https://gist.github.com/4099519](https://gist.github.com/4099519)). Do you have any suggestions on it?  
What can be the problem?

Note: I also tried to use Wikipedia plugin, however I was never able to  
index more than 5gb (indexing usually stops when data folder reach to 5gb  
index and even when I start the server again it reaches at most ~7gb of  
data).

Thanks,  
Pinar

--

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [November 18, 2012, 1:08pm UTC](https://discuss.elastic.co/t/elasticsearch-failed-to-start-shard-and-deletes-all-index-files/9751/2 "2012-11-18T13:08:54Z")

</div>

Are you running the same version of elasticsearch on all 3 nodes?

Wikipedia river runs much better when you use local dump files instead of  
downloading it and indexing it at the same time.

On Saturday, November 17, 2012 4:03:00 PM UTC-5, Pınar Yanardağ wrote:

> Hi,
> 
> I am trying to index Wikipedia dump with a Python script I wrote. I simply  
> read the Wikipedia dump and for each article, I send a curl request to  
> index id, title and text properties of the page. Here is the mapping I use:
> 
> {  
> "article" : {  
> "properties" : {  
> "text" : {  
> "type" : "string"  
> },  
> "title" : {  
> "type" : "string"  
> },  
> "wid" : {  
> "type" : "date",  
> "format" : "dateOptionalTime"  
> }  
> }  
> }
> 
> After my script finishes its work, I am left with a 55GB data index  
> (extracted wikipedia dump is 39GB). However, after I start Elasticsearch  
> again, it deletes all the indexed files and give the following error:  
> "failed to start  
> shard org.elasticsearch.indices.recovery.RecoveryFailedException:" (more  
> log: [elastic · GitHub](https://gist.github.com/4099519)). Do you have any suggestions on it?  
> What can be the problem?
> 
> Note: I also tried to use Wikipedia plugin, however I was never able to  
> index more than 5gb (indexing usually stops when data folder reach to 5gb  
> index and even when I start the server again it reaches at most ~7gb of  
> data).
> 
> Thanks,  
> Pinar

--

---

<div class="post-metadata">

**Author:** ![Pinar\_Yanardag](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/pinar_yanardag/32/2611_2.png) [@Pinar\_Yanardag](https://discuss.elastic.co/u/Pinar_Yanardag)\
**Post date:** [November 18, 2012, 5:55pm UTC](https://discuss.elastic.co/t/elasticsearch-failed-to-start-shard-and-deletes-all-index-files/9751/3 "2012-11-18T17:55:35Z")

</div>

On Sunday, November 18, 2012 8:08:55 AM UTC-5, Igor Motov wrote:

> Are you running the same version of elasticsearch on all 3 nodes?

Yes, I am using the same version (0.19.11).

> Wikipedia river runs much better when you use local dump files instead of  
> downloading it and indexing it at the same time.

Actually I always tried to index with a local dump like the following:

curl -XPUT localhost:9200/\_river/my\_river/\_meta -d '{ "type" : "wikipedia",  
"index" : { "name" : "my\_index", "type" : "my\_type", "url" :  
"/local/enwiki-latest-pages-articles.xml.bz2", "bulk\_size" : 10 } }'

However I am not sure if Elasticsearch is using the local dump I provided,  
because after I sent the curl above, I see the download link in the logs:  
"[wikipedia][my\_river] creating wikipedia stream river for  
[[http://download.wikimedia.org/enwiki/latest/enwiki-latest-pages-articles.xml.bz2](http://download.wikimedia.org/enwiki/latest/enwiki-latest-pages-articles.xml.bz2)]".

Thanks,  
Pinar

> On Saturday, November 17, 2012 4:03:00 PM UTC-5, Pınar Yanardağ wrote:
> 
> > Hi,
> > 
> > I am trying to index Wikipedia dump with a Python script I wrote. I  
> > simply read the Wikipedia dump and for each article, I send a curl request  
> > to index id, title and text properties of the page. Here is the mapping I  
> > use:
> > 
> > {  
> > "article" : {  
> > "properties" : {  
> > "text" : {  
> > "type" : "string"  
> > },  
> > "title" : {  
> > "type" : "string"  
> > },  
> > "wid" : {  
> > "type" : "date",  
> > "format" : "dateOptionalTime"  
> > }  
> > }  
> > }
> > 
> > After my script finishes its work, I am left with a 55GB data index  
> > (extracted wikipedia dump is 39GB). However, after I start Elasticsearch  
> > again, it deletes all the indexed files and give the following error:  
> > "failed to start  
> > shard org.elasticsearch.indices.recovery.RecoveryFailedException:" (more  
> > log: [elastic · GitHub](https://gist.github.com/4099519)). Do you have any suggestions on  
> > it? What can be the problem?
> > 
> > Note: I also tried to use Wikipedia plugin, however I was never able to  
> > index more than 5gb (indexing usually stops when data folder reach to 5gb  
> > index and even when I start the server again it reaches at most ~7gb of  
> > data).
> > 
> > Thanks,  
> > Pinar

--

---

<div class="post-metadata">

**Author:** ![Igor\_Motov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_motov/32/45193_2.png) [@Igor\_Motov](https://discuss.elastic.co/u/Igor_Motov)\
**Post date:** [November 18, 2012, 7:47pm UTC](https://discuss.elastic.co/t/elasticsearch-failed-to-start-shard-and-deletes-all-index-files/9751/4 "2012-11-18T19:47:23Z")

</div>

It actually should be:

{  
"type" : "wikipedia",  
"wikipedia" : {  
"url" : "file:///local/enwiki-latest-pages-articles.xml.bz2"  
},  
"index": {  
"index" : "my\_index",  
"type" : "my\_type"  
},  
"bulk\_size" : 10  
}

On Sunday, November 18, 2012 12:55:35 PM UTC-5, Pınar Yanardağ wrote:

> On Sunday, November 18, 2012 8:08:55 AM UTC-5, Igor Motov wrote:
> 
> > Are you running the same version of elasticsearch on all 3 nodes?
> 
> Yes, I am using the same version (0.19.11).
> 
> > Wikipedia river runs much better when you use local dump files instead of  
> > downloading it and indexing it at the same time.
> 
> Actually I always tried to index with a local dump like the following:
> 
> curl -XPUT localhost:9200/\_river/my\_river/\_meta -d '{ "type" :  
> "wikipedia", "index" : { "name" : "my\_index", "type" : "my\_type",  
> "url" : "/local/enwiki-latest-pages-articles.xml.bz2", "bulk\_size" : 10 }  
> }'
> 
> However I am not sure if Elasticsearch is using the local dump I provided,  
> because after I sent the curl above, I see the download link in the logs:  
> "[wikipedia][my\_river] creating wikipedia stream river for [  
> [http://download.wikimedia.org/enwiki/latest/enwiki-latest-pages-articles.xml.bz2](http://download.wikimedia.org/enwiki/latest/enwiki-latest-pages-articles.xml.bz2)  
> ]".
> 
> Thanks,  
> Pinar
> 
> > On Saturday, November 17, 2012 4:03:00 PM UTC-5, Pınar Yanardağ wrote:
> > 
> > > Hi,
> > > 
> > > I am trying to index Wikipedia dump with a Python script I wrote. I  
> > > simply read the Wikipedia dump and for each article, I send a curl request  
> > > to index id, title and text properties of the page. Here is the mapping I  
> > > use:
> > > 
> > > {  
> > > "article" : {  
> > > "properties" : {  
> > > "text" : {  
> > > "type" : "string"  
> > > },  
> > > "title" : {  
> > > "type" : "string"  
> > > },  
> > > "wid" : {  
> > > "type" : "date",  
> > > "format" : "dateOptionalTime"  
> > > }  
> > > }  
> > > }
> > > 
> > > After my script finishes its work, I am left with a 55GB data index  
> > > (extracted wikipedia dump is 39GB). However, after I start Elasticsearch  
> > > again, it deletes all the indexed files and give the following error:  
> > > "failed to start  
> > > shard org.elasticsearch.indices.recovery.RecoveryFailedException:" (more  
> > > log: [elastic · GitHub](https://gist.github.com/4099519)). Do you have any suggestions on  
> > > it? What can be the problem?
> > > 
> > > Note: I also tried to use Wikipedia plugin, however I was never able to  
> > > index more than 5gb (indexing usually stops when data folder reach to 5gb  
> > > index and even when I start the server again it reaches at most ~7gb of  
> > > data).
> > > 
> > > Thanks,  
> > > Pinar

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:03am UTC](https://discuss.elastic.co/t/elasticsearch-failed-to-start-shard-and-deletes-all-index-files/9751/5 "2017-07-06T03:03:48Z")

</div>


