# Elasticsearch continually fails after upgrade

**URL:** <https://discuss.elastic.co/t/elasticsearch-continually-fails-after-upgrade/185528>\
**Category:** Elasticsearch\
**Created:** [June 12, 2019, 11:20pm UTC](https://discuss.elastic.co/t/elasticsearch-continually-fails-after-upgrade/185528 "2019-06-12T23:20:12Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Emersumbigens](https://avatars.discourse-cdn.com/v4/letter/e/4da419/32.png) [@Emersumbigens](https://discuss.elastic.co/u/Emersumbigens)\
**Post date:** [June 12, 2019, 11:20pm UTC](https://discuss.elastic.co/t/elasticsearch-continually-fails-after-upgrade/185528/1 "2019-06-12T23:20:12Z")

</div>

I recently upgraded from 6.3.2 to 6.7.1 and everything initially worked appropriately. Currently though elasticsearch fails shortly after restarting it and now the Kibana page isn't ready. Not exactly sure where is the best place to start so I ran the cluster health command:  
[root@ip-10-0-1-207 ec2-user]# curl -XGET '[http://localhost:9200/\_cluster/health](http://localhost:9200/_cluster/health)'{"cluster\_name":"elasticsearch","status":"red","timed\_out":false,"number\_of\_nodes":1,"number\_of\_data\_nodes":1,"active\_primary\_shards":1053,"active\_shards":1053,"relocating\_shards":0,"initializing\_shards":4,"unassigned\_shards":1367,"delayed\_unassigned\_shards":0,"number\_of\_pending\_tasks":9,"number\_of\_in\_flight\_fetch":0,"task\_max\_waiting\_in\_queue\_millis":47898,"active\_shards\_percent\_as\_number":43.44059405940594}  
[root@ip-10-0-1-207 ec2-user]#

Any ideas?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [June 12, 2019, 11:25pm UTC](https://discuss.elastic.co/t/elasticsearch-continually-fails-after-upgrade/185528/2 "2019-06-12T23:25:29Z")

</div>

What do the logs show?

> [@Emersumbigens](#):
>
> "number\_of\_nodes":1,"number\_of\_data\_nodes":1,"active\_primary\_shards":1053,"active\_shards":1053,"relocating\_shards":0,"initializing\_shards":4,"unassigned\_shards":1367

That's a lot of shards for a single node. How many indices? How much data?

---

<div class="post-metadata">

**Author:** ![Emersumbigens](https://avatars.discourse-cdn.com/v4/letter/e/4da419/32.png) [@Emersumbigens](https://discuss.elastic.co/u/Emersumbigens)\
**Post date:** [June 13, 2019, 12:14am UTC](https://discuss.elastic.co/t/elasticsearch-continually-fails-after-upgrade/185528/3 "2019-06-13T00:14:15Z")

</div>

This system is a single-host ELK setup. The clients are 75% Redhat Linux 7.6 instances and 25% Windows Server 2016 hosted in AWS. We've used several types of information delivered from various types of beats: filebeat, auditbeat, metricbeat, Winlogbeat, etc. Currently there are a lot of indices listed and I'm wondering if I should be purging those indices and starting over.

---

<div class="post-metadata">

**Author:** ![Emersumbigens](https://avatars.discourse-cdn.com/v4/letter/e/4da419/32.png) [@Emersumbigens](https://discuss.elastic.co/u/Emersumbigens)\
**Post date:** [June 13, 2019, 12:14am UTC](https://discuss.elastic.co/t/elasticsearch-continually-fails-after-upgrade/185528/4 "2019-06-13T00:14:58Z")

</div>

Attached are the logs that I've found:  
Jun 12 19:30:20 ip-10-0-1-207 kibana: {"type":"log","@timestamp":"2019-06-12T23:30:20Z","tags":["error","task\_manager"],"pid":12326,"message":"Failed to poll for work: [search\_phase\_execution\_exception] all shards failed :: {"path":"/.kibana\_task\_manager/\_doc/\_search","query":{"ignore\_unavailable":true},"body":"{\"query\":{\"bool\":{\"must\":[{\"term\":{\"type\":\"task\"}},{\"bool\":{\"must\":[{\"terms\":{\"task.taskType\":[\"maps\_telemetry\",\"vis\_telemetry\"]}},{\"range\":{\"task.attempts\":{\"lte\":3}}},{\"range\":{\"task.runAt\":{\"lte\":\"now\"}}},{\"range\":{\"kibana.apiVersion\":{\"lte\":1}}}]}}]}},\"size\":10,\"sort\":{\"task.runAt\":{\"order\":\"asc\"}},\"seq\_no\_primary\_term\":true}","statusCode":503,"response":"{\"error\":{\"root\_cause\":,\"type\":\"search\_phase\_execution\_exception\",\"reason\":\"all shards failed\",\"phase\":\"query\",\"grouped\":true,\"failed\_shards\":},\"status\":503}"}"}  
Jun 12 19:30:23 ip-10-0-1-207 kibana: {"type":"log","@timestamp":"2019-06-12T23:30:23Z","tags":["error","task\_manager"],"pid":12326,"message":"Failed to poll for work: [search\_phase\_execution\_exception] all shards failed :: {"path":"/.kibana\_task\_manager/\_doc/\_search","query":{"ignore\_unavailable":true},"body":"{\"query\":{\"bool\":{\"must\":[{\"term\":{\"type\":\"task\"}},{\"bool\":{\"must\":[{\"terms\":{\"task.taskType\":[\"maps\_telemetry\",\"vis\_telemetry\"]}},{\"range\":{\"task.attempts\":{\"lte\":3}}},{\"range\":{\"task.runAt\":{\"lte\":\"now\"}}},{\"range\":{\"kibana.apiVersion\":{\"lte\":1}}}]}}]}},\"size\":10,\"sort\":{\"task.runAt\":{\"order\":\"asc\"}},\"seq\_no\_primary\_term\":true}","statusCode":503,"response":"{\"error\":{\"root\_cause\":,\"type\":\"search\_phase\_execution\_exception\",\"reason\":\"all shards failed\",\"phase\":\"query\",\"grouped\":true,\"failed\_shards\":},\"status\":503}"}"}  
Jun 12 19:30:26 ip-10-0-1-207 kibana: {"type":"log","@timestamp":"2019-06-12T23:30:26Z","tags":["error","task\_manager"],"pid":12326,"message":"Failed to poll for work: [search\_phase\_execution\_exception] all shards failed :: {"path":"/.kibana\_task\_manager/\_doc/\_search","query":{"ignore\_unavailable":true},"body":"{\"query\":{\"bool\":{\"must\":[{\"term\":{\"type\":\"task\"}},{\"bool\":{\"must\":[{\"terms\":{\"task.taskType\":[\"maps\_telemetry\",\"vis\_telemetry\"]}},{\"range\":{\"task.attempts\":{\"lte\":3}}},{\"range\":{\"task.runAt\":{\"lte\":\"now\"}}},{\"range\":{\"kibana.apiVersion\":{\"lte\":1}}}]}}]}},\"size\":10,\"sort\":{\"task.runAt\":{\"order\":\"asc\"}},\"seq\_no\_primary\_term\":true}","statusCode":503,"response":"{\"error\":{\"root\_cause\":,\"type\":\"search\_phase\_execution\_exception\",\"reason\":\"all shards failed\",\"phase\":\"query\",\"grouped\":true,\"failed\_shards\":},\"status\":503}"}"}

```
[2019-06-12T19:29:05,544][WARN][logstash.outputs.elasticsearch] Attempted to resurrect connection to dead ES instance, but got an error. {:url=>"http://localhost:9200/", :error_type=>LogStash::Outputs::ElasticSearch::HttpClient::Pool::HostUnreachableError, :error=>"Elasticsearch Unreachable: [http://localhost:9200/][Manticore::SocketException] Connection refused (Connection refused)"}
[2019-06-12T19:29:10,556][WARN][logstash.outputs.elasticsearch] Attempted to resurrect connection to dead ES instance, but got an error. {:url=>"http://localhost:9200/", :error_type=>LogStash::Outputs::ElasticSearch::HttpClient::Pool::HostUnreachableError, :error=>"Elasticsearch Unreachable: [http://localhost:9200/][Manticore::SocketException] Connection refused (Connection refused)"}
[2019-06-12T19:29:15,621][WARN][logstash.outputs.elasticsearch] Attempted to resurrect connection to dead ES instance, but got an error. {:url=>"http://localhost:9200/", :error_type=>LogStash::Outputs::ElasticSearch::HttpClient::Pool::BadResponseCodeError, :error=>"Got response code '503' contacting Elasticsearch at URL 'http://localhost:9200/'"}
[2019-06-12T19:29:20,670][WARN][logstash.outputs.elasticsearch] Restored connection to ES instance {:url=>"http://localhost:9200/"}

Jun 12 18:25:29 ip-10-0-1-207.ec2.internal systemd[1]: Started Elasticsearch.
Jun 12 19:15:28 ip-10-0-1-207.ec2.internal systemd[1]: elasticsearch.service: main process exited, code=exited, status=127/n/a
Jun 12 19:15:28 ip-10-0-1-207.ec2.internal systemd[1]: Unit elasticsearch.service entered failed state.
Jun 12 19:15:28 ip-10-0-1-207.ec2.internal systemd[1]: elasticsearch.service failed.
Jun 12 19:28:48 ip-10-0-1-207.ec2.internal systemd[1]: Started Elasticsearch.
```

---

<div class="post-metadata">

**Author:** ![Emersumbigens](https://avatars.discourse-cdn.com/v4/letter/e/4da419/32.png) [@Emersumbigens](https://discuss.elastic.co/u/Emersumbigens)\
**Post date:** [June 13, 2019, 1:00am UTC](https://discuss.elastic.co/t/elasticsearch-continually-fails-after-upgrade/185528/5 "2019-06-13T01:00:00Z")

</div>

I ran the following command and it outputs 307 indices (101 that contain data ranging from 50 kb to 900 mb):  
curl -X GET '[http://localhost:9200/\_cat/indices?v](http://localhost:9200/_cat/indices?v)'

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [June 13, 2019, 1:09am UTC](https://discuss.elastic.co/t/elasticsearch-continually-fails-after-upgrade/185528/6 "2019-06-13T01:09:29Z")

</div>

Please format your code/logs/config using the `</>` button, or markdown style back ticks. It helps to make things easy to read which helps us help you 🙂

Ok, you definitely need to reduce your shard count. Look at using the `_shrink` API to do that, or use the `_reindex` API to "merge" your daily indices into monthly ones.

---

<div class="post-metadata">

**Author:** ![Emersumbigens](https://avatars.discourse-cdn.com/v4/letter/e/4da419/32.png) [@Emersumbigens](https://discuss.elastic.co/u/Emersumbigens)\
**Post date:** [June 13, 2019, 3:01am UTC](https://discuss.elastic.co/t/elasticsearch-continually-fails-after-upgrade/185528/7 "2019-06-13T03:01:54Z")

</div>

I'm pretty new to ELK so will have to look into the use of the \_shrink API.  
Wanting to do something basic, which I thought could be accomplished by a single index or maybe a single index per filebeat type. Not yet sure where all the indexes came from or what led to such a high shard count. I didn't configure a shard count anywhere. If I do a complete reinstall, is there a way for me to set the shard count?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [June 13, 2019, 3:52am UTC](https://discuss.elastic.co/t/elasticsearch-continually-fails-after-upgrade/185528/8 "2019-06-13T03:52:14Z")

</div>

Have a look at [Kibana only display logo no content -elasticsearch error shard - #11 by warkolm](https://discuss.elastic.co/t/kibana-only-display-logo-no-content-elasticsearch-error-shard/185529/11) to help with reindexing.

> [@Emersumbigens](#):
>
> Wanting to do something basic, which I thought could be accomplished by a single index or maybe a single index per filebeat type. Not yet sure where all the indexes came from or what led to such a high shard count. I didn't configure a shard count anywhere.

Indices are where the data is stored, they are made of shards. In the 6.X release we defaulted to 5 shards per index with 1 replica set, so a total of 10 per index. We changed that in 7.X to be 1 per index with 1 replica (ie 2).

> [@Emersumbigens](#):
>
> If I do a complete reinstall, is there a way for me to set the shard count?

No need to do that! 🙂 Though I would suggest using the latest 7.X release, as it reduces the burden here.

Otherwise you can set the shard count in the filebeat config file.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 11, 2019, 4:01am UTC](https://discuss.elastic.co/t/elasticsearch-continually-fails-after-upgrade/185528/9 "2019-07-11T04:01:59Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
