# Cluster stopped ingesting, with "failed to obtain in-memory shard lock"

**URL:** <https://discuss.elastic.co/t/cluster-stopped-ingesting-with-failed-to-obtain-in-memory-shard-lock/155016>\
**Category:** Elasticsearch\
**Created:** [November 1, 2018, 12:43pm UTC](https://discuss.elastic.co/t/cluster-stopped-ingesting-with-failed-to-obtain-in-memory-shard-lock/155016 "2018-11-01T12:43:36Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![TimWard](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/timward/32/19574_2.png) [@TimWard](https://discuss.elastic.co/u/TimWard)\
**Post date:** [November 1, 2018, 12:43pm UTC](https://discuss.elastic.co/t/cluster-stopped-ingesting-with-failed-to-obtain-in-memory-shard-lock/155016/1 "2018-11-01T12:43:36Z")

</div>

My cluster logged some "failed to obtain in-memory shard lock" messages over a period of the night, finishing at around 04:33 this morning, and there is no data in most of the indexes after 04:33, ie it has stopped indexing data. (There's just one, vary sparsely used, index with data in it past that time.)

All shards are showing as STARTED. retry\_failed didn't do anything. Restarting each node in the cluster didn't do anything.

What else do I need to look at? How do I get my cluster indexing again?

The reason for a probably in the middle of the night may have been that I was reindexing hundreds of gigabytes of data, and at some point in the process one or two of the nodes might have got short of disk space for relocating shards. There is currently no shortage of disk space on any node.

Data should be but isn't coming in from Logstash, from Metricbeat and from some Python scripts. The index that is still being written to comes from a Java application.

---

<div class="post-metadata">

**Author:** ![TimWard](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/timward/32/19574_2.png) [@TimWard](https://discuss.elastic.co/u/TimWard)\
**Post date:** [November 1, 2018, 12:45pm UTC](https://discuss.elastic.co/t/cluster-stopped-ingesting-with-failed-to-obtain-in-memory-shard-lock/155016/2 "2018-11-01T12:45:54Z")

</div>

Ah, and then I find the following in the Python logs. I'd better now try to find out what the concept of a "read only index" is.

```
{
	u'update': {
		u'status': 403,
		u'_type': u'doc',
		u'_index': u'event-2018.07',
		u'error': {
			u'reason': u'blockedby: [FORBIDDEN/12/indexread-only/allowdelete(api)];',
			u'type': u'cluster_block_exception'
		},
		u'_id': u'et3-tim-2.imagiro.ltd_threshold_Datapushgroupcount_2018-07-02T13: 18: 02.265Z',
		u'data': {
			'doc': {
				'alerted-level': 'critical',
				'alerted': True
			}
		}
	}
}
```

---

<div class="post-metadata">

**Author:** ![DavidTurner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/davidturner/32/22453_2.png) [@DavidTurner](https://discuss.elastic.co/u/DavidTurner)\
**Post date:** [November 1, 2018, 12:51pm UTC](https://discuss.elastic.co/t/cluster-stopped-ingesting-with-failed-to-obtain-in-memory-shard-lock/155016/3 "2018-11-01T12:51:47Z")

</div>

> [@TimWard](#):
>
> The reason for a probably in the middle of the night may have been that I was reindexing hundreds of gigabytes of data, and at some point in the process one or two of the nodes might have got short of disk space for relocating shards. There is currently no shortage of disk space on any node.

This would explain the message `blockedby: [FORBIDDEN/12/indexread-only/allowdelete(api)]`: if a node exceeds the flood-stage disk watermark (95% of disk capacity by default) then all indices with shards on that node are marked as read-only. The [documentation on disk-based shard allocation](https://www.elastic.co/guide/en/elasticsearch/reference/6.x/disk-allocator.html#disk-allocator) describes this in more detail and also describes how to recover.

This may or may not be related to the "failed to obtain in-memory shard lock" message, but it sounds like this is your actual problem here.

---

<div class="post-metadata">

**Author:** ![TimWard](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/timward/32/19574_2.png) [@TimWard](https://discuss.elastic.co/u/TimWard)\
**Post date:** [November 1, 2018, 12:59pm UTC](https://discuss.elastic.co/t/cluster-stopped-ingesting-with-failed-to-obtain-in-memory-shard-lock/155016/4 "2018-11-01T12:59:50Z")

</div>

And having got that clue from the Python application logs the fix (after several hours' research) was

```
PUT /_all/_settings
{
  "index": {
    "blocks.read_only_allow_delete": null
  }
}
```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 29, 2018, 12:59pm UTC](https://discuss.elastic.co/t/cluster-stopped-ingesting-with-failed-to-obtain-in-memory-shard-lock/155016/5 "2018-11-29T12:59:55Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
