# Disk storage not increasing despite indexes being growing

**URL:** <https://discuss.elastic.co/t/disk-storage-not-increasing-despite-indexes-being-growing/211523>\
**Category:** Elasticsearch\
**Created:** [December 11, 2019, 9:20pm UTC](https://discuss.elastic.co/t/disk-storage-not-increasing-despite-indexes-being-growing/211523 "2019-12-11T21:20:16Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![viniciof](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/viniciof/32/53454_2.png) [@viniciof](https://discuss.elastic.co/u/viniciof)\
**Post date:** [December 11, 2019, 9:20pm UTC](https://discuss.elastic.co/t/disk-storage-not-increasing-despite-indexes-being-growing/211523/1 "2019-12-11T21:20:16Z")

</div>

Hi all,

Noticed today something very weird, our Logstash workers keep emitting data consistently to two of our ES clusters, rate as expected.

 ![logstash%20events](https://us1.discourse-cdn.com/elastic/original/3X/4/2/428917587951a7680e7790abee775eba94d0710f.png)

Disk storage in one of them gets consumed as expected also, but the other remains flat (and there are rejected events in thread pool in most of the data nodes)

**Problematic cluster**

 ![problem%20cluster](https://us1.discourse-cdn.com/elastic/original/3X/a/5/a5c5ec5e62afd618e465945d6479d832747a8bb8.png)

**Good cluster**

 ![good%20clsuter](https://us1.discourse-cdn.com/elastic/original/3X/b/7/b7100d080a767e24937110bf824b32da0ddc8186.png)

how can I find out what's wrong ?

rgds,

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 12, 2019, 4:45am UTC](https://discuss.elastic.co/t/disk-storage-not-increasing-despite-indexes-being-growing/211523/2 "2019-12-12T04:45:02Z")

</div>

What does your Logstash outputs look like?

---

<div class="post-metadata">

**Author:** ![viniciof](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/viniciof/32/53454_2.png) [@viniciof](https://discuss.elastic.co/u/viniciof)\
**Post date:** [December 12, 2019, 3:13pm UTC](https://discuss.elastic.co/t/disk-storage-not-increasing-despite-indexes-being-growing/211523/3 "2019-12-12T15:13:19Z")

</div>

They send bulk requests, like the one as follows:

> <https://gist.github.com/vinicioflores/a922e3f9de75454203d77c9e2b8b3dbd>

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 12, 2019, 4:59pm UTC](https://discuss.elastic.co/t/disk-storage-not-increasing-despite-indexes-being-growing/211523/4 "2019-12-12T16:59:05Z")

</div>

What does the configuration look like? Do you have a separate output plugin per cluster?

---

<div class="post-metadata">

**Author:** ![viniciof](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/viniciof/32/53454_2.png) [@viniciof](https://discuss.elastic.co/u/viniciof)\
**Post date:** [December 12, 2019, 9:23pm UTC](https://discuss.elastic.co/t/disk-storage-not-increasing-despite-indexes-being-growing/211523/5 "2019-12-12T21:23:39Z")

</div>

I have separate output plugin for each, here the Logstash pipeline:

**Problematic one**

```
output {
  elasticsearch {
   hosts => [
        "arm-or-006.myserver.com:9996",
        "arm-or-007.myserver.com:9996",
        "arm-or-008.myserver.com:9996",
        "arm-or-010.myserver.com:9996"
    ]
    ssl => true
    cacert => "/app/ssl/cert.pem"
    user => "myuser"
    password => "mypass"
    document_type => "arm"
    document_id => "%{ibi_id}"
    index => "%{ibi_target}-%{+YYYY-MM}"
    doc_as_upsert => true
    action => "update"
    retry_max_interval => 5
    retry_on_conflict => 5
    flush_size => 10000
    timeout => 1000000
  }
}

```

**Good one**

```
output {
  elasticsearch {
    hosts => [
        "arm-lc-001.myserver.com:9996",
        "arm-lc-003.myserver.com:9996",
        "arm-lc-004.myserver.com:9996",
        "arm-lc-005.myserver.com:9996"
    ]
    ssl => true
    cacert => "/app/ssl/cert.pem"
    ssl_certificate_verification => false
    user => "myuser"
    password => "mypass"
    document_type => "arm"
    document_id => "%{ibi_id}"
    index => "%{ibi_target}-%{+YYYY-MM}"
    doc_as_upsert => true
    action => "update"
    retry_max_interval => 5
    retry_on_conflict => 5
    flush_size => 10000
    timeout => 1000000
  }
}
```

---

<div class="post-metadata">

**Author:** ![viniciof](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/viniciof/32/53454_2.png) [@viniciof](https://discuss.elastic.co/u/viniciof)\
**Post date:** [December 12, 2019, 11:56pm UTC](https://discuss.elastic.co/t/disk-storage-not-increasing-despite-indexes-being-growing/211523/6 "2019-12-12T23:56:42Z")

</div>

Also, this is what I see in hot threads:

> <https://gist.github.com/vinicioflores/5a150b1a25e094ef16a15ed9690b4b51>

@Christian_Dahlqvist , what does this mean ?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 13, 2019, 6:50am UTC](https://discuss.elastic.co/t/disk-storage-not-increasing-despite-indexes-being-growing/211523/7 "2019-12-13T06:50:58Z")

</div>

I have never run bulk updates so am not sure if errors here would cause the update to be retried from Logstash or simply dropped. You seem to have a lot of time spent on management. Do you have a very large number of shards in the cluster? Are you using dynamic mappings? Does the hardware profiles supporting the cluster s differ, especially with respect to the type of storage used? Is there anything in the Elasticsearch logs?

---

<div class="post-metadata">

**Author:** ![viniciof](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/viniciof/32/53454_2.png) [@viniciof](https://discuss.elastic.co/u/viniciof)\
**Post date:** [December 13, 2019, 5:10pm UTC](https://discuss.elastic.co/t/disk-storage-not-increasing-despite-indexes-being-growing/211523/8 "2019-12-13T17:10:24Z")

</div>

Here is all sharding info in my cluster:

> <https://gist.github.com/vinicioflores/5253cf93a6e53da6fa774d2581539d7d>

> <https://gist.github.com/vinicioflores/495d94a3e20de3d1f8af52283de27b10>

Here's the mapping info of the most problematic index we have in the cluster (I'm using dynamic mappings template for it)

> <https://gist.github.com/vinicioflores/2bc85779dcc1daad33cb9ff7336c5cb0>

We use SAN/LUN based storage (magnitude of TBs of space) and all servers have same specs. How can I find out if it's due to bad disk I/O?

Didn't find anything in ES logs though

---

<div class="post-metadata">

**Author:** ![viniciof](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/viniciof/32/53454_2.png) [@viniciof](https://discuss.elastic.co/u/viniciof)\
**Post date:** [December 13, 2019, 7:49pm UTC](https://discuss.elastic.co/t/disk-storage-not-increasing-despite-indexes-being-growing/211523/9 "2019-12-13T19:49:10Z")

</div>

I see also my thread pools very packed:

> <https://gist.github.com/vinicioflores/846fdf0bebda2ce889430cff86a6a440>

@Christian_Dahlqvist, How can I find out the details of those specific threads taking all available slots in each queue ?

---

<div class="post-metadata">

**Author:** ![viniciof](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/viniciof/32/53454_2.png) [@viniciof](https://discuss.elastic.co/u/viniciof)\
**Post date:** [December 13, 2019, 7:55pm UTC](https://discuss.elastic.co/t/disk-storage-not-increasing-despite-indexes-being-growing/211523/10 "2019-12-13T19:55:30Z")

</div>

And here my global cluster's settings:

```
{
  "persistent": {
    "cluster": {
      "routing": {
        "allocation": {
          "awareness": {
            "attributes": ""
          }
        }
      }
    },
    "indices": {
      "breaker": {
        "fielddata": {
          "limit": "60%"
        },
        "request": {
          "limit": "30%"
        }
      }
    }
  },
  "transient": {
    "indices": {
      "recovery": {
        "max_bytes_per_sec": "256mb"
      }
    }
  }
}
```

---

<div class="post-metadata">

**Author:** ![viniciof](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/viniciof/32/53454_2.png) [@viniciof](https://discuss.elastic.co/u/viniciof)\
**Post date:** [December 17, 2019, 12:20am UTC](https://discuss.elastic.co/t/disk-storage-not-increasing-despite-indexes-being-growing/211523/11 "2019-12-17T00:20:08Z")

</div>

@Christian_Dahlqvist can you help ? Let me know if any more information is needed

---

<div class="post-metadata">

**Author:** ![viniciof](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/viniciof/32/53454_2.png) [@viniciof](https://discuss.elastic.co/u/viniciof)\
**Post date:** [December 17, 2019, 3:23pm UTC](https://discuss.elastic.co/t/disk-storage-not-increasing-despite-indexes-being-growing/211523/12 "2019-12-17T15:23:20Z")

</div>

Hi @Christian_Dahlqvist,

Here's my node stats. I notice one of my nodes (arm-or-009\_data) is 99% memory utilization .

> <https://gist.github.com/vinicioflores/6b0d98b50521f20b0381c5169463db04>

And these are the stats for the most problematic (slow indexing) index in the cluster, "daas-arm-prod-users-2019-12-new"

> <https://gist.github.com/vinicioflores/d58bbfb9cccaf2bd3b075f0601b3ff5d>

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 18, 2019, 7:08am UTC](https://discuss.elastic.co/t/disk-storage-not-increasing-despite-indexes-being-growing/211523/13 "2019-12-18T07:08:49Z")

</div>

Are there any error messages in the Elasticsearch logs? Can you try enabling the [dead-letter queue](https://www.elastic.co/guide/en/logstash/7.5/dead-letter-queues.html) to see if this captures any errors that would otherwise be ignored/dropped?

---

<div class="post-metadata">

**Author:** ![viniciof](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/viniciof/32/53454_2.png) [@viniciof](https://discuss.elastic.co/u/viniciof)\
**Post date:** [December 18, 2019, 10:42pm UTC](https://discuss.elastic.co/t/disk-storage-not-increasing-despite-indexes-being-growing/211523/14 "2019-12-18T22:42:10Z")

</div>

I checked the logs and no errors seem to be there. There's only one that the cluster complains a lot:

```
[2019-12-18T14:39:44,682][ERROR][o.e.x.m.c.c.ClusterStatsCollector] [arm-or-002_master] collector [cluster_stats] timed out when collecting data

```

and

```
[2019-12-18T14:29:24,586][ERROR][o.e.x.m.c.i.IndexStatsCollector] [arm-or-002_master] collector [index-stats] timed out when collecting data

```

Whenever this is logged, it causes a "blank" patch in the overview section of monitoring of the cluster in Kibana (like cluster is unresponsive during that time of exception)

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/3/f/3fc30d7aed697037316c01e13f322b49f1d492d8.png)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 15, 2020, 10:42pm UTC](https://discuss.elastic.co/t/disk-storage-not-increasing-despite-indexes-being-growing/211523/15 "2020-01-15T22:42:11Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
