# Disk usage after indexing much higher then after restart service

**URL:** <https://discuss.elastic.co/t/disk-usage-after-indexing-much-higher-then-after-restart-service/75053>\
**Category:** Elasticsearch\
**Created:** [February 14, 2017, 2:39pm UTC](https://discuss.elastic.co/t/disk-usage-after-indexing-much-higher-then-after-restart-service/75053 "2017-02-14T14:39:58Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![Benjamin\_Gathmann](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/benjamin_gathmann/32/7561_2.png) [@Benjamin\_Gathmann](https://discuss.elastic.co/u/Benjamin_Gathmann)\
**Post date:** [February 14, 2017, 2:39pm UTC](https://discuss.elastic.co/t/disk-usage-after-indexing-much-higher-then-after-restart-service/75053/1 "2017-02-14T14:39:58Z")

</div>

I noticed that if I index a large amount of documents to my cluster (which currently is restricted to one server), I have e.g. a disk usage of 90 GByte. (`du -sh /var/lib/elasticsearch/mycluster`)  
If I then restart elasticsearch service `(sudo service elasticsearch restart`), the size of that directory is sharply reduced to 68 GByte.  
Why is that so?  
What can I do to free the space without restarting the service?

The phenomenon resembles this issue:

> [@Elasticsearch creating files but not releasing them until I restart the service](https://discuss.elastic.co/t/elasticsearch-creating-files-but-not-releasing-them-until-i-restart-the-service/37385/5):
>
> I know, but updating to 1.X or newer is a much larger project and I would like to get this fixed in the interim. Thanks

but I am running ES 2.4

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [February 15, 2017, 9:25pm UTC](https://discuss.elastic.co/t/disk-usage-after-indexing-much-higher-then-after-restart-service/75053/2 "2017-02-15T21:25:16Z")

</div>

Depends, could be merges, translog or other things.  
What do the APIs (eg \_cat) show about disk use?

---

<div class="post-metadata">

**Author:** ![Benjamin\_Gathmann](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/benjamin_gathmann/32/7561_2.png) [@Benjamin\_Gathmann](https://discuss.elastic.co/u/Benjamin_Gathmann)\
**Post date:** [February 16, 2017, 9:44am UTC](https://discuss.elastic.co/t/disk-usage-after-indexing-much-higher-then-after-restart-service/75053/3 "2017-02-16T09:44:40Z")

</div>

Hi Mark,

Which commands should I run in particular? E.g.:

```
curl -GET 'http://localhost:9200/_cat/indices'
curl -GET 'http://localhost:9200/_stats' 

```

indices output is pretty limited, \_stats output is very large, what should I look for?

---

<div class="post-metadata">

**Author:** ![Benjamin\_Gathmann](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/benjamin_gathmann/32/7561_2.png) [@Benjamin\_Gathmann](https://discuss.elastic.co/u/Benjamin_Gathmann)\
**Post date:** [February 16, 2017, 9:47am UTC](https://discuss.elastic.co/t/disk-usage-after-indexing-much-higher-then-after-restart-service/75053/4 "2017-02-16T09:47:36Z")

</div>

OK, I think now I know what you mean - I should use \_cat API instead of using "du -sh" because it will show the size of the index itself, not the directory size.  
I will now re-index and check it out.

---

<div class="post-metadata">

**Author:** ![Benjamin\_Gathmann](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/benjamin_gathmann/32/7561_2.png) [@Benjamin\_Gathmann](https://discuss.elastic.co/u/Benjamin_Gathmann)\
**Post date:** [February 16, 2017, 12:30pm UTC](https://discuss.elastic.co/t/disk-usage-after-indexing-much-higher-then-after-restart-service/75053/5 "2017-02-16T12:30:16Z")

</div>

OK, so I reindexed and then collected all kinds of data before and after restarting the service.

\_cat/indices output was simply

```
health status index pri rep docs.count docs.deleted store.size pri.store.size
yellow open myindex 5 1 96414689 0 97gb 97gb

```

before and:

```
health status index pri rep docs.count docs.deleted store.size pri.store.size
yellow open myindex 5 1 96414689 0 69.9gb 69.9gb

```

afterwards

As you can see, I have the default settings, so five primary nodes in total.  
For each of these, the directory

```
/var/lib/elasticsearch/mycluster/nodes/0/indices/myindex/0/index

```

shows a du -sh of 20 GB before, and 16 GB afterwards.

translog folders are very small in both cases (around 20 KB).

So I ran

```
du /var/lib/elasticsearch/mycluster/nodes/0/indices/myindex/0/index/* | sort -n 

```

on each of the nodes before and afterwards to get an idea of which files change in size.

After the restart, the files in this directory either stayed constant in size, or they completely disappeared.

The largest of the disappeared files for a sample directory were (first column is size in KByte):

```
284632	/var/lib/elasticsearch/mycluster/nodes/0/indices/myindex/0/index/_t0.cfs
370488	/var/lib/elasticsearch/mycluster/nodes/0/indices/myindex/0/index/_3v.nvd
477448	/var/lib/elasticsearch/mycluster/nodes/0/indices/myindex/0/index/_wy.cfs
506380	/var/lib/elasticsearch/mycluster/nodes/0/indices/myindex/0/index/_hl.cfs
608492	/var/lib/elasticsearch/mycluster/nodes/0/indices/myindex/0/index/_pc.cfs
820228	/var/lib/elasticsearch/mycluster/nodes/0/indices/myindex/0/index/_jt.nvd
838104	/var/lib/elasticsearch/mycluster/nodes/0/indices/myindex/0/index/_rl.cfs
1269064	/var/lib/elasticsearch/mycluster/nodes/0/indices/myindex/0/index/_mm.nvd

```

Can you roughly tell me what happens there?

I can also email you a complete list of the files if you like.

Thanks!

---

<div class="post-metadata">

**Author:** ![Benjamin\_Gathmann](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/benjamin_gathmann/32/7561_2.png) [@Benjamin\_Gathmann](https://discuss.elastic.co/u/Benjamin_Gathmann)\
**Post date:** [February 16, 2017, 12:33pm UTC](https://discuss.elastic.co/t/disk-usage-after-indexing-much-higher-then-after-restart-service/75053/6 "2017-02-16T12:33:04Z")

</div>

By the way, this is an index with a lot of nested documents, maybe this has some special effects.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [February 16, 2017, 9:47pm UTC](https://discuss.elastic.co/t/disk-usage-after-indexing-much-higher-then-after-restart-service/75053/7 "2017-02-16T21:47:02Z")

</div>

Does `_cat/segments` change between reboots?

---

<div class="post-metadata">

**Author:** ![Benjamin\_Gathmann](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/benjamin_gathmann/32/7561_2.png) [@Benjamin\_Gathmann](https://discuss.elastic.co/u/Benjamin_Gathmann)\
**Post date:** [February 17, 2017, 9:29am UTC](https://discuss.elastic.co/t/disk-usage-after-indexing-much-higher-then-after-restart-service/75053/8 "2017-02-17T09:29:06Z")

</div>

Before service restart, the following exist in \_cat/segments, but not after:

```
index shard prirep ip segment generation docs.count docs.deleted size size.memory committed searchable version compound     
my_index 0 p 127.0.0.1 _h2 614 1768000 0 1.3gb 0 true false 5.5.2 false    
my_index 0 p 127.0.0.1 _i7 655 410189 0 262.9mb 0 true false 5.5.2 true     
my_index 0 p 127.0.0.1 _kc 732 820267 0 479mb 0 true false 5.5.2 true     
my_index 0 p 127.0.0.1 _mg 808 75 0 36.2kb 0 true false 5.5.2 true     
my_index 0 p 127.0.0.1 _n9 837 1489193 0 1gb 0 true false 5.5.2 false    
my_index 0 p 127.0.0.1 _ou 894 2131694 0 1.5gb 0 true false 5.5.2 false    
my_index 0 p 127.0.0.1 _py 934 166920 0 105mb 0 true false 5.5.2 true     
my_index 0 p 127.0.0.1 _q9 945 490438 0 282.4mb 0 true false 5.5.2 true     
my_index 0 p 127.0.0.1 _r1 973 16775 0 7.9mb 0 true false 5.5.2 true     

```

Plus the following entry changes in the "committed" column (i.e. it is "true" after restart):

```
my_index 0 p 127.0.0.1 _rc 984 7293551 0 5.5gb 1542548 **false** true 5.5.2 false    

```

Cheers,  
Benjamin

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [February 18, 2017, 1:13am UTC](https://discuss.elastic.co/t/disk-usage-after-indexing-much-higher-then-after-restart-service/75053/9 "2017-02-18T01:13:34Z")

</div>

Looks like things are being merged and then compressed then.

---

<div class="post-metadata">

**Author:** ![Benjamin\_Gathmann](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/benjamin_gathmann/32/7561_2.png) [@Benjamin\_Gathmann](https://discuss.elastic.co/u/Benjamin_Gathmann)\
**Post date:** [February 18, 2017, 6:14am UTC](https://discuss.elastic.co/t/disk-usage-after-indexing-much-higher-then-after-restart-service/75053/10 "2017-02-18T06:14:59Z")

</div>

Ok, but will this eventually happen by itself, or do I have to restart the  
service? I am worried if my index grows to a Terabyte, will I have a few  
hundred Gigabytes lying around that only get cleaned up after service  
restart?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [February 18, 2017, 7:11am UTC](https://discuss.elastic.co/t/disk-usage-after-indexing-much-higher-then-after-restart-service/75053/11 "2017-02-18T07:11:33Z")

</div>

It's an automatic process, see [https://www.elastic.co/guide/en/elasticsearch/reference/5.2/index-modules-merge.html](https://www.elastic.co/guide/en/elasticsearch/reference/5.2/index-modules-merge.html)

---

<div class="post-metadata">

**Author:** ![Benjamin\_Gathmann](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/benjamin_gathmann/32/7561_2.png) [@Benjamin\_Gathmann](https://discuss.elastic.co/u/Benjamin_Gathmann)\
**Post date:** [February 20, 2017, 7:53am UTC](https://discuss.elastic.co/t/disk-usage-after-indexing-much-higher-then-after-restart-service/75053/12 "2017-02-20T07:53:48Z")

</div>

OK, thanks.  
I will watch the cluster and see whether the merging happens after some time (but I trust that ES does what the specs say). 🙂

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 20, 2017, 7:54am UTC](https://discuss.elastic.co/t/disk-usage-after-indexing-much-higher-then-after-restart-service/75053/13 "2017-03-20T07:54:33Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
