# Disk usage on benchmarks

**URL:** <https://discuss.elastic.co/t/disk-usage-on-benchmarks/278908>\
**Category:** Elasticsearch\
**Created:** [July 16, 2021, 1:34pm UTC](https://discuss.elastic.co/t/disk-usage-on-benchmarks/278908 "2021-07-16T13:34:50Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![zj27](https://avatars.discourse-cdn.com/v4/letter/z/b4bc9f/32.png) [@zj27](https://discuss.elastic.co/u/zj27)\
**Post date:** [July 16, 2021, 1:34pm UTC](https://discuss.elastic.co/t/disk-usage-on-benchmarks/278908/1 "2021-07-16T13:34:50Z")

</div>

Hi all,

I'm studying the [Elasticsearch Adhoc Benchmarks](https://elasticsearch-benchmarks.elastic.co/no-omit/nyc_taxis/index.html) and am curious about the disk usage metrics. Why there is a huge gap between final index size (25GB) and total bytes written(313GB)?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 16, 2021, 6:16pm UTC](https://discuss.elastic.co/t/disk-usage-on-benchmarks/278908/2 "2021-07-16T18:16:38Z")

</div>

Lucene uses immutable segments to store data and these are created reasonably small and then merged into larger ones. This means that the same data is written to disk multiple times as more data is added to the index and merging creates larger and larger segments.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 13, 2021, 6:17pm UTC](https://discuss.elastic.co/t/disk-usage-on-benchmarks/278908/3 "2021-08-13T18:17:15Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
