# Linear growth of index size

**URL:** <https://discuss.elastic.co/t/linear-growth-of-index-size/170723>\
**Category:** Elasticsearch\
**Created:** [March 4, 2019, 12:39pm UTC](https://discuss.elastic.co/t/linear-growth-of-index-size/170723 "2019-03-04T12:39:21Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![amitavmohanty01](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/amitavmohanty01/32/58017_2.png) [@amitavmohanty01](https://discuss.elastic.co/u/amitavmohanty01)\
**Post date:** [March 4, 2019, 12:39pm UTC](https://discuss.elastic.co/t/linear-growth-of-index-size/170723/1 "2019-03-04T12:39:21Z")

</div>

To test the growth of index compared to ingested data, we made 1 million calls to insert the same document in an index. We observed linear growth in index size.

 ![03%20PM](https://us1.discourse-cdn.com/elastic/original/3X/a/5/a516607626f57d8ee1ac2283f17838679b463443.png)

The query to load the docs was:

> curl -X POST “[http://elkelastic01.myhost.com](http://elkelastic01.myhost.com):\<my\_port\>/tests/\_doc” -H ‘Content-Type: application/json’ -d’  
> { “field1" : “This index has only one sentence, nothing more than that” }  
> '

As there is a reverse document index, the growth we anticipated was non-linear (plateau to be specific). What is missing ? Please explain.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [March 4, 2019, 1:13pm UTC](https://discuss.elastic.co/t/linear-growth-of-index-size/170723/2 "2019-03-04T13:13:11Z")

</div>

Each document gets assigned a unique id, which takes up space on disk. You also need to keep track of the data, which is likely to be reasonably linear. Given that your actual data should be very compact it is possible that this is driving disk usage.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 4, 2019, 1:43pm UTC](https://discuss.elastic.co/t/linear-growth-of-index-size/170723/3 "2019-03-04T13:43:29Z")

</div>

Adding that every document is stored in a `_source` field. Also you did not share the mapping. If you are using default, then a `field1.keyword` is more likely generated. It uses doc\_values data structure which like column oriented data structure.

---

<div class="post-metadata">

**Author:** ![amitavmohanty01](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/amitavmohanty01/32/58017_2.png) [@amitavmohanty01](https://discuss.elastic.co/u/amitavmohanty01)\
**Post date:** [March 4, 2019, 7:26pm UTC](https://discuss.elastic.co/t/linear-growth-of-index-size/170723/4 "2019-03-04T19:26:59Z")

</div>

> [@Christian\_Dahlqvist](#):
>
> Each document gets assigned a unique id, which takes up space on disk.

Is there any alternative to it or any way to reduce this space? Can we use an auto-increment id or something which is an integer and takes lesser space?

---

<div class="post-metadata">

**Author:** ![amitavmohanty01](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/amitavmohanty01/32/58017_2.png) [@amitavmohanty01](https://discuss.elastic.co/u/amitavmohanty01)\
**Post date:** [March 4, 2019, 7:36pm UTC](https://discuss.elastic.co/t/linear-growth-of-index-size/170723/5 "2019-03-04T19:36:11Z")

</div>

> [@dadoonet](#):
>
> Adding that every document is stored in a `_source` field.

The size with \_source field is 39287618 bytes and the size when \_source is removed is 37506437 bytes. So, I don't think \_source is the culprit here. Am I missing something?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 4, 2019, 7:50pm UTC](https://discuss.elastic.co/t/linear-growth-of-index-size/170723/6 "2019-03-04T19:50:42Z")

</div>

Yeah that's some 4% more.  
As `_source` is compressed, and because you have very similar documents that's probably why you don't see a big difference. That won't be the case with your final system though may be.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 4, 2019, 7:51pm UTC](https://discuss.elastic.co/t/linear-growth-of-index-size/170723/7 "2019-03-04T19:51:20Z")

</div>

I don't think so. This is not possible for now.

---

<div class="post-metadata">

**Author:** ![amitavmohanty01](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/amitavmohanty01/32/58017_2.png) [@amitavmohanty01](https://discuss.elastic.co/u/amitavmohanty01)\
**Post date:** [March 4, 2019, 8:06pm UTC](https://discuss.elastic.co/t/linear-growth-of-index-size/170723/9 "2019-03-04T20:06:51Z")

</div>

> [@dadoonet](#):
>
> I don't think so. This is not possible for now.

Is there value in considering this for a future release?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [March 4, 2019, 8:26pm UTC](https://discuss.elastic.co/t/linear-growth-of-index-size/170723/10 "2019-03-04T20:26:56Z")

</div>

You can assign your own document IDs which may be shorter and take up less space. There is always going to be data stored per document, so even though Elasticsearch uses a terms dictionary and compresses the source the index size growth should be linear with the number of documents.

---

<div class="post-metadata">

**Author:** ![amitavmohanty01](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/amitavmohanty01/32/58017_2.png) [@amitavmohanty01](https://discuss.elastic.co/u/amitavmohanty01)\
**Post date:** [March 5, 2019, 1:01pm UTC](https://discuss.elastic.co/t/linear-growth-of-index-size/170723/11 "2019-03-05T13:01:23Z")

</div>

> [@dadoonet](#):
>
> Also you did not share the mapping. If you are using default, then a `field1.keyword` is more likely generated. It uses doc\_values data structure which like column oriented data structure.

The mapping that was used is as follows:

```
{
 “tests” : {
   “mappings” : {
     “_doc” : {
       “_size” : {
         “enabled” : true
       },
       “_source” : {
         “enabled” : false
       },
       “properties” : {
         “field1” : {
           “type” : “text”
         }
       }
     }
   }
 }
}

```

---

<div class="post-metadata">

**Author:** ![amitavmohanty01](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/amitavmohanty01/32/58017_2.png) [@amitavmohanty01](https://discuss.elastic.co/u/amitavmohanty01)\
**Post date:** [March 5, 2019, 1:30pm UTC](https://discuss.elastic.co/t/linear-growth-of-index-size/170723/12 "2019-03-05T13:30:20Z")

</div>

> [@dadoonet](#):
>
> Yeah that's some 4% more.  
> As `_source` is compressed, and because you have very similar documents that's probably why you don't see a big difference. That won't be the case with your final system though may be.

We tried a different source for documents. We pushed 32.3m documents into an index. The mapping used is as follows:

```
“mappings” : {
     “doc” : {
       “_size” : {
         “enabled” : true
       },
       “properties” : {
         “@timestamp” : {
           “type” : “date”
         },
         “@version” : {
           “type” : “text”,
           “fields” : {
             “keyword” : {
               “type” : “keyword”,
               “ignore_above” : 256
             }
           }
         },
         “hostname” : {
           “type” : “text”,
           “analyzer” : “pattern”
         },
         “log_level” : {
           “type” : “keyword”
         },
         “message” : {
           “type” : “text”,
           “analyzer” : “pattern”
         },
         “package” : {
           “type” : “text”,
           “analyzer” : “pattern”
         },
         “source” : {
           “type” : “text”,
           “analyzer” : “pattern”
         },
         “tags” : {
           “type” : “text”,
           “fields” : {
             “keyword” : {
               “type” : “keyword”,
               “ignore_above” : 256
             }
           }
         },
         “thread” : {
           “type” : “text”,
           “analyzer” : “pattern”
         }
       }
     }
   }
 }

```

The size of the index we observed was 8.5 GB. We re-indexed the same index but removed \_source. The size reduced to 4.7 GB. This restricted our search capabilities. So, we enabled `store` for fields. The size bumped up to 6.7 GB.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 2, 2019, 1:30pm UTC](https://discuss.elastic.co/t/linear-growth-of-index-size/170723/13 "2019-04-02T13:30:21Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
