# ElasticSearch index size peculiarity

**URL:** <https://discuss.elastic.co/t/elasticsearch-index-size-peculiarity/14373>\
**Category:** Elasticsearch\
**Created:** [November 13, 2013, 3:34am UTC](https://discuss.elastic.co/t/elasticsearch-index-size-peculiarity/14373 "2013-11-13T03:34:39Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Avleen\_Vig](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/avleen_vig/32/1985_2.png) [@Avleen\_Vig](https://discuss.elastic.co/u/Avleen_Vig)\
**Post date:** [November 13, 2013, 3:34am UTC](https://discuss.elastic.co/t/elasticsearch-index-size-peculiarity/14373/1 "2013-11-13T03:34:39Z")

</div>

A couple of us have been comparing notes on our Logstash installations at  
the larger end of the scale, and something about ElasticSearch has us  
baffled.  
We're hoping someone here can shed some light on this.

Currently I have 228,262,883 documents in an index.  
It's taking up 243.1Gb of space on disk.  
The average size of the messages going into Logstash (which then converted  
to json and put in ES) was only ~500 bytes each.

At 500 bytes, that's about 106Gb of raw logs.  
I'm adding many fields to the json which gets dropped into ES, but still..  
I would expect that with compression that space used would go _down_, not  
_up_.

This is my mapping: [https://gist.github.com/avleen/7440270](https://gist.github.com/avleen/7440270)  
The only field being analyzed, is "message".

And.. we just removed the "message" field from being sent to elasticsearch.  
The docs/index size ratio did not change much at all (if any).  
Still getting ~1k - 1.5k disk space used, per document in elasticsearch.  
It seems odd that for such a small source, that the stored space should be  
so much larger even with LZ4 compression?

We noticed that while store-level compression might be helping some, it  
doesn't seem to be helping as much as it could. Running gzip on the data  
(raw and the index files) seems to provide quite a bit more compression  
that we're getting right now.  
Likewise, enabling compression on ZFS reduced that space taken by almost  
half.

Overall, I'm trying to index several billion log lines per day, and  
multi-Tb indexes add up in cost.  
Does anyone have any suggestions on what we could do?

(big thanks to Jordan who has already gone way out of way to help with  
this!)  
Thanks 🙂

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![radu\_gheorghe](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/radu_gheorghe/32/556_2.png) [@radu\_gheorghe](https://discuss.elastic.co/u/radu_gheorghe)\
**Post date:** [November 14, 2013, 6:19pm UTC](https://discuss.elastic.co/t/elasticsearch-index-size-peculiarity/14373/2 "2013-11-14T18:19:23Z")

</div>

Hi Avleen,

I only have two ideas, but who knows 🙂

First one is that, looking at your mapping, maybe Logstash really makes  
your document a lot bigger than the originals. Even if most of the stuff is  
not\_analyzed. To verify this, you can do an experiment:

- enable the size field, and index a few docs:  
[Elastic — The Search AI Company | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/mapping-size-field.html)

This should get you the size of the JSONs you're indexing. You may even  
want to run a statistical facet to see what's the average size of the logs  
you indexed:

> **[Elastic — The Search AI Company](https://www.elastic.co)**
>
> Power insights and outcomes with The Elastic Search AI Platform. See into your data and find answers that matter with enterprise solutions designed to help you accelerate time to insight. Try Elastic ...

If you get 3k per document, I'm not sure you can do much.

The second idea may be a bit in Captain Obvious territory, but it's worth  
checking. Which ES version are you on? With 0.90 or later, you should have  
good compression by default, at the Lucene level. So, if you're not already  
there, it might be worth upgrading to the latest 0.90.7, if you're using  
Logstash's elasticsearch\_http output. If you're using the "standard"  
elasticsearch output plugin, it might be worth upgrading to the latest  
Logstash (1.2.2, runs on 0.90.3). Latest Logstash should be faster, too. I  
know because I'm using it:

> **[Logstash Tutorial: Getting Started with Logging - Sematext](https://sematext.com/blog/getting-started-with-logstash/)**
>
> Get started shipping logs with this Logstash tutorial. Learn what it is, what it is used for and how it works. Logging configuration examples.

## Best regards, Radu

Performance Monitoring \* Log Analytics \* Search Analytics  
Solr & Elasticsearch Support \* [http://sematext.com/](http://sematext.com/)

On Tue, Nov 12, 2013 at 7:34 PM, Avleen Vig [avleen@gmail.com](mailto:avleen@gmail.com) wrote:

> A couple of us have been comparing notes on our Logstash installations at  
> the larger end of the scale, and something about Elasticsearch has us  
> baffled.  
> We're hoping someone here can shed some light on this.
> 
> Currently I have 228,262,883 documents in an index.  
> It's taking up 243.1Gb of space on disk.  
> The average size of the messages going into Logstash (which then converted  
> to json and put in ES) was only ~500 bytes each.
> 
> At 500 bytes, that's about 106Gb of raw logs.  
> I'm adding many fields to the json which gets dropped into ES, but still..  
> I would expect that with compression that space used would go _down_, not  
> _up_.
> 
> This is my mapping: [gist:7440270 · GitHub](https://gist.github.com/avleen/7440270)  
> The only field being analyzed, is "message".
> 
> And.. we just removed the "message" field from being sent to  
> elasticsearch. The docs/index size ratio did not change much at all (if  
> any).  
> Still getting ~1k - 1.5k disk space used, per document in elasticsearch.  
> It seems odd that for such a small source, that the stored space should be  
> so much larger even with LZ4 compression?
> 
> We noticed that while store-level compression might be helping some, it  
> doesn't seem to be helping as much as it could. Running gzip on the data  
> (raw and the index files) seems to provide quite a bit more compression  
> that we're getting right now.  
> Likewise, enabling compression on ZFS reduced that space taken by almost  
> half.
> 
> Overall, I'm trying to index several billion log lines per day, and  
> multi-Tb indexes add up in cost.  
> Does anyone have any suggestions on what we could do?
> 
> (big thanks to Jordan who has already gone way out of way to help with  
> this!)  
> Thanks 🙂
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:06am UTC](https://discuss.elastic.co/t/elasticsearch-index-size-peculiarity/14373/3 "2017-07-06T02:06:56Z")

</div>


