# Storage Ratios - I my syslog streams are expanding in elastic search to more than 10:1?

**URL:** <https://discuss.elastic.co/t/storage-ratios-i-my-syslog-streams-are-expanding-in-elastic-search-to-more-than-10-1/9215>\
**Category:** Elasticsearch\
**Created:** [October 2, 2012, 6:11pm UTC](https://discuss.elastic.co/t/storage-ratios-i-my-syslog-streams-are-expanding-in-elastic-search-to-more-than-10-1/9215 "2012-10-02T18:11:56Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Dylan\_Johnson](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dylan_johnson/32/2698_2.png) [@Dylan\_Johnson](https://discuss.elastic.co/u/Dylan_Johnson)\
**Post date:** [October 2, 2012, 6:11pm UTC](https://discuss.elastic.co/t/storage-ratios-i-my-syslog-streams-are-expanding-in-elastic-search-to-more-than-10-1/9215/1 "2012-10-02T18:11:56Z")

</div>

I have a directory that has 10GB of data and i used Logstash to parse the  
date to elasticsearch which works fine. I have 2 index and 0 replication  
however after logstash has finished parsing the date to ES the ES /data  
directory is 80GB ?

This is unworkable ? Whats the reason for this ?

Thanks

Dylan

--

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [October 3, 2012, 1:16am UTC](https://discuss.elastic.co/t/storage-ratios-i-my-syslog-streams-are-expanding-in-elastic-search-to-more-than-10-1/9215/2 "2012-10-03T01:16:58Z")

</div>

Hello Dylan,

ES stored the original JSON in the special \_source field. That right there  
means your index will be at least 10GB in size. Additionally, it is  
possible your fields are also stored and not just index - the store part  
would be another 10GB. On top of that is the inverted index. But I'm not  
sure how you get to 80 GB. Maybe you are using ngrams somewhere? Maybe  
you can check with Skywalker plugin what's inside your index?

## Otis

Search Analytics - [Cloud Monitoring Tools & Services | Sematext](http://sematext.com/search-analytics/index.html)  
Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)

On Tuesday, October 2, 2012 2:11:56 PM UTC-4, Dylan Johnson wrote:

> I have a directory that has 10GB of data and i used Logstash to parse the  
> date to elasticsearch which works fine. I have 2 index and 0 replication  
> however after logstash has finished parsing the date to ES the ES /data  
> directory is 80GB ?
> 
> This is unworkable ? Whats the reason for this ?
> 
> Thanks
> 
> Dylan

--

---

<div class="post-metadata">

**Author:** ![Dylan\_Johnson](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dylan_johnson/32/2698_2.png) [@Dylan\_Johnson](https://discuss.elastic.co/u/Dylan_Johnson)\
**Post date:** [October 3, 2012, 5:41pm UTC](https://discuss.elastic.co/t/storage-ratios-i-my-syslog-streams-are-expanding-in-elastic-search-to-more-than-10-1/9215/3 "2012-10-03T17:41:25Z")

</div>

Seems like my replicas were making up the space.

So i start with a couple of GB of data and end up with about x3 , this seems expensive ?

Does anyone have any info on compressing the index's or using some sort of archiving setup ? In other words whats the best way to save space when using ES ?

Does anyone have any recommendations on configuring ES so that the DB size is similar to that of the original data that was parsed into it ?

My CIO wont take me seriously when i tell him for every TB of data we need 3 TB in ES

Thanks

D

On Oct 3, 2012, at 2:16 AM, Otis Gospodnetic wrote:

> Hello Dylan,
> 
> ES stored the original JSON in the special \_source field. That right there means your index will be at least 10GB in size. Additionally, it is possible your fields are also stored and not just index - the store part would be another 10GB. On top of that is the inverted index. But I'm not sure how you get to 80 GB. Maybe you are using ngrams somewhere? Maybe you can check with Skywalker plugin what's inside your index?
> 
> ## Otis
> 
> Search Analytics - [Cloud Monitoring Tools & Services | Sematext](http://sematext.com/search-analytics/index.html)  
> Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> 
> On Tuesday, October 2, 2012 2:11:56 PM UTC-4, Dylan Johnson wrote:  
> I have a directory that has 10GB of data and i used Logstash to parse the date to elasticsearch which works fine. I have 2 index and 0 replication however after logstash has finished parsing the date to ES the ES /data directory is 80GB ?
> 
> This is unworkable ? Whats the reason for this ?
> 
> Thanks
> 
> Dylan
> 
> --

--

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [October 3, 2012, 8:17pm UTC](https://discuss.elastic.co/t/storage-ratios-i-my-syslog-streams-are-expanding-in-elastic-search-to-more-than-10-1/9215/4 "2012-10-03T20:17:33Z")

</div>

Hello,

You could disable \_source. You could compress it -

> **[Elastic — The Search AI Company](https://www.elastic.co)**
>
> Power insights and outcomes with The Elastic Search AI Platform. See into your data and find answers that matter with enterprise solutions designed to help you accelerate time to insight. Try Elastic ...

You could have 0 replicas (how many does the DB have?)

See also  
[Elastic — The Search AI Company | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/store.html) .

## Otis

Search Analytics - [Cloud Monitoring Tools & Services | Sematext](http://sematext.com/search-analytics/index.html)  
Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)

On Wednesday, October 3, 2012 1:41:41 PM UTC-4, Dylan Johnson wrote:

> Seems like my replicas were making up the space.
> 
> So i start with a couple of GB of data and end up with about x3 , this  
> seems expensive ?
> 
> Does anyone have any info on compressing the index's or using some sort of  
> archiving setup ? In other words whats the best way to save space when  
> using ES ?
> 
> Does anyone have any recommendations on configuring ES so that the DB size  
> is similar to that of the original data that was parsed into it ?
> 
> My CIO wont take me seriously when i tell him for every TB of data we need  
> 3 TB in ES
> 
> Thanks
> 
> D
> 
> On Oct 3, 2012, at 2:16 AM, Otis Gospodnetic wrote:
> 
> Hello Dylan,
> 
> ES stored the original JSON in the special \_source field. That right  
> there means your index will be at least 10GB in size. Additionally, it is  
> possible your fields are also stored and not just index - the store part  
> would be another 10GB. On top of that is the inverted index. But I'm not  
> sure how you get to 80 GB. Maybe you are using ngrams somewhere? Maybe  
> you can check with Skywalker plugin what's inside your index?
> 
> ## Otis
> 
> Search Analytics - [Cloud Monitoring Tools & Services | Sematext](http://sematext.com/search-analytics/index.html)  
> Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> 
> On Tuesday, October 2, 2012 2:11:56 PM UTC-4, Dylan Johnson wrote:
> 
> > I have a directory that has 10GB of data and i used Logstash to parse the  
> > date to elasticsearch which works fine. I have 2 index and 0 replication  
> > however after logstash has finished parsing the date to ES the ES /data  
> > directory is 80GB ?
> > 
> > This is unworkable ? Whats the reason for this ?
> > 
> > Thanks
> > 
> > Dylan
> 
> --

--

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [October 4, 2012, 7:25am UTC](https://discuss.elastic.co/t/storage-ratios-i-my-syslog-streams-are-expanding-in-elastic-search-to-more-than-10-1/9215/5 "2012-10-04T07:25:21Z")

</div>

If you use "vanilla" logstash, without changing anything, then you can look here for some hints (specifically, the new stored compression): [GitHub - elastic/logstash: Logstash - transport and process your logs, events, or other data](https://github.com/logstash/logstash/wiki/Elasticsearch-Storage-Optimization). @whack also gisted (but I can't find it) an experiment that he did with storage sizes with ES, can't find it now, you can possibly ping him.

On Oct 3, 2012, at 7:41 PM, Dylan Johnson [dylandjohnson@googlemail.com](mailto:dylandjohnson@googlemail.com) wrote:

> Seems like my replicas were making up the space.
> 
> So i start with a couple of GB of data and end up with about x3 , this seems expensive ?
> 
> Does anyone have any info on compressing the index's or using some sort of archiving setup ? In other words whats the best way to save space when using ES ?
> 
> Does anyone have any recommendations on configuring ES so that the DB size is similar to that of the original data that was parsed into it ?
> 
> My CIO wont take me seriously when i tell him for every TB of data we need 3 TB in ES
> 
> Thanks
> 
> D
> 
> On Oct 3, 2012, at 2:16 AM, Otis Gospodnetic wrote:
> 
> > Hello Dylan,
> > 
> > ES stored the original JSON in the special \_source field. That right there means your index will be at least 10GB in size. Additionally, it is possible your fields are also stored and not just index - the store part would be another 10GB. On top of that is the inverted index. But I'm not sure how you get to 80 GB. Maybe you are using ngrams somewhere? Maybe you can check with Skywalker plugin what's inside your index?
> > 
> > ## Otis
> > 
> > Search Analytics - [Cloud Monitoring Tools & Services | Sematext](http://sematext.com/search-analytics/index.html)  
> > Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> > 
> > On Tuesday, October 2, 2012 2:11:56 PM UTC-4, Dylan Johnson wrote:  
> > I have a directory that has 10GB of data and i used Logstash to parse the date to elasticsearch which works fine. I have 2 index and 0 replication however after logstash has finished parsing the date to ES the ES /data directory is 80GB ?
> > 
> > This is unworkable ? Whats the reason for this ?
> > 
> > Thanks
> > 
> > Dylan
> > 
> > --
> 
> --

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:10am UTC](https://discuss.elastic.co/t/storage-ratios-i-my-syslog-streams-are-expanding-in-elastic-search-to-more-than-10-1/9215/6 "2017-07-06T03:10:11Z")

</div>


