# Compress : true flag

**URL:** <https://discuss.elastic.co/t/compress-true-flag/4983>\
**Category:** Elasticsearch\
**Created:** [July 28, 2011, 1:58am UTC](https://discuss.elastic.co/t/compress-true-flag/4983 "2011-07-28T01:58:22Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Anurag\_Phadke](https://avatars.discourse-cdn.com/v4/letter/a/5f9b8f/32.png) [@Anurag\_Phadke](https://discuss.elastic.co/u/Anurag_Phadke)\
**Post date:** [July 28, 2011, 1:58am UTC](https://discuss.elastic.co/t/compress-true-flag/4983/1 "2011-07-28T01:58:22Z")

</div>

Added the compress flag using the following:

curl -XPUT localhost:9200/\_settings -d '{  
"index" : {  
"\_source" : {"compress" : true}  
}  
}'

However, there wasn't any significant difference in the size, approx.  
320gb/node before and 312gb/node after compression on indexing 10m  
documents.  
Any idea what might be wrong here?

-anurag

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 28, 2011, 6:46am UTC](https://discuss.elastic.co/t/compress-true-flag/4983/2 "2011-07-28T06:46:33Z")

</div>

The flag should be set on the mapping of a type:  
[Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/source-field.html).

On Thu, Jul 28, 2011 at 4:58 AM, Anurag [anurag.phadke@gmail.com](mailto:anurag.phadke@gmail.com) wrote:

> Added the compress flag using the following:
> 
> curl -XPUT localhost:9200/\_settings -d '{  
> "index" : {  
> "\_source" : {"compress" : true}  
> }  
> }'
> 
> However, there wasn't any significant difference in the size, approx.  
> 320gb/node before and 312gb/node after compression on indexing 10m  
> documents.  
> Any idea what might be wrong here?
> 
> -anurag

---

<div class="post-metadata">

**Author:** ![Anurag\_Phadke](https://avatars.discourse-cdn.com/v4/letter/a/5f9b8f/32.png) [@Anurag\_Phadke](https://discuss.elastic.co/u/Anurag_Phadke)\
**Post date:** [July 28, 2011, 5:03pm UTC](https://discuss.elastic.co/t/compress-true-flag/4983/3 "2011-07-28T17:03:30Z")

</div>

Shay,  
That worked, I tried to save some more disk by running:  
curl -XPOST '[http://localhost:9200/\_optimize?max\_num\_segments=5](http://localhost:9200/_optimize?max_num_segments=5)'

However, this one actually added about 1GB of data to all the nodes,  
probably over-optimization on my end?

-anurag

On Wed, Jul 27, 2011 at 11:46 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:

> The flag should be set on the mapping of a  
> type: [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/source-field.html).
> 
> On Thu, Jul 28, 2011 at 4:58 AM, Anurag [anurag.phadke@gmail.com](mailto:anurag.phadke@gmail.com) wrote:
> 
> > Added the compress flag using the following:
> > 
> > curl -XPUT localhost:9200/\_settings -d '{  
> > "index" : {  
> > "\_source" : {"compress" : true}  
> > }  
> > }'
> > 
> > However, there wasn't any significant difference in the size, approx.  
> > 320gb/node before and 312gb/node after compression on indexing 10m  
> > documents.  
> > Any idea what might be wrong here?
> > 
> > -anurag

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 29, 2011, 6:38am UTC](https://discuss.elastic.co/t/compress-true-flag/4983/4 "2011-07-29T06:38:56Z")

</div>

Thats strange..., that the data was not reduced. Anything in the logs? Can  
you try and issue a flush to the index and see if it gets smaller (there  
might still be open indexing "reader" against those old index files).

On Thu, Jul 28, 2011 at 8:03 PM, Anurag [anurag.phadke@gmail.com](mailto:anurag.phadke@gmail.com) wrote:

> Shay,  
> That worked, I tried to save some more disk by running:  
> curl -XPOST '[http://localhost:9200/\_optimize?max\_num\_segments=5](http://localhost:9200/_optimize?max_num_segments=5)'
> 
> However, this one actually added about 1GB of data to all the nodes,  
> probably over-optimization on my end?
> 
> -anurag
> 
> On Wed, Jul 27, 2011 at 11:46 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> 
> > The flag should be set on the mapping of a  
> > type:  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/source-field.html).
> > 
> > On Thu, Jul 28, 2011 at 4:58 AM, Anurag [anurag.phadke@gmail.com](mailto:anurag.phadke@gmail.com) wrote:
> > 
> > > Added the compress flag using the following:
> > > 
> > > curl -XPUT localhost:9200/\_settings -d '{  
> > > "index" : {  
> > > "\_source" : {"compress" : true}  
> > > }  
> > > }'
> > > 
> > > However, there wasn't any significant difference in the size, approx.  
> > > 320gb/node before and 312gb/node after compression on indexing 10m  
> > > documents.  
> > > Any idea what might be wrong here?
> > > 
> > > -anurag

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 29, 2011, 10:36am UTC](https://discuss.elastic.co/t/compress-true-flag/4983/5 "2011-07-29T10:36:30Z")

</div>

Transaction logs are not removed only on restart. They are removed on a  
flush (where a Lucene commit is executed, and a new transaction log is  
created). You can issue an optimize with flush flag set to true, which will  
flush post optimization (assuming you wait for them).

> Once optimized and "sync" returns, are we sure the data are persisted?

What is "sync"? Data is persisted once it has been indexed, it has nothing  
to do with optimization or flushing.

On Fri, Jul 29, 2011 at 1:20 PM, Olivier Favre [olivier@yakaz.com](mailto:olivier@yakaz.com) wrote:

> I saw that the transactions logs are only getting removed when restarting  
> ES.  
> If I remember well, they were deleted after an \_optimize before 0.17.2.
> 
> Once optimized and "sync" returns, are we sure the data are persisted?  
> If so, we could flush old transactions logs?
> 
> --  
> Olivier Favre
> 
> [www.yakaz.com](http://www.yakaz.com)
> 
> 2011/7/29 Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com)
> 
> > Thats strange..., that the data was not reduced. Anything in the logs? Can  
> > you try and issue a flush to the index and see if it gets smaller (there  
> > might still be open indexing "reader" against those old index files).
> > 
> > On Thu, Jul 28, 2011 at 8:03 PM, Anurag [anurag.phadke@gmail.com](mailto:anurag.phadke@gmail.com) wrote:
> > 
> > > Shay,  
> > > That worked, I tried to save some more disk by running:  
> > > curl -XPOST '[http://localhost:9200/\_optimize?max\_num\_segments=5](http://localhost:9200/_optimize?max_num_segments=5)'
> > > 
> > > However, this one actually added about 1GB of data to all the nodes,  
> > > probably over-optimization on my end?
> > > 
> > > -anurag
> > > 
> > > On Wed, Jul 27, 2011 at 11:46 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> > > 
> > > > The flag should be set on the mapping of a  
> > > > type:  
> > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/source-field.html).
> > > > 
> > > > On Thu, Jul 28, 2011 at 4:58 AM, Anurag [anurag.phadke@gmail.com](mailto:anurag.phadke@gmail.com)  
> > > > wrote:
> > > > 
> > > > > Added the compress flag using the following:
> > > > > 
> > > > > curl -XPUT localhost:9200/\_settings -d '{  
> > > > > "index" : {  
> > > > > "\_source" : {"compress" : true}  
> > > > > }  
> > > > > }'
> > > > > 
> > > > > However, there wasn't any significant difference in the size, approx.  
> > > > > 320gb/node before and 312gb/node after compression on indexing 10m  
> > > > > documents.  
> > > > > Any idea what might be wrong here?
> > > > > 
> > > > > -anurag

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 29, 2011, 2:31pm UTC](https://discuss.elastic.co/t/compress-true-flag/4983/6 "2011-07-29T14:31:17Z")

</div>

Yea, it seems like a problem, opened an issue for it:  
[When flushing, old transaction log is not removed · Issue #1180 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/1180).

Still don't understand the sync question. Are you referring to when files  
are fsync'ed? If so, that depends. Lucene fsync files on commit, and  
elasticsearch will fsync the transaction log on changes.

On Fri, Jul 29, 2011 at 4:50 PM, Olivier Favre [olivier@yakaz.com](mailto:olivier@yakaz.com) wrote:

> I issued  
> curl -XPOST ..../index/\_optimize  
> which didn't remove the transaction logs.  
> (I even think I issued \_flush and \_refresh manually).  
> And according to the docs[http://www.elasticsearch.org/guide/reference/api/admin-indices-optimize.html](http://www.elasticsearch.org/guide/reference/api/admin-indices-optimize.html),  
> the default values for refresh, flush and wait\_for\_merge are true.
> 
> By sync, I meant the Unix command.
> 
> --  
> Olivier Favre
> 
> [www.yakaz.com](http://www.yakaz.com)
> 
> 2011/7/29 Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com)
> 
> > Transaction logs are not removed only on restart. They are removed on a  
> > flush (where a Lucene commit is executed, and a new transaction log is  
> > created). You can issue an optimize with flush flag set to true, which will  
> > flush post optimization (assuming you wait for them).
> > 
> > > Once optimized and "sync" returns, are we sure the data are persisted?
> > 
> > What is "sync"? Data is persisted once it has been indexed, it has nothing  
> > to do with optimization or flushing.
> > 
> > On Fri, Jul 29, 2011 at 1:20 PM, Olivier Favre [olivier@yakaz.com](mailto:olivier@yakaz.com) wrote:
> > 
> > > I saw that the transactions logs are only getting removed when restarting  
> > > ES.  
> > > If I remember well, they were deleted after an \_optimize before 0.17.2.
> > > 
> > > Once optimized and "sync" returns, are we sure the data are persisted?  
> > > If so, we could flush old transactions logs?
> > > 
> > > --  
> > > Olivier Favre
> > > 
> > > [www.yakaz.com](http://www.yakaz.com)
> > > 
> > > 2011/7/29 Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com)
> > > 
> > > > Thats strange..., that the data was not reduced. Anything in the logs?  
> > > > Can you try and issue a flush to the index and see if it gets smaller (there  
> > > > might still be open indexing "reader" against those old index files).
> > > > 
> > > > On Thu, Jul 28, 2011 at 8:03 PM, Anurag [anurag.phadke@gmail.com](mailto:anurag.phadke@gmail.com)wrote:
> > > > 
> > > > > Shay,  
> > > > > That worked, I tried to save some more disk by running:  
> > > > > curl -XPOST '[http://localhost:9200/\_optimize?max\_num\_segments=5](http://localhost:9200/_optimize?max_num_segments=5)'
> > > > > 
> > > > > However, this one actually added about 1GB of data to all the nodes,  
> > > > > probably over-optimization on my end?
> > > > > 
> > > > > -anurag
> > > > > 
> > > > > On Wed, Jul 27, 2011 at 11:46 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> > > > > 
> > > > > > The flag should be set on the mapping of a  
> > > > > > type:  
> > > > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/source-field.html)  
> > > > > > .
> > > > > > 
> > > > > > On Thu, Jul 28, 2011 at 4:58 AM, Anurag [anurag.phadke@gmail.com](mailto:anurag.phadke@gmail.com)  
> > > > > > wrote:
> > > > > > 
> > > > > > > Added the compress flag using the following:
> > > > > > > 
> > > > > > > curl -XPUT localhost:9200/\_settings -d '{  
> > > > > > > "index" : {  
> > > > > > > "\_source" : {"compress" : true}  
> > > > > > > }  
> > > > > > > }'
> > > > > > > 
> > > > > > > However, there wasn't any significant difference in the size,  
> > > > > > > approx.  
> > > > > > > 320gb/node before and 312gb/node after compression on indexing 10m  
> > > > > > > documents.  
> > > > > > > Any idea what might be wrong here?
> > > > > > > 
> > > > > > > -anurag

---

<div class="post-metadata">

**Author:** ![Anurag\_Phadke](https://avatars.discourse-cdn.com/v4/letter/a/5f9b8f/32.png) [@Anurag\_Phadke](https://discuss.elastic.co/u/Anurag_Phadke)\
**Post date:** [August 1, 2011, 2:26am UTC](https://discuss.elastic.co/t/compress-true-flag/4983/7 "2011-08-01T02:26:30Z")

</div>

Shay,  
I called the flush API using:

curl -XPOST '[http://localhost:9200/\_flush](http://localhost:9200/_flush)'  
it returned quickly with status "ok" and didn't change the size.

-anurag

On Fri, Jul 29, 2011 at 7:31 AM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:

> Yea, it seems like a problem, opened an issue for it: [When flushing, old transaction log is not removed · Issue #1180 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/1180).  
> Still don't understand the sync question. Are you referring to when files are fsync'ed? If so, that depends. Lucene fsync files on commit, and elasticsearch will fsync the transaction log on changes.
> 
> On Fri, Jul 29, 2011 at 4:50 PM, Olivier Favre [olivier@yakaz.com](mailto:olivier@yakaz.com) wrote:
> 
> > I issued  
> > curl -XPOST ..../index/\_optimize  
> > which didn't remove the transaction logs.  
> > (I even think I issued \_flush and \_refresh manually).  
> > And according to the docs, the default values for refresh, flush and wait\_for\_merge are true.  
> > By sync, I meant the Unix command.
> > 
> > --  
> > Olivier Favre  
> > [www.yakaz.com](http://www.yakaz.com)
> > 
> > 2011/7/29 Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com)
> > 
> > > Transaction logs are not removed only on restart. They are removed on a flush (where a Lucene commit is executed, and a new transaction log is created). You can issue an optimize with flush flag set to true, which will flush post optimization (assuming you wait for them).
> > > 
> > > > Once optimized and "sync" returns, are we sure the data are persisted?  
> > > > What is "sync"? Data is persisted once it has been indexed, it has nothing to do with optimization or flushing.
> > > 
> > > On Fri, Jul 29, 2011 at 1:20 PM, Olivier Favre [olivier@yakaz.com](mailto:olivier@yakaz.com) wrote:
> > > 
> > > > I saw that the transactions logs are only getting removed when restarting ES.  
> > > > If I remember well, they were deleted after an \_optimize before 0.17.2.
> > > > 
> > > > ## Once optimized and "sync" returns, are we sure the data are persisted? If so, we could flush old transactions logs?
> > > > 
> > > > Olivier Favre  
> > > > [www.yakaz.com](http://www.yakaz.com)
> > > > 
> > > > 2011/7/29 Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com)
> > > > 
> > > > > Thats strange..., that the data was not reduced. Anything in the logs? Can you try and issue a flush to the index and see if it gets smaller (there might still be open indexing "reader" against those old index files).
> > > > > 
> > > > > On Thu, Jul 28, 2011 at 8:03 PM, Anurag [anurag.phadke@gmail.com](mailto:anurag.phadke@gmail.com) wrote:
> > > > > 
> > > > > > Shay,  
> > > > > > That worked, I tried to save some more disk by running:  
> > > > > > curl -XPOST '[http://localhost:9200/\_optimize?max\_num\_segments=5](http://localhost:9200/_optimize?max_num_segments=5)'
> > > > > > 
> > > > > > However, this one actually added about 1GB of data to all the nodes,  
> > > > > > probably over-optimization on my end?
> > > > > > 
> > > > > > -anurag
> > > > > > 
> > > > > > On Wed, Jul 27, 2011 at 11:46 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> > > > > > 
> > > > > > > The flag should be set on the mapping of a  
> > > > > > > type: [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/source-field.html).
> > > > > > > 
> > > > > > > On Thu, Jul 28, 2011 at 4:58 AM, Anurag [anurag.phadke@gmail.com](mailto:anurag.phadke@gmail.com) wrote:
> > > > > > > 
> > > > > > > > Added the compress flag using the following:
> > > > > > > > 
> > > > > > > > curl -XPUT localhost:9200/\_settings -d '{  
> > > > > > > > "index" : {  
> > > > > > > > "\_source" : {"compress" : true}  
> > > > > > > > }  
> > > > > > > > }'
> > > > > > > > 
> > > > > > > > However, there wasn't any significant difference in the size, approx.  
> > > > > > > > 320gb/node before and 312gb/node after compression on indexing 10m  
> > > > > > > > documents.  
> > > > > > > > Any idea what might be wrong here?
> > > > > > > > 
> > > > > > > > -anurag

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [August 1, 2011, 6:28am UTC](https://discuss.elastic.co/t/compress-true-flag/4983/8 "2011-08-01T06:28:54Z")

</div>

I think what you see is a problem in 0.17.2 that was found by Olivier, where  
transaction logs don't get removed on flush. Its already fixed in 0.17  
branch and will be part of upcoming 0.17.3 (released this week).

On Mon, Aug 1, 2011 at 5:26 AM, Anurag [anurag.phadke@gmail.com](mailto:anurag.phadke@gmail.com) wrote:

> Shay,  
> I called the flush API using:
> 
> curl -XPOST '[http://localhost:9200/\_flush](http://localhost:9200/_flush)'  
> it returned quickly with status "ok" and didn't change the size.
> 
> -anurag
> 
> On Fri, Jul 29, 2011 at 7:31 AM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com) wrote:
> 
> > Yea, it seems like a problem, opened an issue for it:  
> > [When flushing, old transaction log is not removed · Issue #1180 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/1180).  
> > Still don't understand the sync question. Are you referring to when files  
> > are fsync'ed? If so, that depends. Lucene fsync files on commit, and  
> > elasticsearch will fsync the transaction log on changes.
> > 
> > On Fri, Jul 29, 2011 at 4:50 PM, Olivier Favre [olivier@yakaz.com](mailto:olivier@yakaz.com)  
> > wrote:
> > 
> > > I issued  
> > > curl -XPOST ..../index/\_optimize  
> > > which didn't remove the transaction logs.  
> > > (I even think I issued \_flush and \_refresh manually).  
> > > And according to the docs, the default values for refresh, flush and  
> > > wait\_for\_merge are true.  
> > > By sync, I meant the Unix command.
> > > 
> > > --  
> > > Olivier Favre  
> > > [www.yakaz.com](http://www.yakaz.com)
> > > 
> > > 2011/7/29 Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com)
> > > 
> > > > Transaction logs are not removed only on restart. They are removed on a  
> > > > flush (where a Lucene commit is executed, and a new transaction log is  
> > > > created). You can issue an optimize with flush flag set to true, which will  
> > > > flush post optimization (assuming you wait for them).
> > > > 
> > > > > Once optimized and "sync" returns, are we sure the data are  
> > > > > persisted?  
> > > > > What is "sync"? Data is persisted once it has been indexed, it has  
> > > > > nothing to do with optimization or flushing.
> > > > 
> > > > On Fri, Jul 29, 2011 at 1:20 PM, Olivier Favre [olivier@yakaz.com](mailto:olivier@yakaz.com)  
> > > > wrote:
> > > > 
> > > > > I saw that the transactions logs are only getting removed when  
> > > > > restarting ES.  
> > > > > If I remember well, they were deleted after an \_optimize before  
> > > > > 0.17.2.
> > > > > 
> > > > > ## Once optimized and "sync" returns, are we sure the data are persisted? If so, we could flush old transactions logs?
> > > > > 
> > > > > Olivier Favre  
> > > > > [www.yakaz.com](http://www.yakaz.com)
> > > > > 
> > > > > 2011/7/29 Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com)
> > > > > 
> > > > > > Thats strange..., that the data was not reduced. Anything in the  
> > > > > > logs? Can you try and issue a flush to the index and see if it gets smaller  
> > > > > > (there might still be open indexing "reader" against those old index files).
> > > > > > 
> > > > > > On Thu, Jul 28, 2011 at 8:03 PM, Anurag [anurag.phadke@gmail.com](mailto:anurag.phadke@gmail.com)  
> > > > > > wrote:
> > > > > > 
> > > > > > > Shay,  
> > > > > > > That worked, I tried to save some more disk by running:  
> > > > > > > curl -XPOST '[http://localhost:9200/\_optimize?max\_num\_segments=5](http://localhost:9200/_optimize?max_num_segments=5)'
> > > > > > > 
> > > > > > > However, this one actually added about 1GB of data to all the nodes,  
> > > > > > > probably over-optimization on my end?
> > > > > > > 
> > > > > > > -anurag
> > > > > > > 
> > > > > > > On Wed, Jul 27, 2011 at 11:46 PM, Shay Banon [kimchy@gmail.com](mailto:kimchy@gmail.com)  
> > > > > > > wrote:
> > > > > > > 
> > > > > > > > The flag should be set on the mapping of a  
> > > > > > > > type:  
> > > > > > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/source-field.html).
> > > > > > > > 
> > > > > > > > On Thu, Jul 28, 2011 at 4:58 AM, Anurag [anurag.phadke@gmail.com](mailto:anurag.phadke@gmail.com)  
> > > > > > > > wrote:
> > > > > > > > 
> > > > > > > > > Added the compress flag using the following:
> > > > > > > > > 
> > > > > > > > > curl -XPUT localhost:9200/\_settings -d '{  
> > > > > > > > > "index" : {  
> > > > > > > > > "\_source" : {"compress" : true}  
> > > > > > > > > }  
> > > > > > > > > }'
> > > > > > > > > 
> > > > > > > > > However, there wasn't any significant difference in the size,  
> > > > > > > > > approx.  
> > > > > > > > > 320gb/node before and 312gb/node after compression on indexing  
> > > > > > > > > 10m  
> > > > > > > > > documents.  
> > > > > > > > > Any idea what might be wrong here?
> > > > > > > > > 
> > > > > > > > > -anurag

---

<div class="post-metadata">

**Author:** ![Anita](https://avatars.discourse-cdn.com/v4/letter/a/49beb7/32.png) [@Anita](https://discuss.elastic.co/u/Anita)\
**Post date:** [April 20, 2012, 10:40pm UTC](https://discuss.elastic.co/t/compress-true-flag/4983/9 "2012-04-20T22:40:21Z")

</div>

Compression flag is not working.

ElasticSearch version : 0.19.1  
Gateway : Hadoop  
Hadoop version: 0.20.2.cdh3u3

We are planning to use ElasticSearch and Hadoop gateway to store and search log data.  
There will be few hundred tera bytes of log data.

**TEST1:**  
Index: order

INDEX STATUS  
{  
state: open  
settings: {  
index.number\_of\_shards: 10  
index.number\_of\_replicas: 0  
index.version.created: 190199  
index.\_source.compress: true  
}  
Compression flag is set by running curl request  
curl -XPUT localhost:9200/order/\_settings -d '{"index": {"\_source" : {"compress" : true} }}'  
I ran the optimizer  
curl -XPOST '[http://localhost:9200/order/\_optimize?max\_num\_segments=5](http://localhost:9200/order/_optimize?max_num_segments=5)'

Order  
size: 1.2gb (1.2gb)  
docs: 349345 (349345)

Each log message is about 1K.

I am writing raw data to a file to check how compression works  
Raw data : 381978518 Apr 20 12:40 test\_es.log ( Non compressed 380M)

Data written on the hdfs : 1304717764 bytes (1.2G)

**Raw data size is about 380M, how come elasticsearch is written about 1.2 G of data even with compress flag is enabled.  
ElasticSearch wrote about 3 times of raw data.  
I do not see data getting compressed at all. In addition to log message fields, elasticsearch writes few more fields, but how come elasticsearch data size 3 times of original data size? Am I missing any thing here?**

Each document contains

\_index: order  
\_type: trade  
\_id: 5c10PkonSeOQrBZkH8JVpg  
\_version: 1  
\_score: 1  
\_source: {

```
 // source cntains log message. It has 12 fields. Log message is about 1K.

```

}

**TEST2:**  
I tested with/without compression.

Cluster : 3 node cluster  
Replication : 2  
Number of shards : 10  
Gateway : hadoop (5 node cluster, replication 3)

Compression Test

INDEX META DATA  
{  
state: open  
settings: {  
index.number\_of\_shards: 10  
index.number\_of\_replicas: 2  
index.version.created: 190199  
index.\_source.compress: true  
}

index.\_source.compress: true

order  
size: 2.2gb (6.7gb)  
docs: 627336 (627336)

Same log data writing to a file:

685678405 Apr 20 10:44 test\_es.log ( size 685M)

Raw data is only about 685M

Without Compression  
{  
order  
size: 2.2gb (6.8gb)  
docs: 631361 (631361)

690077730 Apr 20 12:45 test\_es.log (SIZE 690M)

**I do not see any difference in data size w/o compression, when I run "du -sh" on the data directory.**

**is it right way to set compression flag?  
curl -XPUT localhost:9200/order/\_settings -d '{"index": {"\_source" : {"compress" : true} }}'  
curl -XPUT localhost:9200/order/\_settings -d '{"index": {"\_source" : {"compress\_threshold" : 10000} }}**

I really appreciate any help on this.

Please give me some guidelines how to scale large data.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:31am UTC](https://discuss.elastic.co/t/compress-true-flag/4983/10 "2017-07-06T03:31:38Z")

</div>


