# Elasticsearch index MUCH larger then similar lucene index

**URL:** <https://discuss.elastic.co/t/elasticsearch-index-much-larger-then-similar-lucene-index/12056>\
**Category:** Elasticsearch\
**Created:** [May 21, 2013, 3:53pm UTC](https://discuss.elastic.co/t/elasticsearch-index-much-larger-then-similar-lucene-index/12056 "2013-05-21T15:53:30Z")\
**Posts on this page:** 15\
**Page:** 3

<div class="post-metadata">

**Author:** ![shlomivaknin](https://avatars.discourse-cdn.com/v4/letter/s/4da419/32.png) [@shlomivaknin](https://discuss.elastic.co/u/shlomivaknin)\
**Post date:** [May 30, 2013, 10:20am UTC](https://discuss.elastic.co/t/elasticsearch-index-much-larger-then-similar-lucene-index/12056/41 "2013-05-30T10:20:42Z")

</div>

I copied the wrong line before...

ES was actually:

curl -XPOST '[http://host:9200/test/\_optimize?max\_num\_segments=1](http://host:9200/test/_optimize?max_num_segments=1)'  
_{"ok":true,"\_shards":{"total":1,"successful":1,"failed":0}}_  
\*  
\*  
just to not throw you off in the wrong direction...

On Thursday, May 30, 2013 1:02:25 PM UTC+3, Shlomi wrote:

> Israel, sorry for any inconvenience my thread has caused you.
> 
> now back to the really annoying results:
> 
> ES version :
> 
> > curl -XPOST '[http://host:9200/test/\_optimize?max\_num\_segments=1](http://host:9200/test/_optimize?max_num_segments=1)'  
> > {"ok":true,"\_shards":{"total":0,"successful":0,"failed":0}}
> 
> > ls -ltra  
> > total 16623176  
> > drwxr-xr-x 5 elasticsearch elasticsearch 4096 May 27 17:34 ..  
> > -rw-r--r-- 1 elasticsearch elasticsearch 31 May 27 18:16 \_1s0.fnm  
> > -rw-r--r-- 1 elasticsearch elasticsearch 240390660 May 27 18:16 \_1s0.fdx  
> > -rw-r--r-- 1 elasticsearch elasticsearch 2178157235 May 27 18:16 \_1s0.fdt  
> > -rw-r--r-- 1 elasticsearch elasticsearch 742546522 May 27 18:17 \_1s0.tis  
> > -rw-r--r-- 1 elasticsearch elasticsearch 7152131 May 27 18:17 \_1s0.tii  
> > -rw-r--r-- 1 elasticsearch elasticsearch 440466009 May 27 18:17 \_1s0.prx  
> > -rw-r--r-- 1 elasticsearch elasticsearch 1017914310 May 27 18:17 \_1s0.frq  
> > -rw-r--r-- 1 elasticsearch elasticsearch 30048836 May 27 18:17 \_1s0.nrm  
> > -rw-r--r-- 1 elasticsearch elasticsearch 31 May 27 18:38 \_2oj.fnm  
> > -rw-r--r-- 1 elasticsearch elasticsearch 2149916547 May 27 18:38 \_2oj.fdt  
> > -rw-r--r-- 1 elasticsearch elasticsearch 238283772 May 27 18:38 \_2oj.fdx  
> > -rw-r--r-- 1 elasticsearch elasticsearch 735613612 May 27 18:39 \_2oj.tis  
> > -rw-r--r-- 1 elasticsearch elasticsearch 7082393 May 27 18:39 \_2oj.tii  
> > -rw-r--r-- 1 elasticsearch elasticsearch 434339734 May 27 18:39 \_2oj.prx  
> > -rw-r--r-- 1 elasticsearch elasticsearch 1005557319 May 27 18:39 \_2oj.frq  
> > -rw-r--r-- 1 elasticsearch elasticsearch 29785475 May 27 18:39 \_2oj.nrm  
> > -rw-r--r-- 1 elasticsearch elasticsearch 0 May 30 11:49 write.lock  
> > -rw-r--r-- 1 elasticsearch elasticsearch 31 May 30 11:50 \_37c.fnm  
> > -rw-r--r-- 1 elasticsearch elasticsearch 402061692 May 30 11:50 \_37c.fdx  
> > -rw-r--r-- 1 elasticsearch elasticsearch 3636925770 May 30 11:50 \_37c.fdt  
> > -rw-r--r-- 1 elasticsearch elasticsearch 1229530031 May 30 11:52 \_37c.tis  
> > -rw-r--r-- 1 elasticsearch elasticsearch 11770457 May 30 11:52 \_37c.tii  
> > -rw-r--r-- 1 elasticsearch elasticsearch 735561692 May 30 11:52 \_37c.prx  
> > -rw-r--r-- 1 elasticsearch elasticsearch 1698617265 May 30 11:52 \_37c.frq  
> > -rw-r--r-- 1 elasticsearch elasticsearch 50257715 May 30 11:53 \_37c.nrm  
> > -rw-r--r-- 1 elasticsearch elasticsearch 828 May 30 11:53  
> > segments\_4n  
> > -rw-r--r-- 1 elasticsearch elasticsearch 20 May 30 11:53  
> > segments.gen  
> > -rw-r--r-- 1 elasticsearch elasticsearch 138 May 30 11:53  
> > \_checksums-1369903982814  
> > drwxr-xr-x 2 elasticsearch elasticsearch 20480 May 30 11:53 .
> 
> java version after optimize to normal file format setting max\_segments=1:
> 
> > ls -ltr
> 
> total 8759876  
> -rw-rw-r-- 1 shlomiv shlomiv 24 May 30 11:34 \_ao.fnm  
> -rw-rw-r-- 1 shlomiv shlomiv 883151756 May 30 11:35 \_ao.fdx  
> -rw-rw-r-- 1 shlomiv shlomiv 4343895906 May 30 11:35 \_ao.fdt  
> -rw-rw-r-- 1 shlomiv shlomiv 14132289 May 30 11:36 \_ao.tis  
> -rw-rw-r-- 1 shlomiv shlomiv 197431 May 30 11:36 \_ao.tii  
> -rw-rw-r-- 1 shlomiv shlomiv 506303552 May 30 11:36 \_ao.prx  
> -rw-rw-r-- 1 shlomiv shlomiv 3111989398 May 30 11:36 \_ao.frq  
> -rw-rw-r-- 1 shlomiv shlomiv 110393973 May 30 11:36 \_ao.nrm  
> -rw-rw-r-- 1 shlomiv shlomiv 285 May 30 11:36 segments\_38  
> -rw-rw-r-- 1 shlomiv shlomiv 20 May 30 11:36 segments.gen
> 
> still twice the size, after optimization. Israel is right, this saga is  
> really annoying 🙂
> 
> thanks a lot for your patience, i think its important to understand the  
> cause of this size increase, and not just for my sake

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [May 30, 2013, 10:26am UTC](https://discuss.elastic.co/t/elasticsearch-index-much-larger-then-similar-lucene-index/12056/42 "2013-05-30T10:26:01Z")

</div>

Your Elasticsearch index seems to have several segments (for example, an  
optimized index would have only one .fdx file), are you sure you listed the  
correct directory?

--  
Adrien Grand

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![shlomivaknin](https://avatars.discourse-cdn.com/v4/letter/s/4da419/32.png) [@shlomivaknin](https://discuss.elastic.co/u/shlomivaknin)\
**Post date:** [May 30, 2013, 10:58am UTC](https://discuss.elastic.co/t/elasticsearch-index-much-larger-then-similar-lucene-index/12056/43 "2013-05-30T10:58:00Z")

</div>

i ran that again

> curl -XPOST '[http://host:9200/test/\_optimize?max\_num\_segments=1](http://host:9200/test/_optimize?max_num_segments=1)  
> {"ok":true,"\_shards":{"total":1,"successful":1,"failed":0}}

and got the same listing.

so i ran optimize with luke, and that gave me back a better listing, still  
of the same size..

> ls -ltr  
> total 16586872  
> -rw-r--r-- 1 elasticsearch elasticsearch 31 May 30 13:44 \_37d.fnm  
> -rw-r--r-- 1 elasticsearch elasticsearch 880736116 May 30 13:46 \_37d.fdx  
> -rw-r--r-- 1 elasticsearch elasticsearch 7964999544 May 30 13:46 \_37d.fdt  
> -rw-r--r-- 1 elasticsearch elasticsearch 2665101249 May 30 13:50 \_37d.tis  
> -rw-r--r-- 1 elasticsearch elasticsearch 25370865 May 30 13:50 \_37d.tii  
> -rw-r--r-- 1 elasticsearch elasticsearch 1610367435 May 30 13:50 \_37d.prx  
> -rw-r--r-- 1 elasticsearch elasticsearch 3728236673 May 30 13:50 \_37d.frq  
> -rw-r--r-- 1 elasticsearch elasticsearch 110092018 May 30 13:50 \_37d.nrm  
> -rw-r--r-- 1 elasticsearch elasticsearch 313 May 30 13:50 segments\_4o  
> -rw-r--r-- 1 elasticsearch elasticsearch 20 May 30 13:50  
> segments.gen  
> -rw-r--r-- 1 elasticsearch elasticsearch 270 May 30 13:50  
> \_checksums-1369911034916

thanks

On Thu, May 30, 2013 at 1:26 PM, Adrien Grand \<  
[adrien.grand@elasticsearch.com](mailto:adrien.grand@elasticsearch.com)\> wrote:

> Your Elasticsearch index seems to have several segments (for example, an  
> optimized index would have only one .fdx file), are you sure you listed the  
> correct directory?
> 
> --  
> Adrien Grand
> 
> --  
> You received this message because you are subscribed to a topic in the  
> Google Groups "elasticsearch" group.  
> To unsubscribe from this topic, visit  
> [https://groups.google.com/d/topic/elasticsearch/6j0E-2pTbWg/unsubscribe?hl=en-US](https://groups.google.com/d/topic/elasticsearch/6j0E-2pTbWg/unsubscribe?hl=en-US)  
> .  
> To unsubscribe from this group and all its topics, send an email to  
> [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [May 30, 2013, 11:34am UTC](https://discuss.elastic.co/t/elasticsearch-index-much-larger-then-similar-lucene-index/12056/44 "2013-05-30T11:34:37Z")

</div>

OK, so the larger files are

- fdt: stored fields
- tis, tii: terms dictionary
- prx: positions

So a few ideas:

- Did you use the same analyzers?
- Did you use the same index options?
- Did you mark all your fields stored in your mapping? This would explain  
why the fdt file is almost exactly 2x larger.

--  
Adrien Grand

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![shlomivaknin](https://avatars.discourse-cdn.com/v4/letter/s/4da419/32.png) [@shlomivaknin](https://discuss.elastic.co/u/shlomivaknin)\
**Post date:** [May 30, 2013, 3:03pm UTC](https://discuss.elastic.co/t/elasticsearch-index-much-larger-then-similar-lucene-index/12056/45 "2013-05-30T15:03:40Z")

</div>

Hey

> - Did you use the same analyzers?

I think so.. in the java code i use WhitespaceTokenizer + LowerCaseFilter

- FilteringTokenFilter

WhitespaceTokenizer(Version.LUCENE\_35, reader);  
LowerCaseFilter lowerCaseFilter = new  
LowerCaseFilter(Version.LUCENE\_35, whitespaceTokenizer);  
return new PunctuationFilter(false, lowerCaseFilter);

where  
public class PunctuationFilter extends FilteringTokenFilter {....}

the first two should be equal to ES's  
"lowercase[http://www.elasticsearch.org/guide/reference/index-modules/analysis/lowercase-tokenizer/](http://www.elasticsearch.org/guide/reference/index-modules/analysis/lowercase-tokenizer/)"  
tokenizer (right?), and our implementation of FilteringTokenFilter excludes  
exactly the same tokens as listed in our stop words filter

"analyzer": {  
"ngram-index": {  
"tokenizer": "lowercase",  
"filter": [  
"myStop" \<--- contains the same exact list as used in our  
custom FilteringTokenFilter, i know, i copy-pasted myself 😉  
],  
"type": "custom"  
}

> - Did you use the same index options?

I dont exactly understand what options you mean by index options...

> - Did you mark all your fields stored in your mapping? This would explain  
> why the fdt file is almost exactly 2x larger.

well, my fields are stored, but my \_source and \_all are not (i figured i  
rather have the long be stored as long than have the original json saved,  
to save space).  
but with other attempts having store="no" and \_source enabled gave me  
pretty much similar results (I think maybe a few gb larger when used  
\_source)

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [May 30, 2013, 3:18pm UTC](https://discuss.elastic.co/t/elasticsearch-index-much-larger-then-similar-lucene-index/12056/46 "2013-05-30T15:18:38Z")

</div>

Hi,

On Thu, May 30, 2013 at 5:03 PM, Shlomi Vaknin [shlomivaknin@gmail.com](mailto:shlomivaknin@gmail.com)wrote:

> I dont exactly understand what options you mean by index options...

Index options are a way to tell Lucene whether positions and offsets should  
be indexed for a given field.

Could you reindex 1% of your data and upload your Lucene and Elasticsearch  
indexes somewhere? If you can, I'd be happy do have a deeper look at them  
to better understand what happens.

--  
Adrien Grand

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![shlomivaknin](https://avatars.discourse-cdn.com/v4/letter/s/4da419/32.png) [@shlomivaknin](https://discuss.elastic.co/u/shlomivaknin)\
**Post date:** [May 30, 2013, 3:33pm UTC](https://discuss.elastic.co/t/elasticsearch-index-much-larger-then-similar-lucene-index/12056/47 "2013-05-30T15:33:46Z")

</div>

Hey,

Thanks, i found index options, checking that out now.

about uploading a portion of my data, ill have to get back to you on monday  
for this one. need to ask first 🙂

Thanks!

On Thu, May 30, 2013 at 6:18 PM, Adrien Grand \<  
[adrien.grand@elasticsearch.com](mailto:adrien.grand@elasticsearch.com)\> wrote:

> Hi,
> 
> On Thu, May 30, 2013 at 5:03 PM, Shlomi Vaknin [shlomivaknin@gmail.com](mailto:shlomivaknin@gmail.com)wrote:
> 
> > I dont exactly understand what options you mean by index options...
> 
> Index options are a way to tell Lucene whether positions and offsets  
> should be indexed for a given field.
> 
> Could you reindex 1% of your data and upload your Lucene and Elasticsearch  
> indexes somewhere? If you can, I'd be happy do have a deeper look at them  
> to better understand what happens.
> 
> --  
> Adrien Grand
> 
> --  
> You received this message because you are subscribed to a topic in the  
> Google Groups "elasticsearch" group.  
> To unsubscribe from this topic, visit  
> [https://groups.google.com/d/topic/elasticsearch/6j0E-2pTbWg/unsubscribe?hl=en-US](https://groups.google.com/d/topic/elasticsearch/6j0E-2pTbWg/unsubscribe?hl=en-US)  
> .  
> To unsubscribe from this group and all its topics, send an email to  
> [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [June 1, 2013, 12:46am UTC](https://discuss.elastic.co/t/elasticsearch-index-much-larger-then-similar-lucene-index/12056/48 "2013-06-01T00:46:25Z")

</div>

hi, I would add that make sure that the mapping that you think you are setting are actually set using the get mapping API on the live ES node.

also, I would try and remove all variable, and simply run it with a simple lucene program that does not use any custom analyzes or similarity, and the same withe ES. if you still see a difference, it will be much simpler to help.

last, simpler to ru optimize in ten lucene code down to a single segment, and same with ES (call optimize with max num segments set to 1)

really last, don't index that much, you can index 100mb and you should still see the difference

On Thu, May 30, 2013 at 5:34 PM, Shlomi Vaknin [shlomivaknin@gmail.com](mailto:shlomivaknin@gmail.com)  
wrote:

> Hey,  
> Thanks, i found index options, checking that out now.  
> about uploading a portion of my data, ill have to get back to you on monday  
> for this one. need to ask first 🙂  
> Thanks!  
> On Thu, May 30, 2013 at 6:18 PM, Adrien Grand \<  
> [adrien.grand@elasticsearch.com](mailto:adrien.grand@elasticsearch.com)\> wrote:
> 
> > Hi,
> > 
> > On Thu, May 30, 2013 at 5:03 PM, Shlomi Vaknin [shlomivaknin@gmail.com](mailto:shlomivaknin@gmail.com)wrote:
> > 
> > > I dont exactly understand what options you mean by index options...
> > 
> > Index options are a way to tell Lucene whether positions and offsets  
> > should be indexed for a given field.
> > 
> > Could you reindex 1% of your data and upload your Lucene and Elasticsearch  
> > indexes somewhere? If you can, I'd be happy do have a deeper look at them  
> > to better understand what happens.
> > 
> > --  
> > Adrien Grand
> > 
> > --  
> > You received this message because you are subscribed to a topic in the  
> > Google Groups "elasticsearch" group.  
> > To unsubscribe from this topic, visit  
> > [https://groups.google.com/d/topic/elasticsearch/6j0E-2pTbWg/unsubscribe?hl=en-US](https://groups.google.com/d/topic/elasticsearch/6j0E-2pTbWg/unsubscribe?hl=en-US)  
> > .  
> > To unsubscribe from this group and all its topics, send an email to  
> > [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).
> 
> --  
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![shlomivaknin](https://avatars.discourse-cdn.com/v4/letter/s/4da419/32.png) [@shlomivaknin](https://discuss.elastic.co/u/shlomivaknin)\
**Post date:** [June 6, 2013, 2:20pm UTC](https://discuss.elastic.co/t/elasticsearch-index-much-larger-then-similar-lucene-index/12056/49 "2013-06-06T14:20:25Z")

</div>

Hi,

I took some time off of this subject, and now i am back 🙂

I took Shay's advice, and removed all special stuff. I now use the  
standardAnalyzer with no special similarity and no compound file and I use  
LogByteSizeMergePolicy.  
in elastic, i used head to check the index, here is what i got:  
{

- state: open
- settings: {
  - index.number\_of\_replicas: 0
  - index.number\_of\_shards: 1
  - index.version.created: 200699  
}

- mappings: {
  - test: {
    - \_source: {
      - enabled: false  
}

    - properties: {
      - freq: {
        - store: yes
        - type: long  
}

      - gram: {
        - store: yes
        - type: string  
}  
}

    - \_all: {
      - enabled: false  
}  
}  
}

- aliases: []

}

so i guess the types are as i set them.

you know what, i wont guess, here, i used the mapping api to get it:  
curl -XGET '[http://es:9200/test/\_mapping](http://es:9200/test/_mapping)'

{"test":{"test":{"\_all":{"enabled":false},"\_source":{"enabled":false},"properties":{"freq":{"type":"long","store":"yes"},"gram":{"type":"string","store":"yes"}}}}}

ok, that looks the same. i hope i am not missing anything here..

i didnt specify any special mappings, analyzers, custom similarity or stop  
words. everything is standard.

as suggested, i ran just a small sample, with a final  
IndexWriter.optimize(1) on java and the relevant curl on ES.

here are the results:  
java:

ls -ltra

total 329904  
drwxrwxrwt 35 root root 20480 Jun 6 17:00 ..  
-rw-rw-r-- 1 shlomiv shlomiv 24 Jun 6 17:01 \_d.fnm  
-rw-rw-r-- 1 shlomiv shlomiv 36640292 Jun 6 17:01 \_d.fdx  
-rw-rw-r-- 1 shlomiv shlomiv 165896465 Jun 6 17:01 \_d.fdt  
-rw-rw-r-- 1 shlomiv shlomiv 3341766 Jun 6 17:02 \_d.tis  
-rw-rw-r-- 1 shlomiv shlomiv 45880 Jun 6 17:02 \_d.tii  
-rw-rw-r-- 1 shlomiv shlomiv 11335166 Jun 6 17:02 \_d.prx  
-rw-rw-r-- 1 shlomiv shlomiv 115927867 Jun 6 17:02 \_d.frq  
-rw-rw-r-- 1 shlomiv shlomiv 4580040 Jun 6 17:02 \_d.nrm  
drwxrwxr-x 2 shlomiv shlomiv 4096 Jun 6 17:07 .

and ES:

ls -ltr  
total 657340  
-rw-r--r-- 1 elasticsearch elasticsearch 31 Jun 6 16:06 \_4l.fnm  
-rw-r--r-- 1 elasticsearch elasticsearch 36527308 Jun 6 16:06 \_4l.fdx  
-rw-r--r-- 1 elasticsearch elasticsearch 316074621 Jun 6 16:06 \_4l.fdt  
-rw-r--r-- 1 elasticsearch elasticsearch 117176603 Jun 6 16:07 \_4l.tis  
-rw-r--r-- 1 elasticsearch elasticsearch 1120316 Jun 6 16:07 \_4l.tii  
-rw-r--r-- 1 elasticsearch elasticsearch 56962669 Jun 6 16:07 \_4l.prx  
-rw-r--r-- 1 elasticsearch elasticsearch 140660806 Jun 6 16:07 \_4l.frq  
-rw-r--r-- 1 elasticsearch elasticsearch 4565917 Jun 6 16:07 \_4l.nrm  
-rw-r--r-- 1 elasticsearch elasticsearch 313 Jun 6 16:07 segments\_b  
-rw-r--r-- 1 elasticsearch elasticsearch 20 Jun 6 16:07 segments.gen  
-rw-r--r-- 1 elasticsearch elasticsearch 1994 Jun 6 16:07  
\_checksums-1370524025378  
-rw-r--r-- 1 elasticsearch elasticsearch 0 Jun 6 17:00 write.lock

can anyone kindly make a simple test and let this thread know if its just  
that i am weird or can this spectacle be seen elsewhere?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [June 6, 2013, 4:37pm UTC](https://discuss.elastic.co/t/elasticsearch-index-much-larger-then-similar-lucene-index/12056/50 "2013-06-06T16:37:34Z")

</div>

Note, ES sets omit\_norms and omit\_terms\_freq\_and\_positions to false by  
default. When set to true, this saves some space.

Are you really comparing the correct Lucene versions? It could be you  
are mixing 3.5, 3.6, 3.6.1

From the files in the Lucene version, it does not look like you are  
storing many fields in there.

Jörg

Am 06.06.13 16:20, schrieb Shlomi:

> Hi,
> 
> I took some time off of this subject, and now i am back 🙂

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![shlomivaknin](https://avatars.discourse-cdn.com/v4/letter/s/4da419/32.png) [@shlomivaknin](https://discuss.elastic.co/u/shlomivaknin)\
**Post date:** [June 9, 2013, 10:06am UTC](https://discuss.elastic.co/t/elasticsearch-index-much-larger-then-similar-lucene-index/12056/51 "2013-06-09T10:06:00Z")

</div>

hey Jörg,

I checked and made sure both are using 3.6.2 (ES branch 0.20[https://github.com/elasticsearch/elasticsearch/blob/0.20/pom.xml#L33](https://github.com/elasticsearch/elasticsearch/blob/0.20/pom.xml#L33)),  
and same appeared in my pom.xml .

about omit\_norms etc, i made sure explicitly it would be the same in both  
the java code and elastic mapping:  
{  
"test": {  
"\_all": {  
"enabled": "false"  
},  
"properties": {  
"freq": {  
"store": "yes",  
"compress": "true",  
"index\_options": "docs",  
"omit\_norms": "true",  
"type": "long",  
"index": "not\_analyzed"  
},  
"gram": {  
"store": "yes",  
"compress": "true",  
"index\_options": "docs",  
"omit\_norms": "true",  
"type": "string"  
}  
},  
"\_source": {  
"enabled": "false"  
}  
}  
}

and in the java code i have:

```
    Document document = new Document();

```

Field gram = new Field("ngram", ngram, Field.Store.YES,  
Field.Index.ANALYZED);  
gram.setOmitNorms(false);  
gram.setIndexOptions(FieldInfo.IndexOptions.DOCS\_ONLY);

```
    NumericField frequencyField = new NumericField("frequency", 

```

Field.Store.YES, true);  
frequencyField.setOmitNorms(true);  
frequencyField.setIndexOptions(FieldInfo.IndexOptions.DOCS\_ONLY);  
frequencyField.setLongValue(frequency);

```
    document.add(gram);
    document.add(frequencyField);

```

> From the files in the Lucene version, it does not look like you are  
> storing many fields in there.

I only index two fields, gram and freq, if that is what you meant..

This settings still gives me 2x size on elastic. can anyone confirm this on  
his data?

I think i should write a little something that shows this and put it on  
github, for you to checkout, because i feel we are not getting anywhere..

Thanks for all your patience!

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![shlomivaknin](https://avatars.discourse-cdn.com/v4/letter/s/4da419/32.png) [@shlomivaknin](https://discuss.elastic.co/u/shlomivaknin)\
**Post date:** [June 9, 2013, 4:47pm UTC](https://discuss.elastic.co/t/elasticsearch-index-much-larger-then-similar-lucene-index/12056/52 "2013-06-09T16:47:22Z")

</div>

hey,

I made a repo on github [https://github.com/vadali/es-vs-lucene](https://github.com/vadali/es-vs-lucene), it  
contains two folders, one for ES which is much different then the code i  
currently use, for example this one uses plain REST api instead of  
BulkProcessor, and another folder for lucene which is quite similar to the  
code i currently use.

both of them uses lein[https://github.com/technomancy/leiningen#installation](https://github.com/technomancy/leiningen#installation)as their build tool, and there are full instructions how to run each of  
them in the readme[https://github.com/vadali/es-vs-lucene/blob/master/README.md](https://github.com/vadali/es-vs-lucene/blob/master/README.md)  
.

this repo also contains a randomly generated data file, which exhibits the  
same behavior i was reporting.  
After optimizing with max\_segments = 1, i get about 14mb on lucene and 20.5mb  
on elastic. this is only a small dataset, so it doesnt seem like much, but  
this gets meaningful as the dataset gets larger (and i have HUGE datasets..)

let me know if you had any problems to run this test, and if you have any  
ideas regarding why do we see this size difference.

thanks!

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [June 10, 2013, 9:52pm UTC](https://discuss.elastic.co/t/elasticsearch-index-much-larger-then-similar-lucene-index/12056/53 "2013-06-10T21:52:40Z")

</div>

Your Java source code for Lucene shows IndexOptions.DOCS\_ONLY. This is  
equivalent to ES parameter omit\_term\_freq\_and\_positions = true. This is  
missing in the ES mapping. I think there is a confusion about ES  
index\_options "docs" for Lucene 4, I assume ES does not pick it up.  
Maybe it helps to use the following for Lucene 3.6.1 / ES 0.20

{  
"test": {  
"\_all": {  
"enabled": "false"  
},  
"\_source": {  
"enabled": "false"  
},  
"properties": {  
"freq": {  
"type": "long",  
"store": "yes",  
"omit\_norms": "true",  
"omit\_terms\_freq\_and\_positions" : "true"  
},  
"gram": {  
"type": "string",  
"store": "yes",  
"omit\_norms": "true",  
"omit\_terms\_freq\_and\_positions" : "true"  
}  
}

Jörg

Am 09.06.13 12:06, schrieb Shlomi:

> hey Jörg,
> 
> I checked and made sure both are using 3.6.2 (ES branch 0.20  
> [https://github.com/elasticsearch/elasticsearch/blob/0.20/pom.xml#L33](https://github.com/elasticsearch/elasticsearch/blob/0.20/pom.xml#L33)),  
> and same appeared in my pom.xml .
> 
> about omit\_norms etc, i made sure explicitly it would be the same in  
> both the java code and elastic mapping:  
> {  
> "test": {  
> "\_all": {  
> "enabled": "false"  
> },  
> "properties": {  
> "freq": {  
> "store": "yes",  
> "compress": "true",  
> "index\_options": "docs",  
> "omit\_norms": "true",  
> "type": "long",  
> "index": "not\_analyzed"  
> },  
> "gram": {  
> "store": "yes",  
> "compress": "true",  
> "index\_options": "docs",  
> "omit\_norms": "true",  
> "type": "string"  
> }  
> },  
> "\_source": {  
> "enabled": "false"  
> }  
> }  
> }
> 
> and in the java code i have:
> 
> ```
> Document document = new Document();
> 
> ```
> 
> Field gram = new Field("ngram", ngram, Field.Store.YES,  
> Field.Index.ANALYZED);  
> gram.setOmitNorms(false);  
> gram.setIndexOptions(FieldInfo.IndexOptions.DOCS\_ONLY);
> 
> ```
> NumericField frequencyField = new NumericField("frequency", 
> 
> ```
> 
> Field.Store.YES, true);  
> frequencyField.setOmitNorms(true);  
> frequencyField.setIndexOptions(FieldInfo.IndexOptions.DOCS\_ONLY);  
> frequencyField.setLongValue(frequency);
> 
> ```
> document.add(gram);
> document.add(frequencyField);
> 
> From the files in the Lucene version, it does not look like you
> are storing many fields in there.
> 
> ```
> 
> I only index two fields, gram and freq, if that is what you meant..
> 
> This settings still gives me 2x size on elastic. can anyone confirm  
> this on his data?
> 
> I think i should write a little something that shows this and put it  
> on github, for you to checkout, because i feel we are not getting  
> anywhere..
> 
> Thanks for all your patience!
> 
> --  
> You received this message because you are subscribed to the Google  
> Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send  
> an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![shlomivaknin](https://avatars.discourse-cdn.com/v4/letter/s/4da419/32.png) [@shlomivaknin](https://discuss.elastic.co/u/shlomivaknin)\
**Post date:** [June 11, 2013, 12:09pm UTC](https://discuss.elastic.co/t/elasticsearch-index-much-larger-then-similar-lucene-index/12056/54 "2013-06-11T12:09:39Z")

</div>

Hey Yorg,

Thanks for checking this out!

you are right, i got confused by the docs  
about omit\_terms\_freq\_and\_positions, it said it would be deprecated.  
I also found another problem, where i shouldnt have omitted norms on the  
gram field.

I fixed, pushed and ran the test, but the size didnt change..

On Tuesday, June 11, 2013 12:52:40 AM UTC+3, Jörg Prante wrote:

> Your Java source code for Lucene shows IndexOptions.DOCS\_ONLY. This is  
> equivalent to ES parameter omit\_term\_freq\_and\_positions = true. This is  
> missing in the ES mapping. I think there is a confusion about ES  
> index\_options "docs" for Lucene 4, I assume ES does not pick it up.  
> Maybe it helps to use the following for Lucene 3.6.1 / ES 0.20
> 
> {  
> "test": {  
> "\_all": {  
> "enabled": "false"  
> },  
> "\_source": {  
> "enabled": "false"  
> },  
> "properties": {  
> "freq": {  
> "type": "long",  
> "store": "yes",  
> "omit\_norms": "true",  
> "omit\_terms\_freq\_and\_positions" : "true"  
> },  
> "gram": {  
> "type": "string",  
> "store": "yes",  
> "omit\_norms": "true",  
> "omit\_terms\_freq\_and\_positions" : "true"  
> }  
> }
> 
> Jörg
> 
> Am 09.06.13 12:06, schrieb Shlomi:
> 
> > hey Jörg,
> > 
> > I checked and made sure both are using 3.6.2 (ES branch 0.20  
> > [https://github.com/elasticsearch/elasticsearch/blob/0.20/pom.xml#L33](https://github.com/elasticsearch/elasticsearch/blob/0.20/pom.xml#L33)),
> 
> > and same appeared in my pom.xml .
> > 
> > about omit\_norms etc, i made sure explicitly it would be the same in  
> > both the java code and elastic mapping:  
> > {  
> > "test": {  
> > "\_all": {  
> > "enabled": "false"  
> > },  
> > "properties": {  
> > "freq": {  
> > "store": "yes",  
> > "compress": "true",  
> > "index\_options": "docs",  
> > "omit\_norms": "true",  
> > "type": "long",  
> > "index": "not\_analyzed"  
> > },  
> > "gram": {  
> > "store": "yes",  
> > "compress": "true",  
> > "index\_options": "docs",  
> > "omit\_norms": "true",  
> > "type": "string"  
> > }  
> > },  
> > "\_source": {  
> > "enabled": "false"  
> > }  
> > }  
> > }
> > 
> > and in the java code i have:
> > 
> > ```
> > Document document = new Document(); 
> > 
> > ```
> > 
> > Field gram = new Field("ngram", ngram, Field.Store.YES,  
> > Field.Index.ANALYZED);  
> > gram.setOmitNorms(false);  
> > gram.setIndexOptions(FieldInfo.IndexOptions.DOCS\_ONLY);
> > 
> > ```
> > NumericField frequencyField = new NumericField("frequency", 
> > 
> > ```
> > 
> > Field.Store.YES, true);  
> > frequencyField.setOmitNorms(true);  
> > frequencyField.setIndexOptions(FieldInfo.IndexOptions.DOCS\_ONLY);  
> > frequencyField.setLongValue(frequency);
> > 
> > ```
> > document.add(gram); 
> > document.add(frequencyField); 
> > 
> > From the files in the Lucene version, it does not look like you 
> > are storing many fields in there. 
> > 
> > ```
> > 
> > I only index two fields, gram and freq, if that is what you meant..
> > 
> > This settings still gives me 2x size on elastic. can anyone confirm  
> > this on his data?
> > 
> > I think i should write a little something that shows this and put it  
> > on github, for you to checkout, because i feel we are not getting  
> > anywhere..
> > 
> > Thanks for all your patience!
> > 
> > --  
> > You received this message because you are subscribed to the Google  
> > Groups "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send  
> > an email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:31am UTC](https://discuss.elastic.co/t/elasticsearch-index-much-larger-then-similar-lucene-index/12056/55 "2017-07-06T02:31:49Z")

</div>



[Previous page](https://discuss.elastic.co/t/elasticsearch-index-much-larger-then-similar-lucene-index/12056.md?page=2)
