# NGram Index implications

**URL:** <https://discuss.elastic.co/t/ngram-index-implications/49289>\
**Category:** Elasticsearch\
**Created:** [May 5, 2016, 12:38pm UTC](https://discuss.elastic.co/t/ngram-index-implications/49289 "2016-05-05T12:38:09Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![avibh](https://avatars.discourse-cdn.com/v4/letter/a/9fc29f/32.png) [@avibh](https://discuss.elastic.co/u/avibh)\
**Post date:** [May 5, 2016, 12:38pm UTC](https://discuss.elastic.co/t/ngram-index-implications/49289/1 "2016-05-05T12:38:09Z")

</div>

Hi,

I'd like to use NGram for my indexes using the autocomplete analyzer.  
my only concern is the implications on index ram/disk size.  
is there a rule of thumb of how much the index consumes in case of -

1. "min\_gram": 1,  
"max\_gram": 5

2. "min\_gram": 1,  
"max\_gram": 10

3. "min\_gram": 1,  
"max\_gram": 15

4. "min\_gram": 1,  
"max\_gram": 20

is it a multiplication or something else?

---

<div class="post-metadata">

**Author:** ![Bruce\_Ritchie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bruce_ritchie/32/9370_2.png) [@Bruce\_Ritchie](https://discuss.elastic.co/u/Bruce_Ritchie)\
**Post date:** [May 5, 2016, 1:35pm UTC](https://discuss.elastic.co/t/ngram-index-implications/49289/2 "2016-05-05T13:35:08Z")

</div>

I'd suggest the edge ngram as it would produce less tokens and is generally good enough for autocomplete. As for size with plain ngram it's not double going from 5 to 10 grams as most words are not 10 characters long. The min\_gram actually has a greater impact in the # of tokens than the max typically.

This is a test.

T  
Th  
Thi  
This  
h  
hi  
his  
i  
is  
s  
...

As you can see the # of tokens explodes with a min\_gram of 1. Using the edgengram instead gives you

T  
Th  
Thi  
This

The downside is you won't match inside of words. Typically that's not a huge deal for autocomplete. There is a site plugin (inquisitor) that let's you play with tokenizers and see the results of tokenization.

---

<div class="post-metadata">

**Author:** ![avibh](https://avatars.discourse-cdn.com/v4/letter/a/9fc29f/32.png) [@avibh](https://discuss.elastic.co/u/avibh)\
**Post date:** [May 5, 2016, 2:04pm UTC](https://discuss.elastic.co/t/ngram-index-implications/49289/3 "2016-05-05T14:04:37Z")

</div>

> [@Bruce\_Ritchie](#):
>
> edgengram

thanks for the quick response Brice and the heads up with the "edgengram", i'll check it out.  
now same question with "edgengram", regarding index size?

---

<div class="post-metadata">

**Author:** ![Bruce\_Ritchie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bruce_ritchie/32/9370_2.png) [@Bruce\_Ritchie](https://discuss.elastic.co/u/Bruce_Ritchie)\
**Post date:** [May 5, 2016, 2:31pm UTC](https://discuss.elastic.co/t/ngram-index-implications/49289/4 "2016-05-05T14:31:03Z")

</div>

It's not nearly as bad because the # of tokens doesn't increase anywhere as much. With min=1, max=10 you could see a max of 10x increase in tokens for each field indexed that way but that may not correspond to 10x increase in index size. That is because at least for English many words are \< 10 characters long and it's likely you wouldn't use this tokenizer for every field in your index. I can't give exact #'s without testing with your data.

---

<div class="post-metadata">

**Author:** ![avibh](https://avatars.discourse-cdn.com/v4/letter/a/9fc29f/32.png) [@avibh](https://discuss.elastic.co/u/avibh)\
**Post date:** [May 5, 2016, 4:13pm UTC](https://discuss.elastic.co/t/ngram-index-implications/49289/5 "2016-05-05T16:13:16Z")

</div>

thanks.  
I guess I'll have to run some benchmarks and see for myself 🙂

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:53pm UTC](https://discuss.elastic.co/t/ngram-index-implications/49289/6 "2017-07-05T22:53:40Z")

</div>


