# Better effective substring query idea?

**URL:** <https://discuss.elastic.co/t/better-effective-substring-query-idea/7853>\
**Category:** Elasticsearch\
**Created:** [May 25, 2012, 12:12am UTC](https://discuss.elastic.co/t/better-effective-substring-query-idea/7853 "2012-05-25T00:12:07Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![arta](https://avatars.discourse-cdn.com/v4/letter/a/aca169/32.png) [@arta](https://discuss.elastic.co/u/arta)\
**Post date:** [May 25, 2012, 12:12am UTC](https://discuss.elastic.co/t/better-effective-substring-query-idea/7853/1 "2012-05-25T00:12:07Z")

</div>

Hi,  
I have a string field that is analyzed by NGram analyzer.  
Let's say the name of the field is 'id' and its value is [a-z]+  
I use NGram analyzer because I want to search any substring of ids.  
Here's an example:  
curl -XPUT '[http://localhost:9200/test/](http://localhost:9200/test/)' -d   
'{"index":{"analysis":{  
"analyzer":{"ngram":{"type":"custom","tokenizer":"myngram","filter":["lowercase"]}},  
"tokenizer":{"myngram":{"type":"ngram","min\_gram":1,"max\_gram":3}}}}}'  
curl -XPUT '[http://localhost:9200/test/type1/\_mapping](http://localhost:9200/test/type1/_mapping)' -d   
'{"type1":{"properties":{"id":{"type":"string","index":"analyzed","analyzer":"ngram"}}}}'  
curl -XPUT localhost:9200/test/type1/doc1 -d '{"id":"aaaaaabbbcc"}'  
curl -XPUT localhost:9200/test/type1/doc2 -d '{"id":"aaaabbbbbcc"}'  
curl -XPUT localhost:9200/test/type1/doc3 -d '{"id":"aaaaaaaabcc"}'

I use constant\_score with text query for the search against these docs.  
I use constant\_score because I don't need score and I think it speeds up the search.  
curl -XGET localhost:9200/test/type1/\_search -d   
'{"query":{"constant\_score":{"query":{"text":{"id":{"query":"bbb","operator":"and"}}}}}}'

This search works fine and gets doc1 and doc2.  
However, if I search for a series of the same letters whose length is more than max\_gram,  
I get false hits.  
For example (4 'b's),  
curl -XGET localhost:9200/test/type1/\_search -d   
'{"query":{"constant\_score":{"query":{"text":{"id":{"query":"bbbb","operator":"and"}}}}}}'  
returns doc1 (3 'b's, false hit) and doc2 (4 'b's).  
This is understandable by nature of NGram analyzer.

So, I came up with an idea to apply a filter to filter out false hits.

I changed the mapping so that the id field is a multi\_field and has "id" (NGram'ed) and "id.raw" (not analyzed).  
curl -XDELETE '[http://localhost:9200/test](http://localhost:9200/test)'  
curl -XPUT '[http://localhost:9200/test/](http://localhost:9200/test/)' -d   
'{"index":{"analysis":{  
"analyzer":{"ngram":{"type":"custom", "tokenizer":"myngram","filter":["lowercase"]}},  
"tokenizer":{"myngram":{"type":"ngram","min\_gram":1,"max\_gram":3}}}}}'  
curl -XPUT '[http://localhost:9200/test/type1/\_mapping](http://localhost:9200/test/type1/_mapping)' -d   
'{"type1":{"properties":{  
"id":{"type":"multi\_field","fields":{  
"id":{"type":"string","index":"analyzed","analyzer":"ngram"},  
"raw":{"type":"string","index":"not\_analyzed"}}}}}}'  
curl -XPUT localhost:9200/test/type1/doc1 -d '{"id":"aaaaaabbbcc"}'  
curl -XPUT localhost:9200/test/type1/doc2 -d '{"id":"aaaabbbbbcc"}'  
curl -XPUT localhost:9200/test/type1/doc3 -d '{"id":"aaaaaaaabcc"}'

Then I can get desired hits with specifying wildcard query together:  
curl -XGET localhost:9200/test/type1/\_search -d   
'{"query":{"constant\_score":{"filter":{"and":[  
{"query":{"text":{"id":{"query":"bbbb","operator":"and"}}}},  
{"query":{"wildcard":{"id.raw":"_bbbb_"}}}  
]}}}}'

So far so good. But questions came up:

Question 1)  
I used NGram analyzer for effective substring search.  
However, I ended up with using wildcard filter to get rid of undesired hits.  
Does this approach ruin my original intention, which is 'fast search'?

Question 2)  
The final query seems to be redundant because using solely wildcard filter can do the job.  
Is having the first text query contrubutes the performance?  
In other words, is the wildcard query filter applied only to the results of the first text query?

Question 3)  
Is there better way to get the substring search done?

Thank you for your help.

---

<div class="post-metadata">

**Author:** ![arta](https://avatars.discourse-cdn.com/v4/letter/a/aca169/32.png) [@arta](https://discuss.elastic.co/u/arta)\
**Post date:** [May 28, 2012, 5:49pm UTC](https://discuss.elastic.co/t/better-effective-substring-query-idea/7853/2 "2012-05-28T17:49:37Z")

</div>

I did some experiment.  
I compared search time of following queries.  
A)  
curl -XGET localhost:9200/test/type1/\_search -d \  
'{"query":{"constant\_score":{  
"query":{"wildcard":{"id.raw":"_bbbb_"}}  
}}}'  
B)  
curl -XGET localhost:9200/test/type1/\_search -d \  
'{"query":{"constant\_score":{"filter":{"and":[  
{"query":{"text":{"id":{"query":"bbbb","operator":"and"}}}},  
{"query":{"wildcard":{"id.raw":"_bbbb_"}}}  
]}}}}'  
C)  
curl -XGET localhost:9200/test/type1/\_search -d \  
'{"query":{"filtered":{  
"query":{"constant\_score":{"query":{"text":{"id":{"query":"bbbb","operator":"and"}}}}},

"filter":{"query":{"wildcard":{"id.raw":"_bbbb_"}}}  
}}}'

(Recap: "id" is NGram'ed and "id.raw" is not\_analyzed)

I expected C to be the fastest, but I observed A was the fastest.  
It was probably because the number of documents I have is only a couple of  
thousand or so.

In theory, with a large amount of documents like millions, can I expect C  
to be the fastest?  
Thanks for your help.

---

<div class="post-metadata">

**Author:** ![arta](https://avatars.discourse-cdn.com/v4/letter/a/aca169/32.png) [@arta](https://discuss.elastic.co/u/arta)\
**Post date:** [May 28, 2012, 5:59pm UTC](https://discuss.elastic.co/t/better-effective-substring-query-idea/7853/3 "2012-05-28T17:59:36Z")

</div>

I did some experiment.  
I compared search time of following queries.  
A)  
curl -XGET localhost:9200/test/type1/\_search -d   
'{"query":{"constant\_score":{  
"query":{"wildcard":{"id.raw":"_bbbb_"}}  
}}}'  
B)  
curl -XGET localhost:9200/test/type1/\_search -d   
'{"query":{"constant\_score":{"filter":{"and":[  
{"query":{"text":{"id":{"query":"bbbb","operator":"and"}}}},  
{"query":{"wildcard":{"id.raw":"_bbbb_"}}}  
]}}}}'  
C)  
curl -XGET localhost:9200/test/type1/\_search -d   
'{"query":{"filtered":{  
"query":{"constant\_score":{"query":{"text":{"id":{"query":"bbbb","operator":"and"}}}}},  
"filter":{"query":{"wildcard":{"id.raw":"_bbbb_"}}}  
}}}'

(Recap: "id" is NGram'ed and "id.raw" is not\_analyzed)

I expected C to be the fastest, but I observed A was the fastest.  
It was probably because the number of documents I have is only a couple of thousand or so.

In theory, with a large amount of documents like millions, can I expect C to be the fastest?  
Thanks for your help.

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [May 29, 2012, 8:10am UTC](https://discuss.elastic.co/t/better-effective-substring-query-idea/7853/4 "2012-05-29T08:10:02Z")

</div>

Hi arta

On Mon, 2012-05-28 at 10:59 -0700, arta wrote:

> I did some experiment.  
> I compared search time of following queries.  
> A)  
> curl -XGET localhost:9200/test/type1/\_search -d   
> '{"query":{"constant\_score":{  
> "query":{"wildcard":{"id.raw":"_bbbb_"}}  
> }}}'  
> B)  
> curl -XGET localhost:9200/test/type1/\_search -d   
> '{"query":{"constant\_score":{"filter":{"and":[  
> {"query":{"text":{"id":{"query":"bbbb","operator":"and"}}}},  
> {"query":{"wildcard":{"id.raw":"_bbbb_"}}}  
> ]}}}}'  
> C)  
> curl -XGET localhost:9200/test/type1/\_search -d   
> '{"query":{"filtered":{  
> "query":{"constant\_score":{"query":{"text":{"id":{"query":"bbbb","operator":"and"}}}}},  
> "filter":{"query":{"wildcard":{"id.raw":"_bbbb_"}}}  
> }}}'
> 
> (Recap: "id" is NGram'ed and "id.raw" is not\_analyzed)

If 'id' is ngram'ed, then you don't need to use wildcard queries on it.  
Just a standard text query will suffice.

Have a look at this post:

> <https://stackoverflow.com/questions/9421358/filename-search-with-elasticsearch>

It goes into a fair bit of detail about how to use ngrams

clint

---

<div class="post-metadata">

**Author:** ![arta](https://avatars.discourse-cdn.com/v4/letter/a/aca169/32.png) [@arta](https://discuss.elastic.co/u/arta)\
**Post date:** [May 29, 2012, 5:26pm UTC](https://discuss.elastic.co/t/better-effective-substring-query-idea/7853/5 "2012-05-29T17:26:51Z")

</div>

Thank you clint,  
The post you mentioned is very helpful.

However, the situation I have here is a bit different.  
In your post, you assume "Your users will probably search for words, numbers or dates, but they probably won't expect 'ile' to match 'file'."  
In my case, I need any-substring search capability.

Another reason I came up with applying wildcard filter idea was that I will eventually have a huge number of documents. So I did not want to set max\_gram to be a big number. I believe increasing it increases index size exponentially. I would like to keep max\_gram == 2 or 3 if it gets the job done. The length of 'id' is unlimited (well, practically it will be less than a couple of hundred chars), this is another reason I cannot set big-enough max\_gram.

Thank you for your further help.

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [May 29, 2012, 6:54pm UTC](https://discuss.elastic.co/t/better-effective-substring-query-idea/7853/6 "2012-05-29T18:54:47Z")

</div>

Hi Arta

> However, the situation I have here is a bit different.  
> In your post, you assume "Your users will probably search for words, numbers  
> or dates, but they probably won't expect 'ile' to match 'file'."  
> In my case, I need any-substring search capability.

Then use ngrams instead of edge-ngrams.

> Another reason I came up with applying wildcard filter idea was that I will  
> eventually have a huge number of documents. So I did not want to set  
> max\_gram to be a big number. I believe increasing it increases index size  
> exponentially. I would like to keep max\_gram == 2 or 3 if it gets the job  
> done. The length of 'id' is unlimited (well, practically it will be less  
> than a couple of hundred chars), this is another reason I cannot set  
> big-enough max\_gram.

Wildcards are almost never the answer. They are inefficient and heavy.  
ES has to load all your terms into memory, find which ones match the  
pattern of your wildcard on the fly, then run the query.

Ngrams do this work (or similar) once at index time. They are much more  
efficient and much faster. But yes, they do take up more space. But  
frankly, space is seldom an issue in comparison to memory and CPU. And  
they don't take up as much space as you think. It's not that it is  
storing 'f', 'fo','foo' for every document. It just adds the document  
ID to the document list for each relevant term.

As a worked example, take the two docs 'foobar' and 'boobaz', with a  
min-gram of 2 and a max-gram of 4.

'foobar' gives you:

- fo, oo, ob, ba, ar
- foo, oob, oba, bar
- foob, ooba, obar

'boobaz' gives you:

- bo, oo, ob, ba, az
- boo, oob, oba, baz
- boob, ooba, obaz

Now let's search on 'obaz' - the search term is also broken down into  
ngrams to give us:  
foobar boobaz

- ob X X
- ba X X
- az O X
- oba X X
- baz O X
- obaz O X

'boobaz' is clearly the winner - it is more relevant

Also, if you do a "phrase" search, then it will make sure that the  
matching term has the same ngrams in the SAME ORDER.

So you don't really need a max-gram of 500 characters to achieve what  
you want.

Read the section entitled "UPDATE" on the post I linked to for more on  
the phrase search with ngrams:

> <https://stackoverflow.com/questions/9421358/filename-search-with-elasticsearch>

clint

---

<div class="post-metadata">

**Author:** ![arta](https://avatars.discourse-cdn.com/v4/letter/a/aca169/32.png) [@arta](https://discuss.elastic.co/u/arta)\
**Post date:** [May 29, 2012, 9:24pm UTC](https://discuss.elastic.co/t/better-effective-substring-query-idea/7853/7 "2012-05-29T21:24:44Z")

</div>

Thank you so much clint,  
I will consider having bigger max\_gram and staying away from wildcard.  
I'm convinced that's the better direction.

---

<div class="post-metadata">

**Author:** ![arta](https://avatars.discourse-cdn.com/v4/letter/a/aca169/32.png) [@arta](https://discuss.elastic.co/u/arta)\
**Post date:** [June 1, 2012, 12:50am UTC](https://discuss.elastic.co/t/better-effective-substring-query-idea/7853/8 "2012-06-01T00:50:50Z")

</div>

Now I tried min\_gram=1 max\_gram=100, recreated index, and did following query:  
"text":{"id":{"operator":"and","query":"aaaa \<snip. total 100 a's\> aaa"}}

It got an error saying:  
TooManyClauses[maxClauseCount is set to 1024]

So I googled a bit and added next line in config/elasticsearch.yml  
index.query.bool.max\_clause\_count: 10000

After restarting ES, the query succeeds.  
My interpretation of this experiment is that ES converts the text query into bool query with each analyzed words as its should elements. I think 100 character with 100-gram analyzer creates 5050 segments. (100 + 99 + ... + 1) therefore 5050 should elements.

My question here is, what will be the impact by increasing max\_clause\_count?  
The default is 1024. There must be some reason the default value is that small.

Thanks for your help.

---

<div class="post-metadata">

**Author:** ![arta](https://avatars.discourse-cdn.com/v4/letter/a/aca169/32.png) [@arta](https://discuss.elastic.co/u/arta)\
**Post date:** [June 1, 2012, 2:01am UTC](https://discuss.elastic.co/t/better-effective-substring-query-idea/7853/9 "2012-06-01T02:01:13Z")

</div>

Oh, maybe I'm doing wong thing.  
As Clint mentioned in his reply to another question here:  
[http://elasticsearch-users.115913.n3.nabble.com/help-needed-with-the-query-tt3177477.html#a3178856](http://elasticsearch-users.115913.n3.nabble.com/help-needed-with-the-query-tt3177477.html#a3178856)  
I should have needed to apply different analyzer for search.  
So the max clause count should not affect.  
Is this the right way to go?

Thanks for your help.

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [June 1, 2012, 12:21pm UTC](https://discuss.elastic.co/t/better-effective-substring-query-idea/7853/10 "2012-06-01T12:21:36Z")

</div>

Hi arta

On Thu, 2012-05-31 at 19:01 -0700, arta wrote:

> Oh, maybe I'm doing wong thing.  
> As Clint mentioned in his reply to another question here:  
> [http://elasticsearch-users.115913.n3.nabble.com/help-needed-with-the-query-tt3177477.html#a3178856](http://elasticsearch-users.115913.n3.nabble.com/help-needed-with-the-query-tt3177477.html#a3178856)  
> I should have needed to apply different analyzer for search.  
> So the max clause count should not affect.  
> Is this the right way to go?

Have a read of this:

> <https://stackoverflow.com/questions/9421358/filename-search-with-elasticsearch>

especially the part starting with UPDATE

Also, you don't need a max gram of 100 - have a look at this post:  
[https://groups.google.com/d/msg/elasticsearch/O8iF7vUf3mA/EOqyapWA5VAJ](https://groups.google.com/d/msg/elasticsearch/O8iF7vUf3mA/EOqyapWA5VAJ)

clint

---

<div class="post-metadata">

**Author:** ![arta](https://avatars.discourse-cdn.com/v4/letter/a/aca169/32.png) [@arta](https://discuss.elastic.co/u/arta)\
**Post date:** [June 1, 2012, 5:33pm UTC](https://discuss.elastic.co/t/better-effective-substring-query-idea/7853/11 "2012-06-01T17:33:39Z")

</div>

Thank you very much for the quick response, Clinton!

The reason why I think my max\_gram needs to be a big enough number, is following use case:  
Let's say min\_gram=2 max\_gram=4.  
document #1 has id='foooobar'  
document #2 has id='fooooobar'  
Then here's a search for 'ooooo'  
If the id field uses the same NGram analyzer both for index and search, the search hits both of the two.  
If the id field uses, say, standard analyzer for search, then the search hits none of the two.

Is there a better way (instead of setting big max\_gram) to make this search successful?  
Thanks, again, for your help. I really appreciate your help!

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [June 1, 2012, 5:39pm UTC](https://discuss.elastic.co/t/better-effective-substring-query-idea/7853/12 "2012-06-01T17:39:39Z")

</div>

On Fri, 2012-06-01 at 10:33 -0700, arta wrote:

> Thank you very much for the quick response, Clinton!
> 
> The reason why I think my max\_gram needs to be a big enough number, is  
> following use case:  
> Let's say min\_gram=2 max\_gram=4.  
> document #1 has id='foooobar'  
> document #2 has id='fooooobar'  
> Then here's a search for 'ooooo'  
> If the id field uses the same NGram analyzer both for index and search, the  
> search hits both of the two.  
> If the id field uses, say, standard analyzer for search, then the search  
> hits none of the two.

Sure, that's a problem. What is the likelihood that you're going to  
have such long repeating strings? And how big a problem is it? Depends  
on your data really

clint

---

<div class="post-metadata">

**Author:** ![arta](https://avatars.discourse-cdn.com/v4/letter/a/aca169/32.png) [@arta](https://discuss.elastic.co/u/arta)\
**Post date:** [June 1, 2012, 6:19pm UTC](https://discuss.elastic.co/t/better-effective-substring-query-idea/7853/13 "2012-06-01T18:19:34Z")

</div>

Let me think more deeply about our use cases.  
I think I can come up with a reasonable max\_gram number.  
Thank you again for your help. It really really saved me.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:25am UTC](https://discuss.elastic.co/t/better-effective-substring-query-idea/7853/14 "2017-07-06T03:25:59Z")

</div>


