# Filename search using nGram-tokenizer and span\_near-query

**URL:** <https://discuss.elastic.co/t/filename-search-using-ngram-tokenizer-and-span-near-query/8859>\
**Category:** Elasticsearch\
**Created:** [August 27, 2012, 12:29pm UTC](https://discuss.elastic.co/t/filename-search-using-ngram-tokenizer-and-span-near-query/8859 "2012-08-27T12:29:29Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![da\_mkay](https://avatars.discourse-cdn.com/v4/letter/d/8c91f0/32.png) [@da\_mkay](https://discuss.elastic.co/u/da_mkay)\
**Post date:** [August 27, 2012, 12:29pm UTC](https://discuss.elastic.co/t/filename-search-using-ngram-tokenizer-and-span-near-query/8859/1 "2012-08-27T12:29:29Z")

</div>

Hi,

I try to build a simple filename search. I want that the user can search  
for any part of the name.

Let's say the following filenames are indexed:  
[1] My\_file\_2012.01.12.txt  
[2] My\_file\_2012.01.05.txt  
[3] My\_file\_2012.05.01.txt  
[4] My\_file\_2012.08.27.txt  
[5] My\_file\_2012.12.12.txt  
[6] My\_file\_2011.12.12.txt  
[7] file\_01\_2012.09.09.txt

Then the user might search for:  
"ile\_20" (finds the first six documents)  
"12.txt" (finds 1, 5, 6)  
"12" followed by "01" (finds 1, 2, 3 - NOT 7)  
"2012" followed by "01" (finds 1, 2, 3 - NOT 7)

(Note: Yes, the user might really search for strings like "ile\_20" ... e.g.  
because of copy-and-paste mistakes 🙂 )

Therefore I use a nGram-tokenizer to index each possible input-string. This  
works fine so far.  
To support the "followed by"-search mentioned above I need a query that  
respects the order of the terms, no matter how many text is between these  
two terms (okay let's say max. 100 characters 🙂 ).

Since a "text\_phrase"-query with a "slop" does not respect the ordering of  
the terms correctly, I decided to use a "span\_near" query. This works fine  
in most cases.

See here my full example-index: [https://gist.github.com/3487909](https://gist.github.com/3487909)

As mentioned in the example above the query "'2012' followed by '01'" does  
not work since the nGram tokenizer generates a position-value for each  
token that is not very useful when used by the "span\_near" query. While  
indexing, the term "2012" is assigned to a position value (50) which is  
bigger than the position value for the term "01" (e.g. 10). Since 50 and 10  
are not in order the query will have no results. The in-order-thing works  
only correct for terms which have the same length (e.g. "'12' followed by  
'01'") or if the terms are ordered by length (e.g. "'20' followed by  
'.12'").

So how can I achieve the correct search-behaviour? I just want the ability  
to search for any part(s) of the filename while respecting the order of the  
terms. 🙂  
Maybe there is a way to tell "span\_near" to not use the position but  
instead the "start\_offset"?  
Or is there another query I can use?

Best regards,  
da-mkay

--

---

<div class="post-metadata">

**Author:** ![da\_mkay](https://avatars.discourse-cdn.com/v4/letter/d/8c91f0/32.png) [@da\_mkay](https://discuss.elastic.co/u/da_mkay)\
**Post date:** [September 3, 2012, 2:34pm UTC](https://discuss.elastic.co/t/filename-search-using-ngram-tokenizer-and-span-near-query/8859/2 "2012-09-03T14:34:00Z")

</div>

Does nobody have an idea? 🙂

da-mkay:

> Hi,
> 
> I try to build a simple filename search. I want that the user can search  
> for any part of the name.
> 
> Let's say the following filenames are indexed:  
> [1] My\_file\_2012.01.12.txt  
> [2] My\_file\_2012.01.05.txt  
> [3] My\_file\_2012.05.01.txt  
> [4] My\_file\_2012.08.27.txt  
> [5] My\_file\_2012.12.12.txt  
> [6] My\_file\_2011.12.12.txt  
> [7] file\_01\_2012.09.09.txt
> 
> Then the user might search for:  
> "ile\_20" (finds the first six documents)  
> "12.txt" (finds 1, 5, 6)  
> "12" followed by "01" (finds 1, 2, 3 - NOT 7)  
> "2012" followed by "01" (finds 1, 2, 3 - NOT 7)
> 
> (Note: Yes, the user might really search for strings like "ile\_20" ... e.g.  
> because of copy-and-paste mistakes 🙂 )
> 
> Therefore I use a nGram-tokenizer to index each possible input-string. This  
> works fine so far.  
> To support the "followed by"-search mentioned above I need a query that  
> respects the order of the terms, no matter how many text is between these  
> two terms (okay let's say max. 100 characters 🙂 ).
> 
> Since a "text\_phrase"-query with a "slop" does not respect the ordering of  
> the terms correctly, I decided to use a "span\_near" query. This works fine  
> in most cases.
> 
> See here my full example-index: [ElasticSearch - filename search using nGram · GitHub](https://gist.github.com/3487909)
> 
> As mentioned in the example above the query "'2012' followed by '01'" does  
> not work since the nGram tokenizer generates a position-value for each  
> token that is not very useful when used by the "span\_near" query. While  
> indexing, the term "2012" is assigned to a position value (50) which is  
> bigger than the position value for the term "01" (e.g. 10). Since 50 and 10  
> are not in order the query will have no results. The in-order-thing works  
> only correct for terms which have the same length (e.g. "'12' followed by  
> '01'") or if the terms are ordered by length (e.g. "'20' followed by  
> '.12'").
> 
> So how can I achieve the correct search-behaviour? I just want the ability  
> to search for any part(s) of the filename while respecting the order of the  
> terms. 🙂  
> Maybe there is a way to tell "span\_near" to not use the position but  
> instead the "start\_offset"?  
> Or is there another query I can use?
> 
> Best regards,  
> da-mkay

--

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [September 4, 2012, 10:20am UTC](https://discuss.elastic.co/t/filename-search-using-ngram-tokenizer-and-span-near-query/8859/3 "2012-09-04T10:20:41Z")

</div>

Have a look at this explanation:

> <https://stackoverflow.com/questions/9421358/filename-search-with-elasticsearch/9432450#9432450>

clint

On Mon, Sep 3, 2012 at 4:34 PM, Manuel [manuel@wenns-um-email-geht.de](mailto:manuel@wenns-um-email-geht.de)wrote:

> Does nobody have an idea? 🙂
> 
> da-mkay:
> 
> > Hi,
> > 
> > I try to build a simple filename search. I want that the user can search  
> > for any part of the name.
> > 
> > Let's say the following filenames are indexed:  
> > [1] My\_file\_2012.01.12.txt  
> > [2] My\_file\_2012.01.05.txt  
> > [3] My\_file\_2012.05.01.txt  
> > [4] My\_file\_2012.08.27.txt  
> > [5] My\_file\_2012.12.12.txt  
> > [6] My\_file\_2011.12.12.txt  
> > [7] file\_01\_2012.09.09.txt
> > 
> > Then the user might search for:  
> > "ile\_20" (finds the first six documents)  
> > "12.txt" (finds 1, 5, 6)  
> > "12" followed by "01" (finds 1, 2, 3 - NOT 7)  
> > "2012" followed by "01" (finds 1, 2, 3 - NOT 7)
> > 
> > (Note: Yes, the user might really search for strings like "ile\_20" ...  
> > e.g.  
> > because of copy-and-paste mistakes 🙂 )
> > 
> > Therefore I use a nGram-tokenizer to index each possible input-string.  
> > This  
> > works fine so far.  
> > To support the "followed by"-search mentioned above I need a query that  
> > respects the order of the terms, no matter how many text is between these  
> > two terms (okay let's say max. 100 characters 🙂 ).
> > 
> > Since a "text\_phrase"-query with a "slop" does not respect the ordering  
> > of  
> > the terms correctly, I decided to use a "span\_near" query. This works  
> > fine  
> > in most cases.
> > 
> > See here my full example-index: [ElasticSearch - filename search using nGram · GitHub](https://gist.github.com/3487909)
> > 
> > As mentioned in the example above the query "'2012' followed by '01'"  
> > does  
> > not work since the nGram tokenizer generates a position-value for each  
> > token that is not very useful when used by the "span\_near" query. While  
> > indexing, the term "2012" is assigned to a position value (50) which is  
> > bigger than the position value for the term "01" (e.g. 10). Since 50 and  
> > 10  
> > are not in order the query will have no results. The in-order-thing works  
> > only correct for terms which have the same length (e.g. "'12' followed by  
> > '01'") or if the terms are ordered by length (e.g. "'20' followed by  
> > '.12'").
> > 
> > So how can I achieve the correct search-behaviour? I just want the  
> > ability  
> > to search for any part(s) of the filename while respecting the order of  
> > the  
> > terms. 🙂  
> > Maybe there is a way to tell "span\_near" to not use the position but  
> > instead the "start\_offset"?  
> > Or is there another query I can use?
> > 
> > Best regards,  
> > da-mkay
> 
> --

--

---

<div class="post-metadata">

**Author:** ![da\_mkay](https://avatars.discourse-cdn.com/v4/letter/d/8c91f0/32.png) [@da\_mkay](https://discuss.elastic.co/u/da_mkay)\
**Post date:** [September 4, 2012, 11:52am UTC](https://discuss.elastic.co/t/filename-search-using-ngram-tokenizer-and-span-near-query/8859/4 "2012-09-04T11:52:25Z")

</div>

Thanks, but that is my question at SO which I posted a few month ago 🙂  
In the answer a text\_phrase-query is used which I cannot use. This is  
why I tried using a span-query. See my text below. 😉

Best regards,  
da-mkay

Am 04.09.2012 12:20, schrieb Clinton Gormley:

> Have a look at this explanation:  
> [lucene - Filename search with ElasticSearch - Stack Overflow](http://stackoverflow.com/questions/9421358/filename-search-with-elasticsearch/9432450#9432450)
> 
> clint
> 
> On Mon, Sep 3, 2012 at 4:34 PM, Manuel [manuel@wenns-um-email-geht.de](mailto:manuel@wenns-um-email-geht.de)wrote:
> 
> > Does nobody have an idea? 🙂
> > 
> > da-mkay:
> > 
> > > Hi,
> > > 
> > > I try to build a simple filename search. I want that the user can search  
> > > for any part of the name.
> > > 
> > > Let's say the following filenames are indexed:  
> > > [1] My\_file\_2012.01.12.txt  
> > > [2] My\_file\_2012.01.05.txt  
> > > [3] My\_file\_2012.05.01.txt  
> > > [4] My\_file\_2012.08.27.txt  
> > > [5] My\_file\_2012.12.12.txt  
> > > [6] My\_file\_2011.12.12.txt  
> > > [7] file\_01\_2012.09.09.txt
> > > 
> > > Then the user might search for:  
> > > "ile\_20" (finds the first six documents)  
> > > "12.txt" (finds 1, 5, 6)  
> > > "12" followed by "01" (finds 1, 2, 3 - NOT 7)  
> > > "2012" followed by "01" (finds 1, 2, 3 - NOT 7)
> > > 
> > > (Note: Yes, the user might really search for strings like "ile\_20" ...  
> > > e.g.  
> > > because of copy-and-paste mistakes 🙂 )
> > > 
> > > Therefore I use a nGram-tokenizer to index each possible input-string.  
> > > This  
> > > works fine so far.  
> > > To support the "followed by"-search mentioned above I need a query that  
> > > respects the order of the terms, no matter how many text is between these  
> > > two terms (okay let's say max. 100 characters 🙂 ).
> > > 
> > > Since a "text\_phrase"-query with a "slop" does not respect the ordering  
> > > of  
> > > the terms correctly, I decided to use a "span\_near" query. This works  
> > > fine  
> > > in most cases.
> > > 
> > > See here my full example-index: [ElasticSearch - filename search using nGram · GitHub](https://gist.github.com/3487909)
> > > 
> > > As mentioned in the example above the query "'2012' followed by '01'"  
> > > does  
> > > not work since the nGram tokenizer generates a position-value for each  
> > > token that is not very useful when used by the "span\_near" query. While  
> > > indexing, the term "2012" is assigned to a position value (50) which is  
> > > bigger than the position value for the term "01" (e.g. 10). Since 50 and  
> > > 10  
> > > are not in order the query will have no results. The in-order-thing works  
> > > only correct for terms which have the same length (e.g. "'12' followed by  
> > > '01'") or if the terms are ordered by length (e.g. "'20' followed by  
> > > '.12'").
> > > 
> > > So how can I achieve the correct search-behaviour? I just want the  
> > > ability  
> > > to search for any part(s) of the filename while respecting the order of  
> > > the  
> > > terms. 🙂  
> > > Maybe there is a way to tell "span\_near" to not use the position but  
> > > instead the "start\_offset"?  
> > > Or is there another query I can use?
> > > 
> > > Best regards,  
> > > da-mkay
> > 
> > --

--

---

<div class="post-metadata">

**Author:** ![da\_mkay](https://avatars.discourse-cdn.com/v4/letter/d/8c91f0/32.png) [@da\_mkay](https://discuss.elastic.co/u/da_mkay)\
**Post date:** [September 5, 2012, 1:55pm UTC](https://discuss.elastic.co/t/filename-search-using-ngram-tokenizer-and-span-near-query/8859/5 "2012-09-05T13:55:27Z")

</div>

I found a solution. Due to the NGram-tokenizer each possible keyword is  
indexed as a separate token. So I can use a "query\_string"-query with a  
wildcard on the filename-field: e.g. "2012\*01".  
This takes the order of the terms into account. But I think that those  
queries can get a bit slow with heavy wildcard-usage, right? Especially  
with many requests in parallel.

Best regards,  
da-mkay

Am 04.09.2012 13:52, schrieb Manuel:

> Thanks, but that is my question at SO which I posted a few month ago 🙂  
> In the answer a text\_phrase-query is used which I cannot use. This is  
> why I tried using a span-query. See my text below. 😉
> 
> Best regards,  
> da-mkay
> 
> Am 04.09.2012 12:20, schrieb Clinton Gormley:
> 
> > Have a look at this explanation:  
> > [lucene - Filename search with ElasticSearch - Stack Overflow](http://stackoverflow.com/questions/9421358/filename-search-with-elasticsearch/9432450#9432450)
> > 
> > clint
> > 
> > On Mon, Sep 3, 2012 at 4:34 PM, Manuel [manuel@wenns-um-email-geht.de](mailto:manuel@wenns-um-email-geht.de)wrote:
> > 
> > > Does nobody have an idea? 🙂
> > > 
> > > da-mkay:
> > > 
> > > > Hi,
> > > > 
> > > > I try to build a simple filename search. I want that the user can search  
> > > > for any part of the name.
> > > > 
> > > > Let's say the following filenames are indexed:  
> > > > [1] My\_file\_2012.01.12.txt  
> > > > [2] My\_file\_2012.01.05.txt  
> > > > [3] My\_file\_2012.05.01.txt  
> > > > [4] My\_file\_2012.08.27.txt  
> > > > [5] My\_file\_2012.12.12.txt  
> > > > [6] My\_file\_2011.12.12.txt  
> > > > [7] file\_01\_2012.09.09.txt
> > > > 
> > > > Then the user might search for:  
> > > > "ile\_20" (finds the first six documents)  
> > > > "12.txt" (finds 1, 5, 6)  
> > > > "12" followed by "01" (finds 1, 2, 3 - NOT 7)  
> > > > "2012" followed by "01" (finds 1, 2, 3 - NOT 7)
> > > > 
> > > > (Note: Yes, the user might really search for strings like "ile\_20" ...  
> > > > e.g.  
> > > > because of copy-and-paste mistakes 🙂 )
> > > > 
> > > > Therefore I use a nGram-tokenizer to index each possible input-string.  
> > > > This  
> > > > works fine so far.  
> > > > To support the "followed by"-search mentioned above I need a query that  
> > > > respects the order of the terms, no matter how many text is between these  
> > > > two terms (okay let's say max. 100 characters 🙂 ).
> > > > 
> > > > Since a "text\_phrase"-query with a "slop" does not respect the ordering  
> > > > of  
> > > > the terms correctly, I decided to use a "span\_near" query. This works  
> > > > fine  
> > > > in most cases.
> > > > 
> > > > See here my full example-index: [ElasticSearch - filename search using nGram · GitHub](https://gist.github.com/3487909)
> > > > 
> > > > As mentioned in the example above the query "'2012' followed by '01'"  
> > > > does  
> > > > not work since the nGram tokenizer generates a position-value for each  
> > > > token that is not very useful when used by the "span\_near" query. While  
> > > > indexing, the term "2012" is assigned to a position value (50) which is  
> > > > bigger than the position value for the term "01" (e.g. 10). Since 50 and  
> > > > 10  
> > > > are not in order the query will have no results. The in-order-thing works  
> > > > only correct for terms which have the same length (e.g. "'12' followed by  
> > > > '01'") or if the terms are ordered by length (e.g. "'20' followed by  
> > > > '.12'").
> > > > 
> > > > So how can I achieve the correct search-behaviour? I just want the  
> > > > ability  
> > > > to search for any part(s) of the filename while respecting the order of  
> > > > the  
> > > > terms. 🙂  
> > > > Maybe there is a way to tell "span\_near" to not use the position but  
> > > > instead the "start\_offset"?  
> > > > Or is there another query I can use?
> > > > 
> > > > Best regards,  
> > > > da-mkay
> > > 
> > > --

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:14am UTC](https://discuss.elastic.co/t/filename-search-using-ngram-tokenizer-and-span-near-query/8859/6 "2017-07-06T03:14:02Z")

</div>


