# Inverse edge back-Ngram (or making it "fuzzy" at the end of a word)?

**URL:** <https://discuss.elastic.co/t/inverse-edge-back-ngram-or-making-it-fuzzy-at-the-end-of-a-word/10903>\
**Category:** Elasticsearch\
**Created:** [February 26, 2013, 10:45am UTC](https://discuss.elastic.co/t/inverse-edge-back-ngram-or-making-it-fuzzy-at-the-end-of-a-word/10903 "2013-02-26T10:45:42Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Per\_Ekman](https://avatars.discourse-cdn.com/v4/letter/p/e9bcb4/32.png) [@Per\_Ekman](https://discuss.elastic.co/u/Per_Ekman)\
**Post date:** [February 26, 2013, 10:45am UTC](https://discuss.elastic.co/t/inverse-edge-back-ngram-or-making-it-fuzzy-at-the-end-of-a-word/10903/1 "2013-02-26T10:45:42Z")

</div>

Hi

We are discussing building an index where possible misspellings at the end  
of a word are getting hits.

We were looking at using the EdgeNGram and making ngrams of the last two  
characters, but that gives us an index of just the 2-character variations  
of the word endings.

How would we best do this? Is it possible to configure the inverse of that?  
Should we tokenize it with a regexp? Any other ideas?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [February 26, 2013, 11:02am UTC](https://discuss.elastic.co/t/inverse-edge-back-ngram-or-making-it-fuzzy-at-the-end-of-a-word/10903/2 "2013-02-26T11:02:19Z")

</div>

On Tue, 2013-02-26 at 02:45 -0800, Per Ekman wrote:

> Hi
> 
> We are discussing building an index where possible misspellings at the  
> end of a word are getting hits.
> 
> We were looking at using the EdgeNGram and making ngrams of the last  
> two characters, but that gives us an index of just the 2-character  
> variations of the word endings.
> 
> How would we best do this? Is it possible to configure the inverse of  
> that? Should we tokenize it with a regexp? Any other ideas?

curl -XPUT '[http://127.0.0.1:9200/test/?pretty=1](http://127.0.0.1:9200/test/?pretty=1)' -d '  
{  
"settings" : {  
"analysis" : {  
"filter" : {  
"end\_grams" : {  
"max\_gram" : 2,  
"side" : "back",  
"min\_gram" : 2,  
"type" : "edge\_ngram"  
}  
},  
"analyzer" : {  
"end\_grams" : {  
"filter" : [  
"standard",  
"lowercase",  
"stop",  
"end\_grams"  
],  
"tokenizer" : "standard"  
}  
}  
}  
}  
}  
'

curl -XGET '[http://127.0.0.1:9200/test/\_analyze?pretty=1&text=The+quick](http://127.0.0.1:9200/test/_analyze?pretty=1&text=The+quick)  
+brown+fox+jumped+over+the+lazy+dog&analyzer=end\_grams'

# {

# "tokens" : [

# {

# "end\_offset" : 9,

# "position" : 1,

# "start\_offset" : 7,

# "type" : "word",

# "token" : "ck"

# },

# {

# "end\_offset" : 15,

# "position" : 2,

# "start\_offset" : 13,

# "type" : "word",

# "token" : "wn"

# },

# {

# "end\_offset" : 19,

# "position" : 3,

# "start\_offset" : 17,

# "type" : "word",

# "token" : "ox"

# },

# {

# "end\_offset" : 26,

# "position" : 4,

# "start\_offset" : 24,

# "type" : "word",

# "token" : "ed"

# },

# {

# "end\_offset" : 31,

# "position" : 5,

# "start\_offset" : 29,

# "type" : "word",

# "token" : "er"

# },

# {

# "end\_offset" : 40,

# "position" : 6,

# "start\_offset" : 38,

# "type" : "word",

# "token" : "zy"

# },

# {

# "end\_offset" : 44,

# "position" : 7,

# "start\_offset" : 42,

# "type" : "word",

# "token" : "og"

# }

# ]

# }

clint

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Per\_Ekman](https://avatars.discourse-cdn.com/v4/letter/p/e9bcb4/32.png) [@Per\_Ekman](https://discuss.elastic.co/u/Per_Ekman)\
**Post date:** [February 26, 2013, 11:09am UTC](https://discuss.elastic.co/t/inverse-edge-back-ngram-or-making-it-fuzzy-at-the-end-of-a-word/10903/3 "2013-02-26T11:09:19Z")

</div>

Alright, that is pretty much what we've done so far, but I'm looking at  
getting "bro", "f", "jump"..... into the index, instead of the endings, And  
possibly the original words as well.

On Tue, Feb 26, 2013 at 12:02 PM, Clinton Gormley [clint@traveljury.com](mailto:clint@traveljury.com)wrote:

> On Tue, 2013-02-26 at 02:45 -0800, Per Ekman wrote:
> 
> > Hi
> > 
> > We are discussing building an index where possible misspellings at the  
> > end of a word are getting hits.
> > 
> > We were looking at using the EdgeNGram and making ngrams of the last  
> > two characters, but that gives us an index of just the 2-character  
> > variations of the word endings.
> > 
> > How would we best do this? Is it possible to configure the inverse of  
> > that? Should we tokenize it with a regexp? Any other ideas?
> 
> curl -XPUT '[http://127.0.0.1:9200/test/?pretty=1](http://127.0.0.1:9200/test/?pretty=1)' -d '  
> {  
> "settings" : {  
> "analysis" : {  
> "filter" : {  
> "end\_grams" : {  
> "max\_gram" : 2,  
> "side" : "back",  
> "min\_gram" : 2,  
> "type" : "edge\_ngram"  
> }  
> },  
> "analyzer" : {  
> "end\_grams" : {  
> "filter" : [  
> "standard",  
> "lowercase",  
> "stop",  
> "end\_grams"  
> ],  
> "tokenizer" : "standard"  
> }  
> }  
> }  
> }  
> }  
> '
> 
> curl -XGET '[http://127.0.0.1:9200/test/\_analyze?pretty=1&text=The+quick](http://127.0.0.1:9200/test/_analyze?pretty=1&text=The+quick)  
> +brown+fox+jumped+over+the+lazy+dog&analyzer=end\_grams'
> 
> # {
> 
> # "tokens" : [
> 
> # {
> 
> # "end\_offset" : 9,
> 
> # "position" : 1,
> 
> # "start\_offset" : 7,
> 
> # "type" : "word",
> 
> # "token" : "ck"
> 
> # },
> 
> # {
> 
> # "end\_offset" : 15,
> 
> # "position" : 2,
> 
> # "start\_offset" : 13,
> 
> # "type" : "word",
> 
> # "token" : "wn"
> 
> # },
> 
> # {
> 
> # "end\_offset" : 19,
> 
> # "position" : 3,
> 
> # "start\_offset" : 17,
> 
> # "type" : "word",
> 
> # "token" : "ox"
> 
> # },
> 
> # {
> 
> # "end\_offset" : 26,
> 
> # "position" : 4,
> 
> # "start\_offset" : 24,
> 
> # "type" : "word",
> 
> # "token" : "ed"
> 
> # },
> 
> # {
> 
> # "end\_offset" : 31,
> 
> # "position" : 5,
> 
> # "start\_offset" : 29,
> 
> # "type" : "word",
> 
> # "token" : "er"
> 
> # },
> 
> # {
> 
> # "end\_offset" : 40,
> 
> # "position" : 6,
> 
> # "start\_offset" : 38,
> 
> # "type" : "word",
> 
> # "token" : "zy"
> 
> # },
> 
> # {
> 
> # "end\_offset" : 44,
> 
> # "position" : 7,
> 
> # "start\_offset" : 42,
> 
> # "type" : "word",
> 
> # "token" : "og"
> 
> # }
> 
> # ]
> 
> # }
> 
> clint
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Per\_Ekman](https://avatars.discourse-cdn.com/v4/letter/p/e9bcb4/32.png) [@Per\_Ekman](https://discuss.elastic.co/u/Per_Ekman)\
**Post date:** [February 26, 2013, 11:13am UTC](https://discuss.elastic.co/t/inverse-edge-back-ngram-or-making-it-fuzzy-at-the-end-of-a-word/10903/4 "2013-02-26T11:13:57Z")

</div>

I guess I was really unclear in my original text. I want to know how to  
strip the last couple of characters in a word, and also keep the original

On Tuesday, February 26, 2013 12:09:19 PM UTC+1, Per Ekman wrote:

> Alright, that is pretty much what we've done so far, but I'm looking at  
> getting "bro", "f", "jump"..... into the index, instead of the endings, And  
> possibly the original words as well.
> 
> On Tue, 2013-02-26 at 02:45 -0800, Per Ekman wrote:
> 
> > > Hi
> > > 
> > > We are discussing building an index where possible misspellings at the  
> > > end of a word are getting hits.
> > > 
> > > We were looking at using the EdgeNGram and making ngrams of the last  
> > > two characters, but that gives us an index of just the 2-character  
> > > variations of the word endings.
> > > 
> > > How would we best do this? Is it possible to configure the inverse of  
> > > that? Should we tokenize it with a regexp? Any other ideas?
> > 
> > curl -XPUT '[http://127.0.0.1:9200/test/?pretty=1](http://127.0.0.1:9200/test/?pretty=1)' -d '  
> > {  
> > "settings" : {  
> > "analysis" : {  
> > "filter" : {  
> > "end\_grams" : {  
> > "max\_gram" : 2,  
> > "side" : "back",  
> > "min\_gram" : 2,  
> > "type" : "edge\_ngram"  
> > }  
> > },  
> > "analyzer" : {  
> > "end\_grams" : {  
> > "filter" : [  
> > "standard",  
> > "lowercase",  
> > "stop",  
> > "end\_grams"  
> > ],  
> > "tokenizer" : "standard"  
> > }  
> > }  
> > }  
> > }  
> > }  
> > '
> > 
> > curl -XGET '[http://127.0.0.1:9200/test/\_analyze?pretty=1&text=The+quick](http://127.0.0.1:9200/test/_analyze?pretty=1&text=The+quick)  
> > +brown+fox+jumped+over+the+lazy+dog&analyzer=end\_grams[http://127.0.0.1:9200/test/\_analyze?pretty=1&text=The+quick+brown+fox+jumped+over+the+lazy+dog&analyzer=end\_grams](http://127.0.0.1:9200/test/_analyze?pretty=1&text=The+quick+brown+fox+jumped+over+the+lazy+dog&analyzer=end_grams)  
> > '
> > 
> > # {
> > 
> > # "tokens" : [
> > 
> > # {
> > 
> > # "end\_offset" : 9,
> > 
> > # "position" : 1,
> > 
> > # "start\_offset" : 7,
> > 
> > # "type" : "word",
> > 
> > # "token" : "ck"
> > 
> > # },
> > 
> > # {
> > 
> > # "end\_offset" : 15,
> > 
> > # "position" : 2,
> > 
> > # "start\_offset" : 13,
> > 
> > # "type" : "word",
> > 
> > # "token" : "wn"
> > 
> > # },
> > 
> > # {
> > 
> > # "end\_offset" : 19,
> > 
> > # "position" : 3,
> > 
> > # "start\_offset" : 17,
> > 
> > # "type" : "word",
> > 
> > # "token" : "ox"
> > 
> > # },
> > 
> > # {
> > 
> > # "end\_offset" : 26,
> > 
> > # "position" : 4,
> > 
> > # "start\_offset" : 24,
> > 
> > # "type" : "word",
> > 
> > # "token" : "ed"
> > 
> > # },
> > 
> > # {
> > 
> > # "end\_offset" : 31,
> > 
> > # "position" : 5,
> > 
> > # "start\_offset" : 29,
> > 
> > # "type" : "word",
> > 
> > # "token" : "er"
> > 
> > # },
> > 
> > # {
> > 
> > # "end\_offset" : 40,
> > 
> > # "position" : 6,
> > 
> > # "start\_offset" : 38,
> > 
> > # "type" : "word",
> > 
> > # "token" : "zy"
> > 
> > # },
> > 
> > # {
> > 
> > # "end\_offset" : 44,
> > 
> > # "position" : 7,
> > 
> > # "start\_offset" : 42,
> > 
> > # "type" : "word",
> > 
> > # "token" : "og"
> > 
> > # }
> > 
> > # ]
> > 
> > # }
> > 
> > clint
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [February 26, 2013, 11:30am UTC](https://discuss.elastic.co/t/inverse-edge-back-ngram-or-making-it-fuzzy-at-the-end-of-a-word/10903/5 "2013-02-26T11:30:01Z")

</div>

On Tue, 2013-02-26 at 12:09 +0100, Per Ekman wrote:

> Alright, that is pretty much what we've done so far, but I'm looking  
> at getting "bro", "f", "jump"..... into the index, instead of the  
> endings,

You specified that you wanted ngrams of the last two characters, which  
is why I set "side" to "back".

> And possibly the original words as well.

Just make the edge ngrams long enough.

You may want to use a multi-field to have one field indexed with (eg)  
the standard analyzer, and another indexed with edge-ngrams, and you can  
query both of them in a single query, giving different boosts to each  
clause

clint

> On Tue, Feb 26, 2013 at 12:02 PM, Clinton Gormley  
> [clint@traveljury.com](mailto:clint@traveljury.com) wrote:  
> On Tue, 2013-02-26 at 02:45 -0800, Per Ekman wrote:  
> \> Hi  
> \>  
> \>  
> \> We are discussing building an index where possible  
> misspellings at the  
> \> end of a word are getting hits.  
> \>  
> \>  
> \> We were looking at using the EdgeNGram and making ngrams of  
> the last  
> \> two characters, but that gives us an index of just the  
> 2-character  
> \> variations of the word endings.  
> \>  
> \>  
> \> How would we best do this? Is it possible to configure the  
> inverse of  
> \> that? Should we tokenize it with a regexp? Any other ideas?
> 
> ```
> curl -XPUT 'http://127.0.0.1:9200/test/?pretty=1' -d '
> {
> "settings" : {
> "analysis" : {
> "filter" : {
> "end_grams" : {
> "max_gram" : 2,
> "side" : "back",
> "min_gram" : 2,
> "type" : "edge_ngram"
> }
> },
> "analyzer" : {
> "end_grams" : {
> "filter" : [
> "standard",
> "lowercase",
> "stop",
> "end_grams"
> ],
> "tokenizer" : "standard"
> }
> }
> }
> }
> }
> '
>     
> curl -XGET
> 'http://127.0.0.1:9200/test/_analyze?pretty=1&text=The+quick
> +brown+fox+jumped+over+the+lazy+dog&analyzer=end_grams'
>     
> # {
> # "tokens" : [
> # {
> # "end_offset" : 9,
> # "position" : 1,
> # "start_offset" : 7,
> # "type" : "word",
> # "token" : "ck"
> # },
> # {
> # "end_offset" : 15,
> # "position" : 2,
> # "start_offset" : 13,
> # "type" : "word",
> # "token" : "wn"
> # },
> # {
> # "end_offset" : 19,
> # "position" : 3,
> # "start_offset" : 17,
> # "type" : "word",
> # "token" : "ox"
> # },
> # {
> # "end_offset" : 26,
> # "position" : 4,
> # "start_offset" : 24,
> # "type" : "word",
> # "token" : "ed"
> # },
> # {
> # "end_offset" : 31,
> # "position" : 5,
> # "start_offset" : 29,
> # "type" : "word",
> # "token" : "er"
> # },
> # {
> # "end_offset" : 40,
> # "position" : 6,
> # "start_offset" : 38,
> # "type" : "word",
> # "token" : "zy"
> # },
> # {
> # "end_offset" : 44,
> # "position" : 7,
> # "start_offset" : 42,
> # "type" : "word",
> # "token" : "og"
> # }
> # ]
> # }
>     
>     
> clint
>     
> --
> You received this message because you are subscribed to the
> Google Groups "elasticsearch" group.
> To unsubscribe from this group and stop receiving emails from
> it, send an email to elasticsearch
> +unsubscribe@googlegroups.com.
> For more options, visit
> https://groups.google.com/groups/opt_out.
> 
> ```
> 
> --  
> You received this message because you are subscribed to the Google  
> Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send  
> an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [February 27, 2013, 2:17pm UTC](https://discuss.elastic.co/t/inverse-edge-back-ngram-or-making-it-fuzzy-at-the-end-of-a-word/10903/6 "2013-02-27T14:17:13Z")

</div>

On Tue, 2013-02-26 at 03:13 -0800, Per Ekman wrote:

> I guess I was really unclear in my original text. I want to know how  
> to strip the last couple of characters in a word, and also keep the  
> original

Ah right

Currently you can't do that in the same field - you can have one field  
with the full word, and another field which uses the pattern tokenizer  
to drop the last two letters.

I'm hoping to get a token filter accepted which does allow multiple  
captures per position in the same field:  
[https://issues.apache.org/jira/browse/LUCENE-4766](https://issues.apache.org/jira/browse/LUCENE-4766)  
but it'll be a while before that happens

clint

> On Tuesday, February 26, 2013 12:09:19 PM UTC+1, Per Ekman wrote:  
> Alright, that is pretty much what we've done so far, but I'm  
> looking at getting "bro", "f", "jump"..... into the index,  
> instead of the endings, And possibly the original words as  
> well.
> 
> ```
> On Tue, 2013-02-26 at 02:45 -0800, Per Ekman wrote:
> > Hi
> >
> >
> > We are discussing building an index where possible
> misspellings at the
> > end of a word are getting hits.
> >
> >
> > We were looking at using the EdgeNGram and making
> ngrams of the last
> > two characters, but that gives us an index of just
> the 2-character
> > variations of the word endings.
> >
> >
> > How would we best do this? Is it possible to
> configure the inverse of
> > that? Should we tokenize it with a regexp? Any other
> ideas?
>             
>             
> curl -XPUT 'http://127.0.0.1:9200/test/?pretty=1' -d
> '
> {
> "settings" : {
> "analysis" : {
> "filter" : {
> "end_grams" : {
> "max_gram" : 2,
> "side" : "back",
> "min_gram" : 2,
> "type" : "edge_ngram"
> }
> },
> "analyzer" : {
> "end_grams" : {
> "filter" : [
> "standard",
> "lowercase",
> "stop",
> "end_grams"
> ],
> "tokenizer" : "standard"
> }
> }
> }
> }
> }
> '
>             
> curl -XGET
> 'http://127.0.0.1:9200/test/_analyze?pretty=1&text=The
> +quick
> +brown+fox+jumped+over+the+lazy
> +dog&analyzer=end_grams'
>             
> # {
> # "tokens" : [
> # {
> # "end_offset" : 9,
> # "position" : 1,
> # "start_offset" : 7,
> # "type" : "word",
> # "token" : "ck"
> # },
> # {
> # "end_offset" : 15,
> # "position" : 2,
> # "start_offset" : 13,
> # "type" : "word",
> # "token" : "wn"
> # },
> # {
> # "end_offset" : 19,
> # "position" : 3,
> # "start_offset" : 17,
> # "type" : "word",
> # "token" : "ox"
> # },
> # {
> # "end_offset" : 26,
> # "position" : 4,
> # "start_offset" : 24,
> # "type" : "word",
> # "token" : "ed"
> # },
> # {
> # "end_offset" : 31,
> # "position" : 5,
> # "start_offset" : 29,
> # "type" : "word",
> # "token" : "er"
> # },
> # {
> # "end_offset" : 40,
> # "position" : 6,
> # "start_offset" : 38,
> # "type" : "word",
> # "token" : "zy"
> # },
> # {
> # "end_offset" : 44,
> # "position" : 7,
> # "start_offset" : 42,
> # "type" : "word",
> # "token" : "og"
> # }
> # ]
> # }
>             
>             
> clint
>             
> --
> You received this message because you are subscribed
> to the Google Groups "elasticsearch" group.
> To unsubscribe from this group and stop receiving
> emails from it, send an email to elasticsearch
> +unsubscribe@googlegroups.com.
> For more options, visit
> https://groups.google.com/groups/opt_out.
> 
> ```
> 
> --  
> You received this message because you are subscribed to the Google  
> Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send  
> an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Per\_Ekman](https://avatars.discourse-cdn.com/v4/letter/p/e9bcb4/32.png) [@Per\_Ekman](https://discuss.elastic.co/u/Per_Ekman)\
**Post date:** [February 27, 2013, 3:14pm UTC](https://discuss.elastic.co/t/inverse-edge-back-ngram-or-making-it-fuzzy-at-the-end-of-a-word/10903/7 "2013-02-27T15:14:36Z")

</div>

Cool. Yeah, we were playing around with the pattern tokenizer to achieve  
this

On Wed, Feb 27, 2013 at 3:17 PM, Clinton Gormley [clint@traveljury.com](mailto:clint@traveljury.com)wrote:

> On Tue, 2013-02-26 at 03:13 -0800, Per Ekman wrote:
> 
> > I guess I was really unclear in my original text. I want to know how  
> > to strip the last couple of characters in a word, and also keep the  
> > original
> 
> Ah right
> 
> Currently you can't do that in the same field - you can have one field  
> with the full word, and another field which uses the pattern tokenizer  
> to drop the last two letters.
> 
> I'm hoping to get a token filter accepted which does allow multiple  
> captures per position in the same field:  
> [[LUCENE-4766] Pattern token filter which emits a token for every capturing group - ASF JIRA](https://issues.apache.org/jira/browse/LUCENE-4766)  
> but it'll be a while before that happens
> 
> clint
> 
> > On Tuesday, February 26, 2013 12:09:19 PM UTC+1, Per Ekman wrote:  
> > Alright, that is pretty much what we've done so far, but I'm  
> > looking at getting "bro", "f", "jump"..... into the index,  
> > instead of the endings, And possibly the original words as  
> > well.
> > 
> > ```
> > On Tue, 2013-02-26 at 02:45 -0800, Per Ekman wrote:
> > > Hi
> > >
> > >
> > > We are discussing building an index where possible
> > misspellings at the
> > > end of a word are getting hits.
> > >
> > >
> > > We were looking at using the EdgeNGram and making
> > ngrams of the last
> > > two characters, but that gives us an index of just
> > the 2-character
> > > variations of the word endings.
> > >
> > >
> > > How would we best do this? Is it possible to
> > configure the inverse of
> > > that? Should we tokenize it with a regexp? Any other
> > ideas?
> > 
> > curl -XPUT 'http://127.0.0.1:9200/test/?pretty=1' -d
> > '
> > {
> > "settings" : {
> > "analysis" : {
> > "filter" : {
> > "end_grams" : {
> > "max_gram" : 2,
> > "side" : "back",
> > "min_gram" : 2,
> > "type" : "edge_ngram"
> > }
> > },
> > "analyzer" : {
> > "end_grams" : {
> > "filter" : [
> > "standard",
> > "lowercase",
> > "stop",
> > "end_grams"
> > ],
> > "tokenizer" : "standard"
> > }
> > }
> > }
> > }
> > }
> > '
> > 
> > curl -XGET
> > 'http://127.0.0.1:9200/test/_analyze?pretty=1&text=The
> > +quick
> > +brown+fox+jumped+over+the+lazy
> > +dog&analyzer=end_grams'
> > 
> > # {
> > # "tokens" : [
> > # {
> > # "end_offset" : 9,
> > # "position" : 1,
> > # "start_offset" : 7,
> > # "type" : "word",
> > # "token" : "ck"
> > # },
> > # {
> > # "end_offset" : 15,
> > # "position" : 2,
> > # "start_offset" : 13,
> > # "type" : "word",
> > # "token" : "wn"
> > # },
> > # {
> > # "end_offset" : 19,
> > # "position" : 3,
> > # "start_offset" : 17,
> > # "type" : "word",
> > # "token" : "ox"
> > # },
> > # {
> > # "end_offset" : 26,
> > # "position" : 4,
> > # "start_offset" : 24,
> > # "type" : "word",
> > # "token" : "ed"
> > # },
> > # {
> > # "end_offset" : 31,
> > # "position" : 5,
> > # "start_offset" : 29,
> > # "type" : "word",
> > # "token" : "er"
> > # },
> > # {
> > # "end_offset" : 40,
> > # "position" : 6,
> > # "start_offset" : 38,
> > # "type" : "word",
> > # "token" : "zy"
> > # },
> > # {
> > # "end_offset" : 44,
> > # "position" : 7,
> > # "start_offset" : 42,
> > # "type" : "word",
> > # "token" : "og"
> > # }
> > # ]
> > # }
> > 
> > clint
> > 
> > --
> > You received this message because you are subscribed
> > to the Google Groups "elasticsearch" group.
> > To unsubscribe from this group and stop receiving
> > emails from it, send an email to elasticsearch
> > +unsubscribe@googlegroups.com.
> > For more options, visit
> > https://groups.google.com/groups/opt_out.
> > 
> > ```
> > 
> > --  
> > You received this message because you are subscribed to the Google  
> > Groups "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send  
> > an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).
> 
> --  
> You received this message because you are subscribed to a topic in the  
> Google Groups "elasticsearch" group.  
> To unsubscribe from this topic, visit  
> [https://groups.google.com/d/topic/elasticsearch/d85geUwu1WM/unsubscribe?hl=en-US](https://groups.google.com/d/topic/elasticsearch/d85geUwu1WM/unsubscribe?hl=en-US)  
> .  
> To unsubscribe from this group and all its topics, send an email to  
> [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:49am UTC](https://discuss.elastic.co/t/inverse-edge-back-ngram-or-making-it-fuzzy-at-the-end-of-a-word/10903/8 "2017-07-06T02:49:18Z")

</div>


