# ES 0.17.0, Lucene 3.3 finite state

**URL:** <https://discuss.elastic.co/t/es-0-17-0-lucene-3-3-finite-state/4891>\
**Category:** Elasticsearch\
**Created:** [July 19, 2011, 12:22am UTC](https://discuss.elastic.co/t/es-0-17-0-lucene-3-3-finite-state/4891 "2011-07-19T00:22:24Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Sebastian\_Gavarini](https://avatars.discourse-cdn.com/v4/letter/s/db5fbb/32.png) [@Sebastian\_Gavarini](https://discuss.elastic.co/u/Sebastian_Gavarini)\
**Post date:** [July 19, 2011, 12:22am UTC](https://discuss.elastic.co/t/es-0-17-0-lucene-3-3-finite-state/4891/1 "2011-07-19T00:22:24Z")

</div>

Hi all,

If I understood correctly from the docs, finite state spellchecker is now  
backported to Lucene 3.3 and so in ElasticSearch 0.17.0. I remember a  
discussion in the mailing list some time ago about spellchecking and the  
idea of not using a different index but just the same data ones with a  
finite state machine doing fuzzy matching in an efficient way, that was  
going to be available in Lucene 4. At the time I was going to use the  
n-grams spellchecker in an incremental new index, but I held it because this  
other FST spellchecker seemed better.

What's the state of all this in ElasticSearch 0.17.0? is it possible to  
access that API somehow? Is it possible (and faster) to use the new fuzzy  
queries with FST?

Thanks,  
Sebastian.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [July 19, 2011, 12:30am UTC](https://discuss.elastic.co/t/es-0-17-0-lucene-3-3-finite-state/4891/2 "2011-07-19T00:30:26Z")

</div>

FST for this is not part of Lucene 3.3, only in trunk (upcoming 4.0).

On Tue, Jul 19, 2011 at 3:22 AM, Sebastian Gavarini [sgavarini@gmail.com](mailto:sgavarini@gmail.com)wrote:

> Hi all,
> 
> If I understood correctly from the docs, finite state spellchecker is now  
> backported to Lucene 3.3 and so in Elasticsearch 0.17.0. I remember a  
> discussion in the mailing list some time ago about spellchecking and the  
> idea of not using a different index but just the same data ones with a  
> finite state machine doing fuzzy matching in an efficient way, that was  
> going to be available in Lucene 4. At the time I was going to use the  
> n-grams spellchecker in an incremental new index, but I held it because this  
> other FST spellchecker seemed better.
> 
> What's the state of all this in Elasticsearch 0.17.0? is it possible to  
> access that API somehow? Is it possible (and faster) to use the new fuzzy  
> queries with FST?
> 
> Thanks,  
> Sebastian.

---

<div class="post-metadata">

**Author:** ![Sebastian\_Gavarini](https://avatars.discourse-cdn.com/v4/letter/s/db5fbb/32.png) [@Sebastian\_Gavarini](https://discuss.elastic.co/u/Sebastian_Gavarini)\
**Post date:** [July 19, 2011, 1:01am UTC](https://discuss.elastic.co/t/es-0-17-0-lucene-3-3-finite-state/4891/3 "2011-07-19T01:01:11Z")

</div>

Shay,

I have checked Lucene's 3.3 changelist and it reports two new features, one  
for the spellchecker, but also another one called FST backport:  
[https://issues.apache.org/jira/browse/LUCENE-3140](https://issues.apache.org/jira/browse/LUCENE-3140)  
[https://issues.apache.org/jira/browse/LUCENE-3135](https://issues.apache.org/jira/browse/LUCENE-3135)

Isn't the first one what's needed to make fuzzy fst? (granted that I haven't  
seen any "_query_" class in the 3140 patch, which isn't very encouraging).

Is the backported code only good for the separate spellchecker?

Thanks,  
Sebastian.

On Mon, Jul 18, 2011 at 9:30 PM, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:

> FST for this is not part of Lucene 3.3, only in trunk (upcoming 4.0).
> 
> On Tue, Jul 19, 2011 at 3:22 AM, Sebastian Gavarini [sgavarini@gmail.com](mailto:sgavarini@gmail.com)wrote:
> 
> > Hi all,
> > 
> > If I understood correctly from the docs, finite state spellchecker is now  
> > backported to Lucene 3.3 and so in Elasticsearch 0.17.0. I remember a  
> > discussion in the mailing list some time ago about spellchecking and the  
> > idea of not using a different index but just the same data ones with a  
> > finite state machine doing fuzzy matching in an efficient way, that was  
> > going to be available in Lucene 4. At the time I was going to use the  
> > n-grams spellchecker in an incremental new index, but I held it because this  
> > other FST spellchecker seemed better.
> > 
> > What's the state of all this in Elasticsearch 0.17.0? is it possible to  
> > access that API somehow? Is it possible (and faster) to use the new fuzzy  
> > queries with FST?
> > 
> > Thanks,  
> > Sebastian.

---

<div class="post-metadata">

**Author:** ![rmuir](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rmuir/32/44949_2.png) [@rmuir](https://discuss.elastic.co/u/rmuir)\
**Post date:** [July 19, 2011, 1:33am UTC](https://discuss.elastic.co/t/es-0-17-0-lucene-3-3-finite-state/4891/4 "2011-07-19T01:33:33Z")

</div>

On Mon, Jul 18, 2011 at 9:01 PM, Sebastian Gavarini [sgavarini@gmail.com](mailto:sgavarini@gmail.com) wrote:

> Shay,  
> I have checked Lucene's 3.3 changelist and it reports two new features, one  
> for the spellchecker, but also another one called FST backport:  
> [[LUCENE-3140] Backport FSTs to 3.x - ASF JIRA](https://issues.apache.org/jira/browse/LUCENE-3140)  
> [[LUCENE-3135] backport suggest module to branch 3.x - ASF JIRA](https://issues.apache.org/jira/browse/LUCENE-3135)  
> Isn't the first one what's needed to make fuzzy fst? (granted that I haven't  
> seen any "_query_" class in the 3140 patch, which isn't very encouraging).  
> Is the backported code only good for the separate spellchecker?

Hi, there are two different pieces of finite-state functionality in lucene:

- FSA (finite state automaton). think of this as HashSet: this is  
totally in 4.0-only  
This is used mostly for things that go directly against the index:  
It implements AutomatonQuery, WildcardQuery, RegexpQuery,  
FuzzyQuery, and DirectSpellChecker in 4.0
- FST (finite state transducer). this of this as HashMap: this is in  
3.x, and 4.0  
This is used as a general datastructure, to hold things (e.g. terms  
-\> something).  
It implements the terms index and memory codec in 4.0-only  
But, it implements auto-suggest and soon also, synonyms in 3.x and 4.0

So, the fst backport here only applies to suggest, which fills a  
separate FST from the lucene index.  
The spellchecker stuff in 4.0-only is different, it turns the users  
query into an FSA and runs it against the lucene index directly.

for more information, see Dawid Weiss' presentation here:  
[http://www.lucidimagination.com/sites/default/files/Weiss%20Dawid%20-%20Finite%20State%20Automata%20in%20Lucene.pdf](http://www.lucidimagination.com/sites/default/files/Weiss%20Dawid%20-%20Finite%20State%20Automata%20in%20Lucene.pdf)

--

> **[Home](https://lucidworks.com/)**
>
> Lucidworks' Fusion platform uses industry-leading search technology to power search & discovery for the largest & most successful companies. Request a demo today.

---

<div class="post-metadata">

**Author:** ![Sebastian\_Gavarini](https://avatars.discourse-cdn.com/v4/letter/s/db5fbb/32.png) [@Sebastian\_Gavarini](https://discuss.elastic.co/u/Sebastian_Gavarini)\
**Post date:** [July 19, 2011, 2:02am UTC](https://discuss.elastic.co/t/es-0-17-0-lucene-3-3-finite-state/4891/5 "2011-07-19T02:02:11Z")

</div>

Thanks for the explanation Robert, I'll wait for 4.0 then.

On Mon, Jul 18, 2011 at 10:33 PM, Robert Muir [rcmuir@gmail.com](mailto:rcmuir@gmail.com) wrote:

> On Mon, Jul 18, 2011 at 9:01 PM, Sebastian Gavarini [sgavarini@gmail.com](mailto:sgavarini@gmail.com)  
> wrote:
> 
> > Shay,  
> > I have checked Lucene's 3.3 changelist and it reports two new features,  
> > one  
> > for the spellchecker, but also another one called FST backport:  
> > [[LUCENE-3140] Backport FSTs to 3.x - ASF JIRA](https://issues.apache.org/jira/browse/LUCENE-3140)  
> > [[LUCENE-3135] backport suggest module to branch 3.x - ASF JIRA](https://issues.apache.org/jira/browse/LUCENE-3135)  
> > Isn't the first one what's needed to make fuzzy fst? (granted that I  
> > haven't  
> > seen any "_query_" class in the 3140 patch, which isn't very  
> > encouraging).  
> > Is the backported code only good for the separate spellchecker?
> 
> Hi, there are two different pieces of finite-state functionality in lucene:
> 
> - FSA (finite state automaton). think of this as HashSet: this is  
> totally in 4.0-only  
> This is used mostly for things that go directly against the index:  
> It implements AutomatonQuery, WildcardQuery, RegexpQuery,  
> FuzzyQuery, and DirectSpellChecker in 4.0
> - FST (finite state transducer). this of this as HashMap: this is in  
> 3.x, and 4.0  
> This is used as a general datastructure, to hold things (e.g. terms  
> -\> something).  
> It implements the terms index and memory codec in 4.0-only  
> But, it implements auto-suggest and soon also, synonyms in 3.x and 4.0
> 
> So, the fst backport here only applies to suggest, which fills a  
> separate FST from the lucene index.  
> The spellchecker stuff in 4.0-only is different, it turns the users  
> query into an FSA and runs it against the lucene index directly.
> 
> for more information, see Dawid Weiss' presentation here:
> 
> [http://www.lucidimagination.com/sites/default/files/Weiss%20Dawid%20-%20Finite%20State%20Automata%20in%20Lucene.pdf](http://www.lucidimagination.com/sites/default/files/Weiss%20Dawid%20-%20Finite%20State%20Automata%20in%20Lucene.pdf)
> 
> --  
> [lucidimagination.com](http://lucidimagination.com)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:00am UTC](https://discuss.elastic.co/t/es-0-17-0-lucene-3-3-finite-state/4891/6 "2017-07-06T04:00:11Z")

</div>


