0.21.0.Beta Status (Lucene 4.x)

kimchy · February 12, 2013, 11:01pm

Heya fellows,

Wanted to give a quick update regarding 0.21.0.Beta release with Lucene 4.x. As I mentioned, we have already fully upgraded to Luceen 4.1, and practically done with exposing the most important features we can get for it. We also concluded our major refactoring for a field data code that has significant reduction in memory used for field data (specifically multi valued fields) for faceting/sorting/scripts.

The last major bit we have left is additional refactoring on our faceting code, to allow for more pluggable faceting implementation and different execution modes. The work is progressing nicely, and I hope that by next week we will be able to release 0.21.0.Beta with it in it.

Once the Beta is out, we will have a short release cycle gearing towards GA, hopefully it should't take long.

--
You received this message because you are subscribed to the Google Groups "elasticsearch" group.
To unsubscribe from this group and stop receiving emails from it, send an email to elasticsearch+unsubscribe@googlegroups.com.
For more options, visit https://groups.google.com/groups/opt_out.

Justin_Treher · February 13, 2013, 1:45pm

Thanks for the communication!

On Tuesday, February 12, 2013 6:01:07 PM UTC-5, kimchy wrote:

Heya fellows,

Wanted to give a quick update regarding 0.21.0.Beta release with Lucene
4.x. As I mentioned, we have already fully upgraded to Luceen 4.1, and
practically done with exposing the most important features we can get for
it. We also concluded our major refactoring for a field data code that has
significant reduction in memory used for field data (specifically multi
valued fields) for faceting/sorting/scripts.

The last major bit we have left is additional refactoring on our faceting
code, to allow for more pluggable faceting implementation and different
execution modes. The work is progressing nicely, and I hope that by next
week we will be able to release 0.21.0.Beta with it in it.

Once the Beta is out, we will have a short release cycle gearing towards
GA, hopefully it should't take long.

--
You received this message because you are subscribed to the Google Groups "elasticsearch" group.
To unsubscribe from this group and stop receiving emails from it, send an email to elasticsearch+unsubscribe@googlegroups.com.
For more options, visit https://groups.google.com/groups/opt_out.

jagdeep · February 13, 2013, 1:55pm

Thanks Shay for the update. Eagerly looking forward for the new release.

On Wednesday, February 13, 2013 4:31:07 AM UTC+5:30, kimchy wrote:

Heya fellows,

Wanted to give a quick update regarding 0.21.0.Beta release with Lucene
4.x. As I mentioned, we have already fully upgraded to Luceen 4.1, and
practically done with exposing the most important features we can get for
it. We also concluded our major refactoring for a field data code that has
significant reduction in memory used for field data (specifically multi
valued fields) for faceting/sorting/scripts.

The last major bit we have left is additional refactoring on our faceting
code, to allow for more pluggable faceting implementation and different
execution modes. The work is progressing nicely, and I hope that by next
week we will be able to release 0.21.0.Beta with it in it.

Once the Beta is out, we will have a short release cycle gearing towards
GA, hopefully it should't take long.

--
You received this message because you are subscribed to the Google Groups "elasticsearch" group.
To unsubscribe from this group and stop receiving emails from it, send an email to elasticsearch+unsubscribe@googlegroups.com.
For more options, visit https://groups.google.com/groups/opt_out.

Leonardo_Menezes · February 14, 2013, 8:25am

Hey Shay,
that's great news. For the new code in faceting, should we also expect
performance improvement, or just memory? thanks

Leonardo Menezes
(+34) 688907766
http://lmenezes.com

On Wed, Feb 13, 2013 at 2:55 PM, jagdeep reach.jagdeep@gmail.com wrote:

Thanks Shay for the update. Eagerly looking forward for the new release.

On Wednesday, February 13, 2013 4:31:07 AM UTC+5:30, kimchy wrote:

Heya fellows,

Wanted to give a quick update regarding 0.21.0.Beta release with Lucene
4.x. As I mentioned, we have already fully upgraded to Luceen 4.1, and
practically done with exposing the most important features we can get for
it. We also concluded our major refactoring for a field data code that has
significant reduction in memory used for field data (specifically multi
valued fields) for faceting/sorting/scripts.

The last major bit we have left is additional refactoring on our faceting
code, to allow for more pluggable faceting implementation and different
execution modes. The work is progressing nicely, and I hope that by next
week we will be able to release 0.21.0.Beta with it in it.

Once the Beta is out, we will have a short release cycle gearing towards
GA, hopefully it should't take long.

--
You received this message because you are subscribed to the Google Groups
"elasticsearch" group.
To unsubscribe from this group and stop receiving emails from it, send an
email to elasticsearch+unsubscribe@googlegroups.com.
For more options, visit https://groups.google.com/groups/opt_out.

--
You received this message because you are subscribed to the Google Groups "elasticsearch" group.
To unsubscribe from this group and stop receiving emails from it, send an email to elasticsearch+unsubscribe@googlegroups.com.
For more options, visit https://groups.google.com/groups/opt_out.

Marcin_Dojwa · February 14, 2013, 8:41am

"We also concluded our major refactoring for a field data code that has
significant reduction in memory used for field data" - this is great news
for me I can't wait for the release

2013/2/14 Leonardo Menezes leonardo.menezess@gmail.com

Hey Shay,
that's great news. For the new code in faceting, should we also expect
performance improvement, or just memory? thanks

Leonardo Menezes
(+34) 688907766
http://lmenezes.com

On Wed, Feb 13, 2013 at 2:55 PM, jagdeep reach.jagdeep@gmail.com wrote:

Thanks Shay for the update. Eagerly looking forward for the new release.

On Wednesday, February 13, 2013 4:31:07 AM UTC+5:30, kimchy wrote:

Heya fellows,

Wanted to give a quick update regarding 0.21.0.Beta release with Lucene
4.x. As I mentioned, we have already fully upgraded to Luceen 4.1, and
practically done with exposing the most important features we can get for
it. We also concluded our major refactoring for a field data code that has
significant reduction in memory used for field data (specifically multi
valued fields) for faceting/sorting/scripts.

The last major bit we have left is additional refactoring on our
faceting code, to allow for more pluggable faceting implementation and
different execution modes. The work is progressing nicely, and I hope that
by next week we will be able to release 0.21.0.Beta with it in it.

Once the Beta is out, we will have a short release cycle gearing towards
GA, hopefully it should't take long.

--
You received this message because you are subscribed to the Google Groups
"elasticsearch" group.
To unsubscribe from this group and stop receiving emails from it, send an
email to elasticsearch+unsubscribe@googlegroups.com.
For more options, visit https://groups.google.com/groups/opt_out.

--
You received this message because you are subscribed to the Google Groups
"elasticsearch" group.
To unsubscribe from this group and stop receiving emails from it, send an
email to elasticsearch+unsubscribe@googlegroups.com.
For more options, visit https://groups.google.com/groups/opt_out.

--
You received this message because you are subscribed to the Google Groups "elasticsearch" group.
To unsubscribe from this group and stop receiving emails from it, send an email to elasticsearch+unsubscribe@googlegroups.com.
For more options, visit https://groups.google.com/groups/opt_out.

jprante · February 14, 2013, 8:51am

I think the facet framework is going to be overhauled to add custom
facet features more easily. For example, right now, there can't be
plugins that add new capabilities to faceting.

So I hope I can use collation based sorting on facet elements soon, or
facet value mappings, where facet values can be augmented with other
values for display (like sort keys in collations).

Jörg

Am 14.02.13 09:25, schrieb Leonardo Menezes:

Hey Shay,
that's great news. For the new code in faceting, should we also
expect performance improvement, or just memory? thanks

--
You received this message because you are subscribed to the Google Groups "elasticsearch" group.
To unsubscribe from this group and stop receiving emails from it, send an email to elasticsearch+unsubscribe@googlegroups.com.
For more options, visit https://groups.google.com/groups/opt_out.

Nicolas_Blanc · February 14, 2013, 10:15am

As we need one boring guy... Is this means that i am near to trash my
nightmare code about grouping ? Mainstream ES will be able to group soon ?

--
Nicolas BLANC.

Le mercredi 13 février 2013 00:01:07 UTC+1, kimchy a écrit :

Heya fellows,

Wanted to give a quick update regarding 0.21.0.Beta release with Lucene
4.x. As I mentioned, we have already fully upgraded to Luceen 4.1, and
practically done with exposing the most important features we can get for
it. We also concluded our major refactoring for a field data code that has
significant reduction in memory used for field data (specifically multi
valued fields) for faceting/sorting/scripts.

The last major bit we have left is additional refactoring on our faceting
code, to allow for more pluggable faceting implementation and different
execution modes. The work is progressing nicely, and I hope that by next
week we will be able to release 0.21.0.Beta with it in it.

Once the Beta is out, we will have a short release cycle gearing towards
GA, hopefully it should't take long.

--
You received this message because you are subscribed to the Google Groups "elasticsearch" group.
To unsubscribe from this group and stop receiving emails from it, send an email to elasticsearch+unsubscribe@googlegroups.com.
For more options, visit https://groups.google.com/groups/opt_out.

Ivan · February 14, 2013, 3:49pm

Exciting news. I have an idea for a new facet type, but the
current framework is too much work for a feature I do not fully need.

New fuzzy queries means spell checking soon as well?

--
Ivan

On Thu, Feb 14, 2013 at 12:51 AM, Jörg Prante joergprante@gmail.com wrote:

I think the facet framework is going to be overhauled to add custom facet
features more easily. For example, right now, there can't be plugins that
add new capabilities to faceting.

So I hope I can use collation based sorting on facet elements soon, or
facet value mappings, where facet values can be augmented with other values
for display (like sort keys in collations).

Jörg

Am 14.02.13 09:25, schrieb Leonardo Menezes:

Hey Shay,
that's great news. For the new code in faceting, should we also
expect performance improvement, or just memory? thanks
--
You received this message because you are subscribed to the Google Groups
"elasticsearch" group.
To unsubscribe from this group and stop receiving emails from it, send an
email to elasticsearch+unsubscribe@**googlegroups.com elasticsearch%2Bunsubscribe@googlegroups.com
.
For more options, visit https://groups.google.com/**groups/opt_out https://groups.google.com/groups/opt_out
.

--
You received this message because you are subscribed to the Google Groups "elasticsearch" group.
To unsubscribe from this group and stop receiving emails from it, send an email to elasticsearch+unsubscribe@googlegroups.com.
For more options, visit https://groups.google.com/groups/opt_out.

mattweber · February 14, 2013, 5:23pm

Spell checking/suggestions is already in master....

github.com/elastic/elasticsearch

Add suggest api

opened 02:35PM - 24 Jan 13 UTC

closed 02:41PM - 24 Jan 13 UTC

martijnvg

>feature v0.90.0.Beta1

# Suggest feature The suggest feature suggests similar looking terms based on a… provided text by using a suggester. At the moment there the only supported suggester is `fuzzy`. The suggest feature is available from version `0.21.0`. # Fuzzy suggester The `fuzzy` suggester suggests terms based on edit distance. The provided suggest text is analyzed before terms are suggested. The suggested terms are provided per analyzed suggest text token. The `fuzzy` suggester doesn't take the query into account that is part of request. # Suggest API The suggest request part is defined along side the query part as top field in the json request. ``` curl -s -XPOST 'localhost:9200/_search' -d '{ "query" : { ... }, "suggest" : { ... } }' ``` Several suggestions can be specified per request. Each suggestion is identified with an arbitary name. In the example below two suggestions are requested. Both `my-suggest-1` and `my-suggest-2` suggestions use the `fuzzy` suggester, but have a different `text`. ``` "suggest" : { "my-suggest-1" : { "text" : "the amsterdma meetpu", "fuzzy" : { "field" : "body" } }, "my-suggest-2" : { "text" : "the rottredam meetpu", "fuzzy" : { "field" : "title", } } } ``` The below suggest response example includes the suggestion response for `my-suggest-1` and `my-suggest-2`. Each suggestion part contains entries. Each entry is effectively a token from the suggest text and contains the suggestion entry text, the original start offset and length in the suggest text and if found an arbitary number of options. ``` { ... "suggest": { "my-suggest-1": [ { "text" : "amsterdma", "offset": 4, "length": 9, "options": [ ... ] }, ... ], "my-suggest-2" : [ ... ] } ... } ``` Each options array contains a option object that includes the suggested text, its document frequency and score compared to the suggest entry text. The meaning of the score depends on the used suggester. The fuzzy suggester's score is based on the edit distance. ``` "options": [ { "text": "amsterdam", "freq": 77, "score": 0.8888889 }, ... ] ``` # Global suggest text To avoid repitition of the suggest text, it is possible to define a global text. In the example below the suggest text is defined globally and applies to the `my-suggest-1` and `my-suggest-2` suggestions. ``` "suggest" : { "text" : "the amsterdma meetpu" "my-suggest-1" : { "fuzzy" : { "field" : "title" } }, "my-suggest-2" : { "fuzzy" : { "field" : "body" } } } ``` The suggest text can in the above example also be specied as suggestion specific option. The suggest text specified on suggestion level override the suggest text on the global level. # Other suggest example. In the below example we request suggestions for the following suggest text: `devloping distibutd saerch engies` on the `title` field with a maximum of 3 suggestions per term inside the suggest text. Note that in this example we use the `count` search type. This isn't required, but a nice optimalization. The suggestions are gather in the `query` phase and in the case that we only care about suggestions (so no hits) we don't need to execute the `fetch` phase. ``` curl -s -XPOST 'localhost:9200/_search?search_type=count' -d '{ "suggest" : { "my-title-suggestions-1" : { "text" : "devloping distibutd saerch engies", "fuzzy" : { "size" : 3, "field" : "title" } } } }' ``` The above request could yield the response as stated in the code example below. As you can see if we take the first suggested options of each suggestion entry we get `developing distributed search engines` as result. ``` { ... "suggest": { "my-title-suggestions-1": [ { "text": "devloping", "offset": 0, "length": 9, "options": [ { "text": "developing", "freq": 77, "score": 0.8888889 }, { "text": "deloping", "freq": 1, "score": 0.875 }, { "text": "deploying", "freq": 2, "score": 0.7777778 } ] }, { "text": "distibutd", "offset": 10, "length": 9, "options": [ { "text": "distributed", "freq": 217, "score": 0.7777778 }, { "text": "disributed", "freq": 1, "score": 0.7777778 }, { "text": "distribute", "freq": 1, "score": 0.7777778 } ] }, { "text": "saerch", "offset": 20, "length": 6, "options": [ { "text": "search", "freq": 1038, "score": 0.8333333 }, { "text": "smerch", "freq": 3, "score": 0.8333333 }, { "text": "serch", "freq": 2, "score": 0.8 } ] }, { "text": "engies", "offset": 27, "length": 6, "options": [ { "text": "engines", "freq": 568, "score": 0.8333333 }, { "text": "engles", "freq": 3, "score": 0.8333333 }, { "text": "eggies", "freq": 1, "score": 0.8333333 } ] } ] } ... } ``` # Common suggest options: - `text` - The suggest text. The suggest text is a required option that needs to be set globally or per suggestion. # Common fuzzy suggest options - `field` - The field to fetch the candidate suggestions from. This is an required option that either needs to be set globally or per suggestion. - `analyzer` - The analyzer to analyse the suggest text with. Defaults to the search analyzer of the suggest field. - `size` - The maximum corrections to be returned per suggest text token. - `sort` - Defines how suggestions should be sorted per suggest text term. Two possible value: *\* `score` - Sort by sore first, then document frequency and then the term itself. *\* `frequency` - Sort by document frequency first, then simlarity score and then the term itself. - `suggest_mode` - The suggest mode controls what suggestions are included or controls for what suggest text terms, suggestions should be suggested. Three possible values can be specified: *\* `missing` - Only suggest terms in the suggest text that aren't in the index. This is the default. *\* `popular` - Only suggest suggestions that occur in more docs then the original suggest text term. *\* `always` - Suggest any matching suggestions based on terms in the suggest text. # Other fuzzy suggest options: - `lowercase_terms` - Lower cases the suggest text terms after text analyzation. - `max_edits` - The maximum edit distance candidate suggestions can have in order to be considered as a suggestion. Can only be a value between 1 and 2. Any other value result in an bad request error being thrown. Defaults to 2. - `min_prefix` - The number of minimal prefix characters that must match in order be a candidate suggestions. Defaults to 1. Increasing this number improves spellcheck performance. Usually misspellings don't occur in the beginning of terms. - `min_query_length` - The minimum length a suggest text term must have in order to be included. Defaults to 4. - `shard_size` - Sets the maximum number of suggestions to be retrieved from each individual shard. During the reduce phase only the top N suggestions are returned based on the `size` option. Defaults to the `size` option. Setting this to a value higher than the `size` can be useful in order to get a more accurate document frequency for spelling corrections at the cost of performance. Due to the fact that terms are partitioned amongst shards, the shard level document frequencies of spelling corrections may not be precise. Increasing this will make these document frequencies more precise. - `max_inspections` - A factor that is used to multiply with the `shards_size` in order to inspect more candidate spell corrections on the shard level. Can improve accuracy at the cost of performance. Defaults to 5. - `threshold_frequency` - The minimal threshold in number of documents a suggestion should appear in. This can be specified as an absolute number or as a relative percentage of number of documents. This can improve quality by only suggesting high frequency terms. Defaults to 0f and is not enabled. If a value higher than 1 is specified then the number cannot be fractional. The shard level document frequencies are used for this option. - `max_query_frequency` - The maximum threshold in number of documents a sugges text token can exist in order to be included. Can be a relative percentage number (e.g 0.4) or an absolute number to represent document frequencies. If an value higher than 1 is specified then fractional can not be specified. Defaults to 0.01f. This can be used to exclude high frequency terms from being spellchecked. High frequency terms are usually spelled correctly on top of this this also improves the spellcheck performance. The shard level document frequencies are used for this option.

On Thu, Feb 14, 2013 at 7:49 AM, Ivan Brusic ivan@brusic.com wrote:

Exciting news. I have an idea for a new facet type, but the current
framework is too much work for a feature I do not fully need.

New fuzzy queries means spell checking soon as well?

--
Ivan

On Thu, Feb 14, 2013 at 12:51 AM, Jörg Prante joergprante@gmail.com wrote:

I think the facet framework is going to be overhauled to add custom facet
features more easily. For example, right now, there can't be plugins that
add new capabilities to faceting.

So I hope I can use collation based sorting on facet elements soon, or
facet value mappings, where facet values can be augmented with other values
for display (like sort keys in collations).

Jörg

Am 14.02.13 09:25, schrieb Leonardo Menezes:

Hey Shay,
that's great news. For the new code in faceting, should we also
expect performance improvement, or just memory? thanks

--
You received this message because you are subscribed to the Google Groups
"elasticsearch" group.
To unsubscribe from this group and stop receiving emails from it, send an
email to elasticsearch+unsubscribe@googlegroups.com.
For more options, visit https://groups.google.com/groups/opt_out.

--
You received this message because you are subscribed to the Google Groups
"elasticsearch" group.
To unsubscribe from this group and stop receiving emails from it, send an
email to elasticsearch+unsubscribe@googlegroups.com.
For more options, visit https://groups.google.com/groups/opt_out.

--
You received this message because you are subscribed to the Google Groups "elasticsearch" group.
To unsubscribe from this group and stop receiving emails from it, send an email to elasticsearch+unsubscribe@googlegroups.com.
For more options, visit https://groups.google.com/groups/opt_out.

Topic		Replies	Views
Release with Lucene 4.1.x? Elasticsearch	10	468	January 30, 2013
ElasticSearch 1.0 features? Elasticsearch	6	410	May 8, 2013
Faceting high cardinality fields - idea review Elasticsearch	8	681	October 14, 2012
Performance killed when faceting on high cardinality fields Elasticsearch	25	2934	January 27, 2013
How to improve performance of facet queries? Elasticsearch	6	1443	September 12, 2013

0.21.0.Beta Status (Lucene 4.x)

Related topics