# Analyzer for football scores

**URL:** <https://discuss.elastic.co/t/analyzer-for-football-scores/5691>\
**Category:** Elasticsearch\
**Created:** [October 26, 2011, 2:06pm UTC](https://discuss.elastic.co/t/analyzer-for-football-scores/5691 "2011-10-26T14:06:14Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ridvan\_Gyundogan](https://avatars.discourse-cdn.com/v4/letter/r/e495f1/32.png) [@Ridvan\_Gyundogan](https://discuss.elastic.co/u/Ridvan_Gyundogan)\
**Post date:** [October 26, 2011, 2:06pm UTC](https://discuss.elastic.co/t/analyzer-for-football-scores/5691/1 "2011-10-26T14:06:14Z")

</div>

Hi I know that this might sound funny, but I try to extract football  
scores from text fields.  
For example I have Man United - Man City 1:6. I want to make a term  
facet which groups the documents by scores:  
1 : 1 (339 documents)  
2 : 1 (564 documents)  
...  
....

So I want "1 :1" to be analyzed as a single term. Anyone having idea  
how to do this with the analyzer, tokenizer, filters?

The regular expression for terms would be something like \d\s:\s\d.  
The thing is that the pattern analyzer expects a regular expression  
for the separator not for the term.

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [October 26, 2011, 2:21pm UTC](https://discuss.elastic.co/t/analyzer-for-football-scores/5691/2 "2011-10-26T14:21:40Z")

</div>

Hi Ridvan

> So I want "1 :1" to be analyzed as a single term. Anyone having idea  
> how to do this with the analyzer, tokenizer, filters?
> 
> The regular expression for terms would be something like \d\s:\s\d.  
> The thing is that the pattern analyzer expects a regular expression  
> for the separator not for the term.

You can build a custom analyzer using the pattern TOKENIZER, which  
allows you to specify a 'group' number, so that you can capture tokens,  
instead of matching on separators

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

clint

---

<div class="post-metadata">

**Author:** ![Ridvan\_Gyundogan](https://avatars.discourse-cdn.com/v4/letter/r/e495f1/32.png) [@Ridvan\_Gyundogan](https://discuss.elastic.co/u/Ridvan_Gyundogan)\
**Post date:** [October 26, 2011, 3:16pm UTC](https://discuss.elastic.co/t/analyzer-for-football-scores/5691/3 "2011-10-26T15:16:58Z")

</div>

Hi Clint, thanks for the answer.

I think I get it, but what confuses me is the comment at the bottom,  
of the link you provided:  
"IMPORTANT: The regular expression should match the token separators,  
not the tokens themselves."

On the other side if I look at the following example from Shay :

> <https://github.com/elastic/elasticsearch/issues/928>
>
> Pattern tokenizer allows to define a tokenizer that uses regex to break text int…o tokens. The \`pattern\` parameter accepts the regex expression (and flags the common ES level regex flags).
> 
> It also accepts \`group\` (defaults to -1), from teh docs:
> 
> group=-1 (the default) is equivalent to "split". In this case, the tokens will be equivalent to the output from (without empty tokens):String#split(java.lang.String)
> 
> Using group \>= 0 selects the matching group as the token. For example, if you have:
> 
> \`\`\`
> pattern = \\'(\[^\\'\]+)\\'
> group = 0
> 
> input = aaa 'bbb' 'ccc'
> \`\`\`
> 
> the output will be two tokens: 'bbb' and 'ccc' (including the ' marks). With the same input but using group=1, the output would be: bbb and ccc (no ' marks).

It looks like the pattern is exactly for the tokens, not for the  
separators?

On Oct 26, 5:21 pm, Clinton Gormley [cl...@traveljury.com](mailto:cl...@traveljury.com) wrote:

> Hi Ridvan
> 
> > So I want "1 :1" to be analyzed as a single term. Anyone having idea  
> > how to do this with the analyzer, tokenizer, filters?
> 
> > The regular expression for terms would be something like \d\s:\s\d.  
> > The thing is that the pattern analyzer expects a regular expression  
> > for the separator not for the term.
> 
> You can build a custom analyzer using the pattern TOKENIZER, which  
> allows you to specify a 'group' number, so that you can capture tokens,  
> instead of matching on separators
> 
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/index-modules/analysis/p)...
> 
> clint

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [October 26, 2011, 3:25pm UTC](https://discuss.elastic.co/t/analyzer-for-football-scores/5691/4 "2011-10-26T15:25:54Z")

</div>

Hi Ridavan

> I think I get it, but what confuses me is the comment at the bottom,  
> of the link you provided:  
> "IMPORTANT: The regular expression should match the token separators,  
> not the tokens themselves."

I think that's just a bad copy-paste from the analyzer docs.

> On the other side if I look at the following example from Shay :  
> [Analysis: Pattern Tokenizer · Issue #928 · elastic/elasticsearch · GitHub](https://github.com/elasticsearch/elasticsearch/issues/928)  
> It looks like the pattern is exactly for the tokens, not for the  
> separators?

As it says, if "group" is -1 then it acts as a 'split' on the regex (ie  
your regex should match the token separators), but if group is \> 0 then  
it returns what is matched. For example:

For text 'foobar' and regex /(o(b))a/

## group: tokens:

-1 fo,r  
0 oba  
1 ob  
2 b

clint

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:50am UTC](https://discuss.elastic.co/t/analyzer-for-football-scores/5691/5 "2017-07-06T03:50:52Z")

</div>


