# Using shingle

**URL:** <https://discuss.elastic.co/t/using-shingle/22294>\
**Category:** Elasticsearch\
**Created:** [February 20, 2015, 2:29pm UTC](https://discuss.elastic.co/t/using-shingle/22294 "2015-02-20T14:29:15Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Petr\_Jansky](https://avatars.discourse-cdn.com/v4/letter/p/ec9cab/32.png) [@Petr\_Jansky](https://discuss.elastic.co/u/Petr_Jansky)\
**Post date:** [February 20, 2015, 2:29pm UTC](https://discuss.elastic.co/t/using-shingle/22294/1 "2015-02-20T14:29:15Z")

</div>

Hi there,

I've tried to use shingle for getting bigrams and trigrams

curl -X POST 'localhost:9200/idnes/' -d '{  
"settings" : {  
"analysis" : {  
"filter": {  
"czech\_stop": {  
"type": "stop",  
"stopwords": "_czech_",  
"ignore\_case": "true",  
"remove\_trailing": "false"  
},  
"czech\_stop\_ngram": {  
"type": "stop",  
"stopwords" : ["a", "i", "k", "o", "s", "u", "v", "z", "do",  
"co", "by", "do", "je", "mu", "mi", "mě", "mně", "mne", "na", "ne", "ní,  
"si", "se", "ta", "to", "té", "ti", "ty", "už", "ve", "za", "že", "aby",  
"ani", "ale", "byl", "jak", "jen", "jde", "kdo", "kdy", "kde", "něm",  
"nich", "něj", "než", "pro", "tak", "ten", "tam", "tady", "těch", "jsou",  
"jsem", "není", "nyní", "nimi", "jako", "jaká", "jaké", "jaká", "právě",  
"který", "která", "které", "jeho", "její", "nebo", "jako", "toho", "kdyby",  
"takový", "taková", "takové", "_czech_" ],  
"ignore\_case": "true",  
"remove\_trailing": "false"  
},  
"czech\_keywords": {  
"type": "keyword\_marker",  
"keywords": ["že"]  
},  
"czech\_stemmer": {  
"type": "stemmer",  
"language": "czech"  
},  
"shingle2\_filter": {  
"type": "shingle",  
"min\_shingle\_size": 2,  
"max\_shingle\_size": 2,  
"output\_unigrams": false  
},  
"shingle3\_filter": {  
"type": "shingle",  
"min\_shingle\_size": 3,  
"max\_shingle\_size": 3,  
\*"output\_unigrams": false \*  
}  
},  
"analyzer": {  
....  
"shingle2s\_analyzer": {  
"type": "custom",  
"tokenizer": "standard",  
"filter": ["standard", "lowercase", "czech\_stop\_ngram",  
"shingle2\_filter"]  
},  
"shingle3s\_analyzer": {  
"type": "custom",  
"tokenizer": "standard",  
"filter": ["czech\_stop\_ngram", "shingle3\_filter"]  
}  
}  
}  
},

"mappings" : {  
"article" : {  
"\_id" : {  
"path" : "reference"  
},

```
"properties" : {
    .....
    "content2" : { "type":"string", "analyzer": "shingle2_analyzer"},
    "content3" : { "type":"string", "analyzer": "shingle3_analyzer"},
    "content4" : { "type":"string", "analyzer": "shingle2s_analyzer"},
    "content5" : { "type":"string", "analyzer": "shingle3s_analyzer"},
    ......

```

If I try my analysers using by calling:

curl -X GET  
'localhost:9200/idnes/\_analyze?analyzer=shingle3s\_analyzer&pretty' -d 'a e  
i o u s k z na ke ze nad pod za před Norská strana zatím dostatečně  
nevyhodnotila, jak citlivou otázkou je pro Česko případ synů Evy  
Michalákové. Tak popisuje současnou situaci premiér Bohuslav Sobotka. Ten  
již dostal odpověď na dopis od premiérky Norska Erny Solbergové. S obecnými  
odpověďmi není spokojen a zvažuje do Norska další psaní.' | grep "token"

It works fine. In results there are only trigrams  
"tokens" : [ {  
"token" : "\_ e _",  
"token" : "e \_ ",  
"token" : " \_ Norská",  
"token" : "_ Norská _",  
"token" : "Norská \_ zatím",  
"token" : "_ zatím dostatečně",  
"token" : "zatím dostatečně nevyhodnotila",  
"token" : "dostatečně nevyhodnotila _",  
"token" : "nevyhodnotila \_ citlivou",  
"token" : "_ citlivou otázkou",  
"token" : "citlivou otázkou \_",  
"token" : "otázkou \_ \_",  
....

But there is an issue if I use it on indexed data  
POST idnes/\_search?pretty=true  
{  
"query": {  
"match": {  
"content\_type": "Article"  
}  
},  
"facets" : {  
"tag" : {  
"terms" : {  
"fields" : ["content5"],  
"size" : 20  
}  
}  
}  
}

In the response there are also unigrams.  
"facets": {  
"tag": {  
"\_type": "terms",  
"missing": 452,  
"total": 926077,  
"other": 762645,  
"terms": [  
{  
"term": "a",  
"count": 18150  
},  
{  
"term": "to",  
"count": 17131  
},  
{  
"term": "je",  
"count": 14090  
},  
{  
"term": "se",  
"count": 13621  
},  
{  
"term": "na",  
"count": 12285  
},  
......  
{  
"term": "korun \_ _",  
"count": 551  
},  
{  
"term": "_ \_ případě",  
"count": 499  
},  
{  
"term": "zobrazení videa musíte",  
"count": 449  
}  
.....

1. Why does it happen?
2. Is there any other way how to skip "\_" from stopword than [http://www.elasticsearch.org/blog/searching-with-shingles/](http://www.elasticsearch.org/blog/searching-with-shingles/)  
that doesn't work for Lucene 4.4+?

Thanks  
Petr

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/0d2aa0fb-2a12-404d-bdf4-bb09b970cb5c%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/0d2aa0fb-2a12-404d-bdf4-bb09b970cb5c%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Petr\_Jansky](https://avatars.discourse-cdn.com/v4/letter/p/ec9cab/32.png) [@Petr\_Jansky](https://discuss.elastic.co/u/Petr_Jansky)\
**Post date:** [March 17, 2015, 2:00pm UTC](https://discuss.elastic.co/t/using-shingle/22294/2 "2015-03-17T14:00:10Z")

</div>

Noone? ☹

Petr

Dne pátek 20. února 2015 15:29:15 UTC+1 Petr Janský napsal(a):

> Hi there,
> 
> I've tried to use shingle for getting bigrams and trigrams
> 
> curl -X POST 'localhost:9200/idnes/' -d '{  
> "settings" : {  
> "analysis" : {  
> "filter": {  
> "czech\_stop": {  
> "type": "stop",  
> "stopwords": "_czech_",  
> "ignore\_case": "true",  
> "remove\_trailing": "false"  
> },  
> "czech\_stop\_ngram": {  
> "type": "stop",  
> "stopwords" : ["a", "i", "k", "o", "s", "u", "v", "z", "do",  
> "co", "by", "do", "je", "mu", "mi", "mě", "mně", "mne", "na", "ne", "ní,  
> "si", "se", "ta", "to", "té", "ti", "ty", "už", "ve", "za", "že", "aby",  
> "ani", "ale", "byl", "jak", "jen", "jde", "kdo", "kdy", "kde", "něm",  
> "nich", "něj", "než", "pro", "tak", "ten", "tam", "tady", "těch", "jsou",  
> "jsem", "není", "nyní", "nimi", "jako", "jaká", "jaké", "jaká", "právě",  
> "který", "která", "které", "jeho", "její", "nebo", "jako", "toho", "kdyby",  
> "takový", "taková", "takové", "_czech_" ],  
> "ignore\_case": "true",  
> "remove\_trailing": "false"  
> },  
> "czech\_keywords": {  
> "type": "keyword\_marker",  
> "keywords": ["že"]  
> },  
> "czech\_stemmer": {  
> "type": "stemmer",  
> "language": "czech"  
> },  
> "shingle2\_filter": {  
> "type": "shingle",  
> "min\_shingle\_size": 2,  
> "max\_shingle\_size": 2,  
> "output\_unigrams": false  
> },  
> "shingle3\_filter": {  
> "type": "shingle",  
> "min\_shingle\_size": 3,  
> "max\_shingle\_size": 3,  
> \*"output\_unigrams": false \*  
> }  
> },  
> "analyzer": {  
> ....  
> "shingle2s\_analyzer": {  
> "type": "custom",  
> "tokenizer": "standard",  
> "filter": ["standard", "lowercase", "czech\_stop\_ngram",  
> "shingle2\_filter"]  
> },  
> "shingle3s\_analyzer": {  
> "type": "custom",  
> "tokenizer": "standard",  
> "filter": ["czech\_stop\_ngram", "shingle3\_filter"]  
> }  
> }  
> }  
> },
> 
> "mappings" : {  
> "article" : {  
> "\_id" : {  
> "path" : "reference"  
> },
> 
> ```
> "properties" : {
> .....
> "content2" : { "type":"string", "analyzer": "shingle2_analyzer"},
> "content3" : { "type":"string", "analyzer": "shingle3_analyzer"},
> "content4" : { "type":"string", "analyzer": 
> 
> ```
> 
> "shingle2s\_analyzer"},  
> "content5" : { "type":"string", "analyzer":  
> "shingle3s\_analyzer"},  
> ......
> 
> If I try my analysers using by calling:
> 
> curl -X GET  
> 'localhost:9200/idnes/\_analyze?analyzer=shingle3s\_analyzer&pretty' -d 'a e  
> i o u s k z na ke ze nad pod za před Norská strana zatím dostatečně  
> nevyhodnotila, jak citlivou otázkou je pro Česko případ synů Evy  
> Michalákové. Tak popisuje současnou situaci premiér Bohuslav Sobotka. Ten  
> již dostal odpověď na dopis od premiérky Norska Erny Solbergové. S obecnými  
> odpověďmi není spokojen a zvažuje do Norska další psaní.' | grep "token"
> 
> It works fine. In results there are only trigrams  
> "tokens" : [ {  
> "token" : "\_ e _",  
> "token" : "e \_ ",  
> "token" : " \_ Norská",  
> "token" : "_ Norská _",  
> "token" : "Norská \_ zatím",  
> "token" : "_ zatím dostatečně",  
> "token" : "zatím dostatečně nevyhodnotila",  
> "token" : "dostatečně nevyhodnotila _",  
> "token" : "nevyhodnotila \_ citlivou",  
> "token" : "_ citlivou otázkou",  
> "token" : "citlivou otázkou \_",  
> "token" : "otázkou \_ \_",  
> ....
> 
> But there is an issue if I use it on indexed data  
> POST idnes/\_search?pretty=true  
> {  
> "query": {  
> "match": {  
> "content\_type": "Article"  
> }  
> },  
> "facets" : {  
> "tag" : {  
> "terms" : {  
> "fields" : ["content5"],  
> "size" : 20  
> }  
> }  
> }  
> }
> 
> In the response there are also unigrams.  
> "facets": {  
> "tag": {  
> "\_type": "terms",  
> "missing": 452,  
> "total": 926077,  
> "other": 762645,  
> "terms": [  
> {  
> "term": "a",  
> "count": 18150  
> },  
> {  
> "term": "to",  
> "count": 17131  
> },  
> {  
> "term": "je",  
> "count": 14090  
> },  
> {  
> "term": "se",  
> "count": 13621  
> },  
> {  
> "term": "na",  
> "count": 12285  
> },  
> ......  
> {  
> "term": "korun \_ _",  
> "count": 551  
> },  
> {  
> "term": "_ \_ případě",  
> "count": 499  
> },  
> {  
> "term": "zobrazení videa musíte",  
> "count": 449  
> }  
> .....
> 
> 1. Why does it happen?
> 2. Is there any other way how to skip "\_" from stopword than  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/blog/searching-with-shingles/) that  
> doesn't work for Lucene 4.4+?
> 
> Thanks  
> Petr

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/378228d7-3d93-4248-9728-2d441ecace91%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/378228d7-3d93-4248-9728-2d441ecace91%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:26am UTC](https://discuss.elastic.co/t/using-shingle/22294/3 "2017-07-06T00:26:14Z")

</div>


