# Hunspell russian language. Problem with some grammar cases

**URL:** <https://discuss.elastic.co/t/hunspell-russian-language-problem-with-some-grammar-cases/307575>\
**Category:** Elasticsearch\
**Tags:** docker\
**Created:** [June 18, 2022, 5:16pm UTC](https://discuss.elastic.co/t/hunspell-russian-language-problem-with-some-grammar-cases/307575 "2022-06-18T17:16:36Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![GrigoriyKrasovskiy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/grigoriykrasovskiy/32/103440_2.png) [@GrigoriyKrasovskiy](https://discuss.elastic.co/u/GrigoriyKrasovskiy)\
**Post date:** [June 18, 2022, 5:16pm UTC](https://discuss.elastic.co/t/hunspell-russian-language-problem-with-some-grammar-cases/307575/1 "2022-06-18T17:16:36Z")

</div>

I use Elasticsearch 7.17.4 with docker and hunspell for russian language.  
My settings for index analysis:

```auto
"analysis": {
        "filter": {
          "my_stemmer": {
            "type": "stemmer",
            "language": "russian"
          },
          "ru_RU": {
            "locale": "ru_RU",
            "type": "hunspell"
          }
        },
        "analyzer": {
          "custom_analyzer": {
            "filter": [
              "lowercase",
              "ru_RU",
              "my_stemmer"
            ],
            "char_filter": [
              "html_strip"
            ],
            "tokenizer": "standard"
          }
        }
      }

```

Unfortunately FTS does not work properly with all words: for example I have the following entries:  
новый колодец  
нового колодца  
новому колодцу  
новым колодцем  
новом колодце

When I make the following request:

```auto
GET http://localhost:9200/ingredient/_search?pretty
Content-Type: application/json

{
  "query": {
    "query_string": {
      "query": "колодец",
      "default_field": "name"
    }
  }
}

```

I get the following result:

```auto
{
  "took": 1,
  "timed_out": false,
  "_shards": {
    "total": 1,
    "successful": 1,
    "skipped": 0,
    "failed": 0
  },
  "hits": {
    "total": {
      "value": 4,
      "relation": "eq"
    },
    "max_score": 1.8472799,
    "hits": [
      {
        "_index": "ingredient",
        "_type": "_doc",
        "_id": "35",
        "_score": 1.8472799,
        "_source": {
          "name": "новый колодец",
          "id": 35,
          "_meta": {}
        }
      },
      {
        "_index": "ingredient",
        "_type": "_doc",
        "_id": "36",
        "_score": 1.8472799,
        "_source": {
          "name": "нового колодца",
          "id": 36,
          "_meta": {}
        }
      },
      {
        "_index": "ingredient",
        "_type": "_doc",
        "_id": "37",
        "_score": 1.8472799,
        "_source": {
          "name": "новому колодцу",
          "id": 37,
          "_meta": {}
        }
      },
      {
        "_index": "ingredient",
        "_type": "_doc",
        "_id": "39",
        "_score": 1.8472799,
        "_source": {
          "name": "новом колодце",
          "id": 39,
          "_meta": {}
        }
      }
    ]
  }
}

Response code: 200 (OK); Time: 65ms; Content length: 1235 bytes

```

I never get "новым колодцем".  
My question is: Is it the problem with my hunspell? Should I find any version of it with more data? Or is it a problem with my settings? Maybe I missed something?

---

<div class="post-metadata">

**Author:** ![RabBit\_BR](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rabbit_br/32/82261_2.png) [@RabBit\_BR](https://discuss.elastic.co/u/RabBit_BR)\
**Post date:** [June 20, 2022, 2:13am UTC](https://discuss.elastic.co/t/hunspell-russian-language-problem-with-some-grammar-cases/307575/2 "2022-06-20T02:13:36Z")

</div>

Hi @GrigoriyKrasovskiy

I would track the tokens generated by the custom\_analyzer for each input term. That way, you would be sure that it would be generating the expected tokens.

Anyway I saw this attempt:

> If available, we recommend trying an algorithmic stemmer for your language before using the hunspell token filter. In practice, algorithmic stemmers often outperform dictionary stemmers. See dictionary stemmers.

Have you tried using fuzzines to get the fifth doc in your answers?

```auto
{
  "query": {
    "query_string": {
      "query": "колодец~",
      "default_field": "name",
      "fuzziness": 1
    }
  }
}

```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 18, 2022, 2:13am UTC](https://discuss.elastic.co/t/hunspell-russian-language-problem-with-some-grammar-cases/307575/3 "2022-07-18T02:13:45Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
