# Multiple analyzers with stemmed synonyms

**URL:** <https://discuss.elastic.co/t/multiple-analyzers-with-stemmed-synonyms/237317>\
**Category:** Elasticsearch\
**Created:** [June 16, 2020, 2:27pm UTC](https://discuss.elastic.co/t/multiple-analyzers-with-stemmed-synonyms/237317 "2020-06-16T14:27:28Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Kushikawa](https://avatars.discourse-cdn.com/v4/letter/k/8baadc/32.png) [@Kushikawa](https://discuss.elastic.co/u/Kushikawa)\
**Post date:** [June 16, 2020, 2:27pm UTC](https://discuss.elastic.co/t/multiple-analyzers-with-stemmed-synonyms/237317/1 "2020-06-16T14:27:28Z")

</div>

I want to create an index with a stemmer analyzer to generalize my synonyms and apply it in other analyzers.  
For a simplified example: I want to use all these synonyms `[beautiful, pretty, beauteous, gorgeous]` in multiple analyzers when searching for `beauty`, once beauty and beautiful have the same stem word

```auto
GET /_analyzer
{
  "tokenizer": "standard",
  "filter": ["stemmer"],
  "text": "beautiful beauty"
}
{
  "tokens": [
    {
      "token": "beauti", ...
    },
    {
      "token": "beauti", ...
    }
  ]
}

```

What I have so far is

```auto
PUT /test_synonyms
{
  "settings": {
    "index": {
      "analysis": {
        "filter": {
          "my_synonyms": {
            "type": "synonym",
            "synonyms": ["beautiful, pretty, beauteous, gorgeous"]
          },
          "my_metaphone": {
            "type": "phonetic",
            "encoder": "metaphone",
            "replace": true
          }
        }
      }
    }
  }
}

```

Stemmer analyzer gives me:

```auto
GET /test_synonyms/_analyzer
{
  "tokenizer": "standard",
  "filter": ["stemmer", "my_synonyms"],
  "text": "beauty"
}
{
  "tokens": [
    {
      "token": "beauti", ...
    },
    {
      "token": "pretti", ...
    },
    {
      "token": "beauteo", ...
    },
    {
      "token": "gorgeou", ...
    }
  ]
}

```

Phonetic analyzer gives me:

```auto
GET /test_synonyms/_analyzer
{
  "tokenizer": "standard",
  "filter": ["my_synonyms", "my_metaphone"],
  "text": "beauty"
}
{
  "tokens": [
    {
      "token": "BT", ...
    }
  ]
}

```

Once "BT" doesn't match with any of the tokens:

```auto
GET /test_synonyms/_analyzer
{
  "tokenizer": "standard",
  "filter": ["my_synonyms", "my_phonetic"],
  "text": "beautiful"
}
{
  "tokens": [
    {
      "token": "BTFL", ... /*beautiful*/
    },
    {
      "token": "PRT", ... /*pretty*/
    },
    {
      "token": "BTS", ... /*beauteous*/
    },
    {
      "token": "KRJS", ... /*gorgeous*/
    }
  ]
}

```

I was wondering if there is a way to return the exact synonym words (not their stem), but still use stemmer to find them, and then use this with other analyzers.. Something to give me the response above when searching for `beauty`

I tried to use the stemmer and phonetic filters together, but it gives me:

```auto
GET /test_synonyms/_analyzer
{
  "tokenizer": "standard",
  "filter": ["stemmer", "my_synonyms", "my_phonetic"],
  "text": "beauty" /*or beautiful (equal responses)*/
}
{
  "tokens": [
    {
      "token": "BT", ... /*beauti*/
    },
    {
      "token": "PRT", ... /*pretti*/
    },
    {
      "token": "BT", ... /*beauteou*/
    },
    {
      "token": "KRJ", ... /*gorgeou*/
    }
  ]
}

```

And this isn't what I really want, cuz when I search for "beautiful" and "beauty", the number of documents returned are differents (beautiful score the phonetic matches), and I want them to be the same.

---

<div class="post-metadata">

**Author:** ![cbuescher](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cbuescher/32/60402_2.png) [@cbuescher](https://discuss.elastic.co/u/cbuescher)\
**Post date:** [June 16, 2020, 5:46pm UTC](https://discuss.elastic.co/t/multiple-analyzers-with-stemmed-synonyms/237317/2 "2020-06-16T17:46:48Z")

</div>

> [@Kushikawa](#):
>
> return the exact synonym words (not their stem)

I don't understand why you would you want to do that? If you index your documents using the stemmer, docs with "beauteous" in the input will have the stemmed version written to the index. When you search them later e.g. via synonym expansion you want the same stemmer being aplied to them, otherwise you will not match the intended documents.

Specifically:

"a gorgeous boat" will index "gorgeou" when using a stemmer.  
"beauty" at search time will expand to "gorgeou", otherwise it wouldn't match the document

Am I missing something?

---

<div class="post-metadata">

**Author:** ![Kushikawa](https://avatars.discourse-cdn.com/v4/letter/k/8baadc/32.png) [@Kushikawa](https://discuss.elastic.co/u/Kushikawa)\
**Post date:** [June 17, 2020, 12:48am UTC](https://discuss.elastic.co/t/multiple-analyzers-with-stemmed-synonyms/237317/3 "2020-06-17T00:48:06Z")

</div>

> [@cbuescher](#):
>
> I don't understand why you would you want to do that?

Hi @cbuescher, thank you for your reply. I'm sorry, my final goal was not as simple as I made it look. I'm new to elastic and my problem is related to specific Portuguese cases. I updated my question! Please let me know if it makes a bit more sense now or if I'm going in the wrong direction.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 15, 2020, 12:48am UTC](https://discuss.elastic.co/t/multiple-analyzers-with-stemmed-synonyms/237317/4 "2020-07-15T00:48:14Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
