# Avoid stemming of Acronyms?

**URL:** <https://discuss.elastic.co/t/avoid-stemming-of-acronyms/28464>\
**Category:** Elasticsearch\
**Created:** [September 1, 2015, 7:08pm UTC](https://discuss.elastic.co/t/avoid-stemming-of-acronyms/28464 "2015-09-01T19:08:51Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![apanimesh061](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/apanimesh061/32/927_2.png) [@apanimesh061](https://discuss.elastic.co/u/apanimesh061)\
**Post date:** [September 1, 2015, 7:08pm UTC](https://discuss.elastic.co/t/avoid-stemming-of-acronyms/28464/1 "2015-09-01T19:08:52Z")

</div>

I am using the `pattern_capture` filter to preserve all the acronyms

```
PUT test_index/_settings
{
  "index.analysis.filter": {
    "acronym_en_EN": {
      "type": "pattern_capture",
      "patterns": [
        "(?:[a-zA-Z]\\.)+", 
        "((?:[a-zA-Z]\\.)+[a-zA-Z])",
        "((?:[a-zA-Z]\\.)+[s]$)",
        "((?:[a-zA-Z]\\.)+[s][\\.]$)"
        ],
      "preserve_original": true
    }
  }
}

```

But i noticed that acronyms that end with `s` or `s.` are stemmed as there is one stemmer filter also attached to the analyzer. The regular expressions in the filter above for handling `s` are also not working.

I test the output using this

```
GET test_index/_analyze?tokenizer=standard&filters=lowercase,acronym_en_EN,apostrophe,porter_stemmer_en_EN&text=u.s.a. u.s. s.w.a.t u.t. 

```

this gives me

```
{
   "tokens": [
      {
         "token": "u.s.a",
         "start_offset": 0,
         "end_offset": 5,
         "type": "<ALPHANUM>",
         "position": 1
      },
      {
         "token": "u.",
         "start_offset": 7,
         "end_offset": 10,
         "type": "<ALPHANUM>",
         "position": 2
      },
      {
         "token": "u.",
         "start_offset": 7,
         "end_offset": 10,
         "type": "<ALPHANUM>",
         "position": 2
      },
      {
         "token": "s.w.a.t",
         "start_offset": 12,
         "end_offset": 19,
         "type": "<ALPHANUM>",
         "position": 3
      },
      {
         "token": "u.t",
         "start_offset": 20,
         "end_offset": 23,
         "type": "<ALPHANUM>",
         "position": 4
      }
   ]
}

```

Is there any way I can preserve the acronyms ending with `s` so that for `u.s.` or `u.s` I don't get `u.`?

---

<div class="post-metadata">

**Author:** ![n0othing](https://avatars.discourse-cdn.com/v4/letter/n/b2d939/32.png) [@n0othing](https://discuss.elastic.co/u/n0othing)\
**Post date:** [September 2, 2015, 1:32am UTC](https://discuss.elastic.co/t/avoid-stemming-of-acronyms/28464/2 "2015-09-02T01:32:13Z")

</div>

It looks like the porter stemmer is doing that (though unsure why at the moment). When not used, the acronyms come out as you expect. I was able to use the [keyword marker token filter](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-keyword-marker-tokenfilter.html%22keyword%20marker%20filter%22) to get your current setup to work.

```
PUT test_index
{
  "index.analysis.filter": {
    "acronym_en_EN": {
      "type": "pattern_capture",
      "patterns": [
        "(?:[a-zA-Z]\\.)+", 
        "((?:[a-zA-Z]\\.)+[a-zA-Z])",
        "((?:[a-zA-Z]\\.)+[s]$)",
        "((?:[a-zA-Z]\\.)+[s][\\.]$)"
        ],
      "preserve_original": true
    },
    "porter_stemmer_en_EN" : {
      "type" : "stemmer",
      "name" : "english"
    },
    "no_stem": {
          "type": "keyword_marker",
          "keywords": ["u.s"] 
        }
  }

```

then

```
GET test_index/_analyze?tokenizer=standard&filters=lowercase,acronym_en_EN,apostrophe,no_stem,porter_stemmer_en_EN&text=u.s.a. u.s. s.w.a.t u.t.

```

results in

```
{
   "tokens": [
      {
         "token": "u.s.a",
         "start_offset": 0,
         "end_offset": 5,
         "type": "<ALPHANUM>",
         "position": 1
      },
      {
         "token": "u.s",
         "start_offset": 7,
         "end_offset": 10,
         "type": "<ALPHANUM>",
         "position": 2
      },
      {
         "token": "s.w.a.t",
         "start_offset": 12,
         "end_offset": 19,
         "type": "<ALPHANUM>",
         "position": 3
      },
      {
         "token": "u.t",
         "start_offset": 20,
         "end_offset": 23,
         "type": "<ALPHANUM>",
         "position": 4
      }
   ]
}

```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:52pm UTC](https://discuss.elastic.co/t/avoid-stemming-of-acronyms/28464/3 "2017-07-05T23:52:46Z")

</div>


