# Why does “Richard John” fail to match “Richard Johns” despite only 1 edit distance?

**URL:** <https://discuss.elastic.co/t/why-does-richard-john-fail-to-match-richard-johns-despite-only-1-edit-distance/379860>\
**Category:** Elastic Search\
**Created:** [July 7, 2025, 10:24am UTC](https://discuss.elastic.co/t/why-does-richard-john-fail-to-match-richard-johns-despite-only-1-edit-distance/379860 "2025-07-07T10:24:35Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![yap\_waiyen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yap_waiyen/32/144026_2.png) [@yap\_waiyen](https://discuss.elastic.co/u/yap_waiyen)\
**Post date:** [July 7, 2025, 10:24am UTC](https://discuss.elastic.co/t/why-does-richard-john-fail-to-match-richard-johns-despite-only-1-edit-distance/379860/1 "2025-07-07T10:24:35Z")

</div>

Hi all,  
I’m working on a name matching system using Elasticsearch with fuzziness enabled. I encountered something confusing and would really appreciate some clarification.

#### 🔍 Problem:

- When I input **"Richard John"** , I expected it to match **"Richard Johns"** (just 1 character difference).
- But **it does not return a match**.
- However, when I input **"Elvire Aide"** , it **does match** with **"Elvire Ade"** , which is also just a 1-character difference.

Both scenarios appear to involve a 1-edit distance — so why is one working and not the other?

I’m using a `match` query with fuzziness and token count filtering like so:

```auto
"my_search_analyzer": {
  "tokenizer": "standard",
  "filter": [
    "lowercase",
    "ascii_folding",
    "custom_synonym_unique",
    "stophrase_synonymphrase",
    "custom_stop",
    "custom_synonym_multi"
  ],
  "char_filter": [
    "wordbreaker_filter",
    "punctuation_filter"
  ]
}

```

```auto
{
  "query": {
    "bool": {
      "must": [
        {
          "match": {
            "name_field": {
              "query": "Richard John",
              "analyzer": "my_search_analyzer",
              "fuzziness": "AUTO:1,6",
              "prefix_length": 1,
              "minimum_should_match": "1<-50% 5<75%",
              "operator": "AND"
            }
          }
        }
      ],
      "filter": [
        {
          "term": {
            "name_token_count": 2
          }
        }
      ]
    }
  }
}

```

### Example:

- Indexed name: **"Richard Johns"**
- Search input: **"Richard John"**

These differ by just 1 character (`s`), but no match is returned. However:

- "Elvire Aide" correctly matches "Elvire Ade" ✅
- But "Elvire Adie" does **not** match "Elvire Ade" ❌
- "ADIK EMPIRE" does not match "ADK EMPIRE" either ❌

### My understanding:

- Elasticsearch applies fuzziness **per token**.
- Token count filtering (`name_token_count == 2`) ensures structural matching, but may prevent slightly longer/shorter names from appearing.

#### 🧠 My Questions:

1. Why is **"Richard John" ❌ Richard Johns** , while **"Elvire Aide" ✅ Elvire Ade** , when both differ by only one character?
2. Does Elasticsearch **only apply fuzziness within tokens** , and not across?
3. Is it because of the **token count filter** — and how can I allow near-miss hits like this while still filtering token count?
4. Would using a `script_score`, or per-token `fuzzy` queries inside a `bool`, be a better approach for this type of multi-token fuzzy matching?

Thanks in advance for any suggestions or clarifications!

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 7, 2025, 10:41am UTC](https://discuss.elastic.co/t/why-does-richard-john-fail-to-match-richard-johns-despite-only-1-edit-distance/379860/2 "2025-07-07T10:41:57Z")

</div>

Welcome!

Could you provide a full recreation script as described in [About the Elasticsearch category](https://discuss.elastic.co/t/about-the-elasticsearch-category/21). It will help to better understand what you are doing. Please, try to keep the example as simple as possible.

A full reproduction script is something anyone can **copy and paste in Kibana dev console** , click on the run button to reproduce your use case. It will help readers to understand, reproduce and if needed fix your problem. It will also most likely help to get a faster answer.

Have a look at the [Elastic Stack and Solutions Help · Forums and Slack | Elastic](https://elastic.co/community/help) page. It contains also lot of useful information on how to ask for help.

---

<div class="post-metadata">

**Author:** ![piotrprz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/piotrprz/32/128141_2.png) [@piotrprz](https://discuss.elastic.co/u/piotrprz)\
**Post date:** [July 7, 2025, 11:12am UTC](https://discuss.elastic.co/t/why-does-richard-john-fail-to-match-richard-johns-despite-only-1-edit-distance/379860/3 "2025-07-07T11:12:56Z")

</div>

Hello @yap_waiyen

AFAICT Fuzziness is applied **one token at a time** , but Elasticsearch will only keep the first `fuzzy_max_expansions` ([default = 50](https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-query-string-query)) candidate spellings it finds for each token.  
The token **“john”** has far more than 50 one-edit neighbours in a name index, so _“johns”_ falls outside that cut-off and never makes the query, whereas **“aide”** has only a handful of neighbours, so _“ade”_ is still included and you get the hit. In other words, the miss is caused by the expansion limit, not by the token-count filter or any cross-token rule.

Perhaps you could raise `fuzzy_max_expansions` (or add a simple plural-stripping synonym, or switch to a phonetic/approx-string plugin) and _Richard Johns_ should then match _Richard John_ too.
