# Single letter term matches cause noise in search results

**URL:** <https://discuss.elastic.co/t/single-letter-term-matches-cause-noise-in-search-results/45698>\
**Category:** Elasticsearch\
**Created:** [March 29, 2016, 2:40pm UTC](https://discuss.elastic.co/t/single-letter-term-matches-cause-noise-in-search-results/45698 "2016-03-29T14:40:19Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![jillesvangurp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jillesvangurp/32/3863_2.png) [@jillesvangurp](https://discuss.elastic.co/u/jillesvangurp)\
**Post date:** [March 29, 2016, 2:40pm UTC](https://discuss.elastic.co/t/single-letter-term-matches-cause-noise-in-search-results/45698/1 "2016-03-29T14:40:19Z")

</div>

I'm trying to implement name search using match, match\_phrase and match\_prefix. My data is messy and spread across name fields, first\_name, last\_name, etc.

For the sake of simplicity, I've been trying to narrow down the problem and it seems that one letter names (i.e. initials) are causing me headaches. For example consider these three names:  
"Jilles van Gurp", "Ali G", "G." indexed into the name.value field.

I've simplified everything to the point where I'm using default everything on es 1.7.3 (analyzer, etc.) and the following query

```
GET /tst/_search
{
  "query": {
    "match": {
      "name.value": {
        "query": "jilles g"
      }
    }
  }
}

```

A query like "jilles" will work as expected. However, as soon as I add the letter g "jilles g", it all goes sideways. and "G." ends up on top. It seems it considers this a full token match on G. and that makes that the most important result despite also having a full token match on jilles. However, from my point of view it is actually the weakest result because it does not actually match most of the query. Phrase prefix match does not produce any results here. What would be a good query + analyzer strategy to make this work as expected that does not introduce too many false positives and actually prefers the "jilles van gurp" over "g" for the query "jilles g"?

For reference, I'm indexing contact data and this data is messy. When asked their name, people sometimes just fill in a single letter. So, I have to deal with it one way or another.

---

<div class="post-metadata">

**Author:** ![xavierfacq](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xavierfacq/32/8744_2.png) [@xavierfacq](https://discuss.elastic.co/u/xavierfacq)\
**Post date:** [March 29, 2016, 3:10pm UTC](https://discuss.elastic.co/t/single-letter-term-matches-cause-noise-in-search-results/45698/2 "2016-03-29T15:10:24Z")

</div>

I can try to add :

`"operator" : "and"`

Or you can try also a multi\_match query with :

`"minimum_should_match": "80%"`

---

<div class="post-metadata">

**Author:** ![jillesvangurp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jillesvangurp/32/3863_2.png) [@jillesvangurp](https://discuss.elastic.co/u/jillesvangurp)\
**Post date:** [March 29, 2016, 3:37pm UTC](https://discuss.elastic.co/t/single-letter-term-matches-cause-noise-in-search-results/45698/3 "2016-03-29T15:37:04Z")

</div>

> [@xavierfacq](#):
>
> minimum\_should\_match

Thanks but this does not work with prefix queries. In the end I figured out I need to set "max\_expansions": 10000 (default 10). The problem was that I have so many tokens starting with g that it never expanded to gurp because of the low default max\_expansions. So it sorted itself out as soon as I bumped this to 10k. Yes this makes it slower but at least it is more correct.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:04pm UTC](https://discuss.elastic.co/t/single-letter-term-matches-cause-noise-in-search-results/45698/4 "2017-07-05T23:04:21Z")

</div>


