# Alternatives to wildcard search for "contains" type searching

**URL:** <https://discuss.elastic.co/t/alternatives-to-wildcard-search-for-contains-type-searching/189623>\
**Category:** Elasticsearch\
**Created:** [July 9, 2019, 8:58pm UTC](https://discuss.elastic.co/t/alternatives-to-wildcard-search-for-contains-type-searching/189623 "2019-07-09T20:58:48Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![onearmedscissor](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/onearmedscissor/32/47247_2.png) [@onearmedscissor](https://discuss.elastic.co/u/onearmedscissor)\
**Post date:** [July 9, 2019, 8:58pm UTC](https://discuss.elastic.co/t/alternatives-to-wildcard-search-for-contains-type-searching/189623/1 "2019-07-09T20:58:48Z")

</div>

Hi, I'm dealing with unstructured documents where we're indexing the contents of the text of the document. Generally right now we just have a fairly basic configuration where we use the standard analyzer on the body of text. This works pretty well, however, we're also looking to support "contains" (within a word) type searching for some particular use cases. In particular, we're looking to extract out numbers that are contained within letters e.g.:

ABC1234EFG

The goal is to be able to search by "1234" and get results.

Wondering about different approaches here:

- The first thought here was wildcards, obviously though it's discouraged to use them from a performance standpoint, especially if the wildcard is on the front of the search term.
- Using ngrams. I think the concern here mainly is term explosion (index size + indexing performance - although we're not too concerned about write performance) + deciding the correct min/max size + maybe additional search noise
- Some sort of custom (or built-in, trying to find something?) tokenizer or filter that can take ABC1234EFG and produce ABC 1234 EFG. This is in a body of text though so we would still want the standard tokenizer behavior on other words e.g.:  
"This is my document ABC1234EFG"
- Something with fuzziness?

Then, I guess as a general question what are people generally doing when people want "contains" type searching beyond full word matching when dealing with a full body of text (so generally cannot assume much about the value of the field other than it's a blob of text).

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 9, 2019, 9:11pm UTC](https://discuss.elastic.co/t/alternatives-to-wildcard-search-for-contains-type-searching/189623/2 "2019-07-09T21:11:57Z")

</div>

I didn't test but I thought that the standard analyzer would split such a text by default. Isn't the case?

---

<div class="post-metadata">

**Author:** ![telendt](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/telendt/32/3416_2.png) [@telendt](https://discuss.elastic.co/u/telendt)\
**Post date:** [July 10, 2019, 9:58am UTC](https://discuss.elastic.co/t/alternatives-to-wildcard-search-for-contains-type-searching/189623/3 "2019-07-10T09:58:46Z")

</div>

> Some sort of custom (or built-in, trying to find something?) tokenizer or filter that can take ABC1234EFG and produce ABC 1234 EFG.

That's what Word Delimiter Token Filter is for:

> **[Word delimiter token filter | Elasticsearch Guide \[8.11\] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-word-delimiter-tokenfilter.html)**

example:

```auto
GET _analyze
{
  "tokenizer": "standard",
  "filter": ["word_delimiter"],
  "text": "ABC1234EFG"
}

```

output:

```auto
{
  "tokens" : [
    {
      "token" : "ABC",
      "start_offset" : 0,
      "end_offset" : 3,
      "type" : "<ALPHANUM>",
      "position" : 0
    },
    {
      "token" : "1234",
      "start_offset" : 3,
      "end_offset" : 7,
      "type" : "<ALPHANUM>",
      "position" : 1
    },
    {
      "token" : "EFG",
      "start_offset" : 7,
      "end_offset" : 10,
      "type" : "<ALPHANUM>",
      "position" : 2
    }
  ]
}

```

---

<div class="post-metadata">

**Author:** ![onearmedscissor](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/onearmedscissor/32/47247_2.png) [@onearmedscissor](https://discuss.elastic.co/u/onearmedscissor)\
**Post date:** [July 10, 2019, 2:22pm UTC](https://discuss.elastic.co/t/alternatives-to-wildcard-search-for-contains-type-searching/189623/4 "2019-07-10T14:22:39Z")

</div>

Thanks! Looks like the word delimiter token filter can work for this. Figured there was some standard functionality like this but I missed it!

---

<div class="post-metadata">

**Author:** ![onearmedscissor](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/onearmedscissor/32/47247_2.png) [@onearmedscissor](https://discuss.elastic.co/u/onearmedscissor)\
**Post date:** [July 10, 2019, 2:54pm UTC](https://discuss.elastic.co/t/alternatives-to-wildcard-search-for-contains-type-searching/189623/5 "2019-07-10T14:54:55Z")

</div>

Then also just curious how people tend to handle user's requesting wildcard type search when using Elasticsearch for user facing applications. Basically, this came up because when users don't know exact word matches they then want to reach for wildcards (and ask for it as a feature in the search syntax). A wildcard placed on the front of a term would perform poorly. A wildcard placed on the back of a term (or middle) could perform okay, also it could combined with index prefix feature to improve performance - [https://www.elastic.co/guide/en/elasticsearch/reference/current/index-prefixes.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/index-prefixes.html)

For instance, I was looking into what other applications do and Sharepoint supports wildcards on the end of terms (but not on the front), see [https://docs.microsoft.com/en-us/sharepoint/dev/general-development/keyword-query-language-kql-syntax-reference](https://docs.microsoft.com/en-us/sharepoint/dev/general-development/keyword-query-language-kql-syntax-reference)

So I guess, for user facing search applications do:

- You let users specify wildcard placement? This probably depends on how technical your users are if this makes sense. Do you only allow it on the end of terms (e.g. like Sharepoint above) - that may be confusing as well. Do you combine this with something like index prefixes to keep response times low.
- Don't have users specify wildcard placement but instead just always do a prefix type search (using index prefixes to maintain fast responses) in combination with something like fuzziness to allow for word variation/misspellings? Doesn't allow for searches in the middle e.g. foo\*bar but maybe gets most of the way there in terms of usability?
- Something like reverse filtered edge ngrams to make front loaded wildcards acceptable?
- Using ngrams (what size to pick, potential index overhead?)

---

<div class="post-metadata">

**Author:** ![telendt](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/telendt/32/3416_2.png) [@telendt](https://discuss.elastic.co/u/telendt)\
**Post date:** [July 11, 2019, 8:30am UTC](https://discuss.elastic.co/t/alternatives-to-wildcard-search-for-contains-type-searching/189623/6 "2019-07-11T08:30:43Z")

</div>

I've never had to support user queries that start with wildcards, so I don't know how people do it. But I think that your 3rd guess is right - use Reverse Token Filter to change you "wildcard prefix" problem into "wildcard suffix" one. For wildcard in the middle (`XX*YYY`) you may even use some simple heuristic, like checking whether the wildcard is closer to the start or the end (end either use "regular" or "reversed" tokens). This still does not solve problem of `*XXX*` (wildcard in the front and in the end) - maybe n-grams (not from the edge) can be helpful there?

Good luck!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 8, 2019, 8:31am UTC](https://discuss.elastic.co/t/alternatives-to-wildcard-search-for-contains-type-searching/189623/7 "2019-08-08T08:31:07Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
