# Match terms with characters in a sequence

**URL:** <https://discuss.elastic.co/t/match-terms-with-characters-in-a-sequence/238768>\
**Category:** Elasticsearch\
**Created:** [June 26, 2020, 2:57am UTC](https://discuss.elastic.co/t/match-terms-with-characters-in-a-sequence/238768 "2020-06-26T02:57:53Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![emarthinsen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/emarthinsen/32/23172_2.png) [@emarthinsen](https://discuss.elastic.co/u/emarthinsen)\
**Post date:** [June 26, 2020, 2:57am UTC](https://discuss.elastic.co/t/match-terms-with-characters-in-a-sequence/238768/1 "2020-06-26T02:57:54Z")

</div>

I'm interested in building an autocomplete-style interface where someone can enter characters and the results will be the terms that have those characters in that order, but not necessarily contiguous. This is functionality that I've seen in code editors where, for instance, you can type in `apmou` and it will match " **ap** p/ **mo** dels/ **u** ser.rb" and " **ap** p/ **mo** dels/ **u** ser\_account.rb", for instance. This might be too specialized a search, but I'm curious if there is a way to pull it off with ES.

---

<div class="post-metadata">

**Author:** ![Vinayak\_Sapre](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vinayak_sapre/32/45939_2.png) [@Vinayak\_Sapre](https://discuss.elastic.co/u/Vinayak_Sapre)\
**Post date:** [June 26, 2020, 7:36am UTC](https://discuss.elastic.co/t/match-terms-with-characters-in-a-sequence/238768/2 "2020-06-26T07:36:46Z")

</div>

@emarthinsen  
I will try to work with code editor example you mentioned. Basic idea is

1. Split path in the document on / and for each part create edge ngrams
2. Since we need ordered match, we need to use phrase query.
3. Create multiple phrases to match parts at different levels (apps vs models vs users).

**Index mapping**

```auto
{
  "settings": {
    "index": {
      "max_ngram_diff": 10,
      "analysis": {
        "tokenizer": {
          "path_tokenizer": {
            "type": "simple_pattern_split",
            "pattern": "/"
          }
        },
        "filter": {
          "1_10_edge_ngram": {
            "type": "edge_ngram",
            "min_gram": "1",
            "max_gram": "10"
          }
        },
        "analyzer": {
          "my_analyzer": {
            "filter": [
              "lowercase",
              "1_10_edge_ngram"
            ],
            "tokenizer": "path_tokenizer"
          },
          "my_search_analyzer": {
            "filter": [
              "lowercase"
            ],
            "tokenizer": "path_tokenizer"
          }
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "path_str": {
        "type": "text",
        "analyzer": "my_analyzer",
        "search_analyzer": "my_search_analyzer"
      }
    }
  }
}

```

**Java code to generate query**

```auto
    private static void gen_query(String searchStr) {
        int len = searchStr.length();
        StringBuilder query = new StringBuilder();
        query.append("{\"query\": {\"bool\": {\"should\": [");
        for (int i =0; i < Math.pow(2,len -1); i++) {
            StringBuilder pattern = new StringBuilder();
            for (int j = 0; j < len; j++) {
                pattern.append(searchStr.charAt(j));
                if (j < len -1) {
                    if ((i & (1 << (j))) != 0) {
                        pattern.append("/");
                    }
                }
            }
            query.append("{\"match_phrase\": {\"path_str\": \"" + pattern + "\"}},");
        }
        query.setCharAt(query.length() -1 , ' ');
        query.append("]}},\"sort\": [\"_doc\"]}");
        System.out.println(query);
    }

```

---

<div class="post-metadata">

**Author:** ![emarthinsen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/emarthinsen/32/23172_2.png) [@emarthinsen](https://discuss.elastic.co/u/emarthinsen)\
**Post date:** [June 30, 2020, 7:05am UTC](https://discuss.elastic.co/t/match-terms-with-characters-in-a-sequence/238768/3 "2020-06-30T07:05:49Z")

</div>

@Vinayak_Sapre Very interesting idea using a phrase query. I'm having a little trouble following your Java sample. You lost me a bit with your loop iterations. Are you trying to take the search term and break it down into every permutation of where the slashes could be?

---

<div class="post-metadata">

**Author:** ![Vinayak\_Sapre](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vinayak_sapre/32/45939_2.png) [@Vinayak\_Sapre](https://discuss.elastic.co/u/Vinayak_Sapre)\
**Post date:** [July 1, 2020, 4:23pm UTC](https://discuss.elastic.co/t/match-terms-with-characters-in-a-sequence/238768/4 "2020-07-01T16:23:12Z")

</div>

@emarthinsen

> [@emarthinsen](#):
>
> Are you trying to take the search term and break it down into every permutation of where the slashes could be?

Yes correct.

If you search for "apmou", slash can be at 0 to (length-1) places, where \_ is in "a\_p\_m\_o\_u". At each location you have boolean choice, add or don't add slash. So you can represent all permutations as 0000, 1000, 0100, .... and there will be 2^(length-1) permutations.

The outer for loop goes over permutations. Inner constructs string, add character and check if you need slash after it.

`if ((i & (1 << (j))) != 0) {` checks if you want slash at j-th location in i-th permutation.

Hope this helps.

Btw I noticed I forgot to add lowercase filter. I will update my previous post.

---

<div class="post-metadata">

**Author:** ![Vinayak\_Sapre](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vinayak_sapre/32/45939_2.png) [@Vinayak\_Sapre](https://discuss.elastic.co/u/Vinayak_Sapre)\
**Post date:** [July 10, 2020, 5:16am UTC](https://discuss.elastic.co/t/match-terms-with-characters-in-a-sequence/238768/5 "2020-07-10T05:16:50Z")

</div>

@emarthinsen  
Just curious, did the approach work for you? If not, can you post your findings. If it did, how is the performance. Query must run at typing speed.

---

<div class="post-metadata">

**Author:** ![emarthinsen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/emarthinsen/32/23172_2.png) [@emarthinsen](https://discuss.elastic.co/u/emarthinsen)\
**Post date:** [July 29, 2020, 4:33am UTC](https://discuss.elastic.co/t/match-terms-with-characters-in-a-sequence/238768/6 "2020-07-29T04:33:31Z")

</div>

Hi @Vinayak_Sapre. So sorry, I haven't had a chance to test it. I got pulled in another direction. Hoping to test it out soon though.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 26, 2020, 4:33am UTC](https://discuss.elastic.co/t/match-terms-with-characters-in-a-sequence/238768/7 "2020-08-26T04:33:39Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
