# How can I achieve fast substring search on a 1.5 billion documents index?

**URL:** https://discuss.elastic.co/t/how-can-i-achieve-fast-substring-search-on-a-1-5-billion-documents-index/380836
**Category:** Elastic Search
**Created:** [August 6, 2025, 6:18pm UTC](https://discuss.elastic.co/t/how-can-i-achieve-fast-substring-search-on-a-1-5-billion-documents-index/380836 "2025-08-06T18:18:07Z")
**Posts on this page:** 5
**Page:** 1

<div class="post-metadata">

### Author: ![JoshMax](https://avatars.discourse-cdn.com/v4/letter/j/ebca7d/32.png) [@JoshMax](https://discuss.elastic.co/u/JoshMax)
#### Post date: [August 6, 2025, 6:18pm UTC](https://discuss.elastic.co/t/how-can-i-achieve-fast-substring-search-on-a-1-5-billion-documents-index/380836/1 "2025-08-06T18:18:07Z")

</div>

Hi everyone!

I’m running an Elasticsearch 9.1.0 cluster (6 shards, 0 replicas, refresh disabled) over roughly 1.5 billion email address documents. I need to support fast, case‐insensitive substring searches that match only contiguous character sequences (i.e. exact substrings), for example:

- Should match when searching for `johndoe`:

```bash
xxxjohndoe@example.com
xxxjohndoexxx@example.com
johndoexxx@example.com

```

- Should not match:

```auto
xxxjohnxxxdoexxx@example.com
john.doe@example.com
john_doe@example.com

```

**What I’ve tried**

1. Index mapping

```json
{
    "settings": {
        "index.number_of_shards": 6,
        "index.number_of_replicas": 0,
        "index.refresh_interval": "-1"
    },
    "mappings": {
        "properties": {
            "email": {
                "type": "keyword",
                "ignore_above": 256,
                "normalizer": "lowercase",
                "fields": {
                    "wc": {
                        "type": "wildcard",
                        "ignore_above": 256
                    }
                }
            },
            "gender": {
                "type": "keyword"
            }
        }
    }
}

```

1. Search query

```json
{
  "size": 500,
  "terminate_after": 1000,
  "track_total_hits": false,
  "_source": false,
  "fields": ["email", "gender"],
  "query": {
    "constant_score": {
      "filter": {
        "wildcard": {
          "email.wc": {
            "value": "*johndoe*",
            "case_insensitive": true
          }
        }
      }
    }
  }
}

```

Despite using the specialized **`wildcard`** field and `"rewrite": "constant_score"`, each query still takes **20–30 seconds** , which is far too slow for my needs.

### What I’m looking for

1. Suggestions on index / search structure that would give me fast, exact substring matching at this scale.
2. Alternatives to wildcard queries, are there better ES features or plugins for this use case?

Thanks in advance for any help!

---

<div class="post-metadata">

### Author: ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)
#### Post date: [August 6, 2025, 6:36pm UTC](https://discuss.elastic.co/t/how-can-i-achieve-fast-substring-search-on-a-1-5-billion-documents-index/380836/2 "2025-08-06T18:36:25Z")

</div>

What size are the shards?

What is the average size of the indexed documents?

What is the specification of the cluster in terms of node count, RAM, CPU and type of storage used?

How many matches does a typical search return?

Does latency improve if you reduce the size of the result set?

Do you have any limitations on the length of the substring to search for? Can the substring cover any part of the email address, e.g. `doe@exam`?

Have you used the [profile API](https://www.elastic.co/guide/en/elasticsearch/reference/8.19/search-profile.html) to get some insights into what is taking up time?

---

<div class="post-metadata">

### Author: ![RainTown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/raintown/32/140206_2.png) [@RainTown](https://discuss.elastic.co/u/RainTown)
#### Post date: [August 6, 2025, 7:03pm UTC](https://discuss.elastic.co/t/how-can-i-achieve-fast-substring-search-on-a-1-5-billion-documents-index/380836/3 "2025-08-06T19:03:37Z")

</div>

random suggestion, but why not split [someone@somedomain.com](mailto:someone@somedomain.com) into 2 parts, on the `@` symbol, and just search the before-the-@ part ? Unless you actually want to match [cleverguy@johndoesmath.com](mailto:cleverguy@johndoesmath.com) ?

---

<div class="post-metadata">

### Author: ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)
#### Post date: [August 6, 2025, 7:38pm UTC](https://discuss.elastic.co/t/how-can-i-achieve-fast-substring-search-on-a-1-5-billion-documents-index/380836/4 "2025-08-06T19:38:42Z")

</div>

> [@JoshMax](#):
>
> Despite using the specialized **`wildcard`** field and `"rewrite": "constant_score"`, each query still takes **20–30 seconds** , which is far too slow for my needs.

Even using the `wildcard` datatype, having a query starting with a leading wildcard is one of the worst possible queries that you can have and should be avoided.

You may need to change this and use a n-gram tokenizer as mentioned in the wildcard [documentation notes](https://www.elastic.co/docs/reference/query-languages/query-dsl/query-dsl-wildcard-query#wildcard-query-notes).

You mentioned that your index have 6 shards and 0 replicas, but what is the size of the index? Also, replicas is something that can improve the search speed.

---

<div class="post-metadata">

### Author: ![chongshengdz](https://avatars.discourse-cdn.com/v4/letter/c/e95f7d/32.png) [@chongshengdz](https://discuss.elastic.co/u/chongshengdz)
#### Post date: [August 10, 2025, 12:34pm UTC](https://discuss.elastic.co/t/how-can-i-achieve-fast-substring-search-on-a-1-5-billion-documents-index/380836/5 "2025-08-10T12:34:42Z")

</div>

use infix mapping (n-gram) for the field, you can search and learn from the documentation.
