# How to find docs that contain the exact specified terms

**URL:** <https://discuss.elastic.co/t/how-to-find-docs-that-contain-the-exact-specified-terms/295758>\
**Category:** Elasticsearch\
**Created:** [January 29, 2022, 5:05pm UTC](https://discuss.elastic.co/t/how-to-find-docs-that-contain-the-exact-specified-terms/295758 "2022-01-29T17:05:43Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![Cholley\_Alexandre](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cholley_alexandre/32/90867_2.png) [@Cholley\_Alexandre](https://discuss.elastic.co/u/Cholley_Alexandre)\
**Post date:** [January 29, 2022, 5:05pm UTC](https://discuss.elastic.co/t/how-to-find-docs-that-contain-the-exact-specified-terms/295758/1 "2022-01-29T17:05:43Z")

</div>

Hi, I'm trying to get docs that contain the exact specified terms. Terms query doesn't match because it also returns docs with additional terms.

Create index :

```auto
PUT my-index
{
   "mappings": {
      "properties": {
        "my-field" : {
           "type" : "keyword"
        }
      }
   }
}

```

Index a document:

```auto
PUT my-index/_doc/1
{
   "my-field": ["A", "B"]
}

```

Index another document:

```auto
PUT my-index/_doc/2
{
   "my-field": ["A", "B", "C"]
}

```

Use the terms query to find documents that containing A and B:

```auto
GET my-index/_search?pretty
{
   "query":
   {
      "terms": {
         "my-field": ["A", "B"]
      }
   }
}

```

Results :

```auto
{
  "took" : 178,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 2,
      "relation" : "eq"
    },
    "max_score" : 1.0,
    "hits" : [
      {
        "_index" : "my-index",
        "_type" : "_doc",
        "_id" : "1",
        "_score" : 1.0,
        "_source" : {
          "my-field" : [
            "A",
            "B"
          ]
        }
      },
      {
        "_index" : "my-index",
        "_type" : "_doc",
        "_id" : "2",
        "_score" : 1.0,
        "_source" : {
          "my-field" : [
            "A",
            "B",
            "C"
          ]
        }
      }
    ]
  }
}

```

Elastic returns document 2 and document 1 because both contain A and B in my-field. It is the expected behavior, but I'd like to return only documents containing strictly A and B (if a doc contains A and B and C then it should be excluded from the results).

I've read the documentation (filters, aggregations, ...) but I can't find a way.

Thanks in advance for your help 😉

---

<div class="post-metadata">

**Author:** ![marone](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/marone/32/87145_2.png) [@marone](https://discuss.elastic.co/u/marone)\
**Post date:** [January 29, 2022, 6:18pm UTC](https://discuss.elastic.co/t/how-to-find-docs-that-contain-the-exact-specified-terms/295758/2 "2022-01-29T18:18:07Z")

</div>

Hey @Cholley_Alexandre ,

Why not tell Elasticsearch to not return what you don't want ^^.

I added another document to make sure it doesn't appear in the search result:

```auto
PUT my-index/_doc/3
{
   "my-field": ["A", "C"]
}

```

Then tell Elasticsearch to return documents that **must not** contain `C`; like this:

```auto
GET my-index/_search
{
  "query": {
    "bool": {
      "must_not": [
        {
          "terms": {
            "my-field": [
              "C"
            ]
          }
        }
      ]
    }
  }
}

```

More about [Compound queries](https://www.elastic.co/guide/en/elasticsearch/reference/7.16/query-dsl-bool-query.html).

Hope it helps!

---

<div class="post-metadata">

**Author:** ![Cholley\_Alexandre](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cholley_alexandre/32/90867_2.png) [@Cholley\_Alexandre](https://discuss.elastic.co/u/Cholley_Alexandre)\
**Post date:** [January 29, 2022, 9:20pm UTC](https://discuss.elastic.co/t/how-to-find-docs-that-contain-the-exact-specified-terms/295758/3 "2022-01-29T21:20:36Z")

</div>

Hi marone,

Thank u for your answer. Yes it's a good idea and it works fine 👍.

But in my use case (a research tool in agricultural regulatory data), the users need to combine between 1 and 4 values (chemical substances) and they need to exclude all the others (about 500 possible values). So I'll need to know all possible values to exclude before to compose the request. I could compose a first query (filtered aggregation, see : [Terms aggregation | Elasticsearch Guide [7.16] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-terms-aggregation.html#_filtering_values_4)) to retrieve all the values to exclude and a second query for the search ? But I don't now if it's good for performances... maybe with the aggregation caches... (see: [Aggregations | Elasticsearch Guide [7.16] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations.html#agg-caches))

To finds docs that contain the exact combined terms "A" and "B", in first time I retrieve all values to exclude with a filtered aggregation like this :

```auto
GET my-index/_search
{
  "size": 0,
  "aggs": {
    "my-agg-name": {
      "terms": {
        "field": "my-field",
        "exclude": ["A","B"]
      }
    }
  }
}

```

Result is a bucket with all possibles values (e.g "C", "D", "E", ...) excluding my terms "A" and "B". Programmatically I put these values in an array,  
and in a second time I compose a new query with "must\_not" and my array of values to exclude:

```auto
GET my-index/_search
{
  "query": {
    "bool": {
      "must_not": [
        {
          "terms": {
            "my-field": ["C", "D", "E", ...] 
          }
        }
      ]
    }
  }
}

```

---

<div class="post-metadata">

**Author:** ![marone](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/marone/32/87145_2.png) [@marone](https://discuss.elastic.co/u/marone)\
**Post date:** [January 29, 2022, 10:45pm UTC](https://discuss.elastic.co/t/how-to-find-docs-that-contain-the-exact-specified-terms/295758/4 "2022-01-29T22:45:27Z")

</div>

Nice 🙂

Well I think you don't have performance issue with the `exclude` parameter as long as you don't use wildcard.

> The performance of the `regexp` query can vary based on the regular expression provided. To improve performance, avoid using wildcard patterns, such as `.*` or `.*?+` , without a prefix or suffix.

and

> Avoid beginning patterns with `*` or `?` . This can increase the iterations needed to find matching terms and slow search performance.

See [Regexp](https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-regexp-query.html#regexp-query-field-params) and [Wildcards](https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-wildcard-query.html#wildcard-query-field-params) docs.

Did I answer your question?

---

<div class="post-metadata">

**Author:** ![Tomo\_M](https://avatars.discourse-cdn.com/v4/letter/t/848f3c/32.png) [@Tomo\_M](https://discuss.elastic.co/u/Tomo_M)\
**Post date:** [January 30, 2022, 12:34am UTC](https://discuss.elastic.co/t/how-to-find-docs-that-contain-the-exact-specified-terms/295758/5 "2022-01-30T00:34:00Z")

</div>

Another idea from me is to make something like fingerprint field. You can implement it by runtime field or ingest pipeline. The former is easier and fit when you can reduce documents to reasonable size by filtering those containing both A and B. You may deduplicate and/or sort values and concatenate them to make fingerprint.

---

<div class="post-metadata">

**Author:** ![Cholley\_Alexandre](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cholley_alexandre/32/90867_2.png) [@Cholley\_Alexandre](https://discuss.elastic.co/u/Cholley_Alexandre)\
**Post date:** [January 30, 2022, 1:25pm UTC](https://discuss.elastic.co/t/how-to-find-docs-that-contain-the-exact-specified-terms/295758/6 "2022-01-30T13:25:26Z")

</div>

Great 😉 thank you marone.

---

<div class="post-metadata">

**Author:** ![Cholley\_Alexandre](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cholley_alexandre/32/90867_2.png) [@Cholley\_Alexandre](https://discuss.elastic.co/u/Cholley_Alexandre)\
**Post date:** [January 30, 2022, 1:50pm UTC](https://discuss.elastic.co/t/how-to-find-docs-that-contain-the-exact-specified-terms/295758/7 "2022-01-30T13:50:08Z")

</div>

Hi Tomo\_M, it's a good way too. I think you mean I need to create an additional field in my doc that contains a fingerprint of the concatenated terms (e.g. MD5 hash) , and then I request on this fingerprint.

Create index :

```auto
PUT my-index
{
   "mappings": {
      "properties": {
        "my-field" : {
           "type" : "keyword"
        },
        "fingerprint-of-axact-terms": {
          "type": "keyword"
        }
      }
   }
}

```

Index a document:

```auto
PUT my-index/_doc/1
{
   "my-field": ["A", "B"],
   "fingerprint-of-exact-terms": "5ae395e8ab6a4121fc3445afdde6b13f"
}

```

When a user query on exact terms A and B, I create a hash of these terms and I query on fingerprint field:

```auto
GET my-index/_search?pretty
{
   "query":
   {
      "terms": {
         "fingerprint-of-exact-terms": "5ae395e8ab6a4121fc3445afdde6b13f"
      }
   }
}

```

But the problem is that the order of the terms can give different fingerprint ? (eg. A and B = 5ae395e8ab6a4121fc3445afdde6b13f / B and A = 751bf9b43937fba8970400edf1e28a5d). Otherwise you must set terms in the same order before index and when you query.

---

<div class="post-metadata">

**Author:** ![Tomo\_M](https://avatars.discourse-cdn.com/v4/letter/t/848f3c/32.png) [@Tomo\_M](https://discuss.elastic.co/u/Tomo_M)\
**Post date:** [January 30, 2022, 4:41pm UTC](https://discuss.elastic.co/t/how-to-find-docs-that-contain-the-exact-specified-terms/295758/8 "2022-01-30T16:41:00Z")

</div>

> [@Cholley\_Alexandre](#):
>
> But the problem is that the order of the terms can give different fingerprint ?

Exactly. It depends on the requiements.

If the order has meaning, different fingerprint from different order make sense. If not, order should not affect the fingerprint and you have to sort values before making fingerprint.

In addition, in this case, just using concatenated values (eg 'A\_B') could be acceptable for fingerprint to distinguish identical set.

This is my sample code. Using [runtime mappings](https://www.elastic.co/jp/blog/getting-started-with-elasticsearch-runtime-fields), you need not calculate hash in indexing by yourself.

```auto
PUT /test_runtime/
{
  "mappings": {
    "properties": {
      "my-field":{
        "type": "keyword"
      }
    }
  }
}

POST /test_runtime/_bulk
{"index":{}}
{"my-field":["A", "B", "C"]}
{"index":{}}
{"my-field":["A", "A","B"]}
{"index":{}}
{"my-field":["A", "B"]}
{"index":{}}
{"my-field":["B","A"]}
{"index":{}}
{"my-field":"A"}
{"index":{}}
{"my-field":["A"]}

GET /test_runtime/_search
{
  "query":{
    "bool":{
      "filter":[
        {"term":{"my-field":"A"}},
        {"term":{"my-field":"B"}},
        {"term":{"fingerprint":"A_B"}}
      ]
    }
  },
  "runtime_mappings": {
    "fingerprint": {
      "type": "keyword",
      "script": {
        "source": """
        if (doc['my-field'] instanceof List) {
          Set set = new HashSet(doc['my-field']);
          List list = new ArrayList(set);
          Collections.sort(list);
          String str = String.join('_',list);
          emit(str)
        } else {
          emit(doc['my-field'].value)
        }
        """
      }
    }
  },
  "fields": [
    "fingerprint"
  ]
}

```

Or you can add [ingest script processor](https://www.elastic.co/guide/en/elasticsearch/reference/master/script-processor.html) to calcurate hash. This way has better performance on querying.

---

<div class="post-metadata">

**Author:** ![Tomo\_M](https://avatars.discourse-cdn.com/v4/letter/t/848f3c/32.png) [@Tomo\_M](https://discuss.elastic.co/u/Tomo_M)\
**Post date:** [January 30, 2022, 5:00pm UTC](https://discuss.elastic.co/t/how-to-find-docs-that-contain-the-exact-specified-terms/295758/9 "2022-01-30T17:00:19Z")

</div>

Using ingest pipeline is like this.

```auto
PUT _ingest/pipeline/test_fingerprint_pipeline/
{
  "description":"",
  "processors":[
    {
      "script":{
        "source": """
        ctx['a'] = 'a';
        if (ctx['my-field'] instanceof List) {
        Set set = new HashSet(ctx['my-field']);
        List list = new ArrayList(set);
        Collections.sort(list);
        String str = String.join('_',list);
        ctx['fingerprint'] = str
      } else {
        ctx['fingerprint'] = ctx['my-field']
      }"""
      }
    }
  ]
}

PUT /test_pipeline/
{
  "mappings": {
    "properties": {
      "my-field":{
        "type": "keyword"
      },
      "fingerprint":{
        "type": "keyword"
      }
    }
  },
  "settings":{
    "index.default_pipeline": "test_fingerprint_pipeline"
  }
}
GET test_pipeline

POST /test_pipeline/_bulk
{"index":{}}
{"my-field":["A", "B", "C"]}
{"index":{}}
{"my-field":["A", "A","B"]}
{"index":{}}
{"my-field":["A", "B"]}
{"index":{}}
{"my-field":["B","A"]}
{"index":{}}
{"my-field":"A"}
{"index":{}}
{"my-field":["A"]}

GET /test_pipeline/_search
{
  "track_total_hits": true, 
  "query":{
    "bool":{
      "filter":[
        {"term":{"fingerprint":"A_B"}}
      ]
    }
  }
}

```

---

<div class="post-metadata">

**Author:** ![Cholley\_Alexandre](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cholley_alexandre/32/90867_2.png) [@Cholley\_Alexandre](https://discuss.elastic.co/u/Cholley_Alexandre)\
**Post date:** [January 30, 2022, 8:48pm UTC](https://discuss.elastic.co/t/how-to-find-docs-that-contain-the-exact-specified-terms/295758/10 "2022-01-30T20:48:40Z")

</div>

Hi Tomo\_M, thanks a lot for your samples. It's really interesting 😉 and it helps me to understand the elastic capabilities. In my use case I need to re-index new updated data every week (about 150k docs), so the ingestion pipeline is interesting.

I will try all these ways. Thank you all

---

<div class="post-metadata">

**Author:** ![marone](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/marone/32/87145_2.png) [@marone](https://discuss.elastic.co/u/marone)\
**Post date:** [January 30, 2022, 9:27pm UTC](https://discuss.elastic.co/t/how-to-find-docs-that-contain-the-exact-specified-terms/295758/11 "2022-01-30T21:27:57Z")

</div>

One more thing to add I forgot, since you are using `Keywords` you don't have to worry about performance, because `keywords` fields are not analyzed compared to `text` field. From this perspective you are safe.

---

<div class="post-metadata">

**Author:** ![Cholley\_Alexandre](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cholley_alexandre/32/90867_2.png) [@Cholley\_Alexandre](https://discuss.elastic.co/u/Cholley_Alexandre)\
**Post date:** [January 31, 2022, 8:56am UTC](https://discuss.elastic.co/t/how-to-find-docs-that-contain-the-exact-specified-terms/295758/12 "2022-01-31T08:56:21Z")

</div>

In fact in my real use case, the mapping is like this :

```auto
"my-field": {
   "type": "text",
   "fields": {
      "completion": {
          "analyzer": "default",
          "type": "completion"
        },
        "keyword": {
           "type": "keyword"
      }
   }
}

```

I think that can't change anything for performances as long as I use my-field.keyword ?

Thanks.

---

<div class="post-metadata">

**Author:** ![Tomo\_M](https://avatars.discourse-cdn.com/v4/letter/t/848f3c/32.png) [@Tomo\_M](https://discuss.elastic.co/u/Tomo_M)\
**Post date:** [February 9, 2022, 9:47am UTC](https://discuss.elastic.co/t/how-to-find-docs-that-contain-the-exact-specified-terms/295758/13 "2022-02-09T09:47:58Z")

</div>

Hi @Cholley_Alexandre ,

I found another solution.

> [@Filter terms aggregation on array field](https://discuss.elastic.co/t/filter-terms-aggregation-on-array-field/296627/3):
>
> I was not correct. I found a solution. GET /test-index/\_search { "query":{ "bool": { "filter":[{ "terms": { "authors": ["1", "2"] } }], "must\_not": [{ "regexp": { "authors": "~(1|2)" } }] } } }

If you don't mind duplication of 'A' or 'B',

```auto
GET /test_runtime/_search
{
  "query":{
    "bool": {
      "filter": [
        {"term": {"my-field": "A"}},
        {"term": {"my-field": "B"}}
      ],
      "must_not": [
        {
          "regexp": {
            "my-field": "~(A|B)"
          }
        }
      ]
    }
  }
}

```

will return what you want. If you want to distinguish duplication or order, you may need custom fingerprint strategy as before.

---

<div class="post-metadata">

**Author:** ![Cholley\_Alexandre](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cholley_alexandre/32/90867_2.png) [@Cholley\_Alexandre](https://discuss.elastic.co/u/Cholley_Alexandre)\
**Post date:** [February 11, 2022, 1:07pm UTC](https://discuss.elastic.co/t/how-to-find-docs-that-contain-the-exact-specified-terms/295758/14 "2022-02-11T13:07:11Z")

</div>

Hi @Tomo_M,

Yes it's a perfect solution for my use case. I don't need to distinguish duplication or order.

If I understand correctly, you must also use multiple "term" in query (one "term" for each value) and not a global "terms" query in (with an array of values) :

Works well with multiple "term" in query :

```auto
GET /test_runtime/_search
{
  "query":{
    "bool": {
      "filter": [
        {"term": {"my-field": "A"}},
        {"term": {"my-field": "B"}}
      ],
      "must_not": [
        {
          "regexp": {
            "my-field": "~(A|B)"
          }
        }
      ]
    }
  }
}

```

Doesn't work with a global "terms" in query :

```auto
GET /test_runtime/_search
{
  "query": {
    "bool": {
      "filter": [
        {
          "terms": {
            "my-field": [
              "A",
              "B"
            ]
          }
        }
      ],
      "must_not": [
        {
          "regexp": {
            "my-field": "~(A|B)"
          }
        }
      ]
    }
  }
}

```

And "regexp" with "must\_not" is use to make sure that all other terms are excluded.

Thanks 😉

---

<div class="post-metadata">

**Author:** ![Tomo\_M](https://avatars.discourse-cdn.com/v4/letter/t/848f3c/32.png) [@Tomo\_M](https://discuss.elastic.co/u/Tomo_M)\
**Post date:** [February 11, 2022, 2:29pm UTC](https://discuss.elastic.co/t/how-to-find-docs-that-contain-the-exact-specified-terms/295758/15 "2022-02-11T14:29:10Z")

</div>

Yes, we have to use multiple term queries because terms query means funcs as OR.  
If you satisfied with my post, please mark it as Solution. Thanks 😀

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 11, 2022, 2:29pm UTC](https://discuss.elastic.co/t/how-to-find-docs-that-contain-the-exact-specified-terms/295758/16 "2022-03-11T14:29:52Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
