# Exclude specific terms from term aggregation's buckets list

**URL:** <https://discuss.elastic.co/t/exclude-specific-terms-from-term-aggregations-buckets-list/131122>\
**Category:** Elasticsearch\
**Created:** [May 9, 2018, 7:55am UTC](https://discuss.elastic.co/t/exclude-specific-terms-from-term-aggregations-buckets-list/131122 "2018-05-09T07:55:01Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![randomuser](https://avatars.discourse-cdn.com/v4/letter/r/958977/32.png) [@randomuser](https://discuss.elastic.co/u/randomuser)\
**Post date:** [May 9, 2018, 7:55am UTC](https://discuss.elastic.co/t/exclude-specific-terms-from-term-aggregations-buckets-list/131122/1 "2018-05-09T07:55:02Z")

</div>

Hello,

Is there any way to filter the searched term from the results?

I'm using a standard tokeniser, the field "word" usually contains values like "green table", "big table", "yellow tables".. my aggregation will put the words "table" or "tables" at top as they're the most frequent.

In this example, I don't want the word "table" in the buckets list results.

By the way, regarding the query\_string, whats the most efficient way (performance wise) to search for words that contain a word?

```
GET newindex/_search
    {
      "query": {
        "query_string": {
          "default_field": "word", 
          "query" : "*table*"
        }
      },
      "aggs" : {
          "tables" : {
              "terms" : { 
                "field" : "word.s",
                "size" : 100
              }
          }
      },
      "size" : 0
    }
```

---

<div class="post-metadata">

**Author:** ![randomuser](https://avatars.discourse-cdn.com/v4/letter/r/958977/32.png) [@randomuser](https://discuss.elastic.co/u/randomuser)\
**Post date:** [May 9, 2018, 8:30am UTC](https://discuss.elastic.co/t/exclude-specific-terms-from-term-aggregations-buckets-list/131122/2 "2018-05-09T08:30:09Z")

</div>

If anyone will have the same problem, the easiest solution is to add Exclude parameter:

```
GET newindex/_search
    {
      "query": {
        "query_string": {
          "default_field": "word", 
          "query" : "*table*"
        }
      },
      "aggs" : {
          "tables" : {
              "terms" : { 
                "field" : "word.s",
                "exclude": ".*table.*"
                "size" : 100
              }
          }
      },
      "size" : 0
    }
```

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 9, 2018, 8:35am UTC](https://discuss.elastic.co/t/exclude-specific-terms-from-term-aggregations-buckets-list/131122/3 "2018-05-09T08:35:33Z")

</div>

There is an exclude parameter in aggs.  
[https://www.elastic.co/guide/en/elasticsearch/reference/6.2/search-aggregations-bucket-terms-aggregation.html#\_filtering\_values\_3](https://www.elastic.co/guide/en/elasticsearch/reference/6.2/search-aggregations-bucket-terms-aggregation.html#_filtering_values_3)

About wildcard performances, prefer using ngrams. You ll pay the price at index time instead of query time.

---

<div class="post-metadata">

**Author:** ![randomuser](https://avatars.discourse-cdn.com/v4/letter/r/958977/32.png) [@randomuser](https://discuss.elastic.co/u/randomuser)\
**Post date:** [May 9, 2018, 8:38am UTC](https://discuss.elastic.co/t/exclude-specific-terms-from-term-aggregations-buckets-list/131122/4 "2018-05-09T08:38:23Z")

</div>

How to get full word tokens with Ngrams?  
With a Ngram tokeniser here, the returned tokens would be "tab", "le " etc, can't aggregate on that as the buckets wouldn't make sense

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 9, 2018, 9:31am UTC](https://discuss.elastic.co/t/exclude-specific-terms-from-term-aggregations-buckets-list/131122/5 "2018-05-09T09:31:04Z")

</div>

Use multi-fields: [https://www.elastic.co/guide/en/elasticsearch/reference/current/multi-fields.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/multi-fields.html)

The same content can use ngrams for search (`text` type) and no transformation as a `keyword` type on which you can compute aggs.

---

<div class="post-metadata">

**Author:** ![randomuser](https://avatars.discourse-cdn.com/v4/letter/r/958977/32.png) [@randomuser](https://discuss.elastic.co/u/randomuser)\
**Post date:** [May 9, 2018, 9:51am UTC](https://discuss.elastic.co/t/exclude-specific-terms-from-term-aggregations-buckets-list/131122/6 "2018-05-09T09:51:27Z")

</div>

Yep there was much dilemma a few days ago regarding that.. but in the example I gave, a keyword type would do buckets on "green table" instead on "green" and "table", or is there a way to achieve the same?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 9, 2018, 10:07am UTC](https://discuss.elastic.co/t/exclude-specific-terms-from-term-aggregations-buckets-list/131122/7 "2018-05-09T10:07:51Z")

</div>

You can use a text field with fielddata on. Might work

---

<div class="post-metadata">

**Author:** ![randomuser](https://avatars.discourse-cdn.com/v4/letter/r/958977/32.png) [@randomuser](https://discuss.elastic.co/u/randomuser)\
**Post date:** [May 9, 2018, 11:13am UTC](https://discuss.elastic.co/t/exclude-specific-terms-from-term-aggregations-buckets-list/131122/8 "2018-05-09T11:13:03Z")

</div>

Wasn't sure if I tried, so I've done it again, it doesn't produce the wanted effect.  
I know that using a standard analyser on a text field isn't really an optimal approach, but I didn't see any other way.  
Thanks for the effort though. 🙂

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 9, 2018, 11:32am UTC](https://discuss.elastic.co/t/exclude-specific-terms-from-term-aggregations-buckets-list/131122/9 "2018-05-09T11:32:54Z")

</div>

Share what you got and what you want.

---

<div class="post-metadata">

**Author:** ![randomuser](https://avatars.discourse-cdn.com/v4/letter/r/958977/32.png) [@randomuser](https://discuss.elastic.co/u/randomuser)\
**Post date:** [May 9, 2018, 12:38pm UTC](https://discuss.elastic.co/t/exclude-specific-terms-from-term-aggregations-buckets-list/131122/10 "2018-05-09T12:38:04Z")

</div>

1. There are lots of special characters in the words field
2. The words field can contain multiple words and every word has to be a separate "entity" (token)

That's pretty much why I need to use a standard analyser on a text field with fielddata turned on.

The goal is to count the number of occurrences of each word within all of the words fields on the cluster (that's the query above). I really gave a lot of consideration for other options, researched a lot, this seems like the only option for my use case.

On the other hand, do you have any ideas how I could convert the query above to count every word within the words field and sort them based on the count number?  
The current query produces a doc\_count, which is the number of documents that contain the word, but some documents contain a word multiple times, so it isn't very precise.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [June 1, 2018, 7:10pm UTC](https://discuss.elastic.co/t/exclude-specific-terms-from-term-aggregations-buckets-list/131122/11 "2018-06-01T19:10:27Z")

</div>

Could you provide a full recreation script as described in [About the Elasticsearch category](https://discuss.elastic.co/t/about-the-elasticsearch-category/21). It will help to better understand what you are doing. Please, try to keep the example as simple as possible.

A full reproduction script will help readers to understand, reproduce and if needed fix your problem. It will also most likely help to get a faster answer.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 29, 2018, 7:10pm UTC](https://discuss.elastic.co/t/exclude-specific-terms-from-term-aggregations-buckets-list/131122/12 "2018-06-29T19:10:41Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
