# Terms Aggregation with include filter

**URL:** <https://discuss.elastic.co/t/terms-aggregation-with-include-filter/50976>\
**Category:** Elasticsearch\
**Created:** [May 25, 2016, 5:31pm UTC](https://discuss.elastic.co/t/terms-aggregation-with-include-filter/50976 "2016-05-25T17:31:06Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![ayushsangani](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ayushsangani/32/9612_2.png) [@ayushsangani](https://discuss.elastic.co/u/ayushsangani)\
**Post date:** [May 25, 2016, 5:31pm UTC](https://discuss.elastic.co/t/terms-aggregation-with-include-filter/50976/1 "2016-05-25T17:31:06Z")

</div>

Hey all,

ES version: 2.3.2 (recently upgraded)

I'm doing terms aggregations on `not_analyzed` string field using include filter.  
[https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-terms-aggregation.html#\_filtering\_values\_2](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-terms-aggregation.html#_filtering_values_2)

Is it possible to add a regex flag which performs "CASE\_INSENSITIVE" terms aggregation on string field?

```auto
PUT /my_index
{
  "mappings": {
    "user": {
      "properties": {
        "name": { 
          "type": "string",
          "fields": {
            "raw": { 
              "type": "string",
              "index": "not_analyzed"
            }
          }
        }
      }
    }
  }
}

```

Aggregation Query:

```auto
POST /my_index/_search
{
 "aggregations": {
    "name_regex_terms_agg": {
      "terms": {
        "field": "name.raw",
        "size": 1000,
        "shard_size": 100000,
        "include": "adam.*|.*\\sadam.*"
      }
    }
  }
}

```

Is it possible to add a regex flag which performs "CASE\_INSENSITIVE" terms aggregation on string field?

Like Terms aggregation should look for both Adam or adam.

Please let me know if there is any other information required.

Thanks for the help.

---

<div class="post-metadata">

**Author:** ![msimos](https://avatars.discourse-cdn.com/v4/letter/m/bb73d2/32.png) [@msimos](https://discuss.elastic.co/u/msimos)\
**Post date:** [May 26, 2016, 12:09am UTC](https://discuss.elastic.co/t/terms-aggregation-with-include-filter/50976/2 "2016-05-26T00:09:29Z")

</div>

Hi,

You could try something like this:

```auto
{
  "size": 0, 
    "aggs" : {
        "buckets" : {
            "terms" : {
                "script" : "doc['name.raw'].value.toLowerCase()"
            }
        }
    }
}

```

Otherwise you could index the field using the keyword analyzer and the lowecase tokenizer to emit 1 token that is all lowercase. Then create an aggregation on that field. That will probably be faster then using a script to lowercase the field at query time.

---

<div class="post-metadata">

**Author:** ![ayushsangani](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ayushsangani/32/9612_2.png) [@ayushsangani](https://discuss.elastic.co/u/ayushsangani)\
**Post date:** [May 26, 2016, 4:21pm UTC](https://discuss.elastic.co/t/terms-aggregation-with-include-filter/50976/3 "2016-05-26T16:21:47Z")

</div>

Thanks for the reply Mike.  
I like your second option, but that would require me to reindex all the documents and plus it will return keys in terms aggregation lowercased(which is not desired).

I'm still not sure why Terms Aggregation `include filter` doesn't have CASE\_INSENSITIVE regex flag?

---

<div class="post-metadata">

**Author:** ![msimos](https://avatars.discourse-cdn.com/v4/letter/m/bb73d2/32.png) [@msimos](https://discuss.elastic.co/u/msimos)\
**Post date:** [May 26, 2016, 5:44pm UTC](https://discuss.elastic.co/t/terms-aggregation-with-include-filter/50976/4 "2016-05-26T17:44:08Z")

</div>

Refer to this breaking change:

[https://www.elastic.co/guide/en/elasticsearch/reference/2.3/breaking\_20\_aggregation\_changes.html#\_including\_excluding\_terms](https://www.elastic.co/guide/en/elasticsearch/reference/2.3/breaking_20_aggregation_changes.html#_including_excluding_terms)

The flags parameter is no longer supported.

---

<div class="post-metadata">

**Author:** ![ayushsangani](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ayushsangani/32/9612_2.png) [@ayushsangani](https://discuss.elastic.co/u/ayushsangani)\
**Post date:** [May 26, 2016, 6:12pm UTC](https://discuss.elastic.co/t/terms-aggregation-with-include-filter/50976/5 "2016-05-26T18:12:47Z")

</div>

Ohh ok thanks!

So I found a workaround to use regex queries on `analyzed` fields and do terms aggregation on `not_analyzed` field.

```auto
{
  "aggregations": {
    "name_regex_query": {
      "filter": {
        "regexp": {
          "name": {
            "value": "aust.*|.*\\saust.*",
            "flags_value": 65535
          }
        }
      },
      "aggregations": {
        "name_raw_terms_agg": {
          "terms": {
            "field": "name.raw",
            "size": 1000,
            "shard_size": 100000
          }
        }
      }
    }
  }
}

```

Curious to know if this has any performance impact?

---

<div class="post-metadata">

**Author:** ![subbu.nv](https://avatars.discourse-cdn.com/v4/letter/s/439d5e/32.png) [@subbu.nv](https://discuss.elastic.co/u/subbu.nv)\
**Post date:** [September 30, 2016, 9:06am UTC](https://discuss.elastic.co/t/terms-aggregation-with-include-filter/50976/6 "2016-09-30T09:06:12Z")

</div>

So are you getting the case insensitive results on the regex query?

---

<div class="post-metadata">

**Author:** ![ayushsangani](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ayushsangani/32/9612_2.png) [@ayushsangani](https://discuss.elastic.co/u/ayushsangani)\
**Post date:** [October 4, 2016, 1:53pm UTC](https://discuss.elastic.co/t/terms-aggregation-with-include-filter/50976/7 "2016-10-04T13:53:49Z")

</div>

@subbu.nv Yeah I'm able to get case insensitive results by using keyword analyzer and the lowecase tokenizer.  
Note that flags in regex query is removed in ES 2.X.

---

<div class="post-metadata">

**Author:** ![GregAtPareto](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gregatpareto/32/15989_2.png) [@GregAtPareto](https://discuss.elastic.co/u/GregAtPareto)\
**Post date:** [March 1, 2017, 9:59pm UTC](https://discuss.elastic.co/t/terms-aggregation-with-include-filter/50976/8 "2017-03-01T21:59:34Z")

</div>

I have a similar situation in that I want to do a case insensitive aggregation on a keyword field, so the idea of keyword analyzer and lowercase tokenizer makes sense. However, the syntax for setting this for a field seems to be elusive for me.

First approach was this:

```
PUT authors
{
  "mappings": {
	"famousbooks": {
	  "properties": {
		"Author": {
		  "type": "text",
		  "fields": {
			"use_lowercase": {
			  "type": "text",
			  "analyzer": "keyword",
			  "tokenizer": "lowercase"
			}
		  }
		}
	  }
	}
  }
}

```

But this fails with

```
"error": {
"root_cause": [
  {
    "type": "mapper_parsing_exception",
    "reason": "Mapping definition for [fields] has unsupported parameters: [tokenizer : lowercase]"
  }
],

```

So next step is a custom analyzer:

```
PUT authors
{
  "settings": {
	"analysis": {
	  "analyzer": {
		"myLowercase": {
		  "type": "custom",
		  "tokenizer": "keyword",
		  "filter" : ["lowercase"]
		}
	  }
	}
  },
  "mappings": {
	"famousbooks": {
	  "properties": {
		"Author": {
		  "type": "text",
		  "fields": {
			"use_lowercase": {
			  "type": "text",
			  "analyzer": "myLowercase"
			}
		  }
		}
	  }
	}
  }
}

```

So the aggregation query:

```
GET authors/famousbooks/_search
{
  "size": 0,
  "aggs": {
	"authors-aggs": {
	  "terms": {
		"field": "Author.use_lowercase"
	  }
	}
  }
}

```

returns an error as follows:

```
"error": {
"root_cause": [
  {
    "type": "illegal_argument_exception",
    "reason": "Fielddata is disabled on text fields by default. Set fielddata=true on [Author.use_lowercase] in order to load fielddata in memory by uninverting the inverted index. Note that this can however use significant memory."
  }
],

```

So to make this work I evidently need to enable Fielddata despite the dire warnings of turning this on... it seems like a big stick to use for what seemingly is a simple thing.

I'm hoping I am missing something obvious here, though.

Thx in advance for the help!

---

<div class="post-metadata">

**Author:** ![GregAtPareto](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gregatpareto/32/15989_2.png) [@GregAtPareto](https://discuss.elastic.co/u/GregAtPareto)\
**Post date:** [March 1, 2017, 10:21pm UTC](https://discuss.elastic.co/t/terms-aggregation-with-include-filter/50976/9 "2017-03-01T22:21:02Z")

</div>

Of course I discover a solution soon after I post...

Normalizers! HelpFound in another post on this forum.

```
PUT authors
{
  "settings": {
    "analysis": {
      "normalizer": {
        "myLowercase": {
          "type": "custom",
          "filter": ["lowercase"]
        }
      }
    }
  },
  "mappings": {
    "famousbooks": {
      "properties": {
        "Author": {
          "type": "keyword",
          "normalizer": "myLowercase"
        }
      }
    }
  }
}
```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:02pm UTC](https://discuss.elastic.co/t/terms-aggregation-with-include-filter/50976/10 "2017-07-05T22:02:45Z")

</div>



---

<div class="post-metadata">

**Author:** ![Martijn\_Laarman](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/martijn_laarman/32/4410_2.png) [@Martijn\_Laarman](https://discuss.elastic.co/u/Martijn_Laarman)\
**Post date:** [September 22, 2017, 10:27am UTC](https://discuss.elastic.co/t/terms-aggregation-with-include-filter/50976/11 "2017-09-22T10:27:31Z")

</div>

For future googlers:

```
"terms": {
    "field": "_index",
    "size": 10,
    "exclude" : " __BAD__",
    "script" : {
        "source" : "if (_value =~ /(^|.*\\s+)Index/i) { return _value } else { return ' __BAD__' }",
        "lang" : "painless"
    }
}

```

Note that you have to enable regex in painless as its disabled by default, read the caveats of doing so here:

[https://www.elastic.co/guide/en/elasticsearch/painless/current/painless-examples.html#modules-scripting-painless-regex](https://www.elastic.co/guide/en/elasticsearch/painless/current/painless-examples.html#modules-scripting-painless-regex)
