# Aggregations within aggregations

**URL:** <https://discuss.elastic.co/t/aggregations-within-aggregations/108399>\
**Category:** Elasticsearch\
**Created:** [November 20, 2017, 1:54pm UTC](https://discuss.elastic.co/t/aggregations-within-aggregations/108399 "2017-11-20T13:54:20Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Karolinebryn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/karolinebryn/32/17095_2.png) [@Karolinebryn](https://discuss.elastic.co/u/Karolinebryn)\
**Post date:** [November 20, 2017, 1:54pm UTC](https://discuss.elastic.co/t/aggregations-within-aggregations/108399/1 "2017-11-20T13:54:20Z")

</div>

I have an index that contains documents about movies and the movies are tagged with their genre and main actors. Like this:

```
{
	"title": "The fast and the furious",
	"tags": [
		{
			"groupName": "Genre",
			"tags": ["Action", "Racing"]
		},
		{
			"groupName": "Actors",
			"tags": ["Vin Diesel", "Paul Walker"]
		}
	]
},
{
	"title": "2 fast 2 furious",
	"tags": [
		{
			"groupName": "Genre",
			"tags": ["Action"]
		},
		{
			"groupName": "Actors",
			"tags": ["Paul Walker", "Ludacris"]
		}
	]
}

```

Now I want to use aggregation to aggregate first on groupName then on the tags-list, in order to create a list like this:

Genre

- Action (2)
- Racing (1)

Actors

- Paul Walker (2)
- Vin Diesel (1)
- Ludacris (1)

I have tried to do the following aggregation:

```
POST mymovies/_search
 {
   "query": {
     "match_all": {}
   },
   "aggs": {
     "movie_groupName": {
       "terms": {
         "field": "tags.groupName",
         "size": 20
       },
       "aggs": {
         "movie_tags": {
           "terms": {
             "field": "tags.tags",
             "size": 20
           }
         }
       }
     }
   }
 }

```

This is the result I expected:

```
"aggregations": {
	"movie_groupName": {
		"doc_count_error_upper_bound": 0,
		"sum_other_doc_count": 0,
		"buckets": [
			{
				"key": "Genre",
				"doc_count": 2,
				"movice_tags": {
					"doc_count_error_upper_bound": 0,
					"sum_other_doc_count": 0,
					"buckets": 
					[
						{
							"key": "Action",
							"doc_count": 2
						},
						{
							"key": "Racing",
							"doc_count": 1
						}
					]
				}
			},
			{
				"key": "Actors",
				"doc_count": 2,
				"movice_tags": {
					"doc_count_error_upper_bound": 0,
					"sum_other_doc_count": 0,
					"buckets": 
					[
						{
							"key": "Paul Walker",
							"doc_count": 2
						},
						{
							"key": "Vin Diesel",
							"doc_count": 1
						},
						{
							"key": "Ludacris",
							"doc_count": 1
						}
					]
				}
			}
		]
	}
}

```

But this is the result I get back:

```
"aggregations": {
	"movie_groupName": {
		"doc_count_error_upper_bound": 0,
		"sum_other_doc_count": 0,
		"buckets": [
			{
				"key": "Genre",
				"doc_count": 2,
				"movice_tags": {
					"doc_count_error_upper_bound": 0,
					"sum_other_doc_count": 0,
					"buckets": 
					[
						{
							"key": "Action",
							"doc_count": 2
						},				
						{
							"key": "Paul Walker",
							"doc_count": 2
						},
						{
							"key": "Racing",
							"doc_count": 1
						},
						{
							"key": "Vin Diesel",
							"doc_count": 1
						},
						{
							"key": "Ludacris",
							"doc_count": 1
						}
					]
				}
			},
			{
				"key": "Actors",
				"doc_count": 2,
				"movice_tags": {
					"doc_count_error_upper_bound": 0,
					"sum_other_doc_count": 0,
					"buckets": 
					[
						{
							"key": "Action",
							"doc_count": 2
						},				
						{
							"key": "Paul Walker",
							"doc_count": 2
						},
						{
							"key": "Racing",
							"doc_count": 1
						},
						{
							"key": "Vin Diesel",
							"doc_count": 1
						},
						{
							"key": "Ludacris",
							"doc_count": 1
						}
					]
				}
			}
		]
	}
}

```

Giving me a list that will look like this (which clearly is not what I want):

Genre

- Paul Walker (2)
- Action (2)
- Racing (1)
- Vin Diesel (1)
- Ludacris (1)

Actors

- Paul Walker (2)
- Action (2)
- Racing (1)
- Vin Diesel (1)
- Ludacris (1)

What am I doing wrong?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [November 20, 2017, 2:02pm UTC](https://discuss.elastic.co/t/aggregations-within-aggregations/108399/2 "2017-11-20T14:02:01Z")

</div>

Life is probably made a lot easier if your actors are listed in a field called "actors" and your genres in a field called "genres". e.g.

```
{
    "title": "2 fast 2 furious",
    "actors": ["Paul Walker", "Ludacris"],
    "genres": ["Action"]
}
```

---

<div class="post-metadata">

**Author:** ![Karolinebryn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/karolinebryn/32/17095_2.png) [@Karolinebryn](https://discuss.elastic.co/u/Karolinebryn)\
**Post date:** [November 20, 2017, 2:04pm UTC](https://discuss.elastic.co/t/aggregations-within-aggregations/108399/3 "2017-11-20T14:04:42Z")

</div>

Yeah, I know how to do it like that. But that is not my question. I want the list to be generic.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 20, 2017, 2:06pm UTC](https://discuss.elastic.co/t/aggregations-within-aggregations/108399/4 "2017-11-20T14:06:53Z")

</div>

Probably because your field `tags` is not a `nested` field.  
But I deeply agree with @Mark_Harwood 's comment!

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [November 20, 2017, 2:10pm UTC](https://discuss.elastic.co/t/aggregations-within-aggregations/108399/5 "2017-11-20T14:10:52Z")

</div>

> [@Karolinebryn](#):
>
> But that is not my question. I want the list to be generic

Prepare yourself for pain then 🙂

- Your queries, aggs and docs will all need to adopt `nested` syntax to overcome the cross-matching problem [1] inherent to Lucene.
- Kibana will be unable to use these fields.
- Your index will bloat in size.

[1] [Proposal for nested document support in Lucene | PPT](https://www.slideshare.net/MarkHarwood/proposal-for-nested-document-support-in-lucene)

---

<div class="post-metadata">

**Author:** ![Karolinebryn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/karolinebryn/32/17095_2.png) [@Karolinebryn](https://discuss.elastic.co/u/Karolinebryn)\
**Post date:** [November 20, 2017, 2:10pm UTC](https://discuss.elastic.co/t/aggregations-within-aggregations/108399/6 "2017-11-20T14:10:56Z")

</div>

@dadoonet How do I make it a nested field?  
The problem with @Mark_Harwood s suggestion is that I do not know what the groupNames might be. A list of movies is just an example to simplify my problem. I want to be able to create a list showing the groupnames and number of documents pr tag without having to do any programming every time a groupName is added or removed.

---

<div class="post-metadata">

**Author:** ![Karolinebryn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/karolinebryn/32/17095_2.png) [@Karolinebryn](https://discuss.elastic.co/u/Karolinebryn)\
**Post date:** [November 20, 2017, 2:15pm UTC](https://discuss.elastic.co/t/aggregations-within-aggregations/108399/7 "2017-11-20T14:15:09Z")

</div>

Well this doesn't sound very good 🙈  
I want to group something within a group... Do you have any other suggestions as how to solve this problem?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [November 20, 2017, 2:24pm UTC](https://discuss.elastic.co/t/aggregations-within-aggregations/108399/8 "2017-11-20T14:24:58Z")

</div>

> [@Karolinebryn](#):
>
> I want to group something within a group

People use elasticsearch to do a lot of of grouping groups within groups and then subgroups etc

However, normally the meaning of the data is expressed on the left-hand side of `"fieldname": "value"` pairs. Life is more complicated if the meaning of "value" depends on a neighbouring field/value pair buried in the same object. We then need you to be explicit about _which_ neighbouring item contains the context etc. using `nested`. Messy.

How many unique "groupnames" do you expect to accumulate over time? You might be able to use dynamic mapping to add new fields automatically but this shouldn't be used for too many unique field names (where "too many" is probably a mapping with thousands of fields).

---

<div class="post-metadata">

**Author:** ![Karolinebryn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/karolinebryn/32/17095_2.png) [@Karolinebryn](https://discuss.elastic.co/u/Karolinebryn)\
**Post date:** [November 20, 2017, 2:40pm UTC](https://discuss.elastic.co/t/aggregations-within-aggregations/108399/9 "2017-11-20T14:40:44Z")

</div>

Not that many. Maybe 10-20 unique "groupNames" at max

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [November 20, 2017, 2:55pm UTC](https://discuss.elastic.co/t/aggregations-within-aggregations/108399/10 "2017-11-20T14:55:19Z")

</div>

> [@Karolinebryn](#):
>
> Not that many. Maybe 10-20 unique "groupNames" at max

Totally manageable then.  
Just throw JSON docs with the new fields at elasticsearch and by default it will automatically expand the internal mapping definition used to manage these fields. For example, given this previously-unseen content:

```
{
   "genre" : ["Action", "Science Fiction"],
    "actors" : ["Vin Diesel", "Ludacris"]
}

```

..elasticsearch will automatcally add to the internal mapping definition as follows:

```
{
  "test": {
	"mappings": {
	  "doc": {
		"properties": {
		  "actors": {
			"type": "text",
			"fields": {
			  "keyword": {
				"type": "keyword",
				"ignore_above": 256
			  }
			}
		  },
		  "genre": {
			"type": "text",
			"fields": {
			  "keyword": {
				"type": "keyword",
				"ignore_above": 256
			  }
			}
 ...

```

You can query on the ".text" indexed versions of fields e.g. `actors.text` so that you would match a search for `vin` regardless of case but would use the `keyword` versions of fields in any aggregation e.g. `actor.keyword` so that when returning popular actors you'd get a single keyword `Vin Diesel` rather than the words `vin` and `diesel` as separate items.  
You can nest `terms` aggregations arbitrarily e.g. break down by `genre.keyword` and then the most popular `actor.keyword` values under that.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 18, 2017, 2:55pm UTC](https://discuss.elastic.co/t/aggregations-within-aggregations/108399/11 "2017-12-18T14:55:21Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
