# Getting Source data in aggregation results

**URL:** <https://discuss.elastic.co/t/getting-source-data-in-aggregation-results/234328>\
**Category:** Elasticsearch\
**Created:** [May 26, 2020, 12:21pm UTC](https://discuss.elastic.co/t/getting-source-data-in-aggregation-results/234328 "2020-05-26T12:21:12Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![pranav24](https://avatars.discourse-cdn.com/v4/letter/p/b19c9b/32.png) [@pranav24](https://discuss.elastic.co/u/pranav24)\
**Post date:** [May 26, 2020, 12:21pm UTC](https://discuss.elastic.co/t/getting-source-data-in-aggregation-results/234328/1 "2020-05-26T12:21:12Z")

</div>

I have an index that consists of nested and normal fields.  
The structure of my index is:

```
{
	"name" : "Walter white",
	"age" : "20",
	"email" : "walter.white@gmail.com",
	"subjects" :[
		{
			"subject_name" : "Computer Science",
			"marks" : "80"
		},
		{
			"subject_name" : "Maths",
			"marks" : "95"
		},
		{
			"subject_name" : "Physics",
			"marks" : "90"
		}
	]
}

```

Now, I want to create a report which contains all the data of the students group by their age.  
I have created a query like this:

```
{
  "_source": false,
  "aggs": {
    "ageGroup": {
      "terms": {
        "field": "age"
      },
      "aggs": {
        "top_sales_hits": {
          "top_hits": {
            "size": 100
          }
        }
      }
    }
  }
}

```

I'm getting the desired result. But it is taking too much time to return the result.  
Is there any other way to do the same?

---

<div class="post-metadata">

**Author:** ![pranav24](https://avatars.discourse-cdn.com/v4/letter/p/b19c9b/32.png) [@pranav24](https://discuss.elastic.co/u/pranav24)\
**Post date:** [May 28, 2020, 11:34am UTC](https://discuss.elastic.co/t/getting-source-data-in-aggregation-results/234328/2 "2020-05-28T11:34:53Z")

</div>

Can anyone please help me on this?

---

<div class="post-metadata">

**Author:** ![Hendrik\_Muhs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hendrik_muhs/32/25802_2.png) [@Hendrik\_Muhs](https://discuss.elastic.co/u/Hendrik_Muhs)\
**Post date:** [May 28, 2020, 12:46pm UTC](https://discuss.elastic.co/t/getting-source-data-in-aggregation-results/234328/3 "2020-05-28T12:46:24Z")

</div>

Hi,

you could create a [transform](https://www.elastic.co/guide/en/elasticsearch/reference/current/transforms.html) which indexes the results in another index. With pivot you can `group_by` age buckets as you described.

To access the source you can use a scripted metric aggregation, the following one would take every input document and store it in an array:

```auto
"all_docs": {
  "scripted_metric": {
    "init_script": "state.docs = []",
    "map_script": "state.docs.add(new HashMap(params['_source']))",
    "combine_script": "return state.docs",
    "reduce_script": "def docs = []; for (s in states) {for (d in s) { docs.add(d);}}return docs"
  }
}

```

This might not be what you want, but I hope I give you something to start with.

---

<div class="post-metadata">

**Author:** ![pranav24](https://avatars.discourse-cdn.com/v4/letter/p/b19c9b/32.png) [@pranav24](https://discuss.elastic.co/u/pranav24)\
**Post date:** [June 1, 2020, 12:51pm UTC](https://discuss.elastic.co/t/getting-source-data-in-aggregation-results/234328/4 "2020-06-01T12:51:26Z")

</div>

Thanks @Hendrik_Muhs  
This is my normal use case every user can do this multiple time a day with different field used for group by.  
Is it feasible to create a index again and again for every request.

---

<div class="post-metadata">

**Author:** ![Hendrik\_Muhs](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/hendrik_muhs/32/25802_2.png) [@Hendrik\_Muhs](https://discuss.elastic.co/u/Hendrik_Muhs)\
**Post date:** [June 3, 2020, 6:25am UTC](https://discuss.elastic.co/t/getting-source-data-in-aggregation-results/234328/5 "2020-06-03T06:25:38Z")

</div>

You can create as many indexes as your cluster can hold, however I wonder if transform is the right choice if you do _not_ reuse the output and if you are only interested in the result _once_. You would need to program against the API and manage the created transforms/indices in the background.

Your original concern was `taking too much time to return the result`. With [Async search](https://www.elastic.co/guide/en/elasticsearch/reference/7.7/async-search.html) you can create the search async and pull the result later. This does not leave indices behind, however results must be retrieved during life-time.

---

<div class="post-metadata">

**Author:** ![pranav24](https://avatars.discourse-cdn.com/v4/letter/p/b19c9b/32.png) [@pranav24](https://discuss.elastic.co/u/pranav24)\
**Post date:** [June 3, 2020, 7:52am UTC](https://discuss.elastic.co/t/getting-source-data-in-aggregation-results/234328/6 "2020-06-03T07:52:02Z")

</div>

Thank you @Hendrik_Muhs  
I was not aware of the Async search feature of Elastic search. I will look into it.

Can you tell me is there any alternative to top\_hits aggregation, to get source data in response from Elasticsearch?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 1, 2020, 7:52am UTC](https://discuss.elastic.co/t/getting-source-data-in-aggregation-results/234328/7 "2020-07-01T07:52:05Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
