# Faceted search grouped by full term

**URL:** https://discuss.elastic.co/t/faceted-search-grouped-by-full-term/4905
**Category:** Elasticsearch
**Created:** [July 20, 2011, 9:42am UTC](https://discuss.elastic.co/t/faceted-search-grouped-by-full-term/4905 "2011-07-20T09:42:20Z")
**Posts on this page:** 4
**Page:** 1

<div class="post-metadata">

### Author: ![Tania](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tania/32/3123_2.png) [@Tania](https://discuss.elastic.co/u/Tania)
#### Post date: [July 20, 2011, 9:42am UTC](https://discuss.elastic.co/t/faceted-search-grouped-by-full-term/4905/1 "2011-07-20T09:42:20Z")

</div>

Hi,  
I've been working with ES in the last two months and each day I discover new fantastic features.

I would like to know if it it possible to implement the following case:

I have many docs indexed in the ES server, with a location field that stores location information (locality, region and more). One example of the location field (type string) is:  
location: 'Brooklyn, New York, USA'

I would like to do a faceted search and have those results grouped by location information. However if I do the following request:

curl -XGET [http://localhost:9200/test/\_search?pretty=true](http://localhost:9200/test/_search?pretty=true) -d '{  
"query": {  
"query\_string" :{  
"fields" : ["title", "description", "location"],  
"query": "xxx"  
}  
},  
"facets": {  
"location": {  
"terms": {  
"field" : "location"  
}  
}  
}  
}'

I do obtain the results grouped by location, but by each of the terms in location field, this is:

{  
"took" : 6,  
"timed\_out" : false,  
"\_shards" : {  
"total" : 5,  
"successful" : 5,  
"failed" : 0  
},  
"hits" : { ...  
},  
"facets" : {  
"location" : {  
"\_type" : "terms",  
"missing" : 15,  
"terms" : [ {  
"term" : "USA",  
"count" : 10  
}, {  
"term" : " **New**",  
"count" : 4  
}, {  
"term" : " **York**",  
"count" : 4  
}, {  
"term" : "Brooklyn",  
"count" : 2  
},]  
}  
}  
}

If I declare in the mapping the location field as **not analyzed** , then the whole field is taken as a facet.  
I would like an intermediate solution, where the field is analyzed but instead of splitting the term by whitespace, splittiing it by a comma, this is, I would like to obtain a reponse similar to:  
{  
"took" : 6,  
"timed\_out" : false,  
"\_shards" : {  
"total" : 5,  
"successful" : 5,  
"failed" : 0  
},  
"hits" : { ...  
},  
"facets" : {  
"location" : {  
"\_type" : "terms",  
"missing" : 15,  
"terms" : [ {  
"term" : "USA",  
"count" : 10  
}, {  
"term" : " **New York**",  
"count" : 4  
} {  
"term" : "Brooklyn",  
"count" : 2  
},]  
}  
}  
}

I have been searching for this in the ES guide, and maybe setting an appropriate analyzer could make it but I cant guess how to do it, even if it is a correct approach.  
I appreciate any help or guidance. Thanks in advance!  
Tania

---

<div class="post-metadata">

### Author: ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)
#### Post date: [July 24, 2011, 2:28am UTC](https://discuss.elastic.co/t/faceted-search-grouped-by-full-term/4905/2 "2011-07-24T02:28:02Z")

</div>

Heya,

You need to create a custom analyzer to split data by commas. One such  
analyzer can be built using the pattern

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

analyzer,  
or maybe the pattern tokenizer, and your own custom token filters.

On Wed, Jul 20, 2011 at 12:42 PM, Tania [yosoythania@hotmail.com](mailto:yosoythania@hotmail.com) wrote:

> Hi,  
> I've been working with ES in the last two months and each day I discover  
> new  
> fantastic features.
> 
> I would like to know if it it possible to implement the following case:
> 
> I have many docs indexed in the ES server, with a location field that  
> stores  
> location information (locality, region and more). One example of the  
> location field (type string) is:  
> location: 'Brooklyn, New York, USA'
> 
> I would like to do a faceted search and have those results grouped by  
> location information. However if I do the following request:
> 
> curl -XGET [http://localhost:9200/test/\_search?pretty=true](http://localhost:9200/test/_search?pretty=true) -d '{  
> "query": {  
> "query\_string" :{  
> "fields" : ["title", "description", "location"],  
> "query": "xxx"  
> }  
> },  
> "facets": {  
> "location": {  
> "terms": {  
> "field" : "location"  
> }  
> }  
> }  
> }'
> 
> I do obtain the results grouped by location, but by each of the terms in  
> location field, this is:
> 
> {  
> "took" : 6,  
> "timed\_out" : false,  
> "\_shards" : {  
> "total" : 5,  
> "successful" : 5,  
> "failed" : 0  
> },  
> "hits" : { ...  
> },  
> "facets" : {  
> "location" : {  
> "\_type" : "terms",  
> "missing" : 15,  
> "terms" : [ {  
> "term" : "USA",  
> "count" : 10  
> }, {  
> "term" : "_New_",  
> "count" : 4  
> }, {  
> "term" : "_York_",  
> "count" : 4  
> }, {  
> "term" : "Brooklyn",  
> "count" : 2  
> },]  
> }  
> }  
> }
> 
> If I declare in the mapping the location field as _not analyzed_, then the  
> whole field is taken as a facet.  
> I would like an intermediate solution, where the field is analyzed but  
> instead of splitting the term by whitespace, splittiing it by a comma, this  
> is, I would like to obtain a reponse similar to:  
> {  
> "took" : 6,  
> "timed\_out" : false,  
> "\_shards" : {  
> "total" : 5,  
> "successful" : 5,  
> "failed" : 0  
> },  
> "hits" : { ...  
> },  
> "facets" : {  
> "location" : {  
> "\_type" : "terms",  
> "missing" : 15,  
> "terms" : [ {  
> "term" : "USA",  
> "count" : 10  
> }, {  
> "term" : "_New York_",  
> "count" : 4  
> } {  
> "term" : "Brooklyn",  
> "count" : 2  
> },]  
> }  
> }  
> }
> 
> I have been searching for this in the ES guide, and maybe setting an  
> appropriate analyzer could make it but I cant guess how to do it, even if  
> it  
> is a correct approach.  
> I appreciate any help or guidance. Thanks in advance!  
> Tania
> 
> --  
> View this message in context:  
> [http://elasticsearch-users.115913.n3.nabble.com/Faceted-search-grouped-by-full-term-tp3185007p3185007.html](http://elasticsearch-users.115913.n3.nabble.com/Faceted-search-grouped-by-full-term-tp3185007p3185007.html)  
> Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).

---

<div class="post-metadata">

### Author: ![Tania](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tania/32/3123_2.png) [@Tania](https://discuss.elastic.co/u/Tania)
#### Post date: [July 26, 2011, 4:17pm UTC](https://discuss.elastic.co/t/faceted-search-grouped-by-full-term/4905/3 "2011-07-26T16:17:32Z")

</div>

Thanks! I did what you explain and it worked!  
In case anyone needs it, I explain briefly how I solved it.  
While the index is being created, to especify the special mappings needed, you can also define the analyzers linked to this index.  
In this case:

curl -POST 'localhost:9200/test' -d '  
{  
"settings":{  
"analysis": {  
"analyzer": {  
"comma":{  
"type": "pattern",  
"pattern":","  
}  
}  
}  
},  
"mappings" : {  
"photos" : {  
"properties" : {  
"tags" : {  
"type" : "string",  
"analyzer": "comma"  
}  
}  
}  
}  
}'

After setting the index mappings and settings specifying that the "tags" field is anayzed using the pattern analyzer called "comma" (consists on a comma), the search and facet tasks and so on consider that the token to analyze this strings is a comma, instead of whitespace.

Many thanks kimchy!

Date: Sat, 23 Jul 2011 19:28:21 -0700

```
Heya,

```

You need to create a custom analyzer to split data by commas. One such analyzer can be built using the pattern [http://www.elasticsearch.org/guide/reference/index-modules/analysis/pattern-analyzer.html](http://www.elasticsearch.org/guide/reference/index-modules/analysis/pattern-analyzer.html) analyzer, or maybe the pattern tokenizer, and your own custom token filters.

On Wed, Jul 20, 2011 at 12:42 PM, Tania \<[hidden email]\> wrote:

Hi,

I've been working with ES in the last two months and each day I discover new

fantastic features.

I would like to know if it it possible to implement the following case:

I have many docs indexed in the ES server, with a location field that stores

location information (locality, region and more). One example of the

location field (type string) is:

location: 'Brooklyn, New York, USA'

I would like to do a faceted search and have those results grouped by

location information. However if I do the following request:

curl -XGET [http://localhost:9200/test/\_search?pretty=true](http://localhost:9200/test/_search?pretty=true) -d '{

"query": {

```
"query_string" :{

  "fields" : ["title", "description", "location"],

  "query": "xxx"

}

```

},

"facets": {

```
"location": {

  "terms": {

    "field" : "location"

  }

}

```

}

}'

I do obtain the results grouped by location, but by each of the terms in

location field, this is:

{

"took" : 6,

"timed\_out" : false,

"\_shards" : {

```
"total" : 5,

"successful" : 5,

"failed" : 0

```

},

"hits" : { ...

},

"facets" : {

```
"location" : {

  "_type" : "terms",

  "missing" : 15,

  "terms" : [ {

    "term" : "USA",

    "count" : 10

  }, {

    "term" : "*New*",

    "count" : 4

  }, {

    "term" : "*York*",

    "count" : 4

  }, {

    "term" : "Brooklyn",

    "count" : 2

  },]

}

```

}

}

If I declare in the mapping the location field as _not analyzed_, then the

whole field is taken as a facet.

I would like an intermediate solution, where the field is analyzed but

instead of splitting the term by whitespace, splittiing it by a comma, this

is, I would like to obtain a reponse similar to:

{

"took" : 6,

"timed\_out" : false,

"\_shards" : {

```
"total" : 5,

"successful" : 5,

"failed" : 0

```

},

"hits" : { ...

},

"facets" : {

```
"location" : {

  "_type" : "terms",

  "missing" : 15,

  "terms" : [ {

    "term" : "USA",

    "count" : 10

  }, {

    "term" : "*New York*",

    "count" : 4

  } {

    "term" : "Brooklyn",

    "count" : 2

  },]

}

```

}

}

I have been searching for this in the ES guide, and maybe setting an

appropriate analyzer could make it but I cant guess how to do it, even if it

is a correct approach.

I appreciate any help or guidance. Thanks in advance!

Tania

--

View this message in context: [http://elasticsearch-users.115913.n3.nabble.com/Faceted-search-grouped-by-full-term-tp3185007p3185007.html](http://elasticsearch-users.115913.n3.nabble.com/Faceted-search-grouped-by-full-term-tp3185007p3185007.html)

Sent from the ElasticSearch Users mailing list archive at [Nabble.com](http://Nabble.com).

```
	If you reply to this email, your message will be added to the discussion below:
	http://elasticsearch-users.115913.n3.nabble.com/Faceted-search-grouped-by-full-term-tp3185007p3194581.html

	
	To unsubscribe from Faceted search grouped by full term, click here.
```

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 6, 2017, 3:59am UTC](https://discuss.elastic.co/t/faceted-search-grouped-by-full-term/4905/4 "2017-07-06T03:59:13Z")

</div>


