# Terms facet is tokenizing a field with special characters

**URL:** <https://discuss.elastic.co/t/terms-facet-is-tokenizing-a-field-with-special-characters/4071>\
**Category:** Elasticsearch\
**Created:** [March 10, 2011, 3:12am UTC](https://discuss.elastic.co/t/terms-facet-is-tokenizing-a-field-with-special-characters/4071 "2011-03-10T03:12:21Z")\
**Posts on this page:** 17\
**Page:** 1

<div class="post-metadata">

**Author:** ![jason](https://avatars.discourse-cdn.com/v4/letter/j/dc4da7/32.png) [@jason](https://discuss.elastic.co/u/jason)\
**Post date:** [March 10, 2011, 3:12am UTC](https://discuss.elastic.co/t/terms-facet-is-tokenizing-a-field-with-special-characters/4071/1 "2011-03-10T03:12:21Z")

</div>

Hi,

I am running a terms facet query on non-numeric field, and it is  
working fine. However, if the field contains % symbol, the resutls in  
a TermsFacet are tokenized into separate terms. For example, if I  
have a json document:

{  
url: http%[3Fwww.myserver.com/someaction/%3Fname%3DBob](http://3Fwww.myserver.com/someaction/%3Fname%3DBob)  
}

and if I run: curl -X GET [http://localhost:9200/\_river/my\_idx/\_search](http://localhost:9200/_river/my_idx/_search)  
-d  
{  
"query" : {  
"match\_all" : {}  
},  
"facets" : {  
"facet1" : {  
"terms" : {  
"field" : "url"  
}  
}  
}  
}

then it is returninng {"terms": [{"term": "[www.myserver.com](http://www.myserver.com)",  
"count":1}, {"term": "http", "count":1}, {"term":"3Fname", "count":1},  
{"term":"3DBob", "count":1}

I think it is supposed to return {"terms": [ {"term":"http  
%[3Fwww.myserver.com/someaction/%3Fname%3DBob","count](http://3Fwww.myserver.com/someaction/%3Fname%3DBob%22,%22count)":1}

Please correct me if I am wrong. It never happens if a field doesn't  
contain %.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [March 10, 2011, 6:18am UTC](https://discuss.elastic.co/t/terms-facet-is-tokenizing-a-field-with-special-characters/4071/2 "2011-03-10T06:18:07Z")

</div>

You need to have the field you run the terms facet on marked as not\_analyzed in its mapping. You will need to reindex the data in order for that to take affect (and of course, create the mappings before you index data).  
On Thursday, March 10, 2011 at 5:12 AM, eugene wrote:

> Hi,
> 
> I am running a terms facet query on non-numeric field, and it is  
> working fine. However, if the field contains % symbol, the resutls in  
> a TermsFacet are tokenized into separate terms. For example, if I  
> have a json document:
> 
> {  
> url: http%[3Fwww.myserver.com/someaction/%3Fname%3DBob](http://3Fwww.myserver.com/someaction/%3Fname%3DBob)  
> }
> 
> and if I run: curl -X GET [http://localhost:9200/\_river/my\_idx/\_search](http://localhost:9200/_river/my_idx/_search)  
> -d  
> {  
> "query" : {  
> "match\_all" : {}  
> },  
> "facets" : {  
> "facet1" : {  
> "terms" : {  
> "field" : "url"  
> }  
> }  
> }  
> }
> 
> then it is returninng {"terms": [{"term": "[www.myserver.com](http://www.myserver.com)",  
> "count":1}, {"term": "http", "count":1}, {"term":"3Fname", "count":1},  
> {"term":"3DBob", "count":1}
> 
> I think it is supposed to return {"terms": [ {"term":"http  
> %[3Fwww.myserver.com/someaction/%3Fname%3DBob","count](http://3Fwww.myserver.com/someaction/%3Fname%3DBob%22,%22count)":1}
> 
> Please correct me if I am wrong. It never happens if a field doesn't  
> contain %.

---

<div class="post-metadata">

**Author:** ![jason](https://avatars.discourse-cdn.com/v4/letter/j/dc4da7/32.png) [@jason](https://discuss.elastic.co/u/jason)\
**Post date:** [March 10, 2011, 9:03am UTC](https://discuss.elastic.co/t/terms-facet-is-tokenizing-a-field-with-special-characters/4071/3 "2011-03-10T09:03:35Z")

</div>

Thank you Shay.

Is there any way to add mappings to existing index? I searched the  
source code and only found the examples for CreateIndexBuilder class.

Eugene.

On Mar 9, 10:18 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> You need to have the field you run the terms facet on marked as not\_analyzed in its mapping. You will need to reindex the data in order for that to take affect (and of course, create the mappings before you index data).
> 
> On Thursday, March 10, 2011 at 5:12 AM, eugene wrote:
> 
> > Hi,
> 
> > I am running a terms facet query on non-numeric field, and it is  
> > working fine. However, if the field contains % symbol, the resutls in  
> > a TermsFacet are tokenized into separate terms. For example, if I  
> > have a json document:
> 
> > {  
> > url: http%[3Fwww.myserver.com/someaction/%3Fname%3DBob](http://3Fwww.myserver.com/someaction/%3Fname%3DBob)  
> > }
> 
> > and if I run: curl -X GEThttp://localhost:9200/\_river/my\_idx/\_search  
> > -d  
> > {  
> > "query" : {  
> > "match\_all" : {}  
> > },  
> > "facets" : {  
> > "facet1" : {  
> > "terms" : {  
> > "field" : "url"  
> > }  
> > }  
> > }  
> > }
> 
> > then it is returninng {"terms": [{"term": "[www.myserver.com](http://www.myserver.com)",  
> > "count":1}, {"term": "http", "count":1}, {"term":"3Fname", "count":1},  
> > {"term":"3DBob", "count":1}
> 
> > I think it is supposed to return {"terms": [ {"term":"http  
> > %[3Fwww.myserver.com/someaction/%3Fname%3DBob","count](http://3Fwww.myserver.com/someaction/%3Fname%3DBob%22,%22count)":1}
> 
> > Please correct me if I am wrong. It never happens if a field doesn't  
> > contain %.- Hide quoted text -
> 
> - Show quoted text -

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [March 10, 2011, 12:16pm UTC](https://discuss.elastic.co/t/terms-facet-is-tokenizing-a-field-with-special-characters/4071/4 "2011-03-10T12:16:39Z")

</div>

There is an API to "putMapping" on an index, but, you can't change a field that is already analyzed to be "not\_analyzed" (as its part of the indexing process, so all current indexed docs will be meaningless).  
On Thursday, March 10, 2011 at 11:03 AM, eugene wrote:

> Thank you Shay.
> 
> Is there any way to add mappings to existing index? I searched the  
> source code and only found the examples for CreateIndexBuilder class.
> 
> Eugene.
> 
> On Mar 9, 10:18 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > You need to have the field you run the terms facet on marked as not\_analyzed in its mapping. You will need to reindex the data in order for that to take affect (and of course, create the mappings before you index data).
> > 
> > On Thursday, March 10, 2011 at 5:12 AM, eugene wrote:
> > 
> > > Hi,
> > 
> > > I am running a terms facet query on non-numeric field, and it is  
> > > working fine. However, if the field contains % symbol, the resutls in  
> > > a TermsFacet are tokenized into separate terms. For example, if I  
> > > have a json document:
> > 
> > > {  
> > > url: http%[3Fwww.myserver.com/someaction/%3Fname%3DBob](http://3Fwww.myserver.com/someaction/%3Fname%3DBob)  
> > > }
> > 
> > > and if I run: curl -X GEThttp://localhost:9200/\_river/my\_idx/\_search  
> > > -d  
> > > {  
> > > "query" : {  
> > > "match\_all" : {}  
> > > },  
> > > "facets" : {  
> > > "facet1" : {  
> > > "terms" : {  
> > > "field" : "url"  
> > > }  
> > > }  
> > > }  
> > > }
> > 
> > > then it is returninng {"terms": [{"term": "[www.myserver.com](http://www.myserver.com)",  
> > > "count":1}, {"term": "http", "count":1}, {"term":"3Fname", "count":1},  
> > > {"term":"3DBob", "count":1}
> > 
> > > I think it is supposed to return {"terms": [ {"term":"http  
> > > %[3Fwww.myserver.com/someaction/%3Fname%3DBob","count](http://3Fwww.myserver.com/someaction/%3Fname%3DBob%22,%22count)":1}
> > 
> > > Please correct me if I am wrong. It never happens if a field doesn't  
> > > contain %.- Hide quoted text -
> > 
> > - Show quoted text -

---

<div class="post-metadata">

**Author:** ![jason](https://avatars.discourse-cdn.com/v4/letter/j/dc4da7/32.png) [@jason](https://discuss.elastic.co/u/jason)\
**Post date:** [March 10, 2011, 11:03pm UTC](https://discuss.elastic.co/t/terms-facet-is-tokenizing-a-field-with-special-characters/4071/5 "2011-03-10T23:03:13Z")

</div>

Ok, I am doing this:

```
        	String mappings =

```

XContentFactory.jsonBuilder().startObject().startObject("type1")  
.startObject("myfield").field("type",  
"string").field("store", "yes").field("index",  
"not\_analyzed").endObject().endObject().endObject().toString();

```
			client.admin().indices().preparePutMapping().setType("type1")
				.setSource(mappings).execute().actionGet();

```

I am getting the following exception (notice: I replaced real ip  
address with "ip\_address" on the first line).  
What do you think it indicates? Thank you!

Eugene.

org.elasticsearch.transport.RemoteTransportException: [Stonecutter]  
[inet[/ip\_address:9300]][indices/mapping/put]  
Caused by: org.elasticsearch.ElasticSearchParseException: Failed to  
derive xcontent from  
org.elasticsearch.common.xcontent.XContentBuilder@1c94b8f  
at  
org.elasticsearch.common.xcontent.XContentFactory.xContent(XContentFactory.java:  
136)  
at  
org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.extractMapping(XContentDocumentMapperParser.java:  
316)  
at  
org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(XContentDocumentMapperParser.java:  
114)  
at  
org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(XContentDocumentMapperParser.java:  
54)  
at  
org.elasticsearch.index.mapper.MapperService.parse(MapperService.java:  
209)  
at org.elasticsearch.cluster.metadata.MetaDataMappingService  
$3.execute(MetaDataMappingService.java:193)  
at org.elasticsearch.cluster.service.InternalClusterService  
$2.run(InternalClusterService.java:175)  
at java.util.concurrent.ThreadPoolExecutor  
$Worker.runTask(ThreadPoolExecutor.java:886)  
at java.util.concurrent.ThreadPoolExecutor  
$Worker.run(ThreadPoolExecutor.java:908)  
at java.lang.Thread.run(Thread.java:662)

On Mar 10, 4:16 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> There is an API to "putMapping" on an index, but, you can't change a field that is already analyzed to be "not\_analyzed" (as its part of the indexing process, so all current indexed docs will be meaningless).
> 
> On Thursday, March 10, 2011 at 11:03 AM, eugene wrote:
> 
> > Thank you Shay.
> 
> > Is there any way to add mappings to existing index? I searched the  
> > source code and only found the examples for CreateIndexBuilder class.
> 
> > Eugene.
> 
> > On Mar 9, 10:18 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > 
> > > You need to have the field you run the terms facet on marked as not\_analyzed in its mapping. You will need to reindex the data in order for that to take affect (and of course, create the mappings before you index data).
> 
> > > On Thursday, March 10, 2011 at 5:12 AM, eugene wrote:
> > > 
> > > > Hi,
> 
> > > > I am running a terms facet query on non-numeric field, and it is  
> > > > working fine. However, if the field contains % symbol, the resutls in  
> > > > a TermsFacet are tokenized into separate terms. For example, if I  
> > > > have a json document:
> 
> > > > {  
> > > > url: http%[3Fwww.myserver.com/someaction/%3Fname%3DBob](http://3Fwww.myserver.com/someaction/%3Fname%3DBob)  
> > > > }
> 
> > > > and if I run: curl -X GEThttp://localhost:9200/\_river/my\_idx/\_search  
> > > > -d  
> > > > {  
> > > > "query" : {  
> > > > "match\_all" : {}  
> > > > },  
> > > > "facets" : {  
> > > > "facet1" : {  
> > > > "terms" : {  
> > > > "field" : "url"  
> > > > }  
> > > > }  
> > > > }  
> > > > }
> 
> > > > then it is returninng {"terms": [{"term": "[www.myserver.com](http://www.myserver.com)",  
> > > > "count":1}, {"term": "http", "count":1}, {"term":"3Fname", "count":1},  
> > > > {"term":"3DBob", "count":1}
> 
> > > > I think it is supposed to return {"terms": [ {"term":"http  
> > > > %[3Fwww.myserver.com/someaction/%3Fname%3DBob","count](http://3Fwww.myserver.com/someaction/%3Fname%3DBob%22,%22count)":1}
> 
> > > > Please correct me if I am wrong. It never happens if a field doesn't  
> > > > contain %.- Hide quoted text -
> 
> > > - Show quoted text -- Hide quoted text -
> 
> - Show quoted text -

---

<div class="post-metadata">

**Author:** ![jason](https://avatars.discourse-cdn.com/v4/letter/j/dc4da7/32.png) [@jason](https://discuss.elastic.co/u/jason)\
**Post date:** [March 11, 2011, 2:14am UTC](https://discuss.elastic.co/t/terms-facet-is-tokenizing-a-field-with-special-characters/4071/6 "2011-03-11T02:14:33Z")

</div>

I solved this problem by creating a new index first, and only then  
putting maps to it. I guess mappings should be done only once on a  
new index/type pair.

-Eugene.

On Mar 10, 3:03 pm, eugene [efur...@gmail.com](mailto:efur...@gmail.com) wrote:

> Ok, I am doing this:
> 
> ```
> String mappings =
> 
> ```
> 
> XContentFactory.jsonBuilder().startObject().startObject("type1")  
> .startObject("myfield").field("type",  
> "string").field("store", "yes").field("index",  
> "not\_analyzed").endObject().endObject().endObject().toString();
> 
> ```
> client.admin().indices().preparePutMapping().setType("type1")
> .setSource(mappings).execute().actionGet();
> 
> ```
> 
> I am getting the following exception (notice: I replaced real ip  
> address with "ip\_address" on the first line).  
> What do you think it indicates? Thank you!
> 
> Eugene.
> 
> org.elasticsearch.transport.RemoteTransportException: [Stonecutter]  
> [inet[/ip\_address:9300]][indices/mapping/put]  
> Caused by: org.elasticsearch.ElasticSearchParseException: Failed to  
> derive xcontent from  
> org.elasticsearch.common.xcontent.XContentBuilder@1c94b8f  
> at  
> org.elasticsearch.common.xcontent.XContentFactory.xContent(XContentFactory.­java:  
> 136)  
> at  
> org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.extrac­tMapping(XContentDocumentMapperParser.java:  
> 316)  
> at  
> org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(­XContentDocumentMapperParser.java:  
> 114)  
> at  
> org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(­XContentDocumentMapperParser.java:  
> 54)  
> at  
> org.elasticsearch.index.mapper.MapperService.parse(MapperService.java:  
> 209)  
> at org.elasticsearch.cluster.metadata.MetaDataMappingService  
> $3.execute(MetaDataMappingService.java:193)  
> at org.elasticsearch.cluster.service.InternalClusterService  
> $2.run(InternalClusterService.java:175)  
> at java.util.concurrent.ThreadPoolExecutor  
> $Worker.runTask(ThreadPoolExecutor.java:886)  
> at java.util.concurrent.ThreadPoolExecutor  
> $Worker.run(ThreadPoolExecutor.java:908)  
> at java.lang.Thread.run(Thread.java:662)
> 
> On Mar 10, 4:16 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > There is an API to "putMapping" on an index, but, you can't change a field that is already analyzed to be "not\_analyzed" (as its part of the indexing process, so all current indexed docs will be meaningless).
> 
> > On Thursday, March 10, 2011 at 11:03 AM, eugene wrote:
> > 
> > > Thank you Shay.
> 
> > > Is there any way to add mappings to existing index? I searched the  
> > > source code and only found the examples for CreateIndexBuilder class.
> 
> > > Eugene.
> 
> > > On Mar 9, 10:18 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > 
> > > > You need to have the field you run the terms facet on marked as not\_analyzed in its mapping. You will need to reindex the data in order for that to take affect (and of course, create the mappings before you index data).
> 
> > > > On Thursday, March 10, 2011 at 5:12 AM, eugene wrote:
> > > > 
> > > > > Hi,
> 
> > > > > I am running a terms facet query on non-numeric field, and it is  
> > > > > working fine. However, if the field contains % symbol, the resutls in  
> > > > > a TermsFacet are tokenized into separate terms. For example, if I  
> > > > > have a json document:
> 
> > > > > {  
> > > > > url: http%[3Fwww.myserver.com/someaction/%3Fname%3DBob](http://3Fwww.myserver.com/someaction/%3Fname%3DBob)  
> > > > > }
> 
> > > > > and if I run: curl -X GEThttp://localhost:9200/\_river/my\_idx/\_search  
> > > > > -d  
> > > > > {  
> > > > > "query" : {  
> > > > > "match\_all" : {}  
> > > > > },  
> > > > > "facets" : {  
> > > > > "facet1" : {  
> > > > > "terms" : {  
> > > > > "field" : "url"  
> > > > > }  
> > > > > }  
> > > > > }  
> > > > > }
> 
> > > > > then it is returninng {"terms": [{"term": "[www.myserver.com](http://www.myserver.com)",  
> > > > > "count":1}, {"term": "http", "count":1}, {"term":"3Fname", "count":1},  
> > > > > {"term":"3DBob", "count":1}
> 
> > > > > I think it is supposed to return {"terms": [ {"term":"http  
> > > > > %[3Fwww.myserver.com/someaction/%3Fname%3DBob","count](http://3Fwww.myserver.com/someaction/%3Fname%3DBob%22,%22count)":1}
> 
> > > > > Please correct me if I am wrong. It never happens if a field doesn't  
> > > > > contain %.- Hide quoted text -
> 
> > > > - Show quoted text -- Hide quoted text -
> 
> > - Show quoted text -- Hide quoted text -
> 
> - Show quoted text -

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [March 11, 2011, 1:11pm UTC](https://discuss.elastic.co/t/terms-facet-is-tokenizing-a-field-with-special-characters/4071/7 "2011-03-11T13:11:02Z")

</div>

Not sure that this solved your problem, but a different change that you did where you don't use toString on the builder that represents the mapping. You can pass the builder directory to the putMapping API.  
On Friday, March 11, 2011 at 4:14 AM, eugene wrote:

> I solved this problem by creating a new index first, and only then  
> putting maps to it. I guess mappings should be done only once on a  
> new index/type pair.
> 
> -Eugene.
> 
> On Mar 10, 3:03 pm, eugene [efur...@gmail.com](mailto:efur...@gmail.com) wrote:
> 
> > Ok, I am doing this:
> > 
> > String mappings =  
> > XContentFactory.jsonBuilder().startObject().startObject("type1")  
> > .startObject("myfield").field("type",  
> > "string").field("store", "yes").field("index",  
> > "not\_analyzed").endObject().endObject().endObject().toString();
> > 
> > client.admin().indices().preparePutMapping().setType("type1")  
> > .setSource(mappings).execute().actionGet();
> > 
> > I am getting the following exception (notice: I replaced real ip  
> > address with "ip\_address" on the first line).  
> > What do you think it indicates? Thank you!
> > 
> > Eugene.
> > 
> > org.elasticsearch.transport.RemoteTransportException: [Stonecutter]  
> > [inet[/ip\_address:9300]][indices/mapping/put]  
> > Caused by: org.elasticsearch.ElasticSearchParseException: Failed to  
> > derive xcontent from  
> > org.elasticsearch.common.xcontent.XContentBuilder@1c94b8f  
> > at  
> > org.elasticsearch.common.xcontent.XContentFactory.xContent(XContentFactory.Â­java:  
> > 136)  
> > at  
> > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.extracÂ­tMapping(XContentDocumentMapperParser.java:  
> > 316)  
> > at  
> > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(Â­XContentDocumentMapperParser.java:  
> > 114)  
> > at  
> > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(Â­XContentDocumentMapperParser.java:  
> > 54)  
> > at  
> > org.elasticsearch.index.mapper.MapperService.parse(MapperService.java:  
> > 209)  
> > at org.elasticsearch.cluster.metadata.MetaDataMappingService  
> > $3.execute(MetaDataMappingService.java:193)  
> > at org.elasticsearch.cluster.service.InternalClusterService  
> > $2.run(InternalClusterService.java:175)  
> > at java.util.concurrent.ThreadPoolExecutor  
> > $Worker.runTask(ThreadPoolExecutor.java:886)  
> > at java.util.concurrent.ThreadPoolExecutor  
> > $Worker.run(ThreadPoolExecutor.java:908)  
> > at java.lang.Thread.run(Thread.java:662)
> > 
> > On Mar 10, 4:16 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > 
> > > There is an API to "putMapping" on an index, but, you can't change a field that is already analyzed to be "not\_analyzed" (as its part of the indexing process, so all current indexed docs will be meaningless).
> > 
> > > On Thursday, March 10, 2011 at 11:03 AM, eugene wrote:
> > > 
> > > > Thank you Shay.
> > 
> > > > Is there any way to add mappings to existing index? I searched the  
> > > > source code and only found the examples for CreateIndexBuilder class.
> > 
> > > > Eugene.
> > 
> > > > On Mar 9, 10:18 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > > 
> > > > > You need to have the field you run the terms facet on marked as not\_analyzed in its mapping. You will need to reindex the data in order for that to take affect (and of course, create the mappings before you index data).
> > 
> > > > > On Thursday, March 10, 2011 at 5:12 AM, eugene wrote:
> > > > > 
> > > > > > Hi,
> > 
> > > > > > I am running a terms facet query on non-numeric field, and it is  
> > > > > > working fine. However, if the field contains % symbol, the resutls in  
> > > > > > a TermsFacet are tokenized into separate terms. For example, if I  
> > > > > > have a json document:
> > 
> > > > > > {  
> > > > > > url: http%[3Fwww.myserver.com/someaction/%3Fname%3DBob](http://3Fwww.myserver.com/someaction/%3Fname%3DBob)  
> > > > > > }
> > 
> > > > > > and if I run: curl -X GEThttp://localhost:9200/\_river/my\_idx/\_search  
> > > > > > -d  
> > > > > > {  
> > > > > > "query" : {  
> > > > > > "match\_all" : {}  
> > > > > > },  
> > > > > > "facets" : {  
> > > > > > "facet1" : {  
> > > > > > "terms" : {  
> > > > > > "field" : "url"  
> > > > > > }  
> > > > > > }  
> > > > > > }  
> > > > > > }
> > 
> > > > > > then it is returninng {"terms": [{"term": "[www.myserver.com](http://www.myserver.com)",  
> > > > > > "count":1}, {"term": "http", "count":1}, {"term":"3Fname", "count":1},  
> > > > > > {"term":"3DBob", "count":1}
> > 
> > > > > > I think it is supposed to return {"terms": [ {"term":"http  
> > > > > > %[3Fwww.myserver.com/someaction/%3Fname%3DBob","count](http://3Fwww.myserver.com/someaction/%3Fname%3DBob%22,%22count)":1}
> > 
> > > > > > Please correct me if I am wrong. It never happens if a field doesn't  
> > > > > > contain %.- Hide quoted text -
> > 
> > > > > - Show quoted text -- Hide quoted text -
> > 
> > > - Show quoted text -- Hide quoted text -
> > 
> > - Show quoted text -

---

<div class="post-metadata">

**Author:** ![jason](https://avatars.discourse-cdn.com/v4/letter/j/dc4da7/32.png) [@jason](https://discuss.elastic.co/u/jason)\
**Post date:** [March 14, 2011, 5:59pm UTC](https://discuss.elastic.co/t/terms-facet-is-tokenizing-a-field-with-special-characters/4071/8 "2011-03-14T17:59:44Z")

</div>

Shay,

It appears I didn't fix the error. I run the following:

```
     "query" : {
"match_all": {}
     },
     "facets" : {
           "url" : {
                     "terms" : {
                               "field" : "url"
                    }
           }
    }

```

I am still getting tokenized terms for "url" field. I checked the  
url field and it is set not\_analyzed.

C:\>curl -X GET [http://localhost:9200/test/type1/\_mapping](http://localhost:9200/test/type1/_mapping)  
{"test":{"type1":{"properties":{"timestamp":  
{"store":"yes","type":"string"},"url":  
{"index":"not\_analyzed","store":"yes","type":"string"}}}}}

Am I using the wrong facet?

Thank you,  
Eugene.

On Mar 11, 6:11 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> Not sure that this solved your problem, but a different change that you did where you don't use toString on the builder that represents the mapping. You can pass the builder directory to the putMapping API.
> 
> On Friday, March 11, 2011 at 4:14 AM, eugene wrote:
> 
> > I solved this problem by creating a new index first, and only then  
> > putting maps to it. I guess mappings should be done only once on a  
> > new index/type pair.
> 
> > -Eugene.
> 
> > On Mar 10, 3:03 pm, eugene [efur...@gmail.com](mailto:efur...@gmail.com) wrote:
> > 
> > > Ok, I am doing this:
> 
> > > String mappings =  
> > > XContentFactory.jsonBuilder().startObject().startObject("type1")  
> > > .startObject("myfield").field("type",  
> > > "string").field("store", "yes").field("index",  
> > > "not\_analyzed").endObject().endObject().endObject().toString();
> 
> > > client.admin().indices().preparePutMapping().setType("type1")  
> > > .setSource(mappings).execute().actionGet();
> 
> > > I am getting the following exception (notice: I replaced real ip  
> > > address with "ip\_address" on the first line).  
> > > What do you think it indicates? Thank you!
> 
> > > Eugene.
> 
> > > org.elasticsearch.transport.RemoteTransportException: [Stonecutter]  
> > > [inet[/ip\_address:9300]][indices/mapping/put]  
> > > Caused by: org.elasticsearch.ElasticSearchParseException: Failed to  
> > > derive xcontent from  
> > > org.elasticsearch.common.xcontent.XContentBuilder@1c94b8f  
> > > at  
> > > org.elasticsearch.common.xcontent.XContentFactory.xContent(XContentFactory.­­java:  
> > > 136)  
> > > at  
> > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.extrac­­tMapping(XContentDocumentMapperParser.java:  
> > > 316)  
> > > at  
> > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(­­XContentDocumentMapperParser.java:  
> > > 114)  
> > > at  
> > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(­­XContentDocumentMapperParser.java:  
> > > 54)  
> > > at  
> > > org.elasticsearch.index.mapper.MapperService.parse(MapperService.java:  
> > > 209)  
> > > at org.elasticsearch.cluster.metadata.MetaDataMappingService  
> > > $3.execute(MetaDataMappingService.java:193)  
> > > at org.elasticsearch.cluster.service.InternalClusterService  
> > > $2.run(InternalClusterService.java:175)  
> > > at java.util.concurrent.ThreadPoolExecutor  
> > > $Worker.runTask(ThreadPoolExecutor.java:886)  
> > > at java.util.concurrent.ThreadPoolExecutor  
> > > $Worker.run(ThreadPoolExecutor.java:908)  
> > > at java.lang.Thread.run(Thread.java:662)
> 
> > > On Mar 10, 4:16 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > > > There is an API to "putMapping" on an index, but, you can't change a field that is already analyzed to be "not\_analyzed" (as its part of the indexing process, so all current indexed docs will be meaningless).
> 
> > > > On Thursday, March 10, 2011 at 11:03 AM, eugene wrote:
> > > > 
> > > > > Thank you Shay.
> 
> > > > > Is there any way to add mappings to existing index? I searched the  
> > > > > source code and only found the examples for CreateIndexBuilder class.
> 
> > > > > Eugene.
> 
> > > > > On Mar 9, 10:18 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > > > 
> > > > > > You need to have the field you run the terms facet on marked as not\_analyzed in its mapping. You will need to reindex the data in order for that to take affect (and of course, create the mappings before you index data).
> 
> > > > > > On Thursday, March 10, 2011 at 5:12 AM, eugene wrote:
> > > > > > 
> > > > > > > Hi,
> 
> > > > > > > I am running a terms facet query on non-numeric field, and it is  
> > > > > > > working fine. However, if the field contains % symbol, the resutls in  
> > > > > > > a TermsFacet are tokenized into separate terms. For example, if I  
> > > > > > > have a json document:
> 
> > > > > > > {  
> > > > > > > url: http%[3Fwww.myserver.com/someaction/%3Fname%3DBob](http://3Fwww.myserver.com/someaction/%3Fname%3DBob)  
> > > > > > > }
> 
> > > > > > > and if I run: curl -X GEThttp://localhost:9200/\_river/my\_idx/\_search  
> > > > > > > -d  
> > > > > > > {  
> > > > > > > "query" : {  
> > > > > > > "match\_all" : {}  
> > > > > > > },  
> > > > > > > "facets" : {  
> > > > > > > "facet1" : {  
> > > > > > > "terms" : {  
> > > > > > > "field" : "url"  
> > > > > > > }  
> > > > > > > }  
> > > > > > > }  
> > > > > > > }
> 
> > > > > > > then it is returninng {"terms": [{"term": "[www.myserver.com](http://www.myserver.com)",  
> > > > > > > "count":1}, {"term": "http", "count":1}, {"term":"3Fname", "count":1},  
> > > > > > > {"term":"3DBob", "count":1}
> 
> > > > > > > I think it is supposed to return {"terms": [ {"term":"http  
> > > > > > > %[3Fwww.myserver.com/someaction/%3Fname%3DBob","count](http://3Fwww.myserver.com/someaction/%3Fname%3DBob%22,%22count)":1}
> 
> > > > > > > Please correct me if I am wrong. It never happens if a field doesn't  
> > > > > > > contain %.- Hide quoted text -
> 
> > > > > > - Show quoted text -- Hide quoted text -
> 
> > > > - Show quoted text -- Hide quoted text -
> 
> > > - Show quoted text -- Hide quoted text -
> 
> - Show quoted text -

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [March 15, 2011, 7:00am UTC](https://discuss.elastic.co/t/terms-facet-is-tokenizing-a-field-with-special-characters/4071/9 "2011-03-15T07:00:55Z")

</div>

Please post a full recreation: [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/help) (including putting the mapping and indexing data), simpler to help then.  
On Monday, March 14, 2011 at 7:59 PM, eugene wrote:

> Shay,
> 
> It appears I didn't fix the error. I run the following:
> 
> "query" : {  
> "match\_all": {}  
> },  
> "facets" : {  
> "url" : {  
> "terms" : {  
> "field" : "url"  
> }  
> }  
> }
> 
> I am still getting tokenized terms for "url" field. I checked the  
> url field and it is set not\_analyzed.
> 
> C:\>curl -X GET [http://localhost:9200/test/type1/\_mapping](http://localhost:9200/test/type1/_mapping)  
> {"test":{"type1":{"properties":{"timestamp":  
> {"store":"yes","type":"string"},"url":  
> {"index":"not\_analyzed","store":"yes","type":"string"}}}}}
> 
> Am I using the wrong facet?
> 
> Thank you,  
> Eugene.
> 
> On Mar 11, 6:11 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > Not sure that this solved your problem, but a different change that you did where you don't use toString on the builder that represents the mapping. You can pass the builder directory to the putMapping API.
> > 
> > On Friday, March 11, 2011 at 4:14 AM, eugene wrote:
> > 
> > > I solved this problem by creating a new index first, and only then  
> > > putting maps to it. I guess mappings should be done only once on a  
> > > new index/type pair.
> > 
> > > -Eugene.
> > 
> > > On Mar 10, 3:03 pm, eugene [efur...@gmail.com](mailto:efur...@gmail.com) wrote:
> > > 
> > > > Ok, I am doing this:
> > 
> > > > String mappings =  
> > > > XContentFactory.jsonBuilder().startObject().startObject("type1")  
> > > > .startObject("myfield").field("type",  
> > > > "string").field("store", "yes").field("index",  
> > > > "not\_analyzed").endObject().endObject().endObject().toString();
> > 
> > > > client.admin().indices().preparePutMapping().setType("type1")  
> > > > .setSource(mappings).execute().actionGet();
> > 
> > > > I am getting the following exception (notice: I replaced real ip  
> > > > address with "ip\_address" on the first line).  
> > > > What do you think it indicates? Thank you!
> > 
> > > > Eugene.
> > 
> > > > org.elasticsearch.transport.RemoteTransportException: [Stonecutter]  
> > > > [inet[/ip\_address:9300]][indices/mapping/put]  
> > > > Caused by: org.elasticsearch.ElasticSearchParseException: Failed to  
> > > > derive xcontent from  
> > > > org.elasticsearch.common.xcontent.XContentBuilder@1c94b8f  
> > > > at  
> > > > org.elasticsearch.common.xcontent.XContentFactory.xContent(XContentFactory.Â­Â­java:  
> > > > 136)  
> > > > at  
> > > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.extracÂ­Â­tMapping(XContentDocumentMapperParser.java:  
> > > > 316)  
> > > > at  
> > > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(Â­Â­XContentDocumentMapperParser.java:  
> > > > 114)  
> > > > at  
> > > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(Â­Â­XContentDocumentMapperParser.java:  
> > > > 54)  
> > > > at  
> > > > org.elasticsearch.index.mapper.MapperService.parse(MapperService.java:  
> > > > 209)  
> > > > at org.elasticsearch.cluster.metadata.MetaDataMappingService  
> > > > $3.execute(MetaDataMappingService.java:193)  
> > > > at org.elasticsearch.cluster.service.InternalClusterService  
> > > > $2.run(InternalClusterService.java:175)  
> > > > at java.util.concurrent.ThreadPoolExecutor  
> > > > $Worker.runTask(ThreadPoolExecutor.java:886)  
> > > > at java.util.concurrent.ThreadPoolExecutor  
> > > > $Worker.run(ThreadPoolExecutor.java:908)  
> > > > at java.lang.Thread.run(Thread.java:662)
> > 
> > > > On Mar 10, 4:16 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > 
> > > > > There is an API to "putMapping" on an index, but, you can't change a field that is already analyzed to be "not\_analyzed" (as its part of the indexing process, so all current indexed docs will be meaningless).
> > 
> > > > > On Thursday, March 10, 2011 at 11:03 AM, eugene wrote:
> > > > > 
> > > > > > Thank you Shay.
> > 
> > > > > > Is there any way to add mappings to existing index? I searched the  
> > > > > > source code and only found the examples for CreateIndexBuilder class.
> > 
> > > > > > Eugene.
> > 
> > > > > > On Mar 9, 10:18 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > > > > 
> > > > > > > You need to have the field you run the terms facet on marked as not\_analyzed in its mapping. You will need to reindex the data in order for that to take affect (and of course, create the mappings before you index data).
> > 
> > > > > > > On Thursday, March 10, 2011 at 5:12 AM, eugene wrote:
> > > > > > > 
> > > > > > > > Hi,
> > 
> > > > > > > > I am running a terms facet query on non-numeric field, and it is  
> > > > > > > > working fine. However, if the field contains % symbol, the resutls in  
> > > > > > > > a TermsFacet are tokenized into separate terms. For example, if I  
> > > > > > > > have a json document:
> > 
> > > > > > > > {  
> > > > > > > > url: http%[3Fwww.myserver.com/someaction/%3Fname%3DBob](http://3Fwww.myserver.com/someaction/%3Fname%3DBob)  
> > > > > > > > }
> > 
> > > > > > > > and if I run: curl -X GEThttp://localhost:9200/\_river/my\_idx/\_search  
> > > > > > > > -d  
> > > > > > > > {  
> > > > > > > > "query" : {  
> > > > > > > > "match\_all" : {}  
> > > > > > > > },  
> > > > > > > > "facets" : {  
> > > > > > > > "facet1" : {  
> > > > > > > > "terms" : {  
> > > > > > > > "field" : "url"  
> > > > > > > > }  
> > > > > > > > }  
> > > > > > > > }  
> > > > > > > > }
> > 
> > > > > > > > then it is returninng {"terms": [{"term": "[www.myserver.com](http://www.myserver.com)",  
> > > > > > > > "count":1}, {"term": "http", "count":1}, {"term":"3Fname", "count":1},  
> > > > > > > > {"term":"3DBob", "count":1}
> > 
> > > > > > > > I think it is supposed to return {"terms": [ {"term":"http  
> > > > > > > > %[3Fwww.myserver.com/someaction/%3Fname%3DBob","count](http://3Fwww.myserver.com/someaction/%3Fname%3DBob%22,%22count)":1}
> > 
> > > > > > > > Please correct me if I am wrong. It never happens if a field doesn't  
> > > > > > > > contain %.- Hide quoted text -
> > 
> > > > > > > - Show quoted text -- Hide quoted text -
> > 
> > > > > - Show quoted text -- Hide quoted text -
> > 
> > > > - Show quoted text -- Hide quoted text -
> > 
> > - Show quoted text -

---

<div class="post-metadata">

**Author:** ![jason](https://avatars.discourse-cdn.com/v4/letter/j/dc4da7/32.png) [@jason](https://discuss.elastic.co/u/jason)\
**Post date:** [March 15, 2011, 7:54am UTC](https://discuss.elastic.co/t/terms-facet-is-tokenizing-a-field-with-special-characters/4071/10 "2011-03-15T07:54:25Z")

</div>

Here is how I created index 'test' and put mappings for 'type1':

client.admin().indices().prepareCreate("test").execute().actionGet();

```
        	String mappings =

```

XContentFactory.jsonBuilder().startObject().startObject("type1").startObject("properties")  
.startObject("url").field("type", "string").field("store",  
"yes").field("index", "not\_analyzed").endObject()  
.startObject("timestamp").field("type",  
"string").field("store", "yes").endObject()  
.endObject().endObject().string();

client.admin().indices().preparePutMapping("test").setType("type1").setSource(mappings).execute().actionGet();

Then, I am writing data to the ES under 'test' index (note: url is  
encrypted, where %3F corresponds to ?, %26 to &, and %3D to =.

```
                                 String url = "http%3A//

```

[someurl.com&nbsp;-&nbsp;This website is for sale!&nbsp;-&nbsp;someurl Resources and Information.](http://www.someurl.com/someservice%3FcityVegas%26State%3DNevada)";  
SimpleDateFormat sdf2 = new  
SimpleDateFormat("yyyy-MM-dd hh:mm:ss");  
Calendar cal =  
Calendar.getInstance();  
IndexResponse response = client.prepareIndex("test",  
"type1").setSource(jsonBuilder()  
.startObject()  
.field("url", url)  
.field("timestamp",  
sdf2.format(cal.getTime()))  
.endObject()  
)  
.execute()  
.actionGet();

Here is my response to :\>curl -X GET [http://localhost:9200/test/type1/\_mapping](http://localhost:9200/test/type1/_mapping)  
(assuming I put more than one documents as above):

..."facets":{"url":{"\_type":"terms","missing":0,"terms":  
[{"term":"[www.someurl.com](http://www.someurl.com)","count":4},{"term":"photoflipper","count":  
4},{"term":"http  
,"count":4},{"term":"3fcity","count":4},{"term":"3a","count":4},  
{"term":"26state","count":4},{"term":"26model","count":4},  
{"term":"3dNevada","count":2}]}}}

On Mar 15, 12:00 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> Please post a full recreation:[Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/help)(including putting the mapping and indexing data), simpler to help then.
> 
> On Monday, March 14, 2011 at 7:59 PM, eugene wrote:
> 
> > Shay,
> 
> > It appears I didn't fix the error. I run the following:
> 
> > "query" : {  
> > "match\_all": {}  
> > },  
> > "facets" : {  
> > "url" : {  
> > "terms" : {  
> > "field" : "url"  
> > }  
> > }  
> > }
> 
> > I am still getting tokenized terms for "url" field. I checked the  
> > url field and it is set not\_analyzed.
> 
> > C:\>curl -X GEThttp://localhost:9200/test/type1/\_mapping  
> > {"test":{"type1":{"properties":{"timestamp":  
> > {"store":"yes","type":"string"},"url":  
> > {"index":"not\_analyzed","store":"yes","type":"string"}}}}}
> 
> > Am I using the wrong facet?
> 
> > Thank you,  
> > Eugene.
> 
> > On Mar 11, 6:11 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > 
> > > Not sure that this solved your problem, but a different change that you did where you don't use toString on the builder that represents the mapping. You can pass the builder directory to the putMapping API.
> 
> > > On Friday, March 11, 2011 at 4:14 AM, eugene wrote:
> 
> > > > I solved this problem by creating a new index first, and only then  
> > > > putting maps to it. I guess mappings should be done only once on a  
> > > > new index/type pair.
> 
> > > > -Eugene.
> 
> > > > On Mar 10, 3:03 pm, eugene [efur...@gmail.com](mailto:efur...@gmail.com) wrote:
> > > > 
> > > > > Ok, I am doing this:
> 
> > > > > String mappings =  
> > > > > XContentFactory.jsonBuilder().startObject().startObject("type1")  
> > > > > .startObject("myfield").field("type",  
> > > > > "string").field("store", "yes").field("index",  
> > > > > "not\_analyzed").endObject().endObject().endObject().toString();
> 
> > > > > client.admin().indices().preparePutMapping().setType("type1")  
> > > > > .setSource(mappings).execute().actionGet();
> 
> > > > > I am getting the following exception (notice: I replaced real ip  
> > > > > address with "ip\_address" on the first line).  
> > > > > What do you think it indicates? Thank you!
> 
> > > > > Eugene.
> 
> > > > > org.elasticsearch.transport.RemoteTransportException: [Stonecutter]  
> > > > > [inet[/ip\_address:9300]][indices/mapping/put]  
> > > > > Caused by: org.elasticsearch.ElasticSearchParseException: Failed to  
> > > > > derive xcontent from  
> > > > > org.elasticsearch.common.xcontent.XContentBuilder@1c94b8f  
> > > > > at  
> > > > > org.elasticsearch.common.xcontent.XContentFactory.xContent(XContentFactory.­­­java:  
> > > > > 136)  
> > > > > at  
> > > > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.extrac­­­tMapping(XContentDocumentMapperParser.java:  
> > > > > 316)  
> > > > > at  
> > > > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(­­­XContentDocumentMapperParser.java:  
> > > > > 114)  
> > > > > at  
> > > > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(­­­XContentDocumentMapperParser.java:  
> > > > > 54)  
> > > > > at  
> > > > > org.elasticsearch.index.mapper.MapperService.parse(MapperService.java:  
> > > > > 209)  
> > > > > at org.elasticsearch.cluster.metadata.MetaDataMappingService  
> > > > > $3.execute(MetaDataMappingService.java:193)  
> > > > > at org.elasticsearch.cluster.service.InternalClusterService  
> > > > > $2.run(InternalClusterService.java:175)  
> > > > > at java.util.concurrent.ThreadPoolExecutor  
> > > > > $Worker.runTask(ThreadPoolExecutor.java:886)  
> > > > > at java.util.concurrent.ThreadPoolExecutor  
> > > > > $Worker.run(ThreadPoolExecutor.java:908)  
> > > > > at java.lang.Thread.run(Thread.java:662)
> 
> > > > > On Mar 10, 4:16 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > > > > > There is an API to "putMapping" on an index, but, you can't change a field that is already analyzed to be "not\_analyzed" (as its part of the indexing process, so all current indexed docs will be meaningless).
> 
> > > > > > On Thursday, March 10, 2011 at 11:03 AM, eugene wrote:
> > > > > > 
> > > > > > > Thank you Shay.
> 
> > > > > > > Is there any way to add mappings to existing index? I searched the  
> > > > > > > source code and only found the examples for CreateIndexBuilder class.
> 
> > > > > > > Eugene.
> 
> > > > > > > On Mar 9, 10:18 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > > > > > 
> > > > > > > > You need to have the field you run the terms facet on marked as not\_analyzed in its mapping. You will need to reindex the data in order for that to take affect (and of course, create the mappings before you index data).
> 
> > > > > > > > On Thursday, March 10, 2011 at 5:12 AM, eugene wrote:
> > > > > > > > 
> > > > > > > > > Hi,
> 
> > > > > > > > > I am running a terms facet query on non-numeric field, and it is  
> > > > > > > > > working fine. However, if the field contains % symbol, the resutls in  
> > > > > > > > > a TermsFacet are tokenized into separate terms. For example, if I  
> > > > > > > > > have a json document:
> 
> > > > > > > > > {  
> > > > > > > > > url: http%[3Fwww.myserver.com/someaction/%3Fname%3DBob](http://3Fwww.myserver.com/someaction/%3Fname%3DBob)  
> > > > > > > > > }
> 
> > > > > > > > > and if I run: curl -X GEThttp://localhost:9200/\_river/my\_idx/\_search  
> > > > > > > > > -d  
> > > > > > > > > {  
> > > > > > > > > "query" : {  
> > > > > > > > > "match\_all" : {}  
> > > > > > > > > },  
> > > > > > > > > "facets" : {  
> > > > > > > > > "facet1" : {  
> > > > > > > > > "terms" : {  
> > > > > > > > > "field" : "url"  
> > > > > > > > > }  
> > > > > > > > > }  
> > > > > > > > > }  
> > > > > > > > > }
> 
> > > > > > > > > then it is returninng {"terms": [{"term": "[www.myserver.com](http://www.myserver.com)",  
> > > > > > > > > "count":1}, {"term": "http", "count":1}, {"term":"3Fname", "count":1},  
> > > > > > > > > {"term":"3DBob", "count":1}
> 
> > > > > > > > > I think it is supposed to return {"terms": [ {"term":"http  
> > > > > > > > > %[3Fwww.myserver.com/someaction/%3Fname%3DBob","count](http://3Fwww.myserver.com/someaction/%3Fname%3DBob%22,%22count)":1}
> 
> > > > > > > > > Please correct me if I am wrong. It never happens if a field doesn't  
> > > > > > > > > contain %.- Hide quoted text -
> 
> > > > > > > > - Show quoted text -- Hide quoted text -
> 
> > > > > > - Show quoted text -- Hide quoted text -
> 
> > > > > - Show quoted text -- Hide quoted text -
> 
> > > - Show quoted text -- Hide quoted text -
> 
> - Show quoted text -

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [March 15, 2011, 7:56am UTC](https://discuss.elastic.co/t/terms-facet-is-tokenizing-a-field-with-special-characters/4071/11 "2011-03-15T07:56:43Z")

</div>

Can you provide a curl recreation?  
On Tuesday, March 15, 2011 at 9:54 AM, eugene wrote:

> Here is how I created index 'test' and put mappings for 'type1':
> 
> client.admin().indices().prepareCreate("test").execute().actionGet();
> 
> String mappings =  
> XContentFactory.jsonBuilder().startObject().startObject("type1").startObject("properties")  
> .startObject("url").field("type", "string").field("store",  
> "yes").field("index", "not\_analyzed").endObject()  
> .startObject("timestamp").field("type",  
> "string").field("store", "yes").endObject()  
> .endObject().endObject().string();
> 
> client.admin().indices().preparePutMapping("test").setType("type1").setSource(mappings).execute().actionGet();
> 
> Then, I am writing data to the ES under 'test' index (note: url is  
> encrypted, where %3F corresponds to ?, %26 to &, and %3D to =.
> 
> String url = "http%3A//  
> [someurl.com&nbsp;-&nbsp;This website is for sale!&nbsp;-&nbsp;someurl Resources and Information.](http://www.someurl.com/someservice%3FcityVegas%26State%3DNevada)";  
> SimpleDateFormat sdf2 = new  
> SimpleDateFormat("yyyy-MM-dd hh:mm:ss");  
> Calendar cal =  
> Calendar.getInstance();  
> IndexResponse response = client.prepareIndex("test",  
> "type1").setSource(jsonBuilder()  
> .startObject()  
> .field("url", url)  
> .field("timestamp",  
> sdf2.format(cal.getTime()))  
> .endObject()  
> )  
> .execute()  
> .actionGet();
> 
> Here is my response to :\>curl -X GET [http://localhost:9200/test/type1/\_mapping](http://localhost:9200/test/type1/_mapping)  
> (assuming I put more than one documents as above):
> 
> ..."facets":{"url":{"\_type":"terms","missing":0,"terms":  
> [{"term":"[www.someurl.com](http://www.someurl.com)","count":4},{"term":"photoflipper","count":  
> 4},{"term":"http  
> ,"count":4},{"term":"3fcity","count":4},{"term":"3a","count":4},  
> {"term":"26state","count":4},{"term":"26model","count":4},  
> {"term":"3dNevada","count":2}]}}}
> 
> On Mar 15, 12:00 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > Please post a full recreation:[Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/help)(including putting the mapping and indexing data), simpler to help then.
> > 
> > On Monday, March 14, 2011 at 7:59 PM, eugene wrote:
> > 
> > > Shay,
> > 
> > > It appears I didn't fix the error. I run the following:
> > 
> > > "query" : {  
> > > "match\_all": {}  
> > > },  
> > > "facets" : {  
> > > "url" : {  
> > > "terms" : {  
> > > "field" : "url"  
> > > }  
> > > }  
> > > }
> > 
> > > I am still getting tokenized terms for "url" field. I checked the  
> > > url field and it is set not\_analyzed.
> > 
> > > C:\>curl -X GEThttp://localhost:9200/test/type1/\_mapping  
> > > {"test":{"type1":{"properties":{"timestamp":  
> > > {"store":"yes","type":"string"},"url":  
> > > {"index":"not\_analyzed","store":"yes","type":"string"}}}}}
> > 
> > > Am I using the wrong facet?
> > 
> > > Thank you,  
> > > Eugene.
> > 
> > > On Mar 11, 6:11 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > 
> > > > Not sure that this solved your problem, but a different change that you did where you don't use toString on the builder that represents the mapping. You can pass the builder directory to the putMapping API.
> > 
> > > > On Friday, March 11, 2011 at 4:14 AM, eugene wrote:
> > 
> > > > > I solved this problem by creating a new index first, and only then  
> > > > > putting maps to it. I guess mappings should be done only once on a  
> > > > > new index/type pair.
> > 
> > > > > -Eugene.
> > 
> > > > > On Mar 10, 3:03 pm, eugene [efur...@gmail.com](mailto:efur...@gmail.com) wrote:
> > > > > 
> > > > > > Ok, I am doing this:
> > 
> > > > > > String mappings =  
> > > > > > XContentFactory.jsonBuilder().startObject().startObject("type1")  
> > > > > > .startObject("myfield").field("type",  
> > > > > > "string").field("store", "yes").field("index",  
> > > > > > "not\_analyzed").endObject().endObject().endObject().toString();
> > 
> > > > > > client.admin().indices().preparePutMapping().setType("type1")  
> > > > > > .setSource(mappings).execute().actionGet();
> > 
> > > > > > I am getting the following exception (notice: I replaced real ip  
> > > > > > address with "ip\_address" on the first line).  
> > > > > > What do you think it indicates? Thank you!
> > 
> > > > > > Eugene.
> > 
> > > > > > org.elasticsearch.transport.RemoteTransportException: [Stonecutter]  
> > > > > > [inet[/ip\_address:9300]][indices/mapping/put]  
> > > > > > Caused by: org.elasticsearch.ElasticSearchParseException: Failed to  
> > > > > > derive xcontent from  
> > > > > > org.elasticsearch.common.xcontent.XContentBuilder@1c94b8f  
> > > > > > at  
> > > > > > org.elasticsearch.common.xcontent.XContentFactory.xContent(XContentFactory.Â­Â­Â­java:  
> > > > > > 136)  
> > > > > > at  
> > > > > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.extracÂ­Â­Â­tMapping(XContentDocumentMapperParser.java:  
> > > > > > 316)  
> > > > > > at  
> > > > > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(Â­Â­Â­XContentDocumentMapperParser.java:  
> > > > > > 114)  
> > > > > > at  
> > > > > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(Â­Â­Â­XContentDocumentMapperParser.java:  
> > > > > > 54)  
> > > > > > at  
> > > > > > org.elasticsearch.index.mapper.MapperService.parse(MapperService.java:  
> > > > > > 209)  
> > > > > > at org.elasticsearch.cluster.metadata.MetaDataMappingService  
> > > > > > $3.execute(MetaDataMappingService.java:193)  
> > > > > > at org.elasticsearch.cluster.service.InternalClusterService  
> > > > > > $2.run(InternalClusterService.java:175)  
> > > > > > at java.util.concurrent.ThreadPoolExecutor  
> > > > > > $Worker.runTask(ThreadPoolExecutor.java:886)  
> > > > > > at java.util.concurrent.ThreadPoolExecutor  
> > > > > > $Worker.run(ThreadPoolExecutor.java:908)  
> > > > > > at java.lang.Thread.run(Thread.java:662)
> > 
> > > > > > On Mar 10, 4:16 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > 
> > > > > > > There is an API to "putMapping" on an index, but, you can't change a field that is already analyzed to be "not\_analyzed" (as its part of the indexing process, so all current indexed docs will be meaningless).
> > 
> > > > > > > On Thursday, March 10, 2011 at 11:03 AM, eugene wrote:
> > > > > > > 
> > > > > > > > Thank you Shay.
> > 
> > > > > > > > Is there any way to add mappings to existing index? I searched the  
> > > > > > > > source code and only found the examples for CreateIndexBuilder class.
> > 
> > > > > > > > Eugene.
> > 
> > > > > > > > On Mar 9, 10:18 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > > > > > > 
> > > > > > > > > You need to have the field you run the terms facet on marked as not\_analyzed in its mapping. You will need to reindex the data in order for that to take affect (and of course, create the mappings before you index data).
> > 
> > > > > > > > > On Thursday, March 10, 2011 at 5:12 AM, eugene wrote:
> > > > > > > > > 
> > > > > > > > > > Hi,
> > 
> > > > > > > > > > I am running a terms facet query on non-numeric field, and it is  
> > > > > > > > > > working fine. However, if the field contains % symbol, the resutls in  
> > > > > > > > > > a TermsFacet are tokenized into separate terms. For example, if I  
> > > > > > > > > > have a json document:
> > 
> > > > > > > > > > {  
> > > > > > > > > > url: http%[3Fwww.myserver.com/someaction/%3Fname%3DBob](http://3Fwww.myserver.com/someaction/%3Fname%3DBob)  
> > > > > > > > > > }
> > 
> > > > > > > > > > and if I run: curl -X GEThttp://localhost:9200/\_river/my\_idx/\_search  
> > > > > > > > > > -d  
> > > > > > > > > > {  
> > > > > > > > > > "query" : {  
> > > > > > > > > > "match\_all" : {}  
> > > > > > > > > > },  
> > > > > > > > > > "facets" : {  
> > > > > > > > > > "facet1" : {  
> > > > > > > > > > "terms" : {  
> > > > > > > > > > "field" : "url"  
> > > > > > > > > > }  
> > > > > > > > > > }  
> > > > > > > > > > }  
> > > > > > > > > > }
> > 
> > > > > > > > > > then it is returninng {"terms": [{"term": "[www.myserver.com](http://www.myserver.com)",  
> > > > > > > > > > "count":1}, {"term": "http", "count":1}, {"term":"3Fname", "count":1},  
> > > > > > > > > > {"term":"3DBob", "count":1}
> > 
> > > > > > > > > > I think it is supposed to return {"terms": [ {"term":"http  
> > > > > > > > > > %[3Fwww.myserver.com/someaction/%3Fname%3DBob","count](http://3Fwww.myserver.com/someaction/%3Fname%3DBob%22,%22count)":1}
> > 
> > > > > > > > > > Please correct me if I am wrong. It never happens if a field doesn't  
> > > > > > > > > > contain %.- Hide quoted text -
> > 
> > > > > > > > > - Show quoted text -- Hide quoted text -
> > 
> > > > > > > - Show quoted text -- Hide quoted text -
> > 
> > > > > > - Show quoted text -- Hide quoted text -
> > 
> > > > - Show quoted text -- Hide quoted text -
> > 
> > - Show quoted text -

---

<div class="post-metadata">

**Author:** ![jason](https://avatars.discourse-cdn.com/v4/letter/j/dc4da7/32.png) [@jason](https://discuss.elastic.co/u/jason)\
**Post date:** [March 15, 2011, 3:57pm UTC](https://discuss.elastic.co/t/terms-facet-is-tokenizing-a-field-with-special-characters/4071/12 "2011-03-15T15:57:19Z")

</div>

I created it using java. Could there a difference if I do this in  
java instead of curl?

On Mar 15, 12:56 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> Can you provide a curl recreation?
> 
> On Tuesday, March 15, 2011 at 9:54 AM, eugene wrote:
> 
> > Here is how I created index 'test' and put mappings for 'type1':
> 
> > client.admin().indices().prepareCreate("test").execute().actionGet();
> 
> > String mappings =  
> > XContentFactory.jsonBuilder().startObject().startObject("type1").startObjec­t("properties")  
> > .startObject("url").field("type", "string").field("store",  
> > "yes").field("index", "not\_analyzed").endObject()  
> > .startObject("timestamp").field("type",  
> > "string").field("store", "yes").endObject()  
> > .endObject().endObject().string();
> 
> > client.admin().indices().preparePutMapping("test").setType("type1").setSour­ce(mappings).execute().actionGet();
> 
> > Then, I am writing data to the ES under 'test' index (note: url is  
> > encrypted, where %3F corresponds to ?, %26 to &, and %3D to =.
> 
> > String url = "http%3A//  
> > [someurl.com&nbsp;-&nbsp;This website is for sale!&nbsp;-&nbsp;someurl Resources and Information.](http://www.someurl.com/someservice%3FcityVegas%26State%3DNevada)";  
> > SimpleDateFormat sdf2 = new  
> > SimpleDateFormat("yyyy-MM-dd hh:mm:ss");  
> > Calendar cal =  
> > Calendar.getInstance();  
> > IndexResponse response = client.prepareIndex("test",  
> > "type1").setSource(jsonBuilder()  
> > .startObject()  
> > .field("url", url)  
> > .field("timestamp",  
> > sdf2.format(cal.getTime()))  
> > .endObject()  
> > )  
> > .execute()  
> > .actionGet();
> 
> > Here is my response to :\>curl -X GEThttp://localhost:9200/test/type1/\_mapping  
> > (assuming I put more than one documents as above):
> 
> > ..."facets":{"url":{"\_type":"terms","missing":0,"terms":  
> > [{"term":"[www.someurl.com](http://www.someurl.com)","count":4},{"term":"photoflipper","count":  
> > 4},{"term":"http  
> > ,"count":4},{"term":"3fcity","count":4},{"term":"3a","count":4},  
> > {"term":"26state","count":4},{"term":"26model","count":4},  
> > {"term":"3dNevada","count":2}]}}}
> 
> > On Mar 15, 12:00 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > 
> > > Please post a full recreation:[Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/help)(includingputting the mapping and indexing data), simpler to help then.
> 
> > > On Monday, March 14, 2011 at 7:59 PM, eugene wrote:
> > > 
> > > > Shay,
> 
> > > > It appears I didn't fix the error. I run the following:
> 
> > > > "query" : {  
> > > > "match\_all": {}  
> > > > },  
> > > > "facets" : {  
> > > > "url" : {  
> > > > "terms" : {  
> > > > "field" : "url"  
> > > > }  
> > > > }  
> > > > }
> 
> > > > I am still getting tokenized terms for "url" field. I checked the  
> > > > url field and it is set not\_analyzed.
> 
> > > > C:\>curl -X GEThttp://localhost:9200/test/type1/\_mapping  
> > > > {"test":{"type1":{"properties":{"timestamp":  
> > > > {"store":"yes","type":"string"},"url":  
> > > > {"index":"not\_analyzed","store":"yes","type":"string"}}}}}
> 
> > > > Am I using the wrong facet?
> 
> > > > Thank you,  
> > > > Eugene.
> 
> > > > On Mar 11, 6:11 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > > 
> > > > > Not sure that this solved your problem, but a different change that you did where you don't use toString on the builder that represents the mapping. You can pass the builder directory to the putMapping API.
> 
> > > > > On Friday, March 11, 2011 at 4:14 AM, eugene wrote:
> 
> > > > > > I solved this problem by creating a new index first, and only then  
> > > > > > putting maps to it. I guess mappings should be done only once on a  
> > > > > > new index/type pair.
> 
> > > > > > -Eugene.
> 
> > > > > > On Mar 10, 3:03 pm, eugene [efur...@gmail.com](mailto:efur...@gmail.com) wrote:
> > > > > > 
> > > > > > > Ok, I am doing this:
> 
> > > > > > > String mappings =  
> > > > > > > XContentFactory.jsonBuilder().startObject().startObject("type1")  
> > > > > > > .startObject("myfield").field("type",  
> > > > > > > "string").field("store", "yes").field("index",  
> > > > > > > "not\_analyzed").endObject().endObject().endObject().toString();
> 
> > > > > > > client.admin().indices().preparePutMapping().setType("type1")  
> > > > > > > .setSource(mappings).execute().actionGet();
> 
> > > > > > > I am getting the following exception (notice: I replaced real ip  
> > > > > > > address with "ip\_address" on the first line).  
> > > > > > > What do you think it indicates? Thank you!
> 
> > > > > > > Eugene.
> 
> > > > > > > org.elasticsearch.transport.RemoteTransportException: [Stonecutter]  
> > > > > > > [inet[/ip\_address:9300]][indices/mapping/put]  
> > > > > > > Caused by: org.elasticsearch.ElasticSearchParseException: Failed to  
> > > > > > > derive xcontent from  
> > > > > > > org.elasticsearch.common.xcontent.XContentBuilder@1c94b8f  
> > > > > > > at  
> > > > > > > org.elasticsearch.common.xcontent.XContentFactory.xContent(XContentFactory.­­­­java:  
> > > > > > > 136)  
> > > > > > > at  
> > > > > > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.extrac­­­­tMapping(XContentDocumentMapperParser.java:  
> > > > > > > 316)  
> > > > > > > at  
> > > > > > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(­­­­XContentDocumentMapperParser.java:  
> > > > > > > 114)  
> > > > > > > at  
> > > > > > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(­­­­XContentDocumentMapperParser.java:  
> > > > > > > 54)  
> > > > > > > at  
> > > > > > > org.elasticsearch.index.mapper.MapperService.parse(MapperService.java:  
> > > > > > > 209)  
> > > > > > > at org.elasticsearch.cluster.metadata.MetaDataMappingService  
> > > > > > > $3.execute(MetaDataMappingService.java:193)  
> > > > > > > at org.elasticsearch.cluster.service.InternalClusterService  
> > > > > > > $2.run(InternalClusterService.java:175)  
> > > > > > > at java.util.concurrent.ThreadPoolExecutor  
> > > > > > > $Worker.runTask(ThreadPoolExecutor.java:886)  
> > > > > > > at java.util.concurrent.ThreadPoolExecutor  
> > > > > > > $Worker.run(ThreadPoolExecutor.java:908)  
> > > > > > > at java.lang.Thread.run(Thread.java:662)
> 
> > > > > > > On Mar 10, 4:16 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > > > > > > > There is an API to "putMapping" on an index, but, you can't change a field that is already analyzed to be "not\_analyzed" (as its part of the indexing process, so all current indexed docs will be meaningless).
> 
> > > > > > > > On Thursday, March 10, 2011 at 11:03 AM, eugene wrote:
> > > > > > > > 
> > > > > > > > > Thank you Shay.
> 
> > > > > > > > > Is there any way to add mappings to existing index? I searched the  
> > > > > > > > > source code and only found the examples for CreateIndexBuilder class.
> 
> > > > > > > > > Eugene.
> 
> > > > > > > > > On Mar 9, 10:18 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > > > > > > > 
> > > > > > > > > > You need to have the field you run the terms facet on marked as not\_analyzed in its mapping. You will need to reindex the data in order for that to take affect (and of course, create the mappings before you index data).
> 
> > > > > > > > > > On Thursday, March 10, 2011 at 5:12 AM, eugene wrote:
> > > > > > > > > > 
> > > > > > > > > > > Hi,
> 
> > > > > > > > > > > I am running a terms facet query on non-numeric field, and it is  
> > > > > > > > > > > working fine. However, if the field contains % symbol, the resutls in  
> > > > > > > > > > > a TermsFacet are tokenized into separate terms. For example, if I  
> > > > > > > > > > > have a json document:
> 
> > > > > > > > > > > {  
> > > > > > > > > > > url: http%[3Fwww.myserver.com/someaction/%3Fname%3DBob](http://3Fwww.myserver.com/someaction/%3Fname%3DBob)  
> > > > > > > > > > > }
> 
> > > > > > > > > > > and if I run: curl -X GEThttp://localhost:9200/\_river/my\_idx/\_search  
> > > > > > > > > > > -d  
> > > > > > > > > > > {  
> > > > > > > > > > > "query" : {  
> > > > > > > > > > > "match\_all" : {}  
> > > > > > > > > > > },  
> > > > > > > > > > > "facets" : {  
> > > > > > > > > > > "facet1" : {  
> > > > > > > > > > > "terms" : {  
> > > > > > > > > > > "field" : "url"  
> > > > > > > > > > > }  
> > > > > > > > > > > }  
> > > > > > > > > > > }  
> > > > > > > > > > > }
> 
> > > > > > > > > > > then it is returninng {"terms": [{"term": "[www.myserver.com](http://www.myserver.com)",  
> > > > > > > > > > > "count":1}, {"term": "http", "count":1}, {"term":"3Fname", "count":1},  
> > > > > > > > > > > {"term":"3DBob", "count":1}
> 
> > > > > > > > > > > I think it is supposed to return {"terms": [ {"term":"http  
> > > > > > > > > > > %[3Fwww.myserver.com/someaction/%3Fname%3DBob","count](http://3Fwww.myserver.com/someaction/%3Fname%3DBob%22,%22count)":1}
> 
> > > > > > > > > > > Please correct me if I am wrong. It never happens if a field doesn't  
> > > > > > > > > > > contain %.- Hide quoted text -
> 
> > > > > > > > > > - Show quoted text -- Hide quoted text -
> 
> > > > > > > > - Show quoted text -- Hide quoted text -
> 
> > > > > > > - Show quoted text -- Hide quoted text -
> 
> > > > > - Show quoted text -- Hide quoted text -
> 
> > > - Show quoted text -- Hide quoted text -
> 
> - Show quoted text -

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [March 15, 2011, 4:12pm UTC](https://discuss.elastic.co/t/terms-facet-is-tokenizing-a-field-with-special-characters/4071/13 "2011-03-15T16:12:38Z")

</div>

No, there isn't a difference, its just simpler to recreate it with curl. The REST API is built on top of the Java API.

If you have problems with curl, then -\>_gist_\<- a simple test case that recreates it with Java, I can give it a go as well.  
On Tuesday, March 15, 2011 at 5:57 PM, eugene wrote:

> I created it using java. Could there a difference if I do this in  
> java instead of curl?
> 
> On Mar 15, 12:56 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > Can you provide a curl recreation?
> > 
> > On Tuesday, March 15, 2011 at 9:54 AM, eugene wrote:
> > 
> > > Here is how I created index 'test' and put mappings for 'type1':
> > 
> > > client.admin().indices().prepareCreate("test").execute().actionGet();
> > 
> > > String mappings =  
> > > XContentFactory.jsonBuilder().startObject().startObject("type1").startObjecÂ­t("properties")  
> > > .startObject("url").field("type", "string").field("store",  
> > > "yes").field("index", "not\_analyzed").endObject()  
> > > .startObject("timestamp").field("type",  
> > > "string").field("store", "yes").endObject()  
> > > .endObject().endObject().string();
> > 
> > > client.admin().indices().preparePutMapping("test").setType("type1").setSourÂ­ce(mappings).execute().actionGet();
> > 
> > > Then, I am writing data to the ES under 'test' index (note: url is  
> > > encrypted, where %3F corresponds to ?, %26 to &, and %3D to =.
> > 
> > > String url = "http%3A//  
> > > [someurl.com&nbsp;-&nbsp;This website is for sale!&nbsp;-&nbsp;someurl Resources and Information.](http://www.someurl.com/someservice%3FcityVegas%26State%3DNevada)";  
> > > SimpleDateFormat sdf2 = new  
> > > SimpleDateFormat("yyyy-MM-dd hh:mm:ss");  
> > > Calendar cal =  
> > > Calendar.getInstance();  
> > > IndexResponse response = client.prepareIndex("test",  
> > > "type1").setSource(jsonBuilder()  
> > > .startObject()  
> > > .field("url", url)  
> > > .field("timestamp",  
> > > sdf2.format(cal.getTime()))  
> > > .endObject()  
> > > )  
> > > .execute()  
> > > .actionGet();
> > 
> > > Here is my response to :\>curl -X GEThttp://localhost:9200/test/type1/\_mapping  
> > > (assuming I put more than one documents as above):
> > 
> > > ..."facets":{"url":{"\_type":"terms","missing":0,"terms":  
> > > [{"term":"[www.someurl.com](http://www.someurl.com)","count":4},{"term":"photoflipper","count":  
> > > 4},{"term":"http  
> > > ,"count":4},{"term":"3fcity","count":4},{"term":"3a","count":4},  
> > > {"term":"26state","count":4},{"term":"26model","count":4},  
> > > {"term":"3dNevada","count":2}]}}}
> > 
> > > On Mar 15, 12:00 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > 
> > > > Please post a full recreation:[Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/help)(includingputting the mapping and indexing data), simpler to help then.
> > 
> > > > On Monday, March 14, 2011 at 7:59 PM, eugene wrote:
> > > > 
> > > > > Shay,
> > 
> > > > > It appears I didn't fix the error. I run the following:
> > 
> > > > > "query" : {  
> > > > > "match\_all": {}  
> > > > > },  
> > > > > "facets" : {  
> > > > > "url" : {  
> > > > > "terms" : {  
> > > > > "field" : "url"  
> > > > > }  
> > > > > }  
> > > > > }
> > 
> > > > > I am still getting tokenized terms for "url" field. I checked the  
> > > > > url field and it is set not\_analyzed.
> > 
> > > > > C:\>curl -X GEThttp://localhost:9200/test/type1/\_mapping  
> > > > > {"test":{"type1":{"properties":{"timestamp":  
> > > > > {"store":"yes","type":"string"},"url":  
> > > > > {"index":"not\_analyzed","store":"yes","type":"string"}}}}}
> > 
> > > > > Am I using the wrong facet?
> > 
> > > > > Thank you,  
> > > > > Eugene.
> > 
> > > > > On Mar 11, 6:11 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > > > 
> > > > > > Not sure that this solved your problem, but a different change that you did where you don't use toString on the builder that represents the mapping. You can pass the builder directory to the putMapping API.
> > 
> > > > > > On Friday, March 11, 2011 at 4:14 AM, eugene wrote:
> > 
> > > > > > > I solved this problem by creating a new index first, and only then  
> > > > > > > putting maps to it. I guess mappings should be done only once on a  
> > > > > > > new index/type pair.
> > 
> > > > > > > -Eugene.
> > 
> > > > > > > On Mar 10, 3:03 pm, eugene [efur...@gmail.com](mailto:efur...@gmail.com) wrote:
> > > > > > > 
> > > > > > > > Ok, I am doing this:
> > 
> > > > > > > > String mappings =  
> > > > > > > > XContentFactory.jsonBuilder().startObject().startObject("type1")  
> > > > > > > > .startObject("myfield").field("type",  
> > > > > > > > "string").field("store", "yes").field("index",  
> > > > > > > > "not\_analyzed").endObject().endObject().endObject().toString();
> > 
> > > > > > > > client.admin().indices().preparePutMapping().setType("type1")  
> > > > > > > > .setSource(mappings).execute().actionGet();
> > 
> > > > > > > > I am getting the following exception (notice: I replaced real ip  
> > > > > > > > address with "ip\_address" on the first line).  
> > > > > > > > What do you think it indicates? Thank you!
> > 
> > > > > > > > Eugene.
> > 
> > > > > > > > org.elasticsearch.transport.RemoteTransportException: [Stonecutter]  
> > > > > > > > [inet[/ip\_address:9300]][indices/mapping/put]  
> > > > > > > > Caused by: org.elasticsearch.ElasticSearchParseException: Failed to  
> > > > > > > > derive xcontent from  
> > > > > > > > org.elasticsearch.common.xcontent.XContentBuilder@1c94b8f  
> > > > > > > > at  
> > > > > > > > org.elasticsearch.common.xcontent.XContentFactory.xContent(XContentFactory.Â­Â­Â­Â­java:  
> > > > > > > > 136)  
> > > > > > > > at  
> > > > > > > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.extracÂ­Â­Â­Â­tMapping(XContentDocumentMapperParser.java:  
> > > > > > > > 316)  
> > > > > > > > at  
> > > > > > > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(Â­Â­Â­Â­XContentDocumentMapperParser.java:  
> > > > > > > > 114)  
> > > > > > > > at  
> > > > > > > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(Â­Â­Â­Â­XContentDocumentMapperParser.java:  
> > > > > > > > 54)  
> > > > > > > > at  
> > > > > > > > org.elasticsearch.index.mapper.MapperService.parse(MapperService.java:  
> > > > > > > > 209)  
> > > > > > > > at org.elasticsearch.cluster.metadata.MetaDataMappingService  
> > > > > > > > $3.execute(MetaDataMappingService.java:193)  
> > > > > > > > at org.elasticsearch.cluster.service.InternalClusterService  
> > > > > > > > $2.run(InternalClusterService.java:175)  
> > > > > > > > at java.util.concurrent.ThreadPoolExecutor  
> > > > > > > > $Worker.runTask(ThreadPoolExecutor.java:886)  
> > > > > > > > at java.util.concurrent.ThreadPoolExecutor  
> > > > > > > > $Worker.run(ThreadPoolExecutor.java:908)  
> > > > > > > > at java.lang.Thread.run(Thread.java:662)
> > 
> > > > > > > > On Mar 10, 4:16 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > 
> > > > > > > > > There is an API to "putMapping" on an index, but, you can't change a field that is already analyzed to be "not\_analyzed" (as its part of the indexing process, so all current indexed docs will be meaningless).
> > 
> > > > > > > > > On Thursday, March 10, 2011 at 11:03 AM, eugene wrote:
> > > > > > > > > 
> > > > > > > > > > Thank you Shay.
> > 
> > > > > > > > > > Is there any way to add mappings to existing index? I searched the  
> > > > > > > > > > source code and only found the examples for CreateIndexBuilder class.
> > 
> > > > > > > > > > Eugene.
> > 
> > > > > > > > > > On Mar 9, 10:18 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > > > > > > > > 
> > > > > > > > > > > You need to have the field you run the terms facet on marked as not\_analyzed in its mapping. You will need to reindex the data in order for that to take affect (and of course, create the mappings before you index data).
> > 
> > > > > > > > > > > On Thursday, March 10, 2011 at 5:12 AM, eugene wrote:
> > > > > > > > > > > 
> > > > > > > > > > > > Hi,
> > 
> > > > > > > > > > > > I am running a terms facet query on non-numeric field, and it is  
> > > > > > > > > > > > working fine. However, if the field contains % symbol, the resutls in  
> > > > > > > > > > > > a TermsFacet are tokenized into separate terms. For example, if I  
> > > > > > > > > > > > have a json document:
> > 
> > > > > > > > > > > > {  
> > > > > > > > > > > > url: http%[3Fwww.myserver.com/someaction/%3Fname%3DBob](http://3Fwww.myserver.com/someaction/%3Fname%3DBob)  
> > > > > > > > > > > > }
> > 
> > > > > > > > > > > > and if I run: curl -X GEThttp://localhost:9200/\_river/my\_idx/\_search  
> > > > > > > > > > > > -d  
> > > > > > > > > > > > {  
> > > > > > > > > > > > "query" : {  
> > > > > > > > > > > > "match\_all" : {}  
> > > > > > > > > > > > },  
> > > > > > > > > > > > "facets" : {  
> > > > > > > > > > > > "facet1" : {  
> > > > > > > > > > > > "terms" : {  
> > > > > > > > > > > > "field" : "url"  
> > > > > > > > > > > > }  
> > > > > > > > > > > > }  
> > > > > > > > > > > > }  
> > > > > > > > > > > > }
> > 
> > > > > > > > > > > > then it is returninng {"terms": [{"term": "[www.myserver.com](http://www.myserver.com)",  
> > > > > > > > > > > > "count":1}, {"term": "http", "count":1}, {"term":"3Fname", "count":1},  
> > > > > > > > > > > > {"term":"3DBob", "count":1}
> > 
> > > > > > > > > > > > I think it is supposed to return {"terms": [ {"term":"http  
> > > > > > > > > > > > %[3Fwww.myserver.com/someaction/%3Fname%3DBob","count](http://3Fwww.myserver.com/someaction/%3Fname%3DBob%22,%22count)":1}
> > 
> > > > > > > > > > > > Please correct me if I am wrong. It never happens if a field doesn't  
> > > > > > > > > > > > contain %.- Hide quoted text -
> > 
> > > > > > > > > > > - Show quoted text -- Hide quoted text -
> > 
> > > > > > > > > - Show quoted text -- Hide quoted text -
> > 
> > > > > > > > - Show quoted text -- Hide quoted text -
> > 
> > > > > > - Show quoted text -- Hide quoted text -
> > 
> > > > - Show quoted text -- Hide quoted text -
> > 
> > - Show quoted text -

---

<div class="post-metadata">

**Author:** ![jason](https://avatars.discourse-cdn.com/v4/letter/j/dc4da7/32.png) [@jason](https://discuss.elastic.co/u/jason)\
**Post date:** [March 15, 2011, 9:31pm UTC](https://discuss.elastic.co/t/terms-facet-is-tokenizing-a-field-with-special-characters/4071/14 "2011-03-15T21:31:56Z")

</div>

Here is how I created it with curl:

C:\>curl -X PUT [http://localhost:9200/test/type1/\_mappings](http://localhost:9200/test/type1/_mappings) -d  
@mappings.json  
{"ok":true,"\_index":"test","\_type":"type2","\_id":"\_mappings","\_version":  
1}

where the contents of mappings.json file are as follows:

{  
"type1" : {  
"properties" : {  
"url" : {"type" : "string", "store" : "yes", "index" :  
"not\_analyzed"},  
"timestamp" : {"type" : "string", "store" : "yes"}  
}  
}  
}

On Mar 15, 9:12 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> No, there isn't a difference, its just simpler to recreate it with curl. The REST API is built on top of the Java API.
> 
> If you have problems with curl, then -\>_gist_\<- a simple test case that recreates it with Java, I can give it a go as well.
> 
> On Tuesday, March 15, 2011 at 5:57 PM, eugene wrote:
> 
> > I created it using java. Could there a difference if I do this in  
> > java instead of curl?
> 
> > On Mar 15, 12:56 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > 
> > > Can you provide a curl recreation?
> 
> > > On Tuesday, March 15, 2011 at 9:54 AM, eugene wrote:
> 
> > > > Here is how I created index 'test' and put mappings for 'type1':
> 
> > > > client.admin().indices().prepareCreate("test").execute().actionGet();
> 
> > > > String mappings =  
> > > > XContentFactory.jsonBuilder().startObject().startObject("type1").startObjec­­t("properties")  
> > > > .startObject("url").field("type", "string").field("store",  
> > > > "yes").field("index", "not\_analyzed").endObject()  
> > > > .startObject("timestamp").field("type",  
> > > > "string").field("store", "yes").endObject()  
> > > > .endObject().endObject().string();
> 
> > > > client.admin().indices().preparePutMapping("test").setType("type1").setSour­­ce(mappings).execute().actionGet();
> 
> > > > Then, I am writing data to the ES under 'test' index (note: url is  
> > > > encrypted, where %3F corresponds to ?, %26 to &, and %3D to =.
> 
> > > > String url = "http%3A//  
> > > > [someurl.com&nbsp;-&nbsp;This website is for sale!&nbsp;-&nbsp;someurl Resources and Information.](http://www.someurl.com/someservice%3FcityVegas%26State%3DNevada)";  
> > > > SimpleDateFormat sdf2 = new  
> > > > SimpleDateFormat("yyyy-MM-dd hh:mm:ss");  
> > > > Calendar cal =  
> > > > Calendar.getInstance();  
> > > > IndexResponse response = client.prepareIndex("test",  
> > > > "type1").setSource(jsonBuilder()  
> > > > .startObject()  
> > > > .field("url", url)  
> > > > .field("timestamp",  
> > > > sdf2.format(cal.getTime()))  
> > > > .endObject()  
> > > > )  
> > > > .execute()  
> > > > .actionGet();
> 
> > > > Here is my response to :\>curl -X GEThttp://localhost:9200/test/type1/\_mapping  
> > > > (assuming I put more than one documents as above):
> 
> > > > ..."facets":{"url":{"\_type":"terms","missing":0,"terms":  
> > > > [{"term":"[www.someurl.com](http://www.someurl.com)","count":4},{"term":"photoflipper","count":  
> > > > 4},{"term":"http  
> > > > ,"count":4},{"term":"3fcity","count":4},{"term":"3a","count":4},  
> > > > {"term":"26state","count":4},{"term":"26model","count":4},  
> > > > {"term":"3dNevada","count":2}]}}}
> 
> > > > On Mar 15, 12:00 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > > 
> > > > > Please post a full recreation:[Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/help)(includingputtingthe mapping and indexing data), simpler to help then.
> 
> > > > > On Monday, March 14, 2011 at 7:59 PM, eugene wrote:
> > > > > 
> > > > > > Shay,
> 
> > > > > > It appears I didn't fix the error. I run the following:
> 
> > > > > > "query" : {  
> > > > > > "match\_all": {}  
> > > > > > },  
> > > > > > "facets" : {  
> > > > > > "url" : {  
> > > > > > "terms" : {  
> > > > > > "field" : "url"  
> > > > > > }  
> > > > > > }  
> > > > > > }
> 
> > > > > > I am still getting tokenized terms for "url" field. I checked the  
> > > > > > url field and it is set not\_analyzed.
> 
> > > > > > C:\>curl -X GEThttp://localhost:9200/test/type1/\_mapping  
> > > > > > {"test":{"type1":{"properties":{"timestamp":  
> > > > > > {"store":"yes","type":"string"},"url":  
> > > > > > {"index":"not\_analyzed","store":"yes","type":"string"}}}}}
> 
> > > > > > Am I using the wrong facet?
> 
> > > > > > Thank you,  
> > > > > > Eugene.
> 
> > > > > > On Mar 11, 6:11 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > > > > 
> > > > > > > Not sure that this solved your problem, but a different change that you did where you don't use toString on the builder that represents the mapping. You can pass the builder directory to the putMapping API.
> 
> > > > > > > On Friday, March 11, 2011 at 4:14 AM, eugene wrote:
> 
> > > > > > > > I solved this problem by creating a new index first, and only then  
> > > > > > > > putting maps to it. I guess mappings should be done only once on a  
> > > > > > > > new index/type pair.
> 
> > > > > > > > -Eugene.
> 
> > > > > > > > On Mar 10, 3:03 pm, eugene [efur...@gmail.com](mailto:efur...@gmail.com) wrote:
> > > > > > > > 
> > > > > > > > > Ok, I am doing this:
> 
> > > > > > > > > String mappings =  
> > > > > > > > > XContentFactory.jsonBuilder().startObject().startObject("type1")  
> > > > > > > > > .startObject("myfield").field("type",  
> > > > > > > > > "string").field("store", "yes").field("index",  
> > > > > > > > > "not\_analyzed").endObject().endObject().endObject().toString();
> 
> > > > > > > > > client.admin().indices().preparePutMapping().setType("type1")  
> > > > > > > > > .setSource(mappings).execute().actionGet();
> 
> > > > > > > > > I am getting the following exception (notice: I replaced real ip  
> > > > > > > > > address with "ip\_address" on the first line).  
> > > > > > > > > What do you think it indicates? Thank you!
> 
> > > > > > > > > Eugene.
> 
> > > > > > > > > org.elasticsearch.transport.RemoteTransportException: [Stonecutter]  
> > > > > > > > > [inet[/ip\_address:9300]][indices/mapping/put]  
> > > > > > > > > Caused by: org.elasticsearch.ElasticSearchParseException: Failed to  
> > > > > > > > > derive xcontent from  
> > > > > > > > > org.elasticsearch.common.xcontent.XContentBuilder@1c94b8f  
> > > > > > > > > at  
> > > > > > > > > org.elasticsearch.common.xcontent.XContentFactory.xContent(XContentFactory.­­­­­java:  
> > > > > > > > > 136)  
> > > > > > > > > at  
> > > > > > > > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.extrac­­­­­tMapping(XContentDocumentMapperParser.java:  
> > > > > > > > > 316)  
> > > > > > > > > at  
> > > > > > > > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(­­­­­XContentDocumentMapperParser.java:  
> > > > > > > > > 114)  
> > > > > > > > > at  
> > > > > > > > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(­­­­­XContentDocumentMapperParser.java:  
> > > > > > > > > 54)  
> > > > > > > > > at  
> > > > > > > > > org.elasticsearch.index.mapper.MapperService.parse(MapperService.java:  
> > > > > > > > > 209)  
> > > > > > > > > at org.elasticsearch.cluster.metadata.MetaDataMappingService  
> > > > > > > > > $3.execute(MetaDataMappingService.java:193)  
> > > > > > > > > at org.elasticsearch.cluster.service.InternalClusterService  
> > > > > > > > > $2.run(InternalClusterService.java:175)  
> > > > > > > > > at java.util.concurrent.ThreadPoolExecutor  
> > > > > > > > > $Worker.runTask(ThreadPoolExecutor.java:886)  
> > > > > > > > > at java.util.concurrent.ThreadPoolExecutor  
> > > > > > > > > $Worker.run(ThreadPoolExecutor.java:908)  
> > > > > > > > > at java.lang.Thread.run(Thread.java:662)
> 
> > > > > > > > > On Mar 10, 4:16 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > > > > > > > > > There is an API to "putMapping" on an index, but, you can't change a field that is already analyzed to be "not\_analyzed" (as its part of the indexing process, so all current indexed docs will be meaningless).
> 
> > > > > > > > > > On Thursday, March 10, 2011 at 11:03 AM, eugene wrote:
> > > > > > > > > > 
> > > > > > > > > > > Thank you Shay.
> 
> > > > > > > > > > > Is there any way to add mappings to existing index? I searched the  
> > > > > > > > > > > source code and only found the examples for CreateIndexBuilder class.
> 
> > > > > > > > > > > Eugene.
> 
> > > > > > > > > > > On Mar 9, 10:18 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > > > > > > > > > 
> > > > > > > > > > > > You need to have the field you run the terms facet on marked as not\_analyzed in its mapping. You will need to reindex the data in order for that to take affect (and of course, create the mappings before you index data).
> 
> > > > > > > > > > > > On Thursday, March 10, 2011 at 5:12 AM, eugene wrote:
> > > > > > > > > > > > 
> > > > > > > > > > > > > Hi,
> 
> > > > > > > > > > > > > I am running a terms facet query on non-numeric field, and it is  
> > > > > > > > > > > > > working fine. However, if the field contains % symbol, the resutls in  
> > > > > > > > > > > > > a TermsFacet are tokenized into separate terms. For example, if I  
> > > > > > > > > > > > > have a json document:
> 
> > > > > > > > > > > > > {  
> > > > > > > > > > > > > url: http%[3Fwww.myserver.com/someaction/%3Fname%3DBob](http://3Fwww.myserver.com/someaction/%3Fname%3DBob)  
> > > > > > > > > > > > > }
> 
> > > > > > > > > > > > > and if I run: curl -X GEThttp://localhost:9200/\_river/my\_idx/\_search  
> > > > > > > > > > > > > -d  
> > > > > > > > > > > > > {  
> > > > > > > > > > > > > "query" : {  
> > > > > > > > > > > > > "match\_all" : {}  
> > > > > > > > > > > > > },  
> > > > > > > > > > > > > "facets" : {  
> > > > > > > > > > > > > "facet1" : {  
> > > > > > > > > > > > > "terms" : {  
> > > > > > > > > > > > > "field" : "url"  
> > > > > > > > > > > > > }  
> > > > > > > > > > > > > }  
> > > > > > > > > > > > > }  
> > > > > > > > > > > > > }
> 
> > > > > > > > > > > > > then it is returninng {"terms": [{"term": "[www.myserver.com](http://www.myserver.com)",  
> > > > > > > > > > > > > "count":1}, {"term": "http", "count":1}, {"term":"3Fname", "count":1},  
> > > > > > > > > > > > > {"term":"3DBob", "count":1}
> 
> > > > > > > > > > > > > I think it is supposed to return {"terms": [ {"term":"http  
> > > > > > > > > > > > > %[3Fwww.myserver.com/someaction/%3Fname%3DBob","count](http://3Fwww.myserver.com/someaction/%3Fname%3DBob%22,%22count)":1}
> 
> > > > > > > > > > > > > Please correct me if I am wrong. It never happens if a field doesn't  
> > > > > > > > > > > > > contain %.- Hide quoted text -
> 
> > > > > > > > > > > > - Show quoted text -- Hide quoted text -
> 
> > > > > > > > > > - Show quoted text -- Hide quoted text -
> 
> > > > > > > > > - Show quoted text -- Hide quoted text -
> 
> > > > > > > - Show quoted text -- Hide quoted text -
> 
> > > > > - Show quoted text -- Hide quoted text -
> 
> > > - Show quoted text -- Hide quoted text -
> 
> - Show quoted text -

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [March 15, 2011, 11:09pm UTC](https://discuss.elastic.co/t/terms-facet-is-tokenizing-a-field-with-special-characters/4071/15 "2011-03-15T23:09:05Z")

</div>

Eugene,

I can see that curl is problematic, just gist a simple (yet complete) Java test case, I will have a look.

-shay.banon  
On Tuesday, March 15, 2011 at 11:31 PM, eugene wrote:

> Here is how I created it with curl:
> 
> C:\>curl -X PUT [http://localhost:9200/test/type1/\_mappings](http://localhost:9200/test/type1/_mappings) -d  
> @mappings.json  
> {"ok":true,"\_index":"test","\_type":"type2","\_id":"\_mappings","\_version":  
> 1}
> 
> where the contents of mappings.json file are as follows:
> 
> {  
> "type1" : {  
> "properties" : {  
> "url" : {"type" : "string", "store" : "yes", "index" :  
> "not\_analyzed"},  
> "timestamp" : {"type" : "string", "store" : "yes"}  
> }  
> }  
> }
> 
> On Mar 15, 9:12 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> 
> > No, there isn't a difference, its just simpler to recreate it with curl. The REST API is built on top of the Java API.
> > 
> > If you have problems with curl, then -\>_gist_\<- a simple test case that recreates it with Java, I can give it a go as well.
> > 
> > On Tuesday, March 15, 2011 at 5:57 PM, eugene wrote:
> > 
> > > I created it using java. Could there a difference if I do this in  
> > > java instead of curl?
> > 
> > > On Mar 15, 12:56 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > 
> > > > Can you provide a curl recreation?
> > 
> > > > On Tuesday, March 15, 2011 at 9:54 AM, eugene wrote:
> > 
> > > > > Here is how I created index 'test' and put mappings for 'type1':
> > 
> > > > > client.admin().indices().prepareCreate("test").execute().actionGet();
> > 
> > > > > String mappings =  
> > > > > XContentFactory.jsonBuilder().startObject().startObject("type1").startObjecÂ­Â­t("properties")  
> > > > > .startObject("url").field("type", "string").field("store",  
> > > > > "yes").field("index", "not\_analyzed").endObject()  
> > > > > .startObject("timestamp").field("type",  
> > > > > "string").field("store", "yes").endObject()  
> > > > > .endObject().endObject().string();
> > 
> > > > > client.admin().indices().preparePutMapping("test").setType("type1").setSourÂ­Â­ce(mappings).execute().actionGet();
> > 
> > > > > Then, I am writing data to the ES under 'test' index (note: url is  
> > > > > encrypted, where %3F corresponds to ?, %26 to &, and %3D to =.
> > 
> > > > > String url = "http%3A//  
> > > > > [someurl.com&nbsp;-&nbsp;This website is for sale!&nbsp;-&nbsp;someurl Resources and Information.](http://www.someurl.com/someservice%3FcityVegas%26State%3DNevada)";  
> > > > > SimpleDateFormat sdf2 = new  
> > > > > SimpleDateFormat("yyyy-MM-dd hh:mm:ss");  
> > > > > Calendar cal =  
> > > > > Calendar.getInstance();  
> > > > > IndexResponse response = client.prepareIndex("test",  
> > > > > "type1").setSource(jsonBuilder()  
> > > > > .startObject()  
> > > > > .field("url", url)  
> > > > > .field("timestamp",  
> > > > > sdf2.format(cal.getTime()))  
> > > > > .endObject()  
> > > > > )  
> > > > > .execute()  
> > > > > .actionGet();
> > 
> > > > > Here is my response to :\>curl -X GEThttp://localhost:9200/test/type1/\_mapping  
> > > > > (assuming I put more than one documents as above):
> > 
> > > > > ..."facets":{"url":{"\_type":"terms","missing":0,"terms":  
> > > > > [{"term":"[www.someurl.com](http://www.someurl.com)","count":4},{"term":"photoflipper","count":  
> > > > > 4},{"term":"http  
> > > > > ,"count":4},{"term":"3fcity","count":4},{"term":"3a","count":4},  
> > > > > {"term":"26state","count":4},{"term":"26model","count":4},  
> > > > > {"term":"3dNevada","count":2}]}}}
> > 
> > > > > On Mar 15, 12:00 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > > > 
> > > > > > Please post a full recreation:[Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/help)(includingputtingthe mapping and indexing data), simpler to help then.
> > 
> > > > > > On Monday, March 14, 2011 at 7:59 PM, eugene wrote:
> > > > > > 
> > > > > > > Shay,
> > 
> > > > > > > It appears I didn't fix the error. I run the following:
> > 
> > > > > > > "query" : {  
> > > > > > > "match\_all": {}  
> > > > > > > },  
> > > > > > > "facets" : {  
> > > > > > > "url" : {  
> > > > > > > "terms" : {  
> > > > > > > "field" : "url"  
> > > > > > > }  
> > > > > > > }  
> > > > > > > }
> > 
> > > > > > > I am still getting tokenized terms for "url" field. I checked the  
> > > > > > > url field and it is set not\_analyzed.
> > 
> > > > > > > C:\>curl -X GEThttp://localhost:9200/test/type1/\_mapping  
> > > > > > > {"test":{"type1":{"properties":{"timestamp":  
> > > > > > > {"store":"yes","type":"string"},"url":  
> > > > > > > {"index":"not\_analyzed","store":"yes","type":"string"}}}}}
> > 
> > > > > > > Am I using the wrong facet?
> > 
> > > > > > > Thank you,  
> > > > > > > Eugene.
> > 
> > > > > > > On Mar 11, 6:11 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > > > > > 
> > > > > > > > Not sure that this solved your problem, but a different change that you did where you don't use toString on the builder that represents the mapping. You can pass the builder directory to the putMapping API.
> > 
> > > > > > > > On Friday, March 11, 2011 at 4:14 AM, eugene wrote:
> > 
> > > > > > > > > I solved this problem by creating a new index first, and only then  
> > > > > > > > > putting maps to it. I guess mappings should be done only once on a  
> > > > > > > > > new index/type pair.
> > 
> > > > > > > > > -Eugene.
> > 
> > > > > > > > > On Mar 10, 3:03 pm, eugene [efur...@gmail.com](mailto:efur...@gmail.com) wrote:
> > > > > > > > > 
> > > > > > > > > > Ok, I am doing this:
> > 
> > > > > > > > > > String mappings =  
> > > > > > > > > > XContentFactory.jsonBuilder().startObject().startObject("type1")  
> > > > > > > > > > .startObject("myfield").field("type",  
> > > > > > > > > > "string").field("store", "yes").field("index",  
> > > > > > > > > > "not\_analyzed").endObject().endObject().endObject().toString();
> > 
> > > > > > > > > > client.admin().indices().preparePutMapping().setType("type1")  
> > > > > > > > > > .setSource(mappings).execute().actionGet();
> > 
> > > > > > > > > > I am getting the following exception (notice: I replaced real ip  
> > > > > > > > > > address with "ip\_address" on the first line).  
> > > > > > > > > > What do you think it indicates? Thank you!
> > 
> > > > > > > > > > Eugene.
> > 
> > > > > > > > > > org.elasticsearch.transport.RemoteTransportException: [Stonecutter]  
> > > > > > > > > > [inet[/ip\_address:9300]][indices/mapping/put]  
> > > > > > > > > > Caused by: org.elasticsearch.ElasticSearchParseException: Failed to  
> > > > > > > > > > derive xcontent from  
> > > > > > > > > > org.elasticsearch.common.xcontent.XContentBuilder@1c94b8f  
> > > > > > > > > > at  
> > > > > > > > > > org.elasticsearch.common.xcontent.XContentFactory.xContent(XContentFactory.Â­Â­Â­Â­Â­java:  
> > > > > > > > > > 136)  
> > > > > > > > > > at  
> > > > > > > > > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.extracÂ­Â­Â­Â­Â­tMapping(XContentDocumentMapperParser.java:  
> > > > > > > > > > 316)  
> > > > > > > > > > at  
> > > > > > > > > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(Â­Â­Â­Â­Â­XContentDocumentMapperParser.java:  
> > > > > > > > > > 114)  
> > > > > > > > > > at  
> > > > > > > > > > org.elasticsearch.index.mapper.xcontent.XContentDocumentMapperParser.parse(Â­Â­Â­Â­Â­XContentDocumentMapperParser.java:  
> > > > > > > > > > 54)  
> > > > > > > > > > at  
> > > > > > > > > > org.elasticsearch.index.mapper.MapperService.parse(MapperService.java:  
> > > > > > > > > > 209)  
> > > > > > > > > > at org.elasticsearch.cluster.metadata.MetaDataMappingService  
> > > > > > > > > > $3.execute(MetaDataMappingService.java:193)  
> > > > > > > > > > at org.elasticsearch.cluster.service.InternalClusterService  
> > > > > > > > > > $2.run(InternalClusterService.java:175)  
> > > > > > > > > > at java.util.concurrent.ThreadPoolExecutor  
> > > > > > > > > > $Worker.runTask(ThreadPoolExecutor.java:886)  
> > > > > > > > > > at java.util.concurrent.ThreadPoolExecutor  
> > > > > > > > > > $Worker.run(ThreadPoolExecutor.java:908)  
> > > > > > > > > > at java.lang.Thread.run(Thread.java:662)
> > 
> > > > > > > > > > On Mar 10, 4:16 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > 
> > > > > > > > > > > There is an API to "putMapping" on an index, but, you can't change a field that is already analyzed to be "not\_analyzed" (as its part of the indexing process, so all current indexed docs will be meaningless).
> > 
> > > > > > > > > > > On Thursday, March 10, 2011 at 11:03 AM, eugene wrote:
> > > > > > > > > > > 
> > > > > > > > > > > > Thank you Shay.
> > 
> > > > > > > > > > > > Is there any way to add mappings to existing index? I searched the  
> > > > > > > > > > > > source code and only found the examples for CreateIndexBuilder class.
> > 
> > > > > > > > > > > > Eugene.
> > 
> > > > > > > > > > > > On Mar 9, 10:18 pm, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:
> > > > > > > > > > > > 
> > > > > > > > > > > > > You need to have the field you run the terms facet on marked as not\_analyzed in its mapping. You will need to reindex the data in order for that to take affect (and of course, create the mappings before you index data).
> > 
> > > > > > > > > > > > > On Thursday, March 10, 2011 at 5:12 AM, eugene wrote:
> > > > > > > > > > > > > 
> > > > > > > > > > > > > > Hi,
> > 
> > > > > > > > > > > > > > I am running a terms facet query on non-numeric field, and it is  
> > > > > > > > > > > > > > working fine. However, if the field contains % symbol, the resutls in  
> > > > > > > > > > > > > > a TermsFacet are tokenized into separate terms. For example, if I  
> > > > > > > > > > > > > > have a json document:
> > 
> > > > > > > > > > > > > > {  
> > > > > > > > > > > > > > url: http%[3Fwww.myserver.com/someaction/%3Fname%3DBob](http://3Fwww.myserver.com/someaction/%3Fname%3DBob)  
> > > > > > > > > > > > > > }
> > 
> > > > > > > > > > > > > > and if I run: curl -X GEThttp://localhost:9200/\_river/my\_idx/\_search  
> > > > > > > > > > > > > > -d  
> > > > > > > > > > > > > > {  
> > > > > > > > > > > > > > "query" : {  
> > > > > > > > > > > > > > "match\_all" : {}  
> > > > > > > > > > > > > > },  
> > > > > > > > > > > > > > "facets" : {  
> > > > > > > > > > > > > > "facet1" : {  
> > > > > > > > > > > > > > "terms" : {  
> > > > > > > > > > > > > > "field" : "url"  
> > > > > > > > > > > > > > }  
> > > > > > > > > > > > > > }  
> > > > > > > > > > > > > > }  
> > > > > > > > > > > > > > }
> > 
> > > > > > > > > > > > > > then it is returninng {"terms": [{"term": "[www.myserver.com](http://www.myserver.com)",  
> > > > > > > > > > > > > > "count":1}, {"term": "http", "count":1}, {"term":"3Fname", "count":1},  
> > > > > > > > > > > > > > {"term":"3DBob", "count":1}
> > 
> > > > > > > > > > > > > > I think it is supposed to return {"terms": [ {"term":"http  
> > > > > > > > > > > > > > %[3Fwww.myserver.com/someaction/%3Fname%3DBob","count](http://3Fwww.myserver.com/someaction/%3Fname%3DBob%22,%22count)":1}
> > 
> > > > > > > > > > > > > > Please correct me if I am wrong. It never happens if a field doesn't  
> > > > > > > > > > > > > > contain %.- Hide quoted text -
> > 
> > > > > > > > > > > > > - Show quoted text -- Hide quoted text -
> > 
> > > > > > > > > > > - Show quoted text -- Hide quoted text -
> > 
> > > > > > > > > > - Show quoted text -- Hide quoted text -
> > 
> > > > > > > > - Show quoted text -- Hide quoted text -
> > 
> > > > > > - Show quoted text -- Hide quoted text -
> > 
> > > > - Show quoted text -- Hide quoted text -
> > 
> > - Show quoted text -

---

<div class="post-metadata">

**Author:** ![SANDPATH7](https://avatars.discourse-cdn.com/v4/letter/s/c37758/32.png) [@SANDPATH7](https://discuss.elastic.co/u/SANDPATH7)\
**Post date:** [February 16, 2012, 9:10pm UTC](https://discuss.elastic.co/t/terms-facet-is-tokenizing-a-field-with-special-characters/4071/16 "2012-02-16T21:10:02Z")

</div>

Hi,

I am also facing the same problem. I am using facet to get all the unique values and their count for a field. And i am getting wrong result.

term: web  
Count: 1191979  
term: misc  
Count: 1191979  
term: passwd  
Count: 1191979  
term: etc  
Count: 1191979

While the actual result should be:  
term: WEB-MISC /etc/passwd  
Count: 1191979

The term is getting truncated. Is some thing else, i need to do?

Thanks,

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:38am UTC](https://discuss.elastic.co/t/terms-facet-is-tokenizing-a-field-with-special-characters/4071/17 "2017-07-06T03:38:59Z")

</div>


