# matchPhraseQuery can not retrieve documents with trailing “’s” even if set word delimiter tokenfilter when created indices

**URL:** <https://discuss.elastic.co/t/matchphrasequery-can-not-retrieve-documents-with-trailing-s-even-if-set-word-delimiter-tokenfilter-when-created-indices/11934>\
**Category:** Elasticsearch\
**Created:** [May 13, 2013, 9:20am UTC](https://discuss.elastic.co/t/matchphrasequery-can-not-retrieve-documents-with-trailing-s-even-if-set-word-delimiter-tokenfilter-when-created-indices/11934 "2013-05-13T09:20:14Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jingang\_Wang](https://avatars.discourse-cdn.com/v4/letter/j/ce7236/32.png) [@Jingang\_Wang](https://discuss.elastic.co/u/Jingang_Wang)\
**Post date:** [May 13, 2013, 9:20am UTC](https://discuss.elastic.co/t/matchphrasequery-can-not-retrieve-documents-with-trailing-s-even-if-set-word-delimiter-tokenfilter-when-created-indices/11934/1 "2013-05-13T09:20:14Z")

</div>

Hi there,

I want to use ES to query some documents mentioning some person names.

For example, when I use Bill Gates to conduct a matchPhraseQuery, I can  
just get the documents exactly mention the name "Bill Gates".  
While a lot of documents mention Bill Gates indirectly, say, they may  
mention "Bill Gates's company".

How should I construct a query to retrieve these documents? Thanks.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![brian\_yoder](https://avatars.discourse-cdn.com/v4/letter/b/f1d935/32.png) [@brian\_yoder](https://discuss.elastic.co/u/brian_yoder)\
**Post date:** [May 13, 2013, 7:20pm UTC](https://discuss.elastic.co/t/matchphrasequery-can-not-retrieve-documents-with-trailing-s-even-if-set-word-delimiter-tokenfilter-when-created-indices/11934/2 "2013-05-13T19:20:31Z")

</div>

You probably want a stemming analyzer (for example, snowball). In that  
case, for example, Debbie's, Debby's, Debby, and Debbie all match each  
other.

Of course, the documents will need to be reloaded to cause the proper  
stemmed terms to be indexed.

Cheers!  
Brian

On Monday, May 13, 2013 5:20:14 AM UTC-4, Jingang Wang wrote:

> Hi there,
> 
> I want to use ES to query some documents mentioning some person names.
> 
> For example, when I use Bill Gates to conduct a matchPhraseQuery, I can  
> just get the documents exactly mention the name "Bill Gates".  
> While a lot of documents mention Bill Gates indirectly, say, they may  
> mention "Bill Gates's company".
> 
> How should I construct a query to retrieve these documents? Thanks.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![brian\_yoder](https://avatars.discourse-cdn.com/v4/letter/b/f1d935/32.png) [@brian\_yoder](https://discuss.elastic.co/u/brian_yoder)\
**Post date:** [May 13, 2013, 10:41pm UTC](https://discuss.elastic.co/t/matchphrasequery-can-not-retrieve-documents-with-trailing-s-even-if-set-word-delimiter-tokenfilter-when-created-indices/11934/3 "2013-05-13T22:41:29Z")

</div>

Hi Jingang,

Here is a full example with the index settings and mappings and a curl  
command to show how various forms of Gates may be indexed so that they  
match each other without any additional work on the part of your query:

Analyzing the string using the "cn" field in the "test" index:

$ curl '[http://localhost:9200/test/\_analyze?field=cn&pretty=true](http://localhost:9200/test/_analyze?field=cn&pretty=true)' -d "_gates  
gate's gates' gates's_" && echo  
{  
"tokens" : [ {  
"token" : "_gate_",  
"start\_offset" : 0,  
"end\_offset" : 5,  
"type" : "",  
"position" : 1  
}, {  
"token" : "_gate_",  
"start\_offset" : 6,  
"end\_offset" : 12,  
"type" : "",  
"position" : 2  
}, {  
"token" : "_gate_",  
"start\_offset" : 13,  
"end\_offset" : 18,  
"type" : "",  
"position" : 3  
}, {  
"token" : "_gate_",  
"start\_offset" : 20,  
"end\_offset" : 27,  
"type" : "",  
"position" : 4  
} ]  
}

And here are a subset of the settings and mappings. Note that I needed to  
fully construct my own filter and analyzers in order to more fully specify  
things such as the language to use.

{  
"settings" : {  
"index" : {  
"number\_of\_shards" : 1,  
"refresh\_interval" : "2s",  
"number\_of\_replicas" : 0,  
"analysis" : {

```
    "filter" : {
      "english_snowball_filter" : {
        "type" : "snowball",
        "language" : "English"
      }
    },
    "analyzer" : {
      "english_stemming_analyzer" : {
        "type" : "custom",
        "tokenizer" : "standard",
        "filter" : [ "standard", "lowercase", "asciifolding", 

```

"english\_snowball\_filter" ]  
},  
"english\_standard\_analyzer" : {  
"type" : "custom",  
"tokenizer" : "standard",  
"filter" : ["standard", "lowercase", "asciifolding"]  
}  
}  
}  
}  
},  
"mappings" : {  
"person" : {  
"\_all" : {  
"enabled" : false  
},  
"properties" : {  
"uid" : {  
"type" : "long"  
},  
"cn" : {  
"type" : "multi\_field",  
"fields" : {  
"cn" : {  
"type" : "string",  
"analyzer" : "english\_stemming\_analyzer"  
},  
"raw" : {  
"type" : "string",  
"index" : "not\_analyzed"  
}  
}  
},  
"location" : {  
"type" : "geo\_point",  
"lat\_lon" : true  
},  
"telno" : {  
"type" : "multi\_field",  
"fields" : {  
"telno" : {  
"type" : "string",  
"analyzer" : "english\_standard\_analyzer"  
},  
"num" : {  
"type" : "long"  
}  
}  
},  
"date" : {  
"type" : "date",  
"format" : "dateOptionalTime"  
}  
}  
}  
}  
}

Hope this helps!

Regards,  
Brian

On Monday, May 13, 2013 5:20:14 AM UTC-4, Jingang Wang wrote:

> Hi there,
> 
> I want to use ES to query some documents mentioning some person names.
> 
> For example, when I use Bill Gates to conduct a matchPhraseQuery, I can  
> just get the documents exactly mention the name "Bill Gates".  
> While a lot of documents mention Bill Gates indirectly, say, they may  
> mention "Bill Gates's company".
> 
> How should I construct a query to retrieve these documents? Thanks.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![brian\_yoder](https://avatars.discourse-cdn.com/v4/letter/b/f1d935/32.png) [@brian\_yoder](https://discuss.elastic.co/u/brian_yoder)\
**Post date:** [May 13, 2013, 10:59pm UTC](https://discuss.elastic.co/t/matchphrasequery-can-not-retrieve-documents-with-trailing-s-even-if-set-word-delimiter-tokenfilter-when-created-indices/11934/4 "2013-05-13T22:59:20Z")

</div>

And here's another example of analyzing the same string as below, but this  
time using the _cn.raw_ field. It's an example of multi-field mapping in  
which a field may be indexed (or not) in two or more ways but its source  
values only need to be stored once. Really awesome!!!

$ curl '[http://localhost:9200/sgen/\_analyze?field=\*cn.raw\*&pretty=true](http://localhost:9200/sgen/_analyze?field=*cn.raw*&pretty=true)' -d _"gates  
gate's gates' gates's"_ && echo  
{  
"tokens" : [ {  
"token" : _"gates gate's gates' gates's"_,  
"start\_offset" : 0,  
"end\_offset" : 27,  
"type" : "word",  
"position" : 1  
} ]  
}

On Monday, May 13, 2013 6:41:29 PM UTC-4, InquiringMind wrote:

> Hi Jingang,
> 
> Here is a full example with the index settings and mappings and a curl  
> command to show how various forms of Gates may be indexed so that they  
> match each other without any additional work on the part of your query:
> 
> Analyzing the string using the "cn" field in the "test" index:
> 
> $ curl '[http://localhost:9200/test/\_analyze?field=cn&pretty=true](http://localhost:9200/test/_analyze?field=cn&pretty=true)' -d "_gates  
> gate's gates' gates's_" && echo  
> {  
> "tokens" : [ {  
> "token" : "_gate_",  
> "start\_offset" : 0,  
> "end\_offset" : 5,  
> "type" : "",  
> "position" : 1  
> }, {  
> "token" : "_gate_",  
> "start\_offset" : 6,  
> "end\_offset" : 12,  
> "type" : "",  
> "position" : 2  
> }, {  
> "token" : "_gate_",  
> "start\_offset" : 13,  
> "end\_offset" : 18,  
> "type" : "",  
> "position" : 3  
> }, {  
> "token" : "_gate_",  
> "start\_offset" : 20,  
> "end\_offset" : 27,  
> "type" : "",  
> "position" : 4  
> } ]  
> }
> 
> And here are a subset of the settings and mappings. Note that I needed to  
> fully construct my own filter and analyzers in order to more fully specify  
> things such as the language to use.
> 
> {  
> "settings" : {  
> "index" : {  
> "number\_of\_shards" : 1,  
> "refresh\_interval" : "2s",  
> "number\_of\_replicas" : 0,  
> "analysis" : {
> 
> ```
> "filter" : {
> "english_snowball_filter" : {
> "type" : "snowball",
> "language" : "English"
> }
> },
> "analyzer" : {
> "english_stemming_analyzer" : {
> "type" : "custom",
> "tokenizer" : "standard",
> "filter" : [ "standard", "lowercase", "asciifolding", 
> 
> ```
> 
> "english\_snowball\_filter" ]  
> },  
> "english\_standard\_analyzer" : {  
> "type" : "custom",  
> "tokenizer" : "standard",  
> "filter" : ["standard", "lowercase", "asciifolding"]  
> }  
> }  
> }  
> }  
> },  
> "mappings" : {  
> "person" : {  
> "\_all" : {  
> "enabled" : false  
> },  
> "properties" : {  
> "uid" : {  
> "type" : "long"  
> },  
> "cn" : {  
> "type" : "multi\_field",  
> "fields" : {  
> "cn" : {  
> "type" : "string",  
> "analyzer" : "english\_stemming\_analyzer"  
> },  
> "raw" : {  
> "type" : "string",  
> "index" : "not\_analyzed"  
> }  
> }  
> },  
> "location" : {  
> "type" : "geo\_point",  
> "lat\_lon" : true  
> },  
> "telno" : {  
> "type" : "multi\_field",  
> "fields" : {  
> "telno" : {  
> "type" : "string",  
> "analyzer" : "english\_standard\_analyzer"  
> },  
> "num" : {  
> "type" : "long"  
> }  
> }  
> },  
> "date" : {  
> "type" : "date",  
> "format" : "dateOptionalTime"  
> }  
> }  
> }  
> }  
> }
> 
> Hope this helps!
> 
> Regards,  
> Brian
> 
> On Monday, May 13, 2013 5:20:14 AM UTC-4, Jingang Wang wrote:
> 
> > Hi there,
> > 
> > I want to use ES to query some documents mentioning some person names.
> > 
> > For example, when I use Bill Gates to conduct a matchPhraseQuery, I can  
> > just get the documents exactly mention the name "Bill Gates".  
> > While a lot of documents mention Bill Gates indirectly, say, they may  
> > mention "Bill Gates's company".
> > 
> > How should I construct a query to retrieve these documents? Thanks.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Jingang\_Wang](https://avatars.discourse-cdn.com/v4/letter/j/ce7236/32.png) [@Jingang\_Wang](https://discuss.elastic.co/u/Jingang_Wang)\
**Post date:** [May 14, 2013, 2:53am UTC](https://discuss.elastic.co/t/matchphrasequery-can-not-retrieve-documents-with-trailing-s-even-if-set-word-delimiter-tokenfilter-when-created-indices/11934/5 "2013-05-14T02:53:31Z")

</div>

Dear  
​ Brian,

Thanks you so much for elaborate examples and explanation.  
​I have changed my settings as your example, while I still can not match  
the phrase in the documents.

Here is my settings of index:  
​  
{  
"20120103": {  
"settings": {  
"index.analysis.filter.my\_delimiter.generate\_word\_parts": "true",  
"index.analysis.filter.my\_delimiter.stem\_english\_possessive": "true",  
"index.analysis.filter.my\_delimiter.preserve\_original": "true",  
"index.analysis.analyzer.my\_analyzer.tokenizer": "standard",  
"index.analysis.filter.english\_stemming\_filter.language": "English",  
"index.analysis.filter.my\_delimiter.split\_on\_numerics": "true",  
"index.analysis.filter.my\_delimiter.catenate\_all": "true",  
"index.number\_of\_shards": "10",  
"index.analysis.filter.my\_delimiter.catenate\_numbers": "true",  
"index.analysis.analyzer.my\_analyzer.type": "custom",  
"index.analysis.filter.english\_stemming\_filter.type": "snowball",  
"index.analysis.filter.my\_delimiter.type": "word\_delimiter",  
"index.analysis.filter.my\_delimiter.catenate\_words": "true",  
"index.number\_of\_replicas": "0",  
"index.analysis.analyzer.my\_analyzer.filter.2": "my\_delimiter",  
"index.analysis.analyzer.my\_analyzer.filter.1": "lowercase",  
"index.analysis.analyzer.my\_analyzer.filter.0": "standard",  
"index.analysis.filter.my\_delimiter.split\_on\_case\_change": "true",  
"index.analysis.analyzer.my\_analyzer.filter.5": "stop",  
"index.analysis.analyzer.my\_analyzer.filter.4":  
"english\_stemming\_filter",  
"index.analysis.analyzer.my\_analyzer.filter.3": "asciifolding",  
"index.version.created": "900001"  
}  
}  
}

And my query is：

QueryBuilder qb = QueryBuilders  
.boolQuery()  
.must(matchPhraseQuery("body\_cleansed", "Aharon  
Barak").analyzer("my\_analyzer"));

There is a documents which contains “Aharon Barak's policy”, but the query  
could not retrieve it.

On Tue, May 14, 2013 at 6:59 AM, InquiringMind [brian.from.fl@gmail.com](mailto:brian.from.fl@gmail.com)wrote:

> And here's another example of analyzing the same string as below, but this  
> time using the _cn.raw_ field. It's an example of multi-field mapping in  
> which a field may be indexed (or not) in two or more ways but its source  
> values only need to be stored once. Really awesome!!!
> 
> $ curl '[http://localhost:9200/sgen/\_analyze?field=\*cn.raw\*&pretty=true](http://localhost:9200/sgen/_analyze?field=*cn.raw*&pretty=true)'  
> -d _"gates gate's gates' gates's"_ && echo  
> {  
> "tokens" : [ {  
> "token" : _"gates gate's gates' gates's"_,  
> "start\_offset" : 0,  
> "end\_offset" : 27,  
> "type" : "word",  
> "position" : 1  
> } ]  
> }
> 
> On Monday, May 13, 2013 6:41:29 PM UTC-4, InquiringMind wrote:
> 
> > Hi Jingang,
> > 
> > Here is a full example with the index settings and mappings and a curl  
> > command to show how various forms of Gates may be indexed so that they  
> > match each other without any additional work on the part of your query:
> > 
> > Analyzing the string using the "cn" field in the "test" index:
> > 
> > $ curl '[http://localhost:9200/test/\_\*\*analyze?field=cn&pretty=true](http://localhost:9200/test/_**analyze?field=cn&pretty=true)[http://localhost:9200/test/\_analyze?field=cn&pretty=true](http://localhost:9200/test/_analyze?field=cn&pretty=true)'  
> > -d "_gates gate's gates' gates's_" && echo  
> > {  
> > "tokens" : [ {  
> > "token" : "_gate_",  
> > "start\_offset" : 0,  
> > "end\_offset" : 5,  
> > "type" : "",  
> > "position" : 1  
> > }, {  
> > "token" : "_gate_",  
> > "start\_offset" : 6,  
> > "end\_offset" : 12,  
> > "type" : "",  
> > "position" : 2  
> > }, {  
> > "token" : "_gate_",  
> > "start\_offset" : 13,  
> > "end\_offset" : 18,  
> > "type" : "",  
> > "position" : 3  
> > }, {  
> > "token" : "_gate_",  
> > "start\_offset" : 20,  
> > "end\_offset" : 27,  
> > "type" : "",  
> > "position" : 4  
> > } ]  
> > }
> > 
> > And here are a subset of the settings and mappings. Note that I needed to  
> > fully construct my own filter and analyzers in order to more fully specify  
> > things such as the language to use.
> > 
> > {  
> > "settings" : {  
> > "index" : {  
> > "number\_of\_shards" : 1,  
> > "refresh\_interval" : "2s",  
> > "number\_of\_replicas" : 0,  
> > "analysis" : {
> > 
> > ```
> > "filter" : {
> > "english_snowball_filter" : {
> > "type" : "snowball",
> > "language" : "English"
> > }
> > },
> > "analyzer" : {
> > "english_stemming_analyzer" : {
> > "type" : "custom",
> > "tokenizer" : "standard",
> > "filter" : [ "standard", "lowercase", "asciifolding",
> > 
> > ```
> > 
> > "english\_snowball\_filter" ]  
> > },  
> > "english\_standard\_analyzer" : {  
> > "type" : "custom",  
> > "tokenizer" : "standard",  
> > "filter" : ["standard", "lowercase", "asciifolding"]  
> > }  
> > }  
> > }  
> > }  
> > },  
> > "mappings" : {  
> > "person" : {  
> > "\_all" : {  
> > "enabled" : false  
> > },  
> > "properties" : {  
> > "uid" : {  
> > "type" : "long"  
> > },  
> > "cn" : {  
> > "type" : "multi\_field",  
> > "fields" : {  
> > "cn" : {  
> > "type" : "string",  
> > "analyzer" : "english\_stemming\_analyzer"  
> > },  
> > "raw" : {  
> > "type" : "string",  
> > "index" : "not\_analyzed"  
> > }  
> > }  
> > },  
> > "location" : {  
> > "type" : "geo\_point",  
> > "lat\_lon" : true  
> > },  
> > "telno" : {  
> > "type" : "multi\_field",  
> > "fields" : {  
> > "telno" : {  
> > "type" : "string",  
> > "analyzer" : "english\_standard\_analyzer"  
> > },  
> > "num" : {  
> > "type" : "long"  
> > }  
> > }  
> > },  
> > "date" : {  
> > "type" : "date",  
> > "format" : "dateOptionalTime"  
> > }  
> > }  
> > }  
> > }  
> > }
> > 
> > Hope this helps!
> > 
> > Regards,  
> > Brian
> > 
> > On Monday, May 13, 2013 5:20:14 AM UTC-4, Jingang Wang wrote:
> > 
> > > Hi there,
> > > 
> > > I want to use ES to query some documents mentioning some person names.
> > > 
> > > For example, when I use Bill Gates to conduct a matchPhraseQuery, I can  
> > > just get the documents exactly mention the name "Bill Gates".  
> > > While a lot of documents mention Bill Gates indirectly, say, they may  
> > > mention "Bill Gates's company".
> > > 
> > > How should I construct a query to retrieve these documents? Thanks.
> > 
> > --  
> > You received this message because you are subscribed to a topic in the  
> > Google Groups "elasticsearch" group.  
> > To unsubscribe from this topic, visit  
> > [https://groups.google.com/d/topic/elasticsearch/SIK3lc215Bk/unsubscribe?hl=en-US](https://groups.google.com/d/topic/elasticsearch/SIK3lc215Bk/unsubscribe?hl=en-US)  
> > .  
> > To unsubscribe from this group and all its topics, send an email to  
> > [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
Wang Jingang(王金刚)  
Ph.D. Candidate at  
Lab of High Volume Language Information Processing & Cloud Computing  
School of Computer Science  
Beijing Institute of Technology  
Beijing 100081  
P.R China

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Jingang\_Wang](https://avatars.discourse-cdn.com/v4/letter/j/ce7236/32.png) [@Jingang\_Wang](https://discuss.elastic.co/u/Jingang_Wang)\
**Post date:** [May 14, 2013, 3:17am UTC](https://discuss.elastic.co/t/matchphrasequery-can-not-retrieve-documents-with-trailing-s-even-if-set-word-delimiter-tokenfilter-when-created-indices/11934/6 "2013-05-14T03:17:25Z")

</div>

My mapping looks like as follows;

XContentBuilder mapping = jsonBuilder()  
.startObject()  
.startObject("kba")  
.startObject("\_all").field("enable", true).field("index\_analyzer",  
"my\_analyzer").field("search\_analyzer", "my\_analyzer").endObject()  
.startObject("properties")  
.startObject("\_source").field("compress", "true").endObject()  
.startObject("stream\_id").field("type",  
"string").field("index","not\_analyzed").endObject()  
.startObject("source").field("type", "string").field("index",  
"not\_analyzed").endObject()  
.startObject("epoch\_ticks").field("type",  
"double").field("index", "not\_analyzed").endObject()  
.startObject("zulu\_timestamp").field("type",  
"string").field("index","not\_analyzed").endObject()  
.startObject("title\_cleansed").field("type","string").field("index",  
"analyzed").field("index\_analyzer","my\_analyzer").field("search\_analyzer","my\_analyzer").endObject()

.startObject("body\_cleansed").field("type","string").field("index",  
"analyzed").field("index\_analyzer","my\_analyzer").field("search\_analyzer","my\_analyzer").endObject()  
.endObject()  
.endObject()  
.endObject();  
PutMappingRequest mappingRequest =  
Requests.putMappingRequest(indexName).type("kba").source(mapping);

On Tue, May 14, 2013 at 10:53 AM, Jingang Wang [bitwjg@gmail.com](mailto:bitwjg@gmail.com) wrote:

> Dear  
> ​ Brian,
> 
> Thanks you so much for elaborate examples and explanation.  
> ​I have changed my settings as your example, while I still can not match  
> the phrase in the documents.
> 
> Here is my settings of index:  
> ​  
> {  
> "20120103": {  
> "settings": {  
> "index.analysis.filter.my\_delimiter.generate\_word\_parts": "true",  
> "index.analysis.filter.my\_delimiter.stem\_english\_possessive": "true",  
> "index.analysis.filter.my\_delimiter.preserve\_original": "true",  
> "index.analysis.analyzer.my\_analyzer.tokenizer": "standard",  
> "index.analysis.filter.english\_stemming\_filter.language": "English",  
> "index.analysis.filter.my\_delimiter.split\_on\_numerics": "true",  
> "index.analysis.filter.my\_delimiter.catenate\_all": "true",  
> "index.number\_of\_shards": "10",  
> "index.analysis.filter.my\_delimiter.catenate\_numbers": "true",  
> "index.analysis.analyzer.my\_analyzer.type": "custom",  
> "index.analysis.filter.english\_stemming\_filter.type": "snowball",  
> "index.analysis.filter.my\_delimiter.type": "word\_delimiter",  
> "index.analysis.filter.my\_delimiter.catenate\_words": "true",  
> "index.number\_of\_replicas": "0",  
> "index.analysis.analyzer.my\_analyzer.filter.2": "my\_delimiter",  
> "index.analysis.analyzer.my\_analyzer.filter.1": "lowercase",  
> "index.analysis.analyzer.my\_analyzer.filter.0": "standard",  
> "index.analysis.filter.my\_delimiter.split\_on\_case\_change": "true",  
> "index.analysis.analyzer.my\_analyzer.filter.5": "stop",  
> "index.analysis.analyzer.my\_analyzer.filter.4":  
> "english\_stemming\_filter",  
> "index.analysis.analyzer.my\_analyzer.filter.3": "asciifolding",  
> "index.version.created": "900001"  
> }  
> }  
> }
> 
> And my query is：
> 
> QueryBuilder qb = QueryBuilders  
> .boolQuery()  
> .must(matchPhraseQuery("body\_cleansed", "Aharon  
> Barak").analyzer("my\_analyzer"));
> 
> There is a documents which contains “Aharon Barak's policy”, but the query  
> could not retrieve it.
> 
> On Tue, May 14, 2013 at 6:59 AM, InquiringMind [brian.from.fl@gmail.com](mailto:brian.from.fl@gmail.com)wrote:
> 
> > And here's another example of analyzing the same string as below, but  
> > this time using the _cn.raw_ field. It's an example of multi-field  
> > mapping in which a field may be indexed (or not) in two or more ways but  
> > its source values only need to be stored once. Really awesome!!!
> > 
> > $ curl '[http://localhost:9200/sgen/\_analyze?field=\*cn.raw\*&pretty=true](http://localhost:9200/sgen/_analyze?field=*cn.raw*&pretty=true)'  
> > -d _"gates gate's gates' gates's"_ && echo  
> > {  
> > "tokens" : [ {  
> > "token" : _"gates gate's gates' gates's"_,  
> > "start\_offset" : 0,  
> > "end\_offset" : 27,  
> > "type" : "word",  
> > "position" : 1  
> > } ]  
> > }
> > 
> > On Monday, May 13, 2013 6:41:29 PM UTC-4, InquiringMind wrote:
> > 
> > > Hi Jingang,
> > > 
> > > Here is a full example with the index settings and mappings and a curl  
> > > command to show how various forms of Gates may be indexed so that they  
> > > match each other without any additional work on the part of your query:
> > > 
> > > Analyzing the string using the "cn" field in the "test" index:
> > > 
> > > $ curl '[http://localhost:9200/test/\_\*\*analyze?field=cn&pretty=true](http://localhost:9200/test/_**analyze?field=cn&pretty=true)[http://localhost:9200/test/\_analyze?field=cn&pretty=true](http://localhost:9200/test/_analyze?field=cn&pretty=true)'  
> > > -d "_gates gate's gates' gates's_" && echo  
> > > {  
> > > "tokens" : [ {  
> > > "token" : "_gate_",  
> > > "start\_offset" : 0,  
> > > "end\_offset" : 5,  
> > > "type" : "",  
> > > "position" : 1  
> > > }, {  
> > > "token" : "_gate_",  
> > > "start\_offset" : 6,  
> > > "end\_offset" : 12,  
> > > "type" : "",  
> > > "position" : 2  
> > > }, {  
> > > "token" : "_gate_",  
> > > "start\_offset" : 13,  
> > > "end\_offset" : 18,  
> > > "type" : "",  
> > > "position" : 3  
> > > }, {  
> > > "token" : "_gate_",  
> > > "start\_offset" : 20,  
> > > "end\_offset" : 27,  
> > > "type" : "",  
> > > "position" : 4  
> > > } ]  
> > > }
> > > 
> > > And here are a subset of the settings and mappings. Note that I needed  
> > > to fully construct my own filter and analyzers in order to more fully  
> > > specify things such as the language to use.
> > > 
> > > {  
> > > "settings" : {  
> > > "index" : {  
> > > "number\_of\_shards" : 1,  
> > > "refresh\_interval" : "2s",  
> > > "number\_of\_replicas" : 0,  
> > > "analysis" : {
> > > 
> > > ```
> > > "filter" : {
> > > "english_snowball_filter" : {
> > > "type" : "snowball",
> > > "language" : "English"
> > > }
> > > },
> > > "analyzer" : {
> > > "english_stemming_analyzer" : {
> > > "type" : "custom",
> > > "tokenizer" : "standard",
> > > "filter" : [ "standard", "lowercase", "asciifolding",
> > > 
> > > ```
> > > 
> > > "english\_snowball\_filter" ]  
> > > },  
> > > "english\_standard\_analyzer" : {  
> > > "type" : "custom",  
> > > "tokenizer" : "standard",  
> > > "filter" : ["standard", "lowercase", "asciifolding"]  
> > > }  
> > > }  
> > > }  
> > > }  
> > > },  
> > > "mappings" : {  
> > > "person" : {  
> > > "\_all" : {  
> > > "enabled" : false  
> > > },  
> > > "properties" : {  
> > > "uid" : {  
> > > "type" : "long"  
> > > },  
> > > "cn" : {  
> > > "type" : "multi\_field",  
> > > "fields" : {  
> > > "cn" : {  
> > > "type" : "string",  
> > > "analyzer" : "english\_stemming\_analyzer"  
> > > },  
> > > "raw" : {  
> > > "type" : "string",  
> > > "index" : "not\_analyzed"  
> > > }  
> > > }  
> > > },  
> > > "location" : {  
> > > "type" : "geo\_point",  
> > > "lat\_lon" : true  
> > > },  
> > > "telno" : {  
> > > "type" : "multi\_field",  
> > > "fields" : {  
> > > "telno" : {  
> > > "type" : "string",  
> > > "analyzer" : "english\_standard\_analyzer"  
> > > },  
> > > "num" : {  
> > > "type" : "long"  
> > > }  
> > > }  
> > > },  
> > > "date" : {  
> > > "type" : "date",  
> > > "format" : "dateOptionalTime"  
> > > }  
> > > }  
> > > }  
> > > }  
> > > }
> > > 
> > > Hope this helps!
> > > 
> > > Regards,  
> > > Brian
> > > 
> > > On Monday, May 13, 2013 5:20:14 AM UTC-4, Jingang Wang wrote:
> > > 
> > > > Hi there,
> > > > 
> > > > I want to use ES to query some documents mentioning some person names.
> > > > 
> > > > For example, when I use Bill Gates to conduct a matchPhraseQuery, I can  
> > > > just get the documents exactly mention the name "Bill Gates".  
> > > > While a lot of documents mention Bill Gates indirectly, say, they may  
> > > > mention "Bill Gates's company".
> > > > 
> > > > How should I construct a query to retrieve these documents? Thanks.
> > > 
> > > --  
> > > You received this message because you are subscribed to a topic in the  
> > > Google Groups "elasticsearch" group.  
> > > To unsubscribe from this topic, visit  
> > > [https://groups.google.com/d/topic/elasticsearch/SIK3lc215Bk/unsubscribe?hl=en-US](https://groups.google.com/d/topic/elasticsearch/SIK3lc215Bk/unsubscribe?hl=en-US)  
> > > .  
> > > To unsubscribe from this group and all its topics, send an email to  
> > > [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).
> 
> --  
> Wang Jingang(王金刚)  
> Ph.D. Candidate at  
> Lab of High Volume Language Information Processing & Cloud Computing  
> School of Computer Science  
> Beijing Institute of Technology  
> Beijing 100081  
> P.R China

--  
Wang Jingang(王金刚)  
Ph.D. Candidate at  
Lab of High Volume Language Information Processing & Cloud Computing  
School of Computer Science  
Beijing Institute of Technology  
Beijing 100081  
P.R China

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![brian\_yoder](https://avatars.discourse-cdn.com/v4/letter/b/f1d935/32.png) [@brian\_yoder](https://discuss.elastic.co/u/brian_yoder)\
**Post date:** [May 14, 2013, 3:27pm UTC](https://discuss.elastic.co/t/matchphrasequery-can-not-retrieve-documents-with-trailing-s-even-if-set-word-delimiter-tokenfilter-when-created-indices/11934/7 "2013-05-14T15:27:26Z")

</div>

Hi Jingang,

The most important debugging tool is to use the \_analyze command I showed  
you.

For example, if you analyze "Foo Bar" and you get two tokens, "Foo" and  
"Bar", then you will immediately know that Foo and Bar can match that  
string but neither foo nor bar can.

The \_analyze function a very useful tool that helped me get through my  
query issues related to analyzers and mappings.

Also, it's an excellent idea to extract the \_mapping for the specified  
index and type. I found in a lot of my early efforts that I had what looked  
to me like a valid mapping, but with a few JSON mistakes Elasticsearch  
didn't recognize it. So it's not my eye that I trust to see if the mappings  
I intend are actually there; I always ask Elasticsearch what it thinks I  
gave it. Because in the end, ES's opinion is the only one that counts!

Regards,  
Brian

On Monday, May 13, 2013 10:53:31 PM UTC-4, Jingang Wang wrote:

> Dear  
> ​ Brian,
> 
> Thanks you so much for elaborate examples and explanation.  
> ​I have changed my settings as your example, while I still can not match  
> the phrase in the documents.
> 
> Here is my settings of index:  
> ​  
> {  
> "20120103": {  
> "settings": {  
> "index.analysis.filter.my\_delimiter.generate\_word\_parts": "true",  
> "index.analysis.filter.my\_delimiter.stem\_english\_possessive": "true",  
> "index.analysis.filter.my\_delimiter.preserve\_original": "true",  
> "index.analysis.analyzer.my\_analyzer.tokenizer": "standard",  
> "index.analysis.filter.english\_stemming\_filter.language": "English",  
> "index.analysis.filter.my\_delimiter.split\_on\_numerics": "true",  
> "index.analysis.filter.my\_delimiter.catenate\_all": "true",  
> "index.number\_of\_shards": "10",  
> "index.analysis.filter.my\_delimiter.catenate\_numbers": "true",  
> "index.analysis.analyzer.my\_analyzer.type": "custom",  
> "index.analysis.filter.english\_stemming\_filter.type": "snowball",  
> "index.analysis.filter.my\_delimiter.type": "word\_delimiter",  
> "index.analysis.filter.my\_delimiter.catenate\_words": "true",  
> "index.number\_of\_replicas": "0",  
> "index.analysis.analyzer.my\_analyzer.filter.2": "my\_delimiter",  
> "index.analysis.analyzer.my\_analyzer.filter.1": "lowercase",  
> "index.analysis.analyzer.my\_analyzer.filter.0": "standard",  
> "index.analysis.filter.my\_delimiter.split\_on\_case\_change": "true",  
> "index.analysis.analyzer.my\_analyzer.filter.5": "stop",  
> "index.analysis.analyzer.my\_analyzer.filter.4":  
> "english\_stemming\_filter",  
> "index.analysis.analyzer.my\_analyzer.filter.3": "asciifolding",  
> "index.version.created": "900001"  
> }  
> }  
> }
> 
> And my query is：
> 
> QueryBuilder qb = QueryBuilders  
> .boolQuery()  
> .must(matchPhraseQuery("body\_cleansed", "Aharon  
> Barak").analyzer("my\_analyzer"));
> 
> There is a documents which contains “Aharon Barak's policy”, but the query  
> could not retrieve it.
> 
> On Tue, May 14, 2013 at 6:59 AM, InquiringMind \<[brian....@gmail.com](mailto:brian....@gmail.com)\<javascript:\>
> 
> > wrote:
> 
> > And here's another example of analyzing the same string as below, but  
> > this time using the _cn.raw_ field. It's an example of multi-field  
> > mapping in which a field may be indexed (or not) in two or more ways but  
> > its source values only need to be stored once. Really awesome!!!
> > 
> > $ curl '[http://localhost:9200/sgen/\_analyze?field=\*cn.raw\*&pretty=true](http://localhost:9200/sgen/_analyze?field=*cn.raw*&pretty=true)'  
> > -d _"gates gate's gates' gates's"_ && echo  
> > {  
> > "tokens" : [ {  
> > "token" : _"gates gate's gates' gates's"_,  
> > "start\_offset" : 0,  
> > "end\_offset" : 27,  
> > "type" : "word",  
> > "position" : 1  
> > } ]  
> > }
> > 
> > On Monday, May 13, 2013 6:41:29 PM UTC-4, InquiringMind wrote:
> > 
> > > Hi Jingang,
> > > 
> > > Here is a full example with the index settings and mappings and a curl  
> > > command to show how various forms of Gates may be indexed so that they  
> > > match each other without any additional work on the part of your query:
> > > 
> > > Analyzing the string using the "cn" field in the "test" index:
> > > 
> > > $ curl '[http://localhost:9200/test/\_\*\*analyze?field=cn&pretty=true](http://localhost:9200/test/_**analyze?field=cn&pretty=true)[http://localhost:9200/test/\_analyze?field=cn&pretty=true](http://localhost:9200/test/_analyze?field=cn&pretty=true)'  
> > > -d "_gates gate's gates' gates's_" && echo  
> > > {  
> > > "tokens" : [ {  
> > > "token" : "_gate_",  
> > > "start\_offset" : 0,  
> > > "end\_offset" : 5,  
> > > "type" : "",  
> > > "position" : 1  
> > > }, {  
> > > "token" : "_gate_",  
> > > "start\_offset" : 6,  
> > > "end\_offset" : 12,  
> > > "type" : "",  
> > > "position" : 2  
> > > }, {  
> > > "token" : "_gate_",  
> > > "start\_offset" : 13,  
> > > "end\_offset" : 18,  
> > > "type" : "",  
> > > "position" : 3  
> > > }, {  
> > > "token" : "_gate_",  
> > > "start\_offset" : 20,  
> > > "end\_offset" : 27,  
> > > "type" : "",  
> > > "position" : 4  
> > > } ]  
> > > }
> > > 
> > > And here are a subset of the settings and mappings. Note that I needed  
> > > to fully construct my own filter and analyzers in order to more fully  
> > > specify things such as the language to use.
> > > 
> > > {  
> > > "settings" : {  
> > > "index" : {  
> > > "number\_of\_shards" : 1,  
> > > "refresh\_interval" : "2s",  
> > > "number\_of\_replicas" : 0,  
> > > "analysis" : {
> > > 
> > > ```
> > > "filter" : {
> > > "english_snowball_filter" : {
> > > "type" : "snowball",
> > > "language" : "English"
> > > }
> > > },
> > > "analyzer" : {
> > > "english_stemming_analyzer" : {
> > > "type" : "custom",
> > > "tokenizer" : "standard",
> > > "filter" : [ "standard", "lowercase", "asciifolding", 
> > > 
> > > ```
> > > 
> > > "english\_snowball\_filter" ]  
> > > },  
> > > "english\_standard\_analyzer" : {  
> > > "type" : "custom",  
> > > "tokenizer" : "standard",  
> > > "filter" : ["standard", "lowercase", "asciifolding"]  
> > > }  
> > > }  
> > > }  
> > > }  
> > > },  
> > > "mappings" : {  
> > > "person" : {  
> > > "\_all" : {  
> > > "enabled" : false  
> > > },  
> > > "properties" : {  
> > > "uid" : {  
> > > "type" : "long"  
> > > },  
> > > "cn" : {  
> > > "type" : "multi\_field",  
> > > "fields" : {  
> > > "cn" : {  
> > > "type" : "string",  
> > > "analyzer" : "english\_stemming\_analyzer"  
> > > },  
> > > "raw" : {  
> > > "type" : "string",  
> > > "index" : "not\_analyzed"  
> > > }  
> > > }  
> > > },  
> > > "location" : {  
> > > "type" : "geo\_point",  
> > > "lat\_lon" : true  
> > > },  
> > > "telno" : {  
> > > "type" : "multi\_field",  
> > > "fields" : {  
> > > "telno" : {  
> > > "type" : "string",  
> > > "analyzer" : "english\_standard\_analyzer"  
> > > },  
> > > "num" : {  
> > > "type" : "long"  
> > > }  
> > > }  
> > > },  
> > > "date" : {  
> > > "type" : "date",  
> > > "format" : "dateOptionalTime"  
> > > }  
> > > }  
> > > }  
> > > }  
> > > }
> > > 
> > > Hope this helps!
> > > 
> > > Regards,  
> > > Brian
> > > 
> > > On Monday, May 13, 2013 5:20:14 AM UTC-4, Jingang Wang wrote:
> > > 
> > > > Hi there,
> > > > 
> > > > I want to use ES to query some documents mentioning some person names.
> > > > 
> > > > For example, when I use Bill Gates to conduct a matchPhraseQuery, I can  
> > > > just get the documents exactly mention the name "Bill Gates".  
> > > > While a lot of documents mention Bill Gates indirectly, say, they may  
> > > > mention "Bill Gates's company".
> > > > 
> > > > How should I construct a query to retrieve these documents? Thanks.
> > > 
> > > --  
> > > You received this message because you are subscribed to a topic in the  
> > > Google Groups "elasticsearch" group.  
> > > To unsubscribe from this topic, visit  
> > > [https://groups.google.com/d/topic/elasticsearch/SIK3lc215Bk/unsubscribe?hl=en-US](https://groups.google.com/d/topic/elasticsearch/SIK3lc215Bk/unsubscribe?hl=en-US)  
> > > .  
> > > To unsubscribe from this group and all its topics, send an email to  
> > > [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).
> 
> --  
> Wang Jingang(王金刚)  
> Ph.D. Candidate at  
> Lab of High Volume Language Information Processing & Cloud Computing  
> School of Computer Science  
> Beijing Institute of Technology  
> Beijing 100081  
> P.R China

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Jingang\_Wang](https://avatars.discourse-cdn.com/v4/letter/j/ce7236/32.png) [@Jingang\_Wang](https://discuss.elastic.co/u/Jingang_Wang)\
**Post date:** [May 15, 2013, 2:07am UTC](https://discuss.elastic.co/t/matchphrasequery-can-not-retrieve-documents-with-trailing-s-even-if-set-word-delimiter-tokenfilter-when-created-indices/11934/8 "2013-05-15T02:07:49Z")

</div>

Hi Brain,

Your suggestions are so much useful to me.

I have resolved this problem via using \_mapping function to check the  
mapping information of specific index and type.

The problem was resulted from unsuccessful mapping operation.

​Thanks again for your help.​  
​

On Tue, May 14, 2013 at 11:27 PM, InquiringMind [brian.from.fl@gmail.com](mailto:brian.from.fl@gmail.com)wrote:

> Hi Jingang,
> 
> The most important debugging tool is to use the \_analyze command I showed  
> you.
> 
> For example, if you analyze "Foo Bar" and you get two tokens, "Foo" and  
> "Bar", then you will immediately know that Foo and Bar can match that  
> string but neither foo nor bar can.
> 
> The \_analyze function a very useful tool that helped me get through my  
> query issues related to analyzers and mappings.
> 
> Also, it's an excellent idea to extract the \_mapping for the specified  
> index and type. I found in a lot of my early efforts that I had what looked  
> to me like a valid mapping, but with a few JSON mistakes Elasticsearch  
> didn't recognize it. So it's not my eye that I trust to see if the mappings  
> I intend are actually there; I always ask Elasticsearch what it thinks I  
> gave it. Because in the end, ES's opinion is the only one that counts!
> 
> Regards,  
> Brian
> 
> On Monday, May 13, 2013 10:53:31 PM UTC-4, Jingang Wang wrote:
> 
> > Dear  
> > ​ Brian,
> > 
> > Thanks you so much for elaborate examples and explanation.  
> > ​I have changed my settings as your example, while I still can not match  
> > the phrase in the documents.
> > 
> > Here is my settings of index:  
> > ​  
> > {  
> > "20120103": {  
> > "settings": {  
> > "index.analysis.filter.my\_ **delimiter.generate\_word\_parts"** :  
> > "true",  
> > "index.analysis.filter.my\_\*\*delimiter.stem\_english\_\*\*possessive":  
> > "true",  
> > "index.analysis.filter.my\_\*\*delimiter.preserve\_original": "true",  
> > "index.analysis.analyzer.my\_\*\*analyzer.tokenizer": "standard",  
> > "index.analysis.filter.\*\*english\_stemming\_filter.\*\*language":  
> > "English",  
> > "index.analysis.filter.my\_\*\*delimiter.split\_on\_numerics": "true",  
> > "index.analysis.filter.my\_\*\*delimiter.catenate\_all": "true",  
> > "index.number\_of\_shards": "10",  
> > "index.analysis.filter.my\_\*\*delimiter.catenate\_numbers": "true",  
> > "index.analysis.analyzer.my\_\*\*analyzer.type": "custom",  
> > "index.analysis.filter.\*\*english\_stemming\_filter.type": "snowball",  
> > "index.analysis.filter.my\_\*\*delimiter.type": "word\_delimiter",  
> > "index.analysis.filter.my\_\*\*delimiter.catenate\_words": "true",  
> > "index.number\_of\_replicas": "0",  
> > "index.analysis.analyzer.my\_\*\*analyzer.filter.2": "my\_delimiter",  
> > "index.analysis.analyzer.my\_\*\*analyzer.filter.1": "lowercase",  
> > "index.analysis.analyzer.my\_\*\*analyzer.filter.0": "standard",  
> > "index.analysis.filter.my\_\*\*delimiter.split\_on\_case\_\*\*change":  
> > "true",  
> > "index.analysis.analyzer.my\_\*\*analyzer.filter.5": "stop",  
> > "index.analysis.analyzer.my\_\*\*analyzer.filter.4":  
> > "english\_stemming\_filter",  
> > "index.analysis.analyzer.my\_\*\*analyzer.filter.3": "asciifolding",  
> > "index.version.created": "900001"  
> > }  
> > }  
> > }
> > 
> > And my query is：
> > 
> > QueryBuilder qb = QueryBuilders  
> > .boolQuery()  
> > .must(matchPhraseQuery("**body\_cleansed", "Aharon  
> > Barak").analyzer("my\_analyzer"**));
> > 
> > There is a documents which contains “Aharon Barak's policy”, but the  
> > query could not retrieve it.
> > 
> > On Tue, May 14, 2013 at 6:59 AM, InquiringMind [brian....@gmail.com](mailto:brian....@gmail.com)wrote:
> > 
> > > And here's another example of analyzing the same string as below, but  
> > > this time using the _cn.raw_ field. It's an example of multi-field  
> > > mapping in which a field may be indexed (or not) in two or more ways but  
> > > its source values only need to be stored once. Really awesome!!!
> > > 
> > > $ curl '[http://localhost:9200/sgen/\_\*\*analyze?field=](http://localhost:9200/sgen/_**analyze?field=)[http://localhost:9200/sgen/\_analyze?field=](http://localhost:9200/sgen/_analyze?field=)  
> > > _cn.raw_&pretty=\*\*true' -d _"gates gate's gates' gates's"_ && echo  
> > > {  
> > > "tokens" : [ {  
> > > "token" : _"gates gate's gates' gates's"_,  
> > > "start\_offset" : 0,  
> > > "end\_offset" : 27,  
> > > "type" : "word",  
> > > "position" : 1  
> > > } ]  
> > > }
> > > 
> > > On Monday, May 13, 2013 6:41:29 PM UTC-4, InquiringMind wrote:
> > > 
> > > > Hi Jingang,
> > > > 
> > > > Here is a full example with the index settings and mappings and a curl  
> > > > command to show how various forms of Gates may be indexed so that they  
> > > > match each other without any additional work on the part of your query:
> > > > 
> > > > Analyzing the string using the "cn" field in the "test" index:
> > > > 
> > > > $ curl '[http://localhost:9200/test/\_\*\*a\*\*nalyze?field=cn&pretty=true](http://localhost:9200/test/_ **a** nalyze?field=cn&pretty=true)[http://localhost:9200/test/\_analyze?field=cn&pretty=true](http://localhost:9200/test/_analyze?field=cn&pretty=true)'  
> > > > -d "_gates gate's gates' gates's_" && echo  
> > > > {  
> > > > "tokens" : [ {  
> > > > "token" : "_gate_",  
> > > > "start\_offset" : 0,  
> > > > "end\_offset" : 5,  
> > > > "type" : "",  
> > > > "position" : 1  
> > > > }, {  
> > > > "token" : "_gate_",  
> > > > "start\_offset" : 6,  
> > > > "end\_offset" : 12,  
> > > > "type" : "",  
> > > > "position" : 2  
> > > > }, {  
> > > > "token" : "_gate_",  
> > > > "start\_offset" : 13,  
> > > > "end\_offset" : 18,  
> > > > "type" : "",  
> > > > "position" : 3  
> > > > }, {  
> > > > "token" : "_gate_",  
> > > > "start\_offset" : 20,  
> > > > "end\_offset" : 27,  
> > > > "type" : "",  
> > > > "position" : 4  
> > > > } ]  
> > > > }
> > > > 
> > > > And here are a subset of the settings and mappings. Note that I needed  
> > > > to fully construct my own filter and analyzers in order to more fully  
> > > > specify things such as the language to use.
> > > > 
> > > > {  
> > > > "settings" : {  
> > > > "index" : {  
> > > > "number\_of\_shards" : 1,  
> > > > "refresh\_interval" : "2s",  
> > > > "number\_of\_replicas" : 0,  
> > > > "analysis" : {
> > > > 
> > > > ```
> > > > "filter" : {
> > > > "english_snowball_filter" : {
> > > > "type" : "snowball",
> > > > "language" : "English"
> > > > }
> > > > },
> > > > "analyzer" : {
> > > > "english_stemming_analyzer" : {
> > > > "type" : "custom",
> > > > "tokenizer" : "standard",
> > > > "filter" : [ "standard", "lowercase", "asciifolding",
> > > > 
> > > > ```
> > > > 
> > > > "english\_snowball\_filter" ]  
> > > > },  
> > > > "english\_standard\_analyzer" : {  
> > > > "type" : "custom",  
> > > > "tokenizer" : "standard",  
> > > > "filter" : ["standard", "lowercase", "asciifolding"]  
> > > > }  
> > > > }  
> > > > }  
> > > > }  
> > > > },  
> > > > "mappings" : {  
> > > > "person" : {  
> > > > "\_all" : {  
> > > > "enabled" : false  
> > > > },  
> > > > "properties" : {  
> > > > "uid" : {  
> > > > "type" : "long"  
> > > > },  
> > > > "cn" : {  
> > > > "type" : "multi\_field",  
> > > > "fields" : {  
> > > > "cn" : {  
> > > > "type" : "string",  
> > > > "analyzer" : "english\_stemming\_analyzer"  
> > > > },  
> > > > "raw" : {  
> > > > "type" : "string",  
> > > > "index" : "not\_analyzed"  
> > > > }  
> > > > }  
> > > > },  
> > > > "location" : {  
> > > > "type" : "geo\_point",  
> > > > "lat\_lon" : true  
> > > > },  
> > > > "telno" : {  
> > > > "type" : "multi\_field",  
> > > > "fields" : {  
> > > > "telno" : {  
> > > > "type" : "string",  
> > > > "analyzer" : "english\_standard\_analyzer"  
> > > > },  
> > > > "num" : {  
> > > > "type" : "long"  
> > > > }  
> > > > }  
> > > > },  
> > > > "date" : {  
> > > > "type" : "date",  
> > > > "format" : "dateOptionalTime"  
> > > > }  
> > > > }  
> > > > }  
> > > > }  
> > > > }
> > > > 
> > > > Hope this helps!
> > > > 
> > > > Regards,  
> > > > Brian
> > > > 
> > > > On Monday, May 13, 2013 5:20:14 AM UTC-4, Jingang Wang wrote:
> > > > 
> > > > > Hi there,
> > > > > 
> > > > > I want to use ES to query some documents mentioning some person names.
> > > > > 
> > > > > For example, when I use Bill Gates to conduct a matchPhraseQuery, I  
> > > > > can just get the documents exactly mention the name "Bill Gates".  
> > > > > While a lot of documents mention Bill Gates indirectly, say, they may  
> > > > > mention "Bill Gates's company".
> > > > > 
> > > > > How should I construct a query to retrieve these documents? Thanks.
> > > > 
> > > > --  
> > > > You received this message because you are subscribed to a topic in the  
> > > > Google Groups "elasticsearch" group.  
> > > > To unsubscribe from this topic, visit [https://groups.google.com/d/](https://groups.google.com/d/)\*\*  
> > > > topic/elasticsearch/\*\*SIK3lc215Bk/unsubscribe?hl=en-\*\*US[https://groups.google.com/d/topic/elasticsearch/SIK3lc215Bk/unsubscribe?hl=en-US](https://groups.google.com/d/topic/elasticsearch/SIK3lc215Bk/unsubscribe?hl=en-US)  
> > > > .  
> > > > To unsubscribe from this group and all its topics, send an email to  
> > > > elasticsearc...@\*\*[googlegroups.com](http://googlegroups.com).
> > > 
> > > For more options, visit [https://groups.google.com/\*\*groups/opt\_out](https://groups.google.com/**groups/opt_out)[https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)  
> > > .
> > 
> > --  
> > Wang Jingang(王金刚)  
> > Ph.D. Candidate at  
> > Lab of High Volume Language Information Processing & Cloud Computing  
> > School of Computer Science  
> > Beijing Institute of Technology  
> > Beijing 100081  
> > P.R China
> > 
> > --  
> > You received this message because you are subscribed to a topic in the  
> > Google Groups "elasticsearch" group.  
> > To unsubscribe from this topic, visit  
> > [https://groups.google.com/d/topic/elasticsearch/SIK3lc215Bk/unsubscribe?hl=en-US](https://groups.google.com/d/topic/elasticsearch/SIK3lc215Bk/unsubscribe?hl=en-US)  
> > .  
> > To unsubscribe from this group and all its topics, send an email to  
> > [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
Wang Jingang(王金刚)  
Ph.D. Candidate at  
Lab of High Volume Language Information Processing & Cloud Computing  
School of Computer Science  
Beijing Institute of Technology  
Beijing 100081  
P.R China

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:36am UTC](https://discuss.elastic.co/t/matchphrasequery-can-not-retrieve-documents-with-trailing-s-even-if-set-word-delimiter-tokenfilter-when-created-indices/11934/9 "2017-07-06T02:36:43Z")

</div>


