# Excluding punctuation from fields

**URL:** <https://discuss.elastic.co/t/excluding-punctuation-from-fields/7772>\
**Category:** Elasticsearch\
**Created:** [May 19, 2012, 3:14pm UTC](https://discuss.elastic.co/t/excluding-punctuation-from-fields/7772 "2012-05-19T15:14:05Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Michael\_Sick](https://avatars.discourse-cdn.com/v4/letter/m/22d042/32.png) [@Michael\_Sick](https://discuss.elastic.co/u/Michael_Sick)\
**Post date:** [May 19, 2012, 3:14pm UTC](https://discuss.elastic.co/t/excluding-punctuation-from-fields/7772/1 "2012-05-19T15:14:05Z")

</div>

Hi All,

What's the best way (or tradeoffs) to exclude punctuation (or specific  
characters) from certain fields during analysis and searching?

i.e.  
In document, "P.F. Changs" would match a search for "P.F. Changs" or "PF  
Changs".

It looks like I could do this with the Synonym filter and the ICU plugin.  
Are there better options? Which is best or what are the tradeoffs?

Overall, if anyone knows of any resources that compare/contrast the various  
analyzers/filters/..., it would be very helpful.

Thanks for any advice/pointers,

--Mike

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [May 20, 2012, 8:35pm UTC](https://discuss.elastic.co/t/excluding-punctuation-from-fields/7772/2 "2012-05-20T20:35:55Z")

</div>

The standard filter should remove punctuation from tokens.

You can use the analysis API to view the differences between analyzers  
(and unfortunately not between tokenizers or filters). The Lucene in  
Action book has a summary of the different classes.

--  
Ivan

On Sat, May 19, 2012 at 8:14 AM, Michael Sick  
[michael.sick@serenesoftware.com](mailto:michael.sick@serenesoftware.com) wrote:

> Hi All,
> 
> What's the best way (or tradeoffs) to exclude punctuation (or specific  
> characters) from certain fields during analysis and searching?
> 
> i.e.  
> In document, "P.F. Changs" would match a search for "P.F. Changs" or "PF  
> Changs".
> 
> It looks like I could do this with the Synonym filter and the ICU plugin.  
> Are there better options? Which is best or what are the tradeoffs?
> 
> Overall, if anyone knows of any resources that compare/contrast the various  
> analyzers/filters/..., it would be very helpful.
> 
> Thanks for any advice/pointers,
> 
> --Mike

---

<div class="post-metadata">

**Author:** ![Michael\_Sick](https://avatars.discourse-cdn.com/v4/letter/m/22d042/32.png) [@Michael\_Sick](https://discuss.elastic.co/u/Michael_Sick)\
**Post date:** [May 22, 2012, 3:27am UTC](https://discuss.elastic.co/t/excluding-punctuation-from-fields/7772/3 "2012-05-22T03:27:48Z")

</div>

Ivan,

Thanks! The Analysis API is priceless. Thanks,

--Mike

On Sun, May 20, 2012 at 4:35 PM, Ivan Brusic [ivan@brusic.com](mailto:ivan@brusic.com) wrote:

> The standard filter should remove punctuation from tokens.
> 
> You can use the analysis API to view the differences between analyzers  
> (and unfortunately not between tokenizers or filters). The Lucene in  
> Action book has a summary of the different classes.
> 
> --  
> Ivan
> 
> On Sat, May 19, 2012 at 8:14 AM, Michael Sick  
> [michael.sick@serenesoftware.com](mailto:michael.sick@serenesoftware.com) wrote:
> 
> > Hi All,
> > 
> > What's the best way (or tradeoffs) to exclude punctuation (or specific  
> > characters) from certain fields during analysis and searching?
> > 
> > i.e.  
> > In document, "P.F. Changs" would match a search for "P.F. Changs" or "PF  
> > Changs".
> > 
> > It looks like I could do this with the Synonym filter and the ICU plugin.  
> > Are there better options? Which is best or what are the tradeoffs?
> > 
> > Overall, if anyone knows of any resources that compare/contrast the  
> > various  
> > analyzers/filters/..., it would be very helpful.
> > 
> > Thanks for any advice/pointers,
> > 
> > --Mike

---

<div class="post-metadata">

**Author:** ![Michael\_Sick](https://avatars.discourse-cdn.com/v4/letter/m/22d042/32.png) [@Michael\_Sick](https://discuss.elastic.co/u/Michael_Sick)\
**Post date:** [May 26, 2012, 6:41am UTC](https://discuss.elastic.co/t/excluding-punctuation-from-fields/7772/4 "2012-05-26T06:41:38Z")

</div>

I'm still having no luck on this. I've created a more self contained  
example for the behavior. In short, I'm storing a document with a field  
containing:

"P.F. Changs Burgers"

Create & Run Test: [Index / Search Document with Punctuation in ElasticSearch · GitHub](https://gist.github.com/2792582)  
Delete Artifacts: [Delete Example Alias, Index, Template · GitHub](https://gist.github.com/2792590)

I'd like ES to provide a match if I search on "p.f.", "p.f", "pf.", "pf"  
with regard to case. Currently only the 1st two work. My approach relies on  
using the Synonym filter for translating all forms above to "pf". I'd be  
happy to fix this approach or, even better, to learn that there's a general  
approach that will not require as much configuration.

Thanks! --Mike

On Sun, May 20, 2012 at 4:35 PM, Ivan Brusic [ivan@brusic.com](mailto:ivan@brusic.com) wrote:

> The standard filter should remove punctuation from tokens.
> 
> You can use the analysis API to view the differences between analyzers  
> (and unfortunately not between tokenizers or filters). The Lucene in  
> Action book has a summary of the different classes.
> 
> --  
> Ivan
> 
> On Sat, May 19, 2012 at 8:14 AM, Michael Sick  
> [michael.sick@serenesoftware.com](mailto:michael.sick@serenesoftware.com) wrote:
> 
> > Hi All,
> > 
> > What's the best way (or tradeoffs) to exclude punctuation (or specific  
> > characters) from certain fields during analysis and searching?
> > 
> > i.e.  
> > In document, "P.F. Changs" would match a search for "P.F. Changs" or "PF  
> > Changs".
> > 
> > It looks like I could do this with the Synonym filter and the ICU plugin.  
> > Are there better options? Which is best or what are the tradeoffs?
> > 
> > Overall, if anyone knows of any resources that compare/contrast the  
> > various  
> > analyzers/filters/..., it would be very helpful.
> > 
> > Thanks for any advice/pointers,
> > 
> > --Mike

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [May 30, 2012, 6:21pm UTC](https://discuss.elastic.co/t/excluding-punctuation-from-fields/7772/5 "2012-05-30T18:21:24Z")

</div>

I just occurred to me as I was testing things that the standard filter  
does NOT remove punctuation. My custom filter in Lucene was stripping  
punctuation, not the standard filter.

I was able to remove punctuation by using a mapping char\_filter. Mine  
simply removes dots '.'

index :  
analysis :  
analyzer :   
unstemmed :  
type : custom  
filter : [unique , standard, asciifolding, lowercase,]  
char\_filter : [punctuation]  
char\_filter :  
punctuation :  
type: mapping  
mappings: [".=\>"]

On Fri, May 25, 2012 at 11:41 PM, Michael Sick  
[michael.sick@serenesoftware.com](mailto:michael.sick@serenesoftware.com) wrote:

> I'm still having no luck on this. I've created a more self contained example  
> for the behavior. In short, I'm storing a document with a field containing:
> 
> "P.F. Changs Burgers"
> 
> Create & Run Test: [Index / Search Document with Punctuation in ElasticSearch · GitHub](https://gist.github.com/2792582)  
> Delete Artifacts: [Delete Example Alias, Index, Template · GitHub](https://gist.github.com/2792590)
> 
> I'd like ES to provide a match if I search on "p.f.", "p.f", "pf.", "pf"  
> with regard to case. Currently only the 1st two work. My approach relies on  
> using the Synonym filter for translating all forms above to "pf". I'd be  
> happy to fix this approach or, even better, to learn that there's a general  
> approach that will not require as much configuration.
> 
> Thanks! --Mike
> 
> On Sun, May 20, 2012 at 4:35 PM, Ivan Brusic [ivan@brusic.com](mailto:ivan@brusic.com) wrote:
> 
> > The standard filter should remove punctuation from tokens.
> > 
> > You can use the analysis API to view the differences between analyzers  
> > (and unfortunately not between tokenizers or filters). The Lucene in  
> > Action book has a summary of the different classes.
> > 
> > --  
> > Ivan
> > 
> > On Sat, May 19, 2012 at 8:14 AM, Michael Sick  
> > [michael.sick@serenesoftware.com](mailto:michael.sick@serenesoftware.com) wrote:
> > 
> > > Hi All,
> > > 
> > > What's the best way (or tradeoffs) to exclude punctuation (or specific  
> > > characters) from certain fields during analysis and searching?
> > > 
> > > i.e.  
> > > In document, "P.F. Changs" would match a search for "P.F. Changs" or "PF  
> > > Changs".
> > > 
> > > It looks like I could do this with the Synonym filter and the ICU  
> > > plugin.  
> > > Are there better options? Which is best or what are the tradeoffs?
> > > 
> > > Overall, if anyone knows of any resources that compare/contrast the  
> > > various  
> > > analyzers/filters/..., it would be very helpful.
> > > 
> > > Thanks for any advice/pointers,
> > > 
> > > --Mike

---

<div class="post-metadata">

**Author:** ![Michael\_Sick](https://avatars.discourse-cdn.com/v4/letter/m/22d042/32.png) [@Michael\_Sick](https://discuss.elastic.co/u/Michael_Sick)\
**Post date:** [June 1, 2012, 4:44pm UTC](https://discuss.elastic.co/t/excluding-punctuation-from-fields/7772/6 "2012-06-01T16:44:15Z")

</div>

Thanks Ivan - I'll give that a shot.

On Wed, May 30, 2012 at 2:21 PM, Ivan Brusic [ivan@brusic.com](mailto:ivan@brusic.com) wrote:

> I just occurred to me as I was testing things that the standard filter  
> does NOT remove punctuation. My custom filter in Lucene was stripping  
> punctuation, not the standard filter.
> 
> I was able to remove punctuation by using a mapping char\_filter. Mine  
> simply removes dots '.'
> 
> index :  
> analysis :  
> analyzer :  
> unstemmed :  
> type : custom  
> filter : [unique , standard, asciifolding, lowercase,]  
> char\_filter : [punctuation]  
> char\_filter :  
> punctuation :  
> type: mapping  
> mappings: [".=\>"]
> 
> On Fri, May 25, 2012 at 11:41 PM, Michael Sick  
> [michael.sick@serenesoftware.com](mailto:michael.sick@serenesoftware.com) wrote:
> 
> > I'm still having no luck on this. I've created a more self contained  
> > example  
> > for the behavior. In short, I'm storing a document with a field  
> > containing:
> > 
> > "P.F. Changs Burgers"
> > 
> > Create & Run Test: [Index / Search Document with Punctuation in ElasticSearch · GitHub](https://gist.github.com/2792582)  
> > Delete Artifacts: [Delete Example Alias, Index, Template · GitHub](https://gist.github.com/2792590)
> > 
> > I'd like ES to provide a match if I search on "p.f.", "p.f", "pf.", "pf"  
> > with regard to case. Currently only the 1st two work. My approach relies  
> > on  
> > using the Synonym filter for translating all forms above to "pf". I'd be  
> > happy to fix this approach or, even better, to learn that there's a  
> > general  
> > approach that will not require as much configuration.
> > 
> > Thanks! --Mike
> > 
> > On Sun, May 20, 2012 at 4:35 PM, Ivan Brusic [ivan@brusic.com](mailto:ivan@brusic.com) wrote:
> > 
> > > The standard filter should remove punctuation from tokens.
> > > 
> > > You can use the analysis API to view the differences between analyzers  
> > > (and unfortunately not between tokenizers or filters). The Lucene in  
> > > Action book has a summary of the different classes.
> > > 
> > > --  
> > > Ivan
> > > 
> > > On Sat, May 19, 2012 at 8:14 AM, Michael Sick  
> > > [michael.sick@serenesoftware.com](mailto:michael.sick@serenesoftware.com) wrote:
> > > 
> > > > Hi All,
> > > > 
> > > > What's the best way (or tradeoffs) to exclude punctuation (or specific  
> > > > characters) from certain fields during analysis and searching?
> > > > 
> > > > i.e.  
> > > > In document, "P.F. Changs" would match a search for "P.F. Changs" or  
> > > > "PF  
> > > > Changs".
> > > > 
> > > > It looks like I could do this with the Synonym filter and the ICU  
> > > > plugin.  
> > > > Are there better options? Which is best or what are the tradeoffs?
> > > > 
> > > > Overall, if anyone knows of any resources that compare/contrast the  
> > > > various  
> > > > analyzers/filters/..., it would be very helpful.
> > > > 
> > > > Thanks for any advice/pointers,
> > > > 
> > > > --Mike

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:26am UTC](https://discuss.elastic.co/t/excluding-punctuation-from-fields/7772/7 "2017-07-06T03:26:00Z")

</div>


