# Filtering for apostrophes and single quotes confusion?

**URL:** <https://discuss.elastic.co/t/filtering-for-apostrophes-and-single-quotes-confusion/12102>\
**Category:** Elasticsearch\
**Created:** [May 23, 2013, 6:29pm UTC](https://discuss.elastic.co/t/filtering-for-apostrophes-and-single-quotes-confusion/12102 "2013-05-23T18:29:53Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![phill](https://avatars.discourse-cdn.com/v4/letter/p/779978/32.png) [@phill](https://discuss.elastic.co/u/phill)\
**Post date:** [May 23, 2013, 6:29pm UTC](https://discuss.elastic.co/t/filtering-for-apostrophes-and-single-quotes-confusion/12102/1 "2013-05-23T18:29:53Z")

</div>

Recently a customer of ours confused himself and us when he discovered  
his corpus of documents contains (at least) two different characters  
used for the apostrophe.  
I didn't recognize the issue at first.

For those who aren't familiar with this problem, depending on the  
history of characters in a document an apostrophe used as a possessive  
in English, i.e  
"customer's report" (one report from one customer) may be one of  
several characters and apparently the 3.x StandardAnalyzer doesn't take  
this into consideration.  
Did I miss a mention of a fix for this? I was conducting tests against  
Lucene 3.4, but didn't see mention in ES either (I'm not running 4.x yet).

While I found some discussion of this issue over the years, I was  
surprised to not find any general solution either in Standard Analyzer  
or in some extra Filter that I might leverage in a filter chain. Am I  
missing something?

I also have to say that not very large test document sets gathered as  
ordinary domain examples from the web have now been shown include  
different apostrophe characters.

My suggested solution matches the one line from the Snowball page (see  
below) "Clearly other codes for apostrophe can be mapped to this  
[apostrophe] code prior to stemming." I'm sure I DO NOT want to mess  
with things too much. I already don't use standard analyzer (so to not  
bother dropping stopwords), so I was thinking a simple filter chainable  
before standard filter that looks for odd Apostrophes followed by s and  
replaces the odd char with U+0027  
[http://www.fileformat.info/info/unicode/char/0027/index.htm](http://www.fileformat.info/info/unicode/char/0027/index.htm) would do  
the trick.

Any thoughts or help?

-Paul

* * *

All the background information I have on the topic.

Smart Editors (MS Word and I believe Adobe Acrobat) convert the ordinary  
single quote/apostrophe key (on the modern English MS keyboard that is  
(the un-shifted key on double quote and single-quote (?) key) to various  
other characters. Meanwhile, neither browser web page entry boxes (by  
default) nor simpler text editors mess with characters typed usually  
resulting in an APOSTROPHE.  
Paul's example of an apostrophe and a 'full quote' generated by the  
Outlook editor.

Paul's example of an apostrophe and a 'full quote' typed into a browser  
field.

Assuming my e-mail editor, my e-mailer, the list mailer, your e-mail  
program and your viewer all preserved the characters along the way, the  
1st line uses 3 different characters the 2nd uses 1. I only typed one  
character in all cases.

Various character that might show up include the following:  
U+0027  
[http://www.fileformat.info/info/unicode/char/0027/index.htm](http://www.fileformat.info/info/unicode/char/0027/index.htm)APOSTROPHE  
Original ASCII character, _probably_ what your keyboard sends, but I  
can't promise anything.  
U+0091 [http://www.fileformat.info/info/unicode/char/0091/index.htm](http://www.fileformat.info/info/unicode/char/0091/index.htm)  
Left single quotation mark ASCII ISO 8859-1 ISO Latin 1(Note 1), but is  
listed as PRIVATE USE ONLY in official Unicode.  
U+0092 [http://www.fileformat.info/info/unicode/char/0092/index.htm](http://www.fileformat.info/info/unicode/char/0092/index.htm)  
Right single quotation mark ASCII ISO 8859-1 ISO Latin 1(Note 1), but is  
listed as PRIVATE USE ONLY in official Unicode.  
U+2018 [http://www.fileformat.info/info/unicode/char/2018/index.htm](http://www.fileformat.info/info/unicode/char/2018/index.htm)  
LEFT SINGLE QUOTATION MARK The official Unicode character. This is what  
I get from the above example generated in 2013.  
U+2019  
[http://www.fileformat.info/info/unicode/char/2019/index.htm](http://www.fileformat.info/info/unicode/char/2019/index.htm)RIGHT  
SINGLE QUOTATION MARK The official Unicode character. This is what I  
get from the above example generated in 2013.  
U+2019 [http://www.fileformat.info/info/unicode/char/201B/index.htm](http://www.fileformat.info/info/unicode/char/201B/index.htm)  
SINGLE HIGH-REVERSED-9 QUOTATION MARK Mentioned as a special case use  
at [Tartarus.org](http://Tartarus.org) in other contexts, eg. O'Reilly (see link below).

Standards are just crazy things in the real world since they are never  
followed fully.  
Typing a single quote from the keyboard into the website  
[http://www.babelstone.co.uk/unicode/whatisit.html](http://www.babelstone.co.uk/unicode/whatisit.html)  
using either Firefox or IE reports back that it got U+0027 - the old  
fashion apostrophe, but Unicode at the page for U+2019  
[http://www.fileformat.info/info/unicode/char/2019/index.htm](http://www.fileformat.info/info/unicode/char/2019/index.htm) says  
[U+2019] "is the preferred character to use for apostrophe".

The Snowball parser folks spotted the problem and summarized it at:  
[http://snowball.tartarus.org/texts/apostrophe.html](http://snowball.tartarus.org/texts/apostrophe.html)  
But I didn't see any Filters there either, but maybe I didn't search  
well enough, but then maybe I used the wrong apostrophe when searching.

-Paul

(1) ISO 8859-1 ISO Latin 1 [http://www.ascii-code.com/](http://www.ascii-code.com/)

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [May 24, 2013, 10:22am UTC](https://discuss.elastic.co/t/filtering-for-apostrophes-and-single-quotes-confusion/12102/2 "2013-05-24T10:22:51Z")

</div>

Hi Paul

Interesting...

From this:

> For these reasons, the English stemmer treats apostrophe as a letter,  
> removing it from the beginning of a word, where it might have stood for an  
> opening quote, from the end of the word, where it might have stood for a  
> closing quote, or been an apostrophe following _s_. The form _’s_ is also  
> treated as an ending.

... it sounds like things will work correctly as long as you normalize all  
single quotes/apostrophes to the same character, which you can do with a  
char filter:

curl -XPUT '[http://127.0.0.1:9200/test/?pretty=1](http://127.0.0.1:9200/test/?pretty=1)' -d '  
{  
"settings" : {  
"analysis" : {  
"analyzer" : {  
"quotes" : {  
"filter" : [  
"standard",  
"lowercase"  
],  
"char\_filter" : [  
"quotes"  
],  
"tokenizer" : "standard"  
}  
},  
"char\_filter" : {  
"quotes" : {  
"mappings" : [  
"\u0091=\>\u0027",  
"\u0092=\>\u0027",  
"\u2018=\>\u0027",  
"\u2019=\>\u0027"  
],  
"type" : "mapping"  
}  
}  
}  
}  
}  
'

curl -XGET '[http://127.0.0.1:9200/test/\_analyze?pretty&analyzer=quotes](http://127.0.0.1:9200/test/_analyze?pretty&analyzer=quotes)' -d '  
Paul’s example of an apostrophe and a ‘full quote’ generated by the Outlook  
editor.  
'

# {

# "tokens" : [

# {

# "end\_offset" : 6,

# "position" : 1,

# "start\_offset" : 0,

# "type" : "",

# "token" : "paul's"

# },

# {

# "end\_offset" : 14,

# "position" : 2,

# "start\_offset" : 7,

# "type" : "",

# "token" : "example"

# },

# {

# "end\_offset" : 17,

# "position" : 3,

# "start\_offset" : 15,

# "type" : "",

# "token" : "of"

# },

# {

# "end\_offset" : 20,

# "position" : 4,

# "start\_offset" : 18,

# "type" : "",

# "token" : "an"

# },

# {

# "end\_offset" : 31,

# "position" : 5,

# "start\_offset" : 21,

# "type" : "",

# "token" : "apostrophe"

# },

# {

# "end\_offset" : 35,

# "position" : 6,

# "start\_offset" : 32,

# "type" : "",

# "token" : "and"

# },

# {

# "end\_offset" : 37,

# "position" : 7,

# "start\_offset" : 36,

# "type" : "",

# "token" : "a"

# },

# {

# "end\_offset" : 43,

# "position" : 8,

# "start\_offset" : 39,

# "type" : "",

# "token" : "full"

# },

# {

# "end\_offset" : 49,

# "position" : 9,

# "start\_offset" : 44,

# "type" : "",

# "token" : "quote"

# },

# {

# "end\_offset" : 60,

# "position" : 10,

# "start\_offset" : 51,

# "type" : "",

# "token" : "generated"

# },

# {

# "end\_offset" : 63,

# "position" : 11,

# "start\_offset" : 61,

# "type" : "",

# "token" : "by"

# },

# {

# "end\_offset" : 67,

# "position" : 12,

# "start\_offset" : 64,

# "type" : "",

# "token" : "the"

# },

# {

# "end\_offset" : 75,

# "position" : 13,

# "start\_offset" : 68,

# "type" : "",

# "token" : "outlook"

# },

# {

# "end\_offset" : 82,

# "position" : 14,

# "start\_offset" : 76,

# "type" : "",

# "token" : "editor"

# }

# ]

# }

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![phill](https://avatars.discourse-cdn.com/v4/letter/p/779978/32.png) [@phill](https://discuss.elastic.co/u/phill)\
**Post date:** [May 24, 2013, 6:27pm UTC](https://discuss.elastic.co/t/filtering-for-apostrophes-and-single-quotes-confusion/12102/3 "2013-05-24T18:27:48Z")

</div>

On 5/24/2013 3:22 AM, Clinton Gormley wrote:

> Hi Paul
> 
> Interesting...
> 
> From this:
> 
> > For these reasons, the English stemmer treats apostrophe as a letter,  
> > removing it from the beginning of a word, where it might have stood  
> > for an opening quote, from the end of the word, where it might have  
> > stood for a closing quote, or been an apostrophe following _/s/_. The  
> > form _/’s/_ is also treated as an ending.
> 
> ... it sounds like things will work correctly as long as you normalize  
> all single quotes/apostrophes to the same character, which you can do  
> with a char filter:

Thanks for the response. Your are right, the Snowball parser will be  
happy with just simple character replacement, so there's no need to try  
to identify only "xxx's" occurrences using a custom _token_ filter. I'm  
a little nervous that I'd throw off some other Filter, but my particular  
configuration is all under my control, so all is good.

Just to complete the record, while looking at the Lucene code, I did  
spot that the simple EnglishPossesiveFilter also thinks its worth  
looking for one more.  
U+FF07. In Unicode this is called FULLWIDTH APOSTROPHE, which seems  
more likely than the one I listed in my original list (note the URL link  
was right the URL text was wrong) U+201B SINGLE HIGH-REVERSED-9  
QUOTATION MARK (Gosh what a name!) even if that is used in some actual  
non-technical published documents as mentioned on the Snowball page, but  
not as a possessive or a quote.

I'd suggest your example char filter should get one more entry for this  
fat or full-width apostrophe.

"char\_filter" : {  
"quotes" : {  
"mappings" : [  
"\u0091=\>\u0027",  
"\u0092=\>\u0027",  
"\u2018=\>\u0027",  
"\u2019=\>\u0027"  
"\uFF07=\>\u0027"  
],  
"type" : "mapping"  
}  
}

I'd leave out all the myriad others characters that look like  
apostrophes which all seem to be special linguistic marks which I hope  
remain part of the term for any linguist processing or maybe further  
filtered away when no one cares.

-Paul

p.s. The fully correct way to spell Hawaii uses one of those really  
special characters -- Hawai?i. see

> **[MODIFIER LETTER TURNED COMMA (U+02BB)](https://www.fileformat.info/info/unicode/char/02BB/index.htm)**
>
> Get the complete details on Unicode character U+02BB on FileFormat.Info

"used in Hawai`ian orthography as `okina (glottal stop)" (but that last  
sentence used what many folks use for glottal stop - a grave accent).  
Now you too can form sentences of the form "Hawai?i's language  
orthography has it's own special characters."

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:34am UTC](https://discuss.elastic.co/t/filtering-for-apostrophes-and-single-quotes-confusion/12102/4 "2017-07-06T02:34:54Z")

</div>


