# How to search exact text?

**URL:** <https://discuss.elastic.co/t/how-to-search-exact-text/9950>\
**Category:** Elasticsearch\
**Created:** [December 5, 2012, 10:46am UTC](https://discuss.elastic.co/t/how-to-search-exact-text/9950 "2012-12-05T10:46:26Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Amy](https://avatars.discourse-cdn.com/v4/letter/a/4da419/32.png) [@Amy](https://discuss.elastic.co/u/Amy)\
**Post date:** [December 5, 2012, 10:46am UTC](https://discuss.elastic.co/t/how-to-search-exact-text/9950/1 "2012-12-05T10:46:26Z")

</div>

Hi,  
I've added the following 2 docs to my index:  
curl -XPUT localhost:9200/testindex/doc/3 -d '{"language":"it"}'  
curl -XPUT localhost:9200/testindex/doc/4 -d '{"language":"pp"}'

I'd like to search for the docs by language.

The following query returns _no_ documents:  
curl -XPOST localhost:9200/testindex/\_search -d  
'{"query":{"bool":{"must":[{"term":{"language":"it"}}]}}}'

Whereas searching for the other "language" (pp) _does_ return documents:  
curl -XPOST localhost:9200/testindex/\_search -d  
'{"query":{"bool":{"must":[{"term":{"language":"pp"}}]}}}'

Why is "it" a special case? How do I search for the exact text and get back  
results every time?  
Regards,  
Amy.

--

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 5, 2012, 11:14am UTC](https://discuss.elastic.co/t/how-to-search-exact-text/9950/2 "2012-12-05T11:14:27Z")

</div>

By default, Elasticsearch applied a standard analyzer (english analyzer).  
The immediate consequence is that common words are ignored during the analyze  
process.

"IT" is a common word in english. So it has not been indexed.

Your use case indicates that you have coded field, "it" instead of italian, I  
suppose.

So, you can either define a mapping for the field language and set your field as  
"index":"not\_analyzed"

See doc here:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

Or, you can define you own analyzer, for example, I often use a custom analyzer  
with a keyword tokenizer with a lowercase filter.  
And apply it to your field.

Does it help?  
David.

Le 5 décembre 2012 à 11:46, Amy [amyblarney@gmail.com](mailto:amyblarney@gmail.com) a écrit :

> Hi,  
> I've added the following 2 docs to my index:  
> curl -XPUT localhost:9200/testindex/doc/3 -d '{"language":"it"}'  
> curl -XPUT localhost:9200/testindex/doc/4 -d '{"language":"pp"}'
> 
> I'd like to search for the docs by language.
> 
> The following query returns no documents:  
> curl -XPOST localhost:9200/testindex/\_search -d  
> '{"query":{"bool":{"must":[{"term":{"language":"it"}}]}}}'
> 
> Whereas searching for the other "language" (pp) does return documents:  
> curl -XPOST localhost:9200/testindex/\_search -d  
> '{"query":{"bool":{"must":[{"term":{"language":"pp"}}]}}}'
> 
> Why is "it" a special case? How do I search for the exact text and get back  
> results every time?  
> Regards,  
> Amy.
> 
> --

--  
David Pilato  
[http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
[http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

--

---

<div class="post-metadata">

**Author:** ![radu\_gheorghe](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/radu_gheorghe/32/556_2.png) [@radu\_gheorghe](https://discuss.elastic.co/u/radu_gheorghe)\
**Post date:** [December 5, 2012, 11:14am UTC](https://discuss.elastic.co/t/how-to-search-exact-text/9950/3 "2012-12-05T11:14:52Z")

</div>

Hello Amy,

On Wed, Dec 5, 2012 at 12:46 PM, Amy [amyblarney@gmail.com](mailto:amyblarney@gmail.com) wrote:

> Hi,  
> I've added the following 2 docs to my index:  
> curl -XPUT localhost:9200/testindex/doc/3 -d '{"language":"it"}'  
> curl -XPUT localhost:9200/testindex/doc/4 -d '{"language":"pp"}'
> 
> I'd like to search for the docs by language.
> 
> The following query returns _no_ documents:  
> curl -XPOST localhost:9200/testindex/\_search -d  
> '{"query":{"bool":{"must":[{"term":{"language":"it"}}]}}}'
> 
> Whereas searching for the other "language" (pp) _does_ return documents:  
> curl -XPOST localhost:9200/testindex/\_search -d  
> '{"query":{"bool":{"must":[{"term":{"language":"pp"}}]}}}'
> 
> Why is "it" a special case?

It's because by default, fields are analyzed using the standard analyzer.  
And that also ignores English stop words from the list of terms. And "it"  
is an English stop word.

> How do I search for the exact text and get back results every time?

If you want exact results of your documents, you can tell ES not to analyze  
your field at all:  
curl -XPOST localhost:9200/testindex/\_search -d  
'{"query":{"bool":{"must":[{"term":{"language":"it"}}]}}}'

That would also improve performance on indexing new docs.

But you can also customize the analysis process, as ES exposes lots of  
options:

> **[Elastic — The Search AI Company](https://www.elastic.co)**
>
> Power insights and outcomes with The Elastic Search AI Platform. See into your data and find answers that matter with enterprise solutions designed to help you accelerate time to insight. Try Elastic ...

## Best regards, Radu

[http://sematext.com/](http://sematext.com/) -- Elasticsearch -- Solr -- Lucene

--

---

<div class="post-metadata">

**Author:** ![radu\_gheorghe](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/radu_gheorghe/32/556_2.png) [@radu\_gheorghe](https://discuss.elastic.co/u/radu_gheorghe)\
**Post date:** [December 5, 2012, 11:16am UTC](https://discuss.elastic.co/t/how-to-search-exact-text/9950/4 "2012-12-05T11:16:29Z")

</div>

David, there should be some mid-air collision detection on this group 🙂

On Wed, Dec 5, 2012 at 1:14 PM, David Pilato [david@pilato.fr](mailto:david@pilato.fr) wrote:

> \*\*  
> By default, Elasticsearch applied a standard analyzer (english analyzer).  
> The immediate consequence is that common words are ignored during the  
> analyze process.
> 
> "IT" is a common word in english. So it has not been indexed.
> 
> Your use case indicates that you have coded field, "it" instead of  
> italian, I suppose.
> 
> So, you can either define a mapping for the field language and set your  
> field as "index":"not\_analyzed"
> 
> See doc here:  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/core-types.html)
> 
> Or, you can define you own analyzer, for example, I often use a custom  
> analyzer with a keyword tokenizer with a lowercase filter.  
> And apply it to your field.
> 
> Does it help?  
> David.
> 
> Le 5 décembre 2012 à 11:46, Amy [amyblarney@gmail.com](mailto:amyblarney@gmail.com) a écrit :
> 
> Hi,  
> I've added the following 2 docs to my index:  
> curl -XPUT localhost:9200/testindex/doc/3 -d '{"language":"it"}'  
> curl -XPUT localhost:9200/testindex/doc/4 -d '{"language":"pp"}'
> 
> I'd like to search for the docs by language.
> 
> The following query returns _no_ documents:  
> curl -XPOST localhost:9200/testindex/\_search -d  
> '{"query":{"bool":{"must":[{"term":{"language":"it"}}]}}}'
> 
> Whereas searching for the other "language" (pp) _does_ return documents:  
> curl -XPOST localhost:9200/testindex/\_search -d  
> '{"query":{"bool":{"must":[{"term":{"language":"pp"}}]}}}'
> 
> Why is "it" a special case? How do I search for the exact text and get  
> back results every time?  
> Regards,  
> Amy.
> 
> --
> 
> --  
> David Pilato  
> [http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
> [http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
> Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs
> 
> --

--  
[http://sematext.com/](http://sematext.com/) -- Elasticsearch -- Solr -- Lucene

--

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 5, 2012, 11:20am UTC](https://discuss.elastic.co/t/how-to-search-exact-text/9950/5 "2012-12-05T11:20:41Z")

</div>

LOL! Right 😉

What about a \_version field on each thread 😉

Cheers

Le 5 décembre 2012 à 12:16, Radu Gheorghe [radu.gheorghe@sematext.com](mailto:radu.gheorghe@sematext.com) a écrit :

> David, there should be some mid-air collision detection on this group 🙂
> 
> On Wed, Dec 5, 2012 at 1:14 PM, David Pilato \<[david@pilato.fr](mailto:david@pilato.fr)  
> [mailto:david@pilato.fr](mailto:david@pilato.fr) \> wrote:
> 
> > > By default, Elasticsearch applied a standard analyzer (english  
> > > analyzer).  
> > > The immediate consequence is that common words are ignored during the  
> > > analyze process.
> > 
> > "IT" is a common word in english. So it has not been indexed.
> > 
> > Your use case indicates that you have coded field, "it" instead of  
> > italian, I suppose.
> > 
> > So, you can either define a mapping for the field language and set your  
> > field as "index":"not\_analyzed"
> > 
> > See doc here:  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/core-types.html)  
> > [http://www.elasticsearch.org/guide/reference/mapping/core-types.html](http://www.elasticsearch.org/guide/reference/mapping/core-types.html)
> > 
> > Or, you can define you own analyzer, for example, I often use a custom  
> > analyzer with a keyword tokenizer with a lowercase filter.  
> > And apply it to your field.
> > 
> > Does it help?  
> > David.
> > 
> > Le 5 décembre 2012 à 11:46, Amy \< [amyblarney@gmail.com](mailto:amyblarney@gmail.com)  
> > [mailto:amyblarney@gmail.com](mailto:amyblarney@gmail.com) \> a écrit :
> > 
> > ```
> > > > > Hi,
> > 
> > ```
> > 
> > > ```
> > > I've added the following 2 docs to my index:
> > > curl -XPUT localhost:9200/testindex/doc/3 -d '{"language":"it"}'
> > > curl -XPUT localhost:9200/testindex/doc/4 -d '{"language":"pp"}'
> > > 
> > > I'd like to search for the docs by language.
> > > 
> > > The following query returns no documents:
> > > curl -XPOST localhost:9200/testindex/_search -d
> > > 
> > > ```
> > > 
> > > '{"query":{"bool":{"must":[{"term":{"language":"it"}}]}}}'
> > > 
> > > ```
> > > Whereas searching for the other "language" (pp) does return documents:
> > > curl -XPOST localhost:9200/testindex/_search -d
> > > 
> > > ```
> > > 
> > > '{"query":{"bool":{"must":[{"term":{"language":"pp"}}]}}}'
> > > 
> > > ```
> > > Why is "it" a special case? How do I search for the exact text and get
> > > 
> > > ```
> > > 
> > > back results every time?  
> > > Regards,  
> > > Amy.
> > > 
> > > ```
> > > --
> > > 
> > > ```
> > > 
> > > > >
> > 
> > --  
> > David Pilato  
> > [http://www.scrutmydocs.org/](http://www.scrutmydocs.org/) [http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
> > [http://dev.david.pilato.fr/](http://dev.david.pilato.fr/) [http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
> > Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs
> > 
> > --
> > 
> > >
> 
> --  
> [http://sematext.com/](http://sematext.com/) [http://sematext.com/](http://sematext.com/) -- Elasticsearch -- Solr --  
> Lucene
> 
> --

--  
David Pilato  
[http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
[http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

--

---

<div class="post-metadata">

**Author:** ![Amy](https://avatars.discourse-cdn.com/v4/letter/a/4da419/32.png) [@Amy](https://discuss.elastic.co/u/Amy)\
**Post date:** [December 5, 2012, 12:38pm UTC](https://discuss.elastic.co/t/how-to-search-exact-text/9950/6 "2012-12-05T12:38:25Z")

</div>

Hi,  
Wow, that was quick! Thanks! That helped.  
I added the standard analyser by adding the following to the  
elasticsearch.yml config file:

#index Settings  
index:  
analysis:  
analyzer:  
# set standard analyzer with no stop words as the default for both  
indexing and searching  
default:  
type: standard  
stopwords: _none_

On Wednesday, December 5, 2012 11:20:41 AM UTC, David Pilato wrote:

> LOL! Right 😉
> 
> What about a \_version field on each thread 😉
> 
> Cheers
> 
> Le 5 décembre 2012 à 12:16, Radu Gheorghe \<[radu.g...@sematext.com](mailto:radu.g...@sematext.com)\<javascript:\>\>  
> a écrit :
> 
> David, there should be some mid-air collision detection on this group 🙂
> 
> On Wed, Dec 5, 2012 at 1:14 PM, David Pilato \<[da...@pilato.fr](mailto:da...@pilato.fr)\<javascript:\>
> 
> > wrote:
> 
> By default, Elasticsearch applied a standard analyzer (english  
> analyzer).  
> The immediate consequence is that common words are ignored during the  
> analyze process.
> 
> "IT" is a common word in english. So it has not been indexed.
> 
> Your use case indicates that you have coded field, "it" instead of  
> italian, I suppose.
> 
> So, you can either define a mapping for the field language and set your  
> field as "index":"not\_analyzed"
> 
> See doc here:  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/mapping/core-types.html)
> 
> Or, you can define you own analyzer, for example, I often use a custom  
> analyzer with a keyword tokenizer with a lowercase filter.  
> And apply it to your field.
> 
> Does it help?  
> David.
> 
> Le 5 décembre 2012 à 11:46, Amy \< [amybl...@gmail.com](mailto:amybl...@gmail.com) \<javascript:\>\> a  
> écrit :
> 
> Hi,  
> I've added the following 2 docs to my index:  
> curl -XPUT localhost:9200/testindex/doc/3 -d '{"language":"it"}'  
> curl -XPUT localhost:9200/testindex/doc/4 -d '{"language":"pp"}'
> 
> I'd like to search for the docs by language.
> 
> The following query returns _no_ documents:  
> curl -XPOST localhost:9200/testindex/\_search -d  
> '{"query":{"bool":{"must":[{"term":{"language":"it"}}]}}}'
> 
> Whereas searching for the other "language" (pp) _does_ return documents:  
> curl -XPOST localhost:9200/testindex/\_search -d  
> '{"query":{"bool":{"must":[{"term":{"language":"pp"}}]}}}'
> 
> Why is "it" a special case? How do I search for the exact text and get  
> back results every time?  
> Regards,  
> Amy.
> 
> --
> 
> --  
> David Pilato  
> [http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
> [http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
> Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs
> 
> --
> 
> --  
> [http://sematext.com/](http://sematext.com/) -- Elasticsearch -- Solr -- Lucene
> 
> --
> 
> --  
> David Pilato  
> [http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
> [http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
> Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:01am UTC](https://discuss.elastic.co/t/how-to-search-exact-text/9950/7 "2017-07-06T03:01:18Z")

</div>


