# To understand the analysis process

**URL:** <https://discuss.elastic.co/t/to-understand-the-analysis-process/26565>\
**Category:** Elasticsearch\
**Created:** [July 30, 2015, 12:09pm UTC](https://discuss.elastic.co/t/to-understand-the-analysis-process/26565 "2015-07-30T12:09:02Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![terrasacer](https://avatars.discourse-cdn.com/v4/letter/t/ecd19e/32.png) [@terrasacer](https://discuss.elastic.co/u/terrasacer)\
**Post date:** [July 30, 2015, 12:09pm UTC](https://discuss.elastic.co/t/to-understand-the-analysis-process/26565/1 "2015-07-30T12:09:02Z")

</div>

Hi all.

I'm trying to understand the analysis process. Especially during query time.

For example we have a field configured to be not analyzed. The value of this field is Albert Einstein.

If I search "Albert" does not match document with the match query but If I use the match\_phrase\_prefix query document is returned.

Why?

---

<div class="post-metadata">

**Author:** ![PatrickKik](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/patrickkik/32/619_2.png) [@PatrickKik](https://discuss.elastic.co/u/PatrickKik)\
**Post date:** [August 3, 2015, 4:57am UTC](https://discuss.elastic.co/t/to-understand-the-analysis-process/26565/2 "2015-08-03T04:57:25Z")

</div>

What I understand from [https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-match-query.html#\_match\_phrase\_prefix](https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-match-query.html#_match_phrase_prefix) is that your query acts as a prefix.

You could test that by querying for "Einstein". Then both the match query and the match\_phrase\_prefix query would return nothing.

Could you post your findings please?

---

<div class="post-metadata">

**Author:** ![terrasacer](https://avatars.discourse-cdn.com/v4/letter/t/ecd19e/32.png) [@terrasacer](https://discuss.elastic.co/u/terrasacer)\
**Post date:** [September 12, 2015, 8:21pm UTC](https://discuss.elastic.co/t/to-understand-the-analysis-process/26565/3 "2015-09-12T20:21:23Z")

</div>

Hi @PatrickKik,

First, sorry for late reply.

I'm trying to understand that: Is not\_analyzed means that without transforming/converting the text to be stored?

If this is true, how match\_phrase\_prefix query returns results?

Does it analyzed text again during the query time?

---

<div class="post-metadata">

**Author:** ![terrasacer](https://avatars.discourse-cdn.com/v4/letter/t/ecd19e/32.png) [@terrasacer](https://discuss.elastic.co/u/terrasacer)\
**Post date:** [September 26, 2015, 9:51am UTC](https://discuss.elastic.co/t/to-understand-the-analysis-process/26565/4 "2015-09-26T09:51:32Z")

</div>

If someone has an idea, how match\_phrase\_prefix query returns results in the above example?

---

<div class="post-metadata">

**Author:** ![softwaredoug](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/softwaredoug/32/22681_2.png) [@softwaredoug](https://discuss.elastic.co/u/softwaredoug)\
**Post date:** [September 26, 2015, 11:58am UTC](https://discuss.elastic.co/t/to-understand-the-analysis-process/26565/5 "2015-09-26T11:58:49Z")

</div>

According to [here](https://www.elastic.co/guide/en/elasticsearch/reference/1.4/query-dsl-match-query.html#_match_phrase_prefix), match phrase prefix does a prefix query on the last term in the query. With not\_analyzed, the whole string is taken as a term. Therefore the "last term" is [Albert Einstein]. Albert is a prefix of this term, therefore it matches.

Most other search queries are exact term matches. A query term needs to match a document's term _exactly_ after analysis is run to match.

This [blog post](http://opensourceconnections.com/blog/2015/09/18/the-simple-power-of-elasticsearch-analyzers/) might be a good primer for you.

---

<div class="post-metadata">

**Author:** ![terrasacer](https://avatars.discourse-cdn.com/v4/letter/t/ecd19e/32.png) [@terrasacer](https://discuss.elastic.co/u/terrasacer)\
**Post date:** [September 26, 2015, 1:16pm UTC](https://discuss.elastic.co/t/to-understand-the-analysis-process/26565/6 "2015-09-26T13:16:29Z")

</div>

Hi @softwaredoug,

Thank you for this informative answer and blog post.

> Therefore the "last term" is [Albert Einstein]. Albert is a prefix of this term, therefore it matches.

I understand that: The string [Albert Einstein] transformed into two sub-string ["Albert", "Einstein"] by the match\_phrase\_prefix query in search time.

Is that correct?

---

<div class="post-metadata">

**Author:** ![softwaredoug](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/softwaredoug/32/22681_2.png) [@softwaredoug](https://discuss.elastic.co/u/softwaredoug)\
**Post date:** [September 26, 2015, 2:37pm UTC](https://discuss.elastic.co/t/to-understand-the-analysis-process/26565/7 "2015-09-26T14:37:16Z")

</div>

The search string is not\_analyzed as well. The search engine takes the whole token [Albert Einstein] and treats it as a word. There's no extra "substring" involved. [Albert] is a prefix of the larger term [Albert Einstein].

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:47pm UTC](https://discuss.elastic.co/t/to-understand-the-analysis-process/26565/8 "2017-07-05T23:47:57Z")

</div>


