# Extract relevant documents when search input may/may not contain white space

**URL:** <https://discuss.elastic.co/t/extract-relevant-documents-when-search-input-may-may-not-contain-white-space/127040>\
**Category:** Elasticsearch\
**Created:** [April 6, 2018, 7:08am UTC](https://discuss.elastic.co/t/extract-relevant-documents-when-search-input-may-may-not-contain-white-space/127040 "2018-04-06T07:08:04Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![sobiya](https://avatars.discourse-cdn.com/v4/letter/s/ee59a6/32.png) [@sobiya](https://discuss.elastic.co/u/sobiya)\
**Post date:** [April 6, 2018, 7:08am UTC](https://discuss.elastic.co/t/extract-relevant-documents-when-search-input-may-may-not-contain-white-space/127040/1 "2018-04-06T07:08:04Z")

</div>

Hi all,  
I can extract the documents with/without space. But if i give input with space (without space in index),the relevant documents are not hitting at the top. Need to score high when compared to other documents.  
For eg., Index contains "tomcat", "tom cat", "bob cat", "tom fish".  
when I search for "tom cat" (tokenized like "tom", "tomcat" and "cat"), it doesn't produce "tom cat" at the first.  
Expected result in order: "tom cat", "tomcat", "tom fish" and "bob cat".  
Help me to hit the documents with highest relevance score.

Thanks in advance.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [April 6, 2018, 7:25am UTC](https://discuss.elastic.co/t/extract-relevant-documents-when-search-input-may-may-not-contain-white-space/127040/2 "2018-04-06T07:25:08Z")

</div>

You can combine different kind of search clauses within multiple should clauses.  
Which means that the exact match will match in both cases, which means a better score.

An example of this here:

> <https://gist.github.com/dadoonet/5179ee72ecbf08f12f53d4bda1b76bab>

---

<div class="post-metadata">

**Author:** ![sobiya](https://avatars.discourse-cdn.com/v4/letter/s/ee59a6/32.png) [@sobiya](https://discuss.elastic.co/u/sobiya)\
**Post date:** [April 7, 2018, 7:25am UTC](https://discuss.elastic.co/t/extract-relevant-documents-when-search-input-may-may-not-contain-white-space/127040/3 "2018-04-07T07:25:03Z")

</div>

I need to search with space but the original value has no space in it. If I did this, it should return the original value at the top. I can use synonym but it is impossible to set for every values in my index.

![](https://us1.discourse-cdn.com/elastic/original/2X/2/2485d26919316b1d80aa60f5ca8e2cb5bc0275fe.jpg)

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [April 7, 2018, 8:40am UTC](https://discuss.elastic.co/t/extract-relevant-documents-when-search-input-may-may-not-contain-white-space/127040/4 "2018-04-07T08:40:25Z")

</div>

May be ngrams would help? But more indexing time, more space on disk...  
Fuzzy queries on a lowercased analyzer field?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 5, 2018, 8:40am UTC](https://discuss.elastic.co/t/extract-relevant-documents-when-search-input-may-may-not-contain-white-space/127040/5 "2018-05-05T08:40:32Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
