# How to Index Words Actual form and Modified form into Elastic Search

**URL:** <https://discuss.elastic.co/t/how-to-index-words-actual-form-and-modified-form-into-elastic-search/154751>\
**Category:** Elasticsearch\
**Created:** [October 31, 2018, 5:57am UTC](https://discuss.elastic.co/t/how-to-index-words-actual-form-and-modified-form-into-elastic-search/154751 "2018-10-31T05:57:28Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Karthik.krishnan](https://avatars.discourse-cdn.com/v4/letter/k/d6d6ee/32.png) [@Karthik.krishnan](https://discuss.elastic.co/u/Karthik.krishnan)\
**Post date:** [October 31, 2018, 5:57am UTC](https://discuss.elastic.co/t/how-to-index-words-actual-form-and-modified-form-into-elastic-search/154751/1 "2018-10-31T05:57:28Z")

</div>

Hi Team,

I want to index the words into elasticsearch with actual form and modified form.

Example:

The Term "F-35" want to index is "F35" and "F-35", when i search the text F35 or F-35 both should return the document.

Note: Here i am using White space analyzer, so it will not be split into two tokens.

Please someone provide me option to achieve this.

---

<div class="post-metadata">

**Author:** ![SaskiaVola](https://avatars.discourse-cdn.com/v4/letter/s/e79b87/32.png) [@SaskiaVola](https://discuss.elastic.co/u/SaskiaVola)\
**Post date:** [November 1, 2018, 11:55am UTC](https://discuss.elastic.co/t/how-to-index-words-actual-form-and-modified-form-into-elastic-search/154751/2 "2018-11-01T11:55:20Z")

</div>

Hi Karthik,

if you want to remove the punctuation aswell, you should consider using the [standard tokenizer](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-standard-tokenizer.html).  
If that removed more than you want, you can consider using a [Pattern-Analyzer](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-pattern-analyzer.html). But be aware that this might be very slow.

Make sure to test your custom analyzers using the [\_analyze API](https://www.elastic.co/guide/en/elasticsearch/reference/6.4/_testing_analyzers.html)

---

<div class="post-metadata">

**Author:** ![Karthik.krishnan](https://avatars.discourse-cdn.com/v4/letter/k/d6d6ee/32.png) [@Karthik.krishnan](https://discuss.elastic.co/u/Karthik.krishnan)\
**Post date:** [November 1, 2018, 3:49pm UTC](https://discuss.elastic.co/t/how-to-index-words-actual-form-and-modified-form-into-elastic-search/154751/3 "2018-11-01T15:49:34Z")

</div>

Hi SaskiaVola,

Thanks for your time to reply this conversation.

When we use standard tokenizer it will split from "F-35" into 2 different tokens as "F" and "35".  
But i am expecting to be a single token like "F35" and "F-35".

Even Pattern-Analyzer will split the token by punctuation i guess?

---

<div class="post-metadata">

**Author:** ![SaskiaVola](https://avatars.discourse-cdn.com/v4/letter/s/e79b87/32.png) [@SaskiaVola](https://discuss.elastic.co/u/SaskiaVola)\
**Post date:** [November 1, 2018, 5:46pm UTC](https://discuss.elastic.co/t/how-to-index-words-actual-form-and-modified-form-into-elastic-search/154751/4 "2018-11-01T17:46:21Z")

</div>

Hi Karthik,

that's correct. So depending on your data, if you can define a proper pattern for the cases you're referring to, you could use a [character filter](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-pattern-replace-charfilter.html) first, that removes the hyphen inside of words that contain numbers.

Then a query for "F35" and "F-35" would match docs containing both variants.

Hope that works for you.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 29, 2018, 5:46pm UTC](https://discuss.elastic.co/t/how-to-index-words-actual-form-and-modified-form-into-elastic-search/154751/5 "2018-11-29T17:46:22Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
