# Store text with hiphens and dot

**URL:** https://discuss.elastic.co/t/store-text-with-hiphens-and-dot/140622
**Category:** Elasticsearch
**Created:** [July 18, 2018, 8:52pm UTC](https://discuss.elastic.co/t/store-text-with-hiphens-and-dot/140622 "2018-07-18T20:52:58Z")
**Posts on this page:** 3
**Page:** 1

<div class="post-metadata">

### Author: ![shashank123hr](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shashank123hr/32/33495_2.png) [@shashank123hr](https://discuss.elastic.co/u/shashank123hr)
#### Post date: [July 18, 2018, 8:52pm UTC](https://discuss.elastic.co/t/store-text-with-hiphens-and-dot/140622/1 "2018-07-18T20:52:58Z")

</div>

I am trying to come up with a query for elastic search to search for strings with hiphen and dots example : "PN-8100.0500-070, PN-8100.0500, PN-8100.0500-180". The search should be able to match exact strings like "PN-8100.0500-070" or "PN-8100.0500" in the above example. Note : The example string is a single field in elasticsearch index.

I tried using the query string with minimum match set to 100%, but it doesn't quite do complete match on the string (it returns entries with PN-8100.0500 as well).

Would it help if I store the comma separated string as string array and search using keyword analyzer.?

Any suggestions would be greatly appreciated. Thanks

---

<div class="post-metadata">

### Author: ![cbuescher](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cbuescher/32/60402_2.png) [@cbuescher](https://discuss.elastic.co/u/cbuescher)
#### Post date: [July 20, 2018, 12:37pm UTC](https://discuss.elastic.co/t/store-text-with-hiphens-and-dot/140622/2 "2018-07-20T12:37:26Z")

</div>

If there is only a single "PN-123814.2314-323" thingy per field in the input document, its enough to map them to the "keyword" type or use the ".raw" multifield if you use the default mapping. If there are several of these Ids in a field of the input document, you need to find a tokenizer that splits them correctly without removing the dots/hyphens in the ids. For example, if the come as "PN-8100.0500-070, PN-8100.0500, PN-8100.0500-180" like you mentioned, you might be able to use the [Pattern Tokenizer](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-pattern-tokenizer.html), the [Simple Pattern Tokenizer](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-simplepattern-tokenizer.html) or the [Simple Pattern Split Tokenizer](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-simplepatternsplit-tokenizer.html) and break on comma and whitespace.

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [August 17, 2018, 12:37pm UTC](https://discuss.elastic.co/t/store-text-with-hiphens-and-dot/140622/3 "2018-08-17T12:37:30Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
