# Search phone number

**URL:** <https://discuss.elastic.co/t/search-phone-number/32771>\
**Category:** Elasticsearch\
**Created:** [October 22, 2015, 12:33pm UTC](https://discuss.elastic.co/t/search-phone-number/32771 "2015-10-22T12:33:07Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![kodermax](https://avatars.discourse-cdn.com/v4/letter/k/9d8465/32.png) [@kodermax](https://discuss.elastic.co/u/kodermax)\
**Post date:** [October 22, 2015, 12:33pm UTC](https://discuss.elastic.co/t/search-phone-number/32771/1 "2015-10-22T12:33:07Z")

</div>

I have type

```
contact : {
   'phone':[79037767523,79037767523]
}

```

How can i search 9037767523?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [October 22, 2015, 12:37pm UTC](https://discuss.elastic.co/t/search-phone-number/32771/2 "2015-10-22T12:37:11Z")

</div>

Check out ngrams [1]. It's a way of indexing parts of words rather than whole ones.  
The definitive guide [2] has a section on this too.

[1][https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-ngram-tokenizer.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-ngram-tokenizer.html)  
[2] [https://www.elastic.co/guide/en/elasticsearch/guide/current/\_ngrams\_for\_partial\_matching.html](https://www.elastic.co/guide/en/elasticsearch/guide/current/_ngrams_for_partial_matching.html)

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [October 22, 2015, 2:16pm UTC](https://discuss.elastic.co/t/search-phone-number/32771/3 "2015-10-22T14:16:33Z")

</div>

I think I'd try to use [pattern capture token filter](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-pattern-capture-tokenfilter.html) to extract the things you want to match. In the case you mention you'd want to strip the leading 7. I don't know Russia's number resolution rules, but for nanpa I'd use something like

```auto
"phone_number" : {
  "type" : "pattern_capture",
  "preserve_original" : 1,
  "patterns" : [
    "1(\\d{3}(\\d+))"
  ]
}

```

You apply phone\_number analyzer as the `index_analyzer` and just use a `keyword` analyzer that strips `+-()` at search time. Or strip in your application. The `index_analyzer` here would index a number like `19195557321` as `19195557321`, `9195557321`, and `5557321` which matches the way phone numbers are resolved in nanpa. A user searching for `5557321` would get all the numbers ending in `5557321` - `19195557321`, `13215557321`, etc.

I'd also strip all the `+-()` stuff from the numbers before indexing them in elasticsearch. You don't want them in the `_source` because they don't add anything.

I once worked for a phone company so I've thought a lot about phone numbers.

BTW - this is a tradeoff. When you get new resolution rules you have to change the mapping and reindex the whole index. If you moved term expansion that the analyzer is doing outside of elasticsearch then you could be more surgical when the patterns change. I'd suggest doing something like that if you had to cover the whole world. So you'd index

```auto
{
  "phone_number": {
    "raw": "19195557321",
    "expansions": ["1919557321", "9195557321", "5557321"]
  }
}

```

and you'd search on phone\_number.expansions.

Which solution you take is all a matter of how big of a deal this is for you. @Mark_Harwood's solution is perfectly reasonable for lots of applications. Its certainly simpler.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [October 22, 2015, 2:24pm UTC](https://discuss.elastic.co/t/search-phone-number/32771/4 "2015-10-22T14:24:00Z")

</div>

Oh one more from my bookmarks - a plugin dedicated to this: [https://github.com/MyPureCloud/elasticsearch-phone](https://github.com/MyPureCloud/elasticsearch-phone)

Not tried it but obviously from someone who's spent some time thinking about this specific problem.  
Let us know if it is any good!

---

<div class="post-metadata">

**Author:** ![kodermax](https://avatars.discourse-cdn.com/v4/letter/k/9d8465/32.png) [@kodermax](https://discuss.elastic.co/u/kodermax)\
**Post date:** [October 22, 2015, 2:35pm UTC](https://discuss.elastic.co/t/search-phone-number/32771/5 "2015-10-22T14:35:52Z")

</div>

thanks all!

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [October 22, 2015, 2:50pm UTC](https://discuss.elastic.co/t/search-phone-number/32771/6 "2015-10-22T14:50:46Z")

</div>

> [@Mark\_Harwood](#):
>
> Oh one more from my bookmarks - a plugin dedicated to this: [GitHub - purecloudlabs/elasticsearch-phone: An Elasticsearch Phone Number Analyzer Plugin](https://github.com/MyPureCloud/elasticsearch-phone)

Very cool!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:43pm UTC](https://discuss.elastic.co/t/search-phone-number/32771/7 "2017-07-05T23:43:12Z")

</div>


