# Serching requested files in Kibana which ends in a number and the file extension

**URL:** <https://discuss.elastic.co/t/serching-requested-files-in-kibana-which-ends-in-a-number-and-the-file-extension/78906>\
**Category:** Elasticsearch\
**Created:** [March 16, 2017, 4:23pm UTC](https://discuss.elastic.co/t/serching-requested-files-in-kibana-which-ends-in-a-number-and-the-file-extension/78906 "2017-03-16T16:23:44Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![YvorL](https://avatars.discourse-cdn.com/v4/letter/y/9fc348/32.png) [@YvorL](https://discuss.elastic.co/u/YvorL)\
**Post date:** [March 16, 2017, 4:23pm UTC](https://discuss.elastic.co/t/serching-requested-files-in-kibana-which-ends-in-a-number-and-the-file-extension/78906/1 "2017-03-16T16:23:45Z")

</div>

I bumped into a strange issue. When I tried to look for the most requested documents (pdf) on one of the analyzed sites, I saw that there are some docs definitely missing. First I used "request:\*.pdf" then I checked "request:pdf". That was the time I noticed that the wildcard request missed ALL documents which request ended in any number before '.pdf' such as 'calendar\_2017.pdf'. Which is odd, because it is a string field so I don't understand how a number can cause this issue. Is there something I can do without reindexing the data?

---

<div class="post-metadata">

**Author:** ![Brandon\_Kobel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/brandon_kobel/32/14829_2.png) [@Brandon\_Kobel](https://discuss.elastic.co/u/Brandon_Kobel)\
**Post date:** [March 16, 2017, 4:43pm UTC](https://discuss.elastic.co/t/serching-requested-files-in-kibana-which-ends-in-a-number-and-the-file-extension/78906/2 "2017-03-16T16:43:23Z")

</div>

@YvorL would you mind posting the mapping for the specific field that you'd having issues searching on?

---

<div class="post-metadata">

**Author:** ![YvorL](https://avatars.discourse-cdn.com/v4/letter/y/9fc348/32.png) [@YvorL](https://discuss.elastic.co/u/YvorL)\
**Post date:** [March 16, 2017, 4:49pm UTC](https://discuss.elastic.co/t/serching-requested-files-in-kibana-which-ends-in-a-number-and-the-file-extension/78906/3 "2017-03-16T16:49:36Z")

</div>

Is this what you're looking for?  
"request": {  
"type": "text",  
"norms": false,  
"fields": {  
"keyword": {  
"type": "keyword"  
}  
}  
}

---

<div class="post-metadata">

**Author:** ![Brandon\_Kobel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/brandon_kobel/32/14829_2.png) [@Brandon\_Kobel](https://discuss.elastic.co/u/Brandon_Kobel)\
**Post date:** [March 16, 2017, 7:16pm UTC](https://discuss.elastic.co/t/serching-requested-files-in-kibana-which-ends-in-a-number-and-the-file-extension/78906/4 "2017-03-16T19:16:13Z")

</div>

@YvorL if you're using `request:*.pdf` in kibana in the querybar, it's translated into a query\_string query against an analyzed text string.

The standard analyzer splits the following text into the following keywords:

1. calendar\_2017.pdf -\> calendar\_2017, pdf
2. 01302017.pdf -\> 01302017, pdf
3. something.pdf -\> something.pdf

The standard analyzer is generally meant for text fields, but it explains why `*.pdf` doesn't return anything for 1 and 2 above.

You should be able to use `request.keyword: *.pdf` which isn't executing the query against the analyzed field and should return what you're looking for.

If you are able to reindex your data, pulling out the extension either using a pattern analyzer or some other mechanism during ingest would be much more performant.

---

<div class="post-metadata">

**Author:** ![YvorL](https://avatars.discourse-cdn.com/v4/letter/y/9fc348/32.png) [@YvorL](https://discuss.elastic.co/u/YvorL)\
**Post date:** [March 16, 2017, 7:35pm UTC](https://discuss.elastic.co/t/serching-requested-files-in-kibana-which-ends-in-a-number-and-the-file-extension/78906/5 "2017-03-16T19:35:59Z")

</div>

@Brandon_Kobel  
I still don't see why the first two are separated to keywords if the last one isn't. My understanding is that if the text is continuous (and in this case, it'll be the URI) then a dot or an underscore won't act as keyword separator. It's a text field, and it should handle numbers as any other character.  
Regardless, it seems that this is the intended way. I was also avoiding searching in a unanalyzed field with a leading asterisk because it won't be a one-time query. It leaves me to reindexing the data.

Thank you for taking your time!

---

<div class="post-metadata">

**Author:** ![Brandon\_Kobel](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/brandon_kobel/32/14829_2.png) [@Brandon\_Kobel](https://discuss.elastic.co/u/Brandon_Kobel)\
**Post date:** [March 16, 2017, 8:02pm UTC](https://discuss.elastic.co/t/serching-requested-files-in-kibana-which-ends-in-a-number-and-the-file-extension/78906/6 "2017-03-16T20:02:12Z")

</div>

@YvorL unfortunately, the details of the tokenizer are out of my expertise, but I'll move this to the Elasticsearch forum and hopefully they can enlighten us both 🙂

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 13, 2017, 8:02pm UTC](https://discuss.elastic.co/t/serching-requested-files-in-kibana-which-ends-in-a-number-and-the-file-extension/78906/7 "2017-04-13T20:02:52Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
