# Querying a file path field by just the basename

**URL:** <https://discuss.elastic.co/t/querying-a-file-path-field-by-just-the-basename/257582>\
**Category:** Elasticsearch\
**Created:** [December 3, 2020, 8:11pm UTC](https://discuss.elastic.co/t/querying-a-file-path-field-by-just-the-basename/257582 "2020-12-03T20:11:47Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![dansei](https://avatars.discourse-cdn.com/v4/letter/d/e19b73/32.png) [@dansei](https://discuss.elastic.co/u/dansei)\
**Post date:** [December 3, 2020, 8:11pm UTC](https://discuss.elastic.co/t/querying-a-file-path-field-by-just-the-basename/257582/1 "2020-12-03T20:11:47Z")

</div>

I have a field that stores the full path of a file on disk. The file's basename is unique and I'd like to be able to query for just the basename as well.

The simplest solution is to add a second field which contains the basename but I tried using an analyzer:

1. I created a `path_tree_rev_tokenizer` of type `path_hierarchy` with delimiter `/` and set it to reverse the order of tokenization, so `/home/bob` is tokenized as `[home/bob, bob]`.

2. I created a `path_tree_rev` analyzer that uses this tokenizer.

3. I made a Text field called `file.path` with the `path_tree_rev` analyzer.

In a simple test separate from my main application code, I created a document with `/home/bob` in the `file.path` field and queried it with

```
"query": {
    "match": {
        "file.path": "bob"
    }
}

```

and it matched my document successfully.

However when I literally copied and pasted the code into my main application and re-created my index with the new field and analyzer definitions, Elasticsearch is unable to find a "real" document by querying for the basename. The only difference I can see is that in my application the field name is `source.file.path`.

Is there something I'm doing wrong here, or is there a way to diagnose why the query is not succeeding?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 8, 2020, 6:57am UTC](https://discuss.elastic.co/t/querying-a-file-path-field-by-just-the-basename/257582/2 "2020-12-08T06:57:43Z")

</div>

Can you share a sample document as well as the mapping of the index?

---

<div class="post-metadata">

**Author:** ![dansei](https://avatars.discourse-cdn.com/v4/letter/d/e19b73/32.png) [@dansei](https://discuss.elastic.co/u/dansei)\
**Post date:** [December 9, 2020, 3:22pm UTC](https://discuss.elastic.co/t/querying-a-file-path-field-by-just-the-basename/257582/3 "2020-12-09T15:22:31Z")

</div>

Hi @Christian_Dahlqvist. Here is a sequence of queries that creates an index, adds a document, and then queries it:

```auto
curl -XDELETE 'http://localhost:9200/foo?pretty' -d ''
curl -H 'Content-Type: application/json' -XPUT 'http://localhost:9200/foo?pretty' -d '{
  "mappings": {
    "properties": {
      "bar": {
        "properties": {
          "name": {
            "type": "text"
          },
          "path": {
            "analyzer": "path_tree_rev",
            "type": "text"
          }
        },
        "type": "object"
      }
    }
  },
  "settings": {
    "analysis": {
      "analyzer": {
        "path_tree_rev": {
          "tokenizer": "path_tree_rev_tokenizer",
          "type": "custom"
        }
      },
      "tokenizer": {
        "path_tree_rev_tokenizer": {
          "delimiter": "/",
          "reverse": true,
          "type": "path_hierarchy"
        }
      }
    }
  }
}'
curl -H 'Content-Type: application/json' -XPOST 'http://localhost:9200/foo/_doc?pretty' -d '{
  "bar": {
    "name": "Bob",
    "path": "/home/bob"
  }
}'
curl -H 'Content-Type: application/json' -XPOST 'http://localhost:9200/foo/_search?pretty' -d '{
  "query": {
    "match": {
      "bar.path": "bob"
    }
  }
}'

```

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 9, 2020, 3:23pm UTC](https://discuss.elastic.co/t/querying-a-file-path-field-by-just-the-basename/257582/4 "2020-12-09T15:23:52Z")

</div>

Please provide an example that can be run from the Kibana console.

---

<div class="post-metadata">

**Author:** ![dansei](https://avatars.discourse-cdn.com/v4/letter/d/e19b73/32.png) [@dansei](https://discuss.elastic.co/u/dansei)\
**Post date:** [December 9, 2020, 3:33pm UTC](https://discuss.elastic.co/t/querying-a-file-path-field-by-just-the-basename/257582/5 "2020-12-09T15:33:31Z")

</div>

I updated the post above with the actual requests made by Python. Hopefully they can be pasted into Kibana easily.

---

<div class="post-metadata">

**Author:** ![dansei](https://avatars.discourse-cdn.com/v4/letter/d/e19b73/32.png) [@dansei](https://discuss.elastic.co/u/dansei)\
**Post date:** [December 9, 2020, 3:57pm UTC](https://discuss.elastic.co/t/querying-a-file-path-field-by-just-the-basename/257582/6 "2020-12-09T15:57:29Z")

</div>

So I discovered that I can dump all of the terms indexed for a given field. So I tried it out on my full application where searching by base name is _not_ working.

```auto
curl -H 'Content-Type: application/json' -XGET 'http://localhost:9200/my_index/_doc/gDEhSHYBoWN8sy6cFN7j/_termvectors' -d '{
  "fields" : ["source.file"],
  "offsets" : true,
  "payloads" : true,
  "positions" : true,
  "term_statistics" : true,
  "field_statistics" : true
}'

```

This _proves_ that the document is properly indexed by its base name:

```auto
        "some_boring_filename" : {
          "doc_freq" : 1,
          "ttf" : 1,
          "term_freq" : 1,
          "tokens" : [
            {
              "position" : 0,
              "start_offset" : 66,
              "end_offset" : 92
            }
          ]
        },

```

Yet when I query it:

```auto
curl -H 'Content-Type: application/json' -XPOST 'http://localhost:9200/my_index/_search?pretty' -d '{
  "query": {
    "match": {
      "source.file": "some_boring_filename"
    }
  }
}'

```

It is not found:

```auto
{
  "took" : 0,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 0,
      "relation" : "eq"
    },
    "max_score" : null,
    "hits" : []
  }
}

```

---

<div class="post-metadata">

**Author:** ![dansei](https://avatars.discourse-cdn.com/v4/letter/d/e19b73/32.png) [@dansei](https://discuss.elastic.co/u/dansei)\
**Post date:** [December 30, 2020, 7:02pm UTC](https://discuss.elastic.co/t/querying-a-file-path-field-by-just-the-basename/257582/7 "2020-12-30T19:02:25Z")

</div>

Bump to prevent this from being closed.

As the previous post shows, the paths are being properly tokenized, but the search is not working.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 27, 2021, 7:02pm UTC](https://discuss.elastic.co/t/querying-a-file-path-field-by-just-the-basename/257582/8 "2021-01-27T19:02:28Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
