# Identifying Attachment that Matches Query

**URL:** <https://discuss.elastic.co/t/identifying-attachment-that-matches-query/78885>\
**Category:** Elasticsearch\
**Created:** [March 16, 2017, 3:05pm UTC](https://discuss.elastic.co/t/identifying-attachment-that-matches-query/78885 "2017-03-16T15:05:53Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![jattenberg](https://avatars.discourse-cdn.com/v4/letter/j/b9e5f3/32.png) [@jattenberg](https://discuss.elastic.co/u/jattenberg)\
**Post date:** [March 16, 2017, 3:05pm UTC](https://discuss.elastic.co/t/identifying-attachment-that-matches-query/78885/1 "2017-03-16T15:05:53Z")

</div>

I am searching over emails that each have attachments that I've ingested using the ingest attachment pipeline. Many emails have multiple attachments, and sometimes it's the content of the attachments themselves- the extracted text that satisfies the query.

My question is: how do I get the filename of the query that matches? I understand how to get the highlighted content of the matched field, but this is different than the filename.

The mapping for attachments i'm using are just the fields that the ingest-attachment plugin provides, using a foreach pipeline:

```
"attachments" : {
            "properties" : {
              "attachment" : {
                "properties" : {
                  "author" : {
                    "type" : "text",
                    "fields" : {
                      "keyword" : {
                        "type" : "keyword",
                        "ignore_above" : 256
                      }
                    }
                  },
                  "content" : {
                    "type" : "text"
                  },
                  "content_length" : {
                    "type" : "long"
                  },
                  "content_type" : {
                    "type" : "keyword"
                  },
                  "date" : {
                    "type" : "date"
                  },
                  "language" : {
                    "type" : "keyword"
                  }
                }
              },
              "data" : {
                "type" : "object",
                "enabled" : false
              },
              "filename" : {
                "type" : "keyword"
              }
            }
          }

```

currently i can get the highlighted content of the attachments like:

```
{
  "query": {
    "bool": {
      "must": [
        {
          "query_string": {
            "default_field": "_all",
            "query": "\"linkedin\"",
            "_name": "all fields"
          }
        }
      ]
    }
  },
  "from": 0,
  "size": 1,
  "highlight": {
    "fields": {
      "attachments.filename": {},
      "attachments.attachment.content": {}
    },
    "require_field_match": false
  },
  "_source": {
    "excludes": [
      "attachments.attachment.content",
      "attachments.data"
    ]
  }
}

```

but as stated above, this isnt what I need.

Thanks!

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 16, 2017, 3:20pm UTC](https://discuss.elastic.co/t/identifying-attachment-that-matches-query/78885/2 "2017-03-16T15:20:35Z")

</div>

Just to make sure I understand. You index multiple attachments within an array like:

```auto
{
  "attachments": [ {
      "attachment": {
        "content": "BASE64"
      },
      "filename": "file1.txt"
    },{
      "attachment": {
        "content": "BASE64"
      },
      "filename": "file2.txt"
    }
  ]
}

```

Am I correct? What is your mapping then?

---

<div class="post-metadata">

**Author:** ![jattenberg](https://avatars.discourse-cdn.com/v4/letter/j/b9e5f3/32.png) [@jattenberg](https://discuss.elastic.co/u/jattenberg)\
**Post date:** [March 16, 2017, 3:33pm UTC](https://discuss.elastic.co/t/identifying-attachment-that-matches-query/78885/3 "2017-03-16T15:33:08Z")

</div>

Not sure I completely understand the question, I just copied and pasted the attachment part of my mapping above.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 16, 2017, 3:42pm UTC](https://discuss.elastic.co/t/identifying-attachment-that-matches-query/78885/4 "2017-03-16T15:42:10Z")

</div>

Are you using `nested` for `attachments` field. I believe you don't but prefer asking.

If you want to have a relationship between the file content and the file name you need to index separated documents in Lucene.

Which means either:

- with separated documents in Elasticsearch (one per attachment, then copy all the email details in this attachment document)
- using nested documents: each attachment will be indexed internally as a separated document in Lucene alongside another document which will contain all the other fields.

---

<div class="post-metadata">

**Author:** ![jattenberg](https://avatars.discourse-cdn.com/v4/letter/j/b9e5f3/32.png) [@jattenberg](https://discuss.elastic.co/u/jattenberg)\
**Post date:** [March 16, 2017, 4:10pm UTC](https://discuss.elastic.co/t/identifying-attachment-that-matches-query/78885/5 "2017-03-16T16:10:59Z")

</div>

It seems the most straightfoward thing would be to use an explicitly nested type instead of just an array of attachments. I'll need to figure how to change my pipeline to handle this.

Given that attachments are nested, how would I change my above query to get the matched filenames?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 16, 2017, 4:24pm UTC](https://discuss.elastic.co/t/identifying-attachment-that-matches-query/78885/6 "2017-03-16T16:24:34Z")

</div>

I believe that if you have nested fields, you have to use a [nested query](https://www.elastic.co/guide/en/elasticsearch/reference/5.2/query-dsl-nested-query.html) and [nested inner hits](https://www.elastic.co/guide/en/elasticsearch/reference/5.2/search-request-inner-hits.html#nested-inner-hits) which will tell you which inner nested object matches so you can easily have access to its filename field.

---

<div class="post-metadata">

**Author:** ![jattenberg](https://avatars.discourse-cdn.com/v4/letter/j/b9e5f3/32.png) [@jattenberg](https://discuss.elastic.co/u/jattenberg)\
**Post date:** [March 16, 2017, 5:05pm UTC](https://discuss.elastic.co/t/identifying-attachment-that-matches-query/78885/7 "2017-03-16T17:05:03Z")

</div>

hmm this may not be sufficient, I was hoping that if I search `_all`, and the result happens to be in an attachment, I could know which one.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 16, 2017, 5:25pm UTC](https://discuss.elastic.co/t/identifying-attachment-that-matches-query/78885/8 "2017-03-16T17:25:31Z")

</div>

No. That's not possible. `_all` is basically a field where every single data has been put in.

You can only do that if you create one single elasticsearch document per attachment (which I'd recommend).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 13, 2017, 5:25pm UTC](https://discuss.elastic.co/t/identifying-attachment-that-matches-query/78885/9 "2017-04-13T17:25:31Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
