# How to ingest/push multiple BIG attachments into indexed document's array attribute in ElasticSearch?

**URL:** <https://discuss.elastic.co/t/how-to-ingest-push-multiple-big-attachments-into-indexed-documents-array-attribute-in-elasticsearch/140781>\
**Category:** Elasticsearch\
**Created:** [July 19, 2018, 5:30pm UTC](https://discuss.elastic.co/t/how-to-ingest-push-multiple-big-attachments-into-indexed-documents-array-attribute-in-elasticsearch/140781 "2018-07-19T17:30:21Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![mbm-rafal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mbm-rafal/32/33539_2.png) [@mbm-rafal](https://discuss.elastic.co/u/mbm-rafal)\
**Post date:** [July 19, 2018, 5:30pm UTC](https://discuss.elastic.co/t/how-to-ingest-push-multiple-big-attachments-into-indexed-documents-array-attribute-in-elasticsearch/140781/1 "2018-07-19T17:30:21Z")

</div>

Hey,

using: ES 5.1

I'm trying to figure out the solution for indexing multiple attachments for single document into Elasticsearch Index.

I have some limitations around my service (using AWS) that limits my HTTP request up to 100MB per single POST.

Basically I have profiles in ES index, and for each profile I want to store multiple searchable attachments, let's say up to 50 x 10MB pdfs  
That requirement limits my approach because I just cannot send to ES of total 500 MB of data.

One of the approach was to make some kind of partial updates, but still how to make 'partial-Update' by pushing NEW attachment to the existing attachments' array?  
Maybe some flatten attachments index and reference to my main index to find profiles out?

I have to also support highliting in result, so the best approach for me is to have mapping like this:

```
{
  "directory.index.v7": {
    "mappings": {
      "profile.event": {
        "properties": {
          "attachments": {
            "properties": {
              "attachment": {
                "properties": {
                  "content": {
                    "type": "text",
                    "fields": {
                      "keyword": {
                        "type": "keyword",
                        "ignore_above": 256
                      }
                    }
                  },
                  "content_length": {
                    "type": "long"
                  },
                  "content_type": {
                    "type": "text",
                    "fields": {
                      "keyword": {
                        "type": "keyword",
                        "ignore_above": 256
                      }
                    }
                  },
                  "date": {
                    "type": "date"
                  },
                  "language": {
                    "type": "text",
                    "fields": {
                      "keyword": {
                        "type": "keyword",
                        "ignore_above": 256
                      }
                    }
                  }
                }
              },
              "data": {
                "type": "text",
                "fields": {
                  "keyword": {
                    "type": "keyword",
                    "ignore_above": 256
                  }
                }
              },
              "filename": {
                "type": "text",
                "fields": {
                  "keyword": {
                    "type": "keyword",
                    "ignore_above": 256
                  }
                }
              }
            }
          },
          "email": {
            "type": "text",
            "fields": {
              "raw": {
                "type": "keyword"
              }
            }
          }
        }
      }
    }
  }
}

```

But as I meanioned before:

- I cannot ingest all attachment at once.
- I don know how and if is possible to make ATTACHMENT PUSH to `attachments` attribute without including older docs (to not reach a limit for POST)

Please advise!

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 19, 2018, 8:28pm UTC](https://discuss.elastic.co/t/how-to-ingest-push-multiple-big-attachments-into-indexed-documents-array-attribute-in-elasticsearch/140781/2 "2018-07-19T20:28:28Z")

</div>

Instead of indexing one array of attachments, why not indexing individual attachments one by one?

---

<div class="post-metadata">

**Author:** ![mbm-rafal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mbm-rafal/32/33539_2.png) [@mbm-rafal](https://discuss.elastic.co/u/mbm-rafal)\
**Post date:** [July 20, 2018, 9:20am UTC](https://discuss.elastic.co/t/how-to-ingest-push-multiple-big-attachments-into-indexed-documents-array-attribute-in-elasticsearch/140781/3 "2018-07-20T09:20:57Z")

</div>

This is not a collection of attachments though, but of "profiles-having-attachments" I want to find PROFILES by searching through attachments array.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 20, 2018, 5:28pm UTC](https://discuss.elastic.co/t/how-to-ingest-push-multiple-big-attachments-into-indexed-documents-array-attribute-in-elasticsearch/140781/4 "2018-07-20T17:28:05Z")

</div>

If for each attachment, you store the profile information, then you can may be retrieve that information...

Otherwise, I'd recommend doing the extraction on your side before sending the data to elasticsearch. That way json documents will not be huge BASE64 content but just extracted text.

Similar to what FSCrawler is doing. BTW you can use it and its REST layer to simulate an upload and get back the extracted text... Might help.

---

<div class="post-metadata">

**Author:** ![mbm-rafal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mbm-rafal/32/33539_2.png) [@mbm-rafal](https://discuss.elastic.co/u/mbm-rafal)\
**Post date:** [July 20, 2018, 5:55pm UTC](https://discuss.elastic.co/t/how-to-ingest-push-multiple-big-attachments-into-indexed-documents-array-attribute-in-elasticsearch/140781/5 "2018-07-20T17:55:30Z")

</div>

Currently I'm trying to PARSE docs on my end, to get text only (40-100MB) docs are producing actually up to 500kB of text. But still having array forces me to try parent-child approach instead of nested field

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 21, 2018, 5:43am UTC](https://discuss.elastic.co/t/how-to-ingest-push-multiple-big-attachments-into-indexed-documents-array-attribute-in-elasticsearch/140781/6 "2018-07-21T05:43:26Z")

</div>

You can change the elasticsearch settings may be to allow more than 100mb per request over the wire?

> I have some limitations around my service (using AWS) that limits my HTTP request up to 100MB per single POST.

Are you using elasticsearch as a service by AWS?  
Or just EC2 instances where you deployed yourself elasticsearch?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 18, 2018, 5:43am UTC](https://discuss.elastic.co/t/how-to-ingest-push-multiple-big-attachments-into-indexed-documents-array-attribute-in-elasticsearch/140781/7 "2018-08-18T05:43:30Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
