# Bulk API via S3

**URL:** <https://discuss.elastic.co/t/bulk-api-via-s3/42926>\
**Category:** Elasticsearch\
**Created:** [February 27, 2016, 7:46pm UTC](https://discuss.elastic.co/t/bulk-api-via-s3/42926 "2016-02-27T19:46:19Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jonathan\_Spooner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jonathan_spooner/32/2672_2.png) [@Jonathan\_Spooner](https://discuss.elastic.co/u/Jonathan_Spooner)\
**Post date:** [February 27, 2016, 7:46pm UTC](https://discuss.elastic.co/t/bulk-api-via-s3/42926/1 "2016-02-27T19:46:19Z")

</div>

I'm using Spark to build text files for the [Bulk API](https://www.elastic.co/guide/en/elasticsearch/reference/1.4/docs-bulk.html) and I staged all 80GB of them in S3. I have 45gb of data and I just realized I can't use curl to POST a file that is on S3!

A) I can write a bash script to download the file and POST it to ES. This is not optimal and I can's ssh into one of the hosted ES nodes and run it from there.

B) I could stream the data with [lambda](http://docs.aws.amazon.com/elasticsearch-service/latest/developerguide/es-aws-integrations.html) but the data is already in bulk format and I don't want to stream at this time.

I want to bulk load my data set and play with different ES configurations to find the correct cluster size.

I know there are many alternatives but I'd like to see if it's possible to use the bulk api from S3.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [February 28, 2016, 3:21am UTC](https://discuss.elastic.co/t/bulk-api-via-s3/42926/2 "2016-02-28T03:21:31Z")

</div>

You typically want to keep the size of your bulk requests to around a few MB in size or smaller, so sending a huge bilk file in a single request is not recommended. Creating a script to read the files and break them up into smaller chunks is one option. You could also use Logstash as it has an [S3 input plugin](https://www.elastic.co/guide/en/logstash/current/plugins-inputs-s3.html), although you would need to process the data and transform it back into simple records as the Elasticsearch output builds bulk requests internally.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:13pm UTC](https://discuss.elastic.co/t/bulk-api-via-s3/42926/3 "2017-07-05T23:13:05Z")

</div>


