# Staging Data for Elasticsearch bulk loading

**URL:** <https://discuss.elastic.co/t/staging-data-for-elasticsearch-bulk-loading/59453>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [August 31, 2016, 5:45pm UTC](https://discuss.elastic.co/t/staging-data-for-elasticsearch-bulk-loading/59453 "2016-08-31T17:45:57Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![jspooner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jspooner/32/12984_2.png) [@jspooner](https://discuss.elastic.co/u/jspooner)\
**Post date:** [August 31, 2016, 5:45pm UTC](https://discuss.elastic.co/t/staging-data-for-elasticsearch-bulk-loading/59453/1 "2016-08-31T17:45:57Z")

</div>

I'm looking for the best storage format to stage data for elasticsearch ingestion. By best I mean the storage format that will require little or no spark processing time. Ideally this Spark job would load data and push directly to elasticsearch.

Why do I want a Spark job that just loads data into elasticsearch? I have a 3TB dataset and I'd like to load a small percentage of it multiple times so we can test different es configurations and performance. After we have a configuration we're pleased with we'll likely have a daily job that summarizes and pushes to elasticsearch.

I currently have an aggregation job that converts raw events into summaries and writes to S3 in Parquet format. I then have a second job to reads that data and transform it into elastic bulk format before using es-hadoop.

Example Job that pushes to elasticsearch:

```
val dataframe = sqlContext.read.parquet(path)
val myDF = dataframe.select("device_id").distinct
val parentRDD = myDF.map( x=>
  (
    Map(ID -> x(0).toString.trim),
    Map()
  )
)
val childrenRDD = dataframe.map( x=>
  (
    Map(
      ID -> x(0),
      PARENT -> x(1)
    ),
    Map(
      "foo" -> x(0),
      "foo" -> x(2)
    )
  )
)
EsSpark.saveToEsWithMeta(parentRDD, "development/parent")
EsSpark.saveToEsWithMeta(childrenRDD, "development/child")

```

My problem with this example is the map action can take a considerable amount of time. Ideally I'd like to stage parentRDD and childrenRDD to file but I'm not sure what the best format would be?

I like using es-hadoop for pushing to elasticsearch because of it's configuration for bulk file size and retry policy but would be open to using something else.

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [August 31, 2016, 11:20pm UTC](https://discuss.elastic.co/t/staging-data-for-elasticsearch-bulk-loading/59453/2 "2016-08-31T23:20:07Z")

</div>

My advice in this scenario would be to store the data already formatted as JSON and to use the `es.input.json` property to load the data in. Make sure that you only have one document per line in the file. There's a [section in our docs](https://www.elastic.co/guide/en/elasticsearch/hadoop/current/spark.html#spark-write-json) that covers how to do this.

Loading with `es.input.json` enabled minimizes the amount of the serialization work ES-Hadoop has to do. For your use case, it allows you to remove some processing time from Spark during indexing.

As for file format, you're really just looking for a format that has good IO performance. Parquet should do just fine.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:23pm UTC](https://discuss.elastic.co/t/staging-data-for-elasticsearch-bulk-loading/59453/3 "2017-07-06T13:23:28Z")

</div>


