# Loading JSON documents to elasticsearch via es-spark connector

**URL:** <https://discuss.elastic.co/t/loading-json-documents-to-elasticsearch-via-es-spark-connector/141616>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [July 25, 2018, 3:41pm UTC](https://discuss.elastic.co/t/loading-json-documents-to-elasticsearch-via-es-spark-connector/141616 "2018-07-25T15:41:14Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![rkalluri](https://avatars.discourse-cdn.com/v4/letter/r/fbc32d/32.png) [@rkalluri](https://discuss.elastic.co/u/rkalluri)\
**Post date:** [July 25, 2018, 3:41pm UTC](https://discuss.elastic.co/t/loading-json-documents-to-elasticsearch-via-es-spark-connector/141616/1 "2018-07-25T15:41:15Z")

</div>

If I have a spark dataframe full of JSON documents with the following schema, does the es-hadoop connector allow indexing them to elastic.

scala\> df.printSchema  
root  
|-- address: array (nullable = true)  
| |-- element: struct (containsNull = true)  
| | |-- location: string (nullable = true)  
| | |-- std\_city: string (nullable = true)  
| | |-- std\_state: string (nullable = true)  
| | |-- std\_street\_name: string (nullable = true)  
| | |-- std\_street\_number: string (nullable = true)  
| | |-- std\_zip: string (nullable = true)  
|-- date\_of\_birth: array (nullable = true)  
| |-- element: struct (containsNull = true)  
| | |-- dob\_day: string (nullable = true)  
| | |-- dob\_full: string (nullable = true)  
| | |-- dob\_month: string (nullable = true)  
| | |-- dob\_year: string (nullable = true)  
|-- names: array (nullable = true)  
| |-- element: struct (containsNull = true)  
| | |-- first\_name: string (nullable = true)  
| | |-- first\_name\_list: string (nullable = true)  
| | |-- fn\_formalname\_assocs: string (nullable = true)  
| | |-- fn\_nickname\_assocs: string (nullable = true)  
| | |-- last\_name: string (nullable = true)  
| | |-- last\_name\_list: string (nullable = true)  
| | |-- ln\_formalname\_assocs: string (nullable = true)  
| | |-- ln\_nickname\_assocs: string (nullable = true)  
| | |-- middle\_name: string (nullable = true)  
| | |-- middle\_name\_list: string (nullable = true)  
| | |-- mn\_formalname\_assocs: string (nullable = true)  
| | |-- mn\_nickname\_assocs: string (nullable = true)  
|-- phone: array (nullable = true)  
| |-- element: string (containsNull = true)  
|-- pm\_id: string (nullable = true)  
|-- id: string (nullable = true)

---

<div class="post-metadata">

**Author:** ![rkalluri](https://avatars.discourse-cdn.com/v4/letter/r/fbc32d/32.png) [@rkalluri](https://discuss.elastic.co/u/rkalluri)\
**Post date:** [July 25, 2018, 3:51pm UTC](https://discuss.elastic.co/t/loading-json-documents-to-elasticsearch-via-es-spark-connector/141616/2 "2018-07-25T15:51:10Z")

</div>

I will have a mapping in elastic, to match the above. Also is there a way to specify the id field from the json document as the \_id for elastic

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [July 25, 2018, 4:04pm UTC](https://discuss.elastic.co/t/loading-json-documents-to-elasticsearch-via-es-spark-connector/141616/3 "2018-07-25T16:04:50Z")

</div>

This doesn't look like it would be hard to ingest with ES-Hadoop. If you haven't had a chance to take a look at our [Spark documentation](https://www.elastic.co/guide/en/elasticsearch/hadoop/current/spark.html#spark-sql) I recommend it.

> is there a way to specify the id field from the json document as the \_id for elastic

You're looking for the `es.mapping.id` setting for that: [Configuration | Elasticsearch for Apache Hadoop [8.11] | Elastic](https://www.elastic.co/guide/en/elasticsearch/hadoop/current/configuration.html#cfg-mapping)

---

<div class="post-metadata">

**Author:** ![rkalluri](https://avatars.discourse-cdn.com/v4/letter/r/fbc32d/32.png) [@rkalluri](https://discuss.elastic.co/u/rkalluri)\
**Post date:** [July 25, 2018, 5:00pm UTC](https://discuss.elastic.co/t/loading-json-documents-to-elasticsearch-via-es-spark-connector/141616/4 "2018-07-25T17:00:50Z")

</div>

Thanks James, It did work fine. I have used the connector for a long time now and a big fan. Thanks for all the hard work.

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [July 25, 2018, 6:22pm UTC](https://discuss.elastic.co/t/loading-json-documents-to-elasticsearch-via-es-spark-connector/141616/5 "2018-07-25T18:22:40Z")

</div>

Happy to hear! Cheers!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 22, 2018, 6:32pm UTC](https://discuss.elastic.co/t/loading-json-documents-to-elasticsearch-via-es-spark-connector/141616/6 "2018-08-22T18:32:30Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
