# Error while trying to insert data to elasticsearch using hive elasticsearch storage handler from spark

**URL:** <https://discuss.elastic.co/t/error-while-trying-to-insert-data-to-elasticsearch-using-hive-elasticsearch-storage-handler-from-spark/147902>\
**Category:** Elasticsearch\
**Created:** [September 10, 2018, 7:28am UTC](https://discuss.elastic.co/t/error-while-trying-to-insert-data-to-elasticsearch-using-hive-elasticsearch-storage-handler-from-spark/147902 "2018-09-10T07:28:48Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Shiva\_Dorai](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/shiva_dorai/32/35354_2.png) [@Shiva\_Dorai](https://discuss.elastic.co/u/Shiva_Dorai)\
**Post date:** [September 10, 2018, 7:28am UTC](https://discuss.elastic.co/t/error-while-trying-to-insert-data-to-elasticsearch-using-hive-elasticsearch-storage-handler-from-spark/147902/1 "2018-09-10T07:28:49Z")

</div>

I have an hive interface to elasticsearch -

```auto
CREATE EXTERNAL TABLE IF NOT EXISTS analytics_metrics_workflow_es(
    workflow_id string,
    workflow_name string)
    ROW FORMAT SERDE 'org.elasticsearch.hadoop.hive.EsSerDe'
    STORED BY 'org.elasticsearch.hadoop.hive.EsStorageHandler'
    TBLPROPERTIES('es.nodes' = 'xxxx','es.port' = '9200','es.resource' = 'analytics_metrics_workflow/data', 'es.index.auto.create' = 'true', 'es.mapping.id' = 'workflow_id');

```

I have another hive table as source -

```auto
CREATE EXTERNAL TABLE IF NOT EXISTS analytics_metrics_workflow(
    workflow_id string,
    workflow_name string)
    ROW FORMAT DELIMITED FIELDS TERMINATED BY '\t'
    LOCATION '/user/sd/analytics_metrics_workflow_sd';

```

When I insert data from hive shell -  
`insert into analytics_metrics_workflow_es select workflow_id,workflow_name from analytics_metrics_workflow`. This works.

But when I use `spark.sql("insert into analytics_metrics_workflow_es select workflow_id,workflow_name from analytics_metrics_workflow")`, it throws an error -

```auto
scala> spark.sql("insert into analytics_metrics_workflow_es select workflow_id,workflow_name from analytics_metrics_workflow");
[Stage 0:> (0 + 2) / 2]18/09/10 07:16:46 ERROR Utils: Aborting task
java.lang.RuntimeException: cannot find field _col0 from [0:workflow_id, 1:workflow_name]
	at org.apache.hadoop.hive.serde2.objectinspector.ObjectInspectorUtils.getStandardStructFieldRef(ObjectInspectorUtils.java:416)
	at org.apache.hadoop.hive.serde2.objectinspector.StandardStructObjectInspector.getStructFieldRef(StandardStructObjectInspector.java:147)
	at org.elasticsearch.hadoop.hive.HiveFieldExtractor.extractField(HiveFieldExtractor.java:52)
	at org.elasticsearch.hadoop.serialization.field.ConstantFieldExtractor.field(ConstantFieldExtractor.java:36)
	at org.elasticsearch.hadoop.serialization.bulk.AbstractBulkFactory$FieldWriter.write(AbstractBulkFactory.java:103)
	at org.elasticsearch.hadoop.serialization.bulk.TemplatedBulk.writeTemplate(TemplatedBulk.java:80)
	at org.elasticsearch.hadoop.serialization.bulk.TemplatedBulk.write(TemplatedBulk.java:56)
	at org.elasticsearch.hadoop.hive.EsSerDe.serialize(EsSerDe.java:163)
	at org.apache.spark.sql.hive.execution.HiveOutputWriter.write(HiveFileFormat.scala:153)
	at org.apache.spark.sql.execution.datasources.FileFormatWriter$SingleDirectoryWriteTask.execute(FileFormatWriter.scala:392)
	at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask$3.apply(FileFormatWriter.scala:269)
	at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask$3.apply(FileFormatWriter.scala:267)
	at org.apache.spark.util.Utils$.tryWithSafeFinallyAndFailureCallbacks(Utils.scala:1411)
	at org.apache.spark.sql.execution.datasources.FileFormatWriter$.org$apache$spark$sql$execution$datasources$FileFormatWriter$$executeTask(FileFormatWriter.scala:272)
	at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:197)
	at org.apache.spark.sql.execution.datasources.FileFormatWriter$$anonfun$write$1.apply(FileFormatWriter.scala:196)
	at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:87)
	at org.apache.spark.scheduler.Task.run(Task.scala:109)
	at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:345)
	at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)
	at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)
	at java.lang.Thread.run(Thread.java:745)

```

Any help would be highly appreciable.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 8, 2018, 7:38am UTC](https://discuss.elastic.co/t/error-while-trying-to-insert-data-to-elasticsearch-using-hive-elasticsearch-storage-handler-from-spark/147902/2 "2018-10-08T07:38:28Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
