# Getting a "RDD element of type java.util.HashMap cannot be used" error

**URL:** <https://discuss.elastic.co/t/getting-a-rdd-element-of-type-java-util-hashmap-cannot-be-used-error/96134>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [August 7, 2017, 4:07pm UTC](https://discuss.elastic.co/t/getting-a-rdd-element-of-type-java-util-hashmap-cannot-be-used-error/96134 "2017-08-07T16:07:26Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Antonio\_Ye](https://avatars.discourse-cdn.com/v4/letter/a/e95f7d/32.png) [@Antonio\_Ye](https://discuss.elastic.co/u/Antonio_Ye)\
**Post date:** [August 7, 2017, 4:07pm UTC](https://discuss.elastic.co/t/getting-a-rdd-element-of-type-java-util-hashmap-cannot-be-used-error/96134/1 "2017-08-07T16:07:27Z")

</div>

I am using pyspark and I have an RDD of complex JSON strings that I converted to JSON using python's json.loads. When I try to save this to elasticsearch using rdd.saveAsNewAPIHadoopFile I get a "RDD element of type java.util.HashMap cannot be used" error.

Here is the code snipped:

jsonRdd = jsonStringRdd.map(lambda x : json.loads(x))  
jsoRdd.saveAsNewAPIHadoopFile(  
path='-',  
outputFormatClass="org.elasticsearch.hadoop.mr.EsOutputFormat",  
keyClass="org.apache.hadoop.io.NullWritable",  
valueClass="org.elasticsearch.hadoop.mr.LinkedMapWritable",  
conf={ "es.resource" : "test-index-rdd/posts" ,

Here is the error:

Traceback (most recent call last):  
File "", line 7, in   
File "/opt/spark/python/pyspark/rdd.py", line 1421, in saveAsNewAPIHadoopFile  
keyConverter, valueConverter, jconf)  
File "/opt/spark/python/lib/py4j-0.10.4-src.zip/py4j/java\_gateway.py", line 1133, in **call**  
File "/opt/spark/python/pyspark/sql/utils.py", line 63, in deco  
return f(\*a, \*\*kw)  
File "/opt/spark/python/lib/py4j-0.10.4-src.zip/py4j/protocol.py", line 319, in get\_return\_value  
py4j.protocol.Py4JJavaError: An error occurred while calling z:org.apache.spark.api.python.PythonRDD.saveAsNewAPIHadoopFile.  
: org.apache.spark.SparkException: RDD element of type java.util.HashMap cannot be used  
at org.apache.spark.api.python.SerDeUtil$.pythonToPairRDD(SerDeUtil.scala:238)  
at org.apache.spark.api.python.PythonRDD$.saveAsNewAPIHadoopFile(PythonRDD.scala:827)  
at org.apache.spark.api.python.PythonRDD.saveAsNewAPIHadoopFile(PythonRDD.scala)

Can anyone help me to solve this issue?

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [August 7, 2017, 4:14pm UTC](https://discuss.elastic.co/t/getting-a-rdd-element-of-type-java-util-hashmap-cannot-be-used-error/96134/2 "2017-08-07T16:14:46Z")

</div>

You may have to convert the hash map into a map writable object to use the NewAPIHadoop calls, since all interaction with the connector through that is handled by the map reduce code, which only works with writable objects.

---

<div class="post-metadata">

**Author:** ![Antonio\_Ye](https://avatars.discourse-cdn.com/v4/letter/a/e95f7d/32.png) [@Antonio\_Ye](https://discuss.elastic.co/u/Antonio_Ye)\
**Post date:** [August 7, 2017, 4:25pm UTC](https://discuss.elastic.co/t/getting-a-rdd-element-of-type-java-util-hashmap-cannot-be-used-error/96134/3 "2017-08-07T16:25:53Z")

</div>

Is there no easier way to do this? Can I somehow pass the string to ES and have it treat it as JSON and properly index it? I tried the following:

jsonStringRdd.saveAsNewAPIHadoopFile(  
path='-',  
outputFormatClass="org.elasticsearch.hadoop.mr.EsOutputFormat",  
keyClass="org.apache.hadoop.io.NullWritable",  
valueClass="org.elasticsearch.hadoop.mr.LinkedMapWritable",  
conf={ "es.resource" : "test-index-rdd/posts" ,  
"es.input.json": "true"  
})

but I got this error:

py4j.protocol.Py4JJavaError: An error occurred while calling z:org.apache.spark.api.python.PythonRDD.saveAsNewAPIHadoopFile.  
: org.apache.spark.SparkException: RDD element of type java.lang.String cannot be used

---

<div class="post-metadata">

**Author:** ![Antonio\_Ye](https://avatars.discourse-cdn.com/v4/letter/a/e95f7d/32.png) [@Antonio\_Ye](https://discuss.elastic.co/u/Antonio_Ye)\
**Post date:** [August 7, 2017, 5:04pm UTC](https://discuss.elastic.co/t/getting-a-rdd-element-of-type-java-util-hashmap-cannot-be-used-error/96134/4 "2017-08-07T17:04:34Z")

</div>

Tried a much simpler example and got the same error.

j = [{'a':1},{'b':2}]  
rdd = sc.parallelize(j)  
rdd.saveAsNewAPIHadoopFile(  
path='-',  
outputFormatClass="org.elasticsearch.hadoop.mr.EsOutputFormat",  
keyClass="org.apache.hadoop.io.NullWritable",  
valueClass="org.elasticsearch.hadoop.mr.LinkedMapWritable",  
conf={ "es.resource" : "test-index-rdd/posts" ,  
"es.input.json": "true"  
})

Traceback (most recent call last):  
File "", line 7, in   
File "/opt/spark/python/pyspark/rdd.py", line 1421, in saveAsNewAPIHadoopFile  
keyConverter, valueConverter, jconf)  
File "/opt/spark/python/lib/py4j-0.10.4-src.zip/py4j/java\_gateway.py", line 1133, in **call**  
File "/opt/spark/python/pyspark/sql/utils.py", line 63, in deco  
return f(\*a, \*\*kw)  
File "/opt/spark/python/lib/py4j-0.10.4-src.zip/py4j/protocol.py", line 319, in get\_return\_value  
py4j.protocol.Py4JJavaError: An error occurred while calling z:org.apache.spark.api.python.PythonRDD.saveAsNewAPIHadoopFile.  
: org.apache.spark.SparkException: RDD element of type java.util.HashMap cannot be used  
at org.apache.spark.api.python.SerDeUtil$.pythonToPairRDD(SerDeUtil.scala:238)  
at org.apache.spark.api.python.PythonRDD$.saveAsNewAPIHadoopFile(PythonRDD.scala:827)  
at org.apache.spark.api.python.PythonRDD.saveAsNewAPIHadoopFile(PythonRDD.scala)

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [August 7, 2017, 7:42pm UTC](https://discuss.elastic.co/t/getting-a-rdd-element-of-type-java-util-hashmap-cannot-be-used-error/96134/5 "2017-08-07T19:42:24Z")

</div>

In the case of using JSON strings, you would need to use a `Text` object instead of a `MapWritable`. The output format will only work with data objects that implement the Hadoop `Writable` contract. To use `String` and `Map` objects you will need to use the more extensive native support available in Scala and Java.

---

<div class="post-metadata">

**Author:** ![Antonio\_Ye](https://avatars.discourse-cdn.com/v4/letter/a/e95f7d/32.png) [@Antonio\_Ye](https://discuss.elastic.co/u/Antonio_Ye)\
**Post date:** [August 8, 2017, 8:16am UTC](https://discuss.elastic.co/t/getting-a-rdd-element-of-type-java-util-hashmap-cannot-be-used-error/96134/6 "2017-08-08T08:16:43Z")

</div>

So there is no pyspark equivalent to rdd.saveJsonToEs(...) ?

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [August 8, 2017, 8:17pm UTC](https://discuss.elastic.co/t/getting-a-rdd-element-of-type-java-util-hashmap-cannot-be-used-error/96134/7 "2017-08-08T20:17:27Z")

</div>

Unfortunately, not at this time. You may be able to tap into the native support by using the Spark SQL functionality with PySpark, and specifying an Elasticsearch datasource (as described later in our documentation about PySpark).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 5, 2017, 8:17pm UTC](https://discuss.elastic.co/t/getting-a-rdd-element-of-type-java-util-hashmap-cannot-be-used-error/96134/8 "2017-09-05T20:17:52Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
