# Spark elasticsearch 5.0.2 scala.MatchError

**URL:** <https://discuss.elastic.co/t/spark-elasticsearch-5-0-2-scala-matcherror/68364>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [December 7, 2016, 9:10pm UTC](https://discuss.elastic.co/t/spark-elasticsearch-5-0-2-scala-matcherror/68364 "2016-12-07T21:10:20Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![markcitizen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/markcitizen/32/117307_2.png) [@markcitizen](https://discuss.elastic.co/u/markcitizen)\
**Post date:** [December 7, 2016, 9:10pm UTC](https://discuss.elastic.co/t/spark-elasticsearch-5-0-2-scala-matcherror/68364/1 "2016-12-07T21:10:20Z")

</div>

Hello,  
I'm using Spark 2.0 with Spark Elasticsearch version 5.0.2 (Scala 2.11):  
"org.elasticsearch" % "elasticsearch-spark-20\_2.11" % "5.0.2"

When trying to read from an index I'm getting the following error:  
scala.MatchError: Buffer() (of class scala.collection.convert.Wrappers$JListWrapper)

This is how I'm getting the data (session is an instance of SparkSession):  
session.sqlContext.read.format("org.elasticsearch.spark.sql").options(opt).load(esIndexName)

opt is a Map of options:  
Map("es.read.field.as.array.include" -\> "fieldNames",  
"es.input.json" -\> "true",  
"es.field.read.empty.as.null" -\> "true",  
"es.index.read.missing.as.empty" -\> "true",  
"es.read.field.exclude" -\> excludedFields)

Field is question is not in an array, it's a nested object field:  
a {  
b {  
c = "string value"  
}  
}

I found a similar issue here:

> <https://github.com/elastic/elasticsearch-hadoop/issues/661>
>
> Hello,
> I'm using the connector with Spark, and I'm trying to read fields that h…as arrays and Strings, like this:
> 
> Field\_name
> \["value"\]
> value
> 
> when I set 
> es.field.read.as.array.include = Field\_name
> It thows me an error that some fields are not arrays, because some of them are String:
> 
> Caused by: scala.MatchError: value (of class java.lang.String)
> 
> And when I set
> es.field.read.as.array.exclude = Field\_name
> It thows me an error that some fields are not Strings, because some of them are arrays:
> 
> Caused by: scala.MatchError: Buffer(\["value"\]) (of class scala.collection.convert.Wrappers$JListWrapper)
> 
> How can I solve this? My goal is to write indexes of ES into parquet or json files using Spark.
> 
> Regards

But after upgrading to the latest spark-elasticsearch library version I'm still seeing the problem.  
I would appreciate any suggestions on how to fix this.  
Thanks,

M

Spark stack trace:

> > >

WARN TaskSetManager: Lost task 1.0 in stage 0.0 (TID 1, host): scala.MatchError: Buffer() (of class scala.collection.convert.Wrappers$JListWrapper)  
at org.apache.spark.sql.catalyst.CatalystTypeConverters$StringConverter$.toCatalystImpl(CatalystTypeConverters.scala:296)  
at org.apache.spark.sql.catalyst.CatalystTypeConverters$StringConverter$.toCatalystImpl(CatalystTypeConverters.scala:295)  
at org.apache.spark.sql.catalyst.CatalystTypeConverters$CatalystTypeConverter.toCatalyst(CatalystTypeConverters.scala:103)  
at org.apache.spark.sql.catalyst.CatalystTypeConverters$StructConverter.toCatalystImpl(CatalystTypeConverters.scala:261)  
at org.apache.spark.sql.catalyst.CatalystTypeConverters$StructConverter.toCatalystImpl(CatalystTypeConverters.scala:251)  
at org.apache.spark.sql.catalyst.CatalystTypeConverters$CatalystTypeConverter.toCatalyst(CatalystTypeConverters.scala:103)  
at org.apache.spark.sql.catalyst.CatalystTypeConverters$StructConverter.toCatalystImpl(CatalystTypeConverters.scala:261)  
at org.apache.spark.sql.catalyst.CatalystTypeConverters$StructConverter.toCatalystImpl(CatalystTypeConverters.scala:251)  
at org.apache.spark.sql.catalyst.CatalystTypeConverters$CatalystTypeConverter.toCatalyst(CatalystTypeConverters.scala:103)  
at org.apache.spark.sql.catalyst.CatalystTypeConverters$ArrayConverter$$anonfun$toCatalystImpl$2.apply(CatalystTypeConverters.scala:164)  
at scala.collection.TraversableLike$$anonfun$map$1.apply(TraversableLike.scala:234)  
\<\<\<

---

<div class="post-metadata">

**Author:** ![markcitizen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/markcitizen/32/117307_2.png) [@markcitizen](https://discuss.elastic.co/u/markcitizen)\
**Post date:** [December 12, 2016, 6:31pm UTC](https://discuss.elastic.co/t/spark-elasticsearch-5-0-2-scala-matcherror/68364/2 "2016-12-12T18:31:28Z")

</div>

Hello,  
I'm going to answer my own question.  
I was not able to find a solution to this problem. I found some answers online but they referred to older versions of spark-elasticsearch library, and they said that the problem was supposed to be fixed in the latest version.  
Seeing how the latest version was still buggy I decided to skip the conversion process altogether and read ES index as JSON, and parse the data myself.

You can do that using code similar to this one (Scala):  
val readCfg = Map("setting" -\> "value")  
val tuples = session.sparkContext.esJsonRDD("myEsIndex", readCfg)

"tuples" is a collection of Tuple2[String, String] items, where the first one is the entry index and the second one is the entry body (JSON text). You can parse JSON text using Playframework JSON library, for example.  
I hope this helps,

M

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 9, 2017, 6:31pm UTC](https://discuss.elastic.co/t/spark-elasticsearch-5-0-2-scala-matcherror/68364/3 "2017-01-09T18:31:40Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
