# Elasticsearch-spark-30 read missing field(double type) error

**URL:** <https://discuss.elastic.co/t/elasticsearch-spark-30-read-missing-field-double-type-error/320112>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [November 30, 2022, 6:56am UTC](https://discuss.elastic.co/t/elasticsearch-spark-30-read-missing-field-double-type-error/320112 "2022-11-30T06:56:27Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![918246180](https://avatars.discourse-cdn.com/v4/letter/9/eb9ed0/32.png) [@918246180](https://discuss.elastic.co/u/918246180)\
**Post date:** [November 30, 2022, 6:56am UTC](https://discuss.elastic.co/t/elasticsearch-spark-30-read-missing-field-double-type-error/320112/1 "2022-11-30T06:56:27Z")

</div>

![144825](https://us1.discourse-cdn.com/elastic/original/3X/7/9/796123293991dcbe0d78b93eea488e1cf62ecf4a.jpeg)  
hi!  
When I use the elasticsearch-spark-30\_2.12-7.16.1.jar to read data, I find that a field of double type is missing, which causes the exception of the Spark application. But I use elasticsearch-spark20\_ 2.11-7.6.1. jar can successfully obtain data. Will this problem be fixed?

---

<div class="post-metadata">

**Author:** ![Keith\_Massey](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/keith_massey/32/83666_2.png) [@Keith\_Massey](https://discuss.elastic.co/u/Keith_Massey)\
**Post date:** [November 30, 2022, 2:15pm UTC](https://discuss.elastic.co/t/elasticsearch-spark-30-read-missing-field-double-type-error/320112/2 "2022-11-30T14:15:08Z")

</div>

You'll have to provide more information. What version of spark are you using (you are using es-spark jars for two completely different versions of spark, so I'm surprised that one of them works at all)? What version of Elasticsearch are you using? Can you provide code (including mappings and data) to reproduce this? And can you paste the stack trace?

---

<div class="post-metadata">

**Author:** ![918246180](https://avatars.discourse-cdn.com/v4/letter/9/eb9ed0/32.png) [@918246180](https://discuss.elastic.co/u/918246180)\
**Post date:** [December 1, 2022, 7:01am UTC](https://discuss.elastic.co/t/elasticsearch-spark-30-read-missing-field-double-type-error/320112/3 "2022-12-01T07:01:12Z")

</div>

### Issue description

First, I set a field value of double type to be an empty string.  
Then when I use Spark to read this field, it will throw an exception(**Cannot parse value [] for field [field name]**)  
I learned that when reading empty, it can be set to null. Therefore, I set **es.field.read.empty.as.null** to **true**. Unfortunately, another exception was thrown( **scala.None$ is not a valid external type for schema of double** ), and I was stuck in the loop  
I realize that this may be a bug. Please check

### Steps to reproduce

ps:index field **"C":{"type":double"}**

```java
 session.read().format("org.elasticsearch.spark.sql")
            .option("es.field.read.empty.as.null","true")
            .load(indexName)

```

### Strack trace:

**es.field.read.empty.as.null = true**  
Strack trace:  
Caused by:java.lang.RuntimeException:Error while encoding: java.lang.RuntimeException: scala.None$ is not a valid external type for schema of double  
at org.apache.spark.sql.error.QueryExecutionErrors$.expressionEncodingError(QueryExecutionErrors.scala:1052)  
at org.apache.spark.sql.catalyst.encoders.ExpressionEncoder$Serializer.apply(ExpressionEncoder.scala:210)  
at org.apache.spark.sql.catalyst.encoders.ExpressionEncoder$Serializer.apply(ExpressionEncoder.scala:193)  
at scala.collection.Iterator$$anon$10.next(Iterator.scala:459)  
... 19 more

**es.field.read.empty.as.null = false**  
Strack trace:  
Caused by: crest.EsHadoopParsingException: Cannot parse value for field[C]  
at org.elasticsearch.hadoop.serialization.ScrollReader.read(ScrollReader.java:903)  
at org.elasticsearch.hadoop.serialization.ScrollReader.map(ScrollReader.java:1047)  
at org.elasticsearch.hadoop.serialization.ScrollReader.read(ScrollReader.java:889)  
at org.elasticsearch.hadoop.serialization.ScrollReader.readHitAsMap(ScrollReader.java:602)  
... more

```auto

### Version Info

OS: : win7 64bit
JVM : jdk1.8
Hadoop/Spark: spark-sql_2.12-3.2.1.jar
ES-Hadoop : elasticsearch-spark-30_2.12-7.16.1.jar（There is no such problem when I use elasticsearch-spark-20_2.11-7.6.1.jar）
ES : 6.8.14(and I made the same mistake when using 7.17.0)
```

---

<div class="post-metadata">

**Author:** ![Keith\_Massey](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/keith_massey/32/83666_2.png) [@Keith\_Massey](https://discuss.elastic.co/u/Keith_Massey)\
**Post date:** [December 1, 2022, 2:56pm UTC](https://discuss.elastic.co/t/elasticsearch-spark-30-read-missing-field-double-type-error/320112/4 "2022-12-01T14:56:40Z")

</div>

I haven't been able to reproduce it. Here's what I did in my [test docker container](https://github.com/masseyke/es-spark-docker). It is using Elasticsearch 8.1.0 and Spark 3.2.1.  
First, I created the mapping and data in Elasticsearch:

```auto
curl -X PUT "localhost:9200/test?pretty" -H 'Content-Type: application/json' -d'
{
  "mappings": {
    "properties": {
      "C":{"type":"double"}
    }
  }
}
curl -X POST localhost:9200/test/_doc/ -H 'Content-Type: application/json' -d'
{
  "C": ""
}
curl -X POST localhost:9200/test/_doc/ -H 'Content-Type: application/json' -d'
{
  "C": "1.5"
}

```

Then ran `es-spark` and did this:

```auto
import org.apache.spark.sql._
val sqc = new SQLContext(sc)
sqc.read.format("es").option("es.field.read.empty.as.null","true").load("test").show
+----+                                                                          
| C|
+----+
|null|
| 1.5|
+----+

```

---

<div class="post-metadata">

**Author:** ![Keith\_Massey](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/keith_massey/32/83666_2.png) [@Keith\_Massey](https://discuss.elastic.co/u/Keith_Massey)\
**Post date:** [December 1, 2022, 3:01pm UTC](https://discuss.elastic.co/t/elasticsearch-spark-30-read-missing-field-double-type-error/320112/5 "2022-12-01T15:01:08Z")

</div>

If I add an empty string to an array of values for the field:

```auto
elastic@localhost:~$ curl -X POST localhost:9200/test/_doc/ -H 'Content-Type: application/json' -d'
{
  "C": ["2.5", ""]
}
'

```

I get this exception:

```auto
if (assertnotnull(input[0, org.apache.spark.sql.Row, true]).isNullAt) null else validateexternaltype(getexternalrowfield(assertnotnull(input[0, org.apache.spark.sql.Row, true]), 0, C), DoubleType) AS C#5
  at org.apache.spark.sql.errors.QueryExecutionErrors$.expressionEncodingError(QueryExecutionErrors.scala:1052)
  at org.apache.spark.sql.catalyst.encoders.ExpressionEncoder$Serializer.apply(ExpressionEncoder.scala:210)
  at org.apache.spark.sql.catalyst.encoders.ExpressionEncoder$Serializer.apply(ExpressionEncoder.scala:193)
  at scala.collection.Iterator$$anon$10.next(Iterator.scala:461)
  at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.processNext(Unknown Source)
  at org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)
  at org.apache.spark.sql.execution.WholeStageCodegenExec$$anon$1.hasNext(WholeStageCodegenExec.scala:759)
  at org.apache.spark.sql.execution.SparkPlan.$anonfun$getByteArrayRdd$1(SparkPlan.scala:349)
  at org.apache.spark.rdd.RDD.$anonfun$mapPartitionsInternal$2(RDD.scala:898)
  at org.apache.spark.rdd.RDD.$anonfun$mapPartitionsInternal$2$adapted(RDD.scala:898)
  at org.apache.spark.rdd.MapPartitionsRDD.compute(MapPartitionsRDD.scala:52)
  at org.apache.spark.rdd.RDD.computeOrReadCheckpoint(RDD.scala:373)
  at org.apache.spark.rdd.RDD.iterator(RDD.scala:337)
  at org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:90)
  at org.apache.spark.scheduler.Task.run(Task.scala:131)
  at org.apache.spark.executor.Executor$TaskRunner.$anonfun$run$3(Executor.scala:506)
  at org.apache.spark.util.Utils$.tryWithSafeFinally(Utils.scala:1462)
  at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:509)
  at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)
  at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)
  at java.lang.Thread.run(Thread.java:748)
Caused by: java.lang.RuntimeException: scala.collection.convert.Wrappers$JListWrapper is not a valid external type for schema of double
  at org.apache.spark.sql.catalyst.expressions.GeneratedClass$SpecificUnsafeProjection.If_0$(Unknown Source)
  at org.apache.spark.sql.catalyst.expressions.GeneratedClass$SpecificUnsafeProjection.apply(Unknown Source)
  at org.apache.spark.sql.catalyst.encoders.ExpressionEncoder$Serializer.apply(ExpressionEncoder.scala:207)
  ... 19 more

```

But that looks different from what you are seeing.

---

<div class="post-metadata">

**Author:** ![918246180](https://avatars.discourse-cdn.com/v4/letter/9/eb9ed0/32.png) [@918246180](https://discuss.elastic.co/u/918246180)\
**Post date:** [December 2, 2022, 1:41am UTC](https://discuss.elastic.co/t/elasticsearch-spark-30-read-missing-field-double-type-error/320112/6 "2022-12-02T01:41:06Z")

</div>

**etest7 C double type**

 ![etest7 C double type](https://us1.discourse-cdn.com/elastic/original/3X/5/b/5bd57b49466921a3dde3b86a467d3a3dd235167b.jpeg)

**C is null and version is 7.17.0**

 ![C is null](https://us1.discourse-cdn.com/elastic/original/3X/f/c/fc663d625fe2187f1a1e938f0b967b190bf3eb74.jpeg)

**I can read successfully when c has no value**

 ![I can read successfully when c has no value](https://us1.discourse-cdn.com/elastic/original/3X/4/d/4d7b1f77f03cb75982f374312e76a55a356911bd.jpeg)

**Put empty string into c**

 ![Put empty string into c](https://us1.discourse-cdn.com/elastic/original/3X/9/6/96e90eda7aeb6daf337c336e159ad9caa98ff9b2.jpeg)

**es.field.read.empty.as.null=false Strack trace:**

 ![es.field.read.empty.as.null=false](https://us1.discourse-cdn.com/elastic/original/3X/3/6/36ee358287aa4b5599bf4c3a91e5eaf4de79ab3b.jpeg)

**es.field.read.empty.as.null=true Strack trace:**

 ![es.field.read.empty.as.null=true](https://us1.discourse-cdn.com/elastic/original/3X/7/e/7e11b94a715f8a51f9514b69569854073fe1553b.jpeg)

---

<div class="post-metadata">

**Author:** ![918246180](https://avatars.discourse-cdn.com/v4/letter/9/eb9ed0/32.png) [@918246180](https://discuss.elastic.co/u/918246180)\
**Post date:** [December 2, 2022, 1:54am UTC](https://discuss.elastic.co/t/elasticsearch-spark-30-read-missing-field-double-type-error/320112/7 "2022-12-02T01:54:45Z")

</div>

> [@Keith\_Massey](#):
>
> 8.1.0

Thanks your prompt reply extremely

first of all .I'm sorry that I can only upload photos, which may be difficult to read. The reason is that my company requires me not to develop programs on the Internet. And I see that the version you are using is different from mine.

This is the maven dependency I used for this verification：  
elasticsearch-spark-30\_2.12-7.16.1  
spark-sql\_2.12-3.2.1

Elasticsearch version 7.17.0

---

<div class="post-metadata">

**Author:** ![Keith\_Massey](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/keith_massey/32/83666_2.png) [@Keith\_Massey](https://discuss.elastic.co/u/Keith_Massey)\
**Post date:** [December 2, 2022, 1:36pm UTC](https://discuss.elastic.co/t/elasticsearch-spark-30-read-missing-field-double-type-error/320112/8 "2022-12-02T13:36:31Z")

</div>

I see you're using Elasticsearch 7.17 and Spark 3.2. Spark 3.2 support was not added until 8.0 -- [https://github.com/elastic/elasticsearch-hadoop/pull/1807](https://github.com/elastic/elasticsearch-hadoop/pull/1807). I don't know that that is related to your current problem but you will run into problems with that combination. Spark made some breaking changes between 3.0 and 3.2.

---

<div class="post-metadata">

**Author:** ![918246180](https://avatars.discourse-cdn.com/v4/letter/9/eb9ed0/32.png) [@918246180](https://discuss.elastic.co/u/918246180)\
**Post date:** [December 3, 2022, 12:32am UTC](https://discuss.elastic.co/t/elasticsearch-spark-30-read-missing-field-double-type-error/320112/9 "2022-12-03T00:32:06Z")

</div>

Thanks for the reminder, I will try

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 31, 2022, 12:33am UTC](https://discuss.elastic.co/t/elasticsearch-spark-30-read-missing-field-double-type-error/320112/10 "2022-12-31T00:33:00Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
