# Spark ES Read Error

**URL:** <https://discuss.elastic.co/t/spark-es-read-error/233845>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [May 22, 2020, 4:34am UTC](https://discuss.elastic.co/t/spark-es-read-error/233845 "2020-05-22T04:34:26Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![dlSpark](https://avatars.discourse-cdn.com/v4/letter/d/a698b9/32.png) [@dlSpark](https://discuss.elastic.co/u/dlSpark)\
**Post date:** [May 22, 2020, 4:34am UTC](https://discuss.elastic.co/t/spark-es-read-error/233845/1 "2020-05-22T04:34:27Z")

</div>

We have several index(s) stored in out ES (7.1.1) cluster. edit-- Spark is 2.3.1, Scala 2.11.8, Hadoop HDP 3.0.1

When reading as a dataframe by  
val reader = spark.read.format("org.elasticsearch.spark.sql").option("es.nodes", nodeList)  
val x = reader.load("small\_index")  
x.show()

for the smaller index (about a million entries) it works ok

for a larger index (3 billion entries) we get the error below after any action on the dataframe. The job has over 300G allocated should not be a Spark resource issue (?).

Any suggestions welcome

```auto
> org.elasticsearch.hadoop.EsHadoopIllegalArgumentException: invalid response
> at org.elasticsearch.hadoop.util.Assert.isTrue(Assert.java:60)
> at org.elasticsearch.hadoop.serialization.ScrollReader.read(ScrollReader.java:271)
> at org.elasticsearch.hadoop.serialization.ScrollReader.read(ScrollReader.java:262)
> at org.elasticsearch.hadoop.rest.RestRepository.scroll(RestRepository.java:313)
> at org.elasticsearch.hadoop.rest.ScrollQuery.hasNext(ScrollQuery.java:93)
> at org.elasticsearch.spark.rdd.AbstractEsRDDIterator.hasNext(AbstractEsRDDIterator.scala:61)
> at scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:408)
> at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.agg_doA ggregateWithoutKey_0$(Unknown Source)
> at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.process Next(Unknown Source)
> at org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)
> at org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$10$$anon$1.hasNext(WholeStageCodegen Exec.scala:614)
> at scala.collection.Iterator$$anon$11.hasNext(Iterator.scala:408)
> at org.apache.spark.shuffle.sort.BypassMergeSortShuffleWriter.write(BypassMergeSortShuffleWriter.java :125)
> at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:96)
> at org.apache.spark.scheduler.ShuffleMapTask.runTask(ShuffleMapTask.scala:53)
> at org.apache.spark.scheduler.Task.run(Task.scala:109)
> at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:345)
> at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)
> at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)
> at java.lang.Thread.run(Thread.java:748)

```

---

<div class="post-metadata">

**Author:** ![Luca\_Belluccini](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/luca_belluccini/32/33239_2.png) [@Luca\_Belluccini](https://discuss.elastic.co/u/Luca_Belluccini)\
**Post date:** [May 22, 2020, 6:31am UTC](https://discuss.elastic.co/t/spark-es-read-error/233845/2 "2020-05-22T06:31:03Z")

</div>

I would suggest to:

- enable the debug logs in the Elasticsearch spark library to show what is the actual "invalid response" [https://www.elastic.co/guide/en/elasticsearch/hadoop/current/logging.html](https://www.elastic.co/guide/en/elasticsearch/hadoop/current/logging.html)
- you can tweak the settings `es.scroll.keepalive` (default 10m) and `es.scroll.size` (default 50). In particular I would increase the size.

---

<div class="post-metadata">

**Author:** ![dlSpark](https://avatars.discourse-cdn.com/v4/letter/d/a698b9/32.png) [@dlSpark](https://discuss.elastic.co/u/dlSpark)\
**Post date:** [May 22, 2020, 7:00am UTC](https://discuss.elastic.co/t/spark-es-read-error/233845/3 "2020-05-22T07:00:03Z")

</div>

Thanks for the response.

just changing those settings (tried size at 1000, 20000, 10000) had no effect.

is there any logging I can change with user permissions, getting the admins to change hadoop config files requires some time and justification

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 19, 2020, 7:08am UTC](https://discuss.elastic.co/t/spark-es-read-error/233845/4 "2020-06-19T07:08:42Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
