# Elastic Search Hadoop Connector - Spark Facing Issues while Saving to ES

**URL:** <https://discuss.elastic.co/t/elastic-search-hadoop-connector-spark-facing-issues-while-saving-to-es/50852>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [May 24, 2016, 2:34pm UTC](https://discuss.elastic.co/t/elastic-search-hadoop-connector-spark-facing-issues-while-saving-to-es/50852 "2016-05-24T14:34:14Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![krishmah](https://avatars.discourse-cdn.com/v4/letter/k/d78d45/32.png) [@krishmah](https://discuss.elastic.co/u/krishmah)\
**Post date:** [May 24, 2016, 2:34pm UTC](https://discuss.elastic.co/t/elastic-search-hadoop-connector-spark-facing-issues-while-saving-to-es/50852/1 "2016-05-24T14:34:14Z")

</div>

I am using 5.0.0 alpha 2 version of Elastic Search along with the corresponding ES Hadoop connector. Version of Spark that I am using is 1.6.1. I am running on EMR.

Here are the configuration of Nodes on ES:

1 Master Node - M4.Large  
10 Data Nodes - C4. 4X Large  
1 Client Node - M4.Large

I was trying to load 7 days worth of log data by Partitioning them into 600 Partitions. Spark Executors: 6 and Spark Executor Cores : 4 and roughly we are loading 127315 records from a single partition.

I get this error and it does not give out additional messages

java.lang.NullPointerException  
at org.elasticsearch.hadoop.rest.RestClient.extractError(RestClient.java:229)  
at org.elasticsearch.hadoop.rest.RestClient.retryFailedEntries(RestClient.java:195)  
at org.elasticsearch.hadoop.rest.RestClient.bulk(RestClient.java:166)  
at org.elasticsearch.hadoop.rest.RestRepository.tryFlush(RestRepository.java:224)  
at org.elasticsearch.hadoop.rest.RestRepository.flush(RestRepository.java:247)  
at org.elasticsearch.hadoop.rest.RestRepository.close(RestRepository.java:266)  
at org.elasticsearch.hadoop.rest.RestService$PartitionWriter.close(RestService.java:130)  
at org.elasticsearch.spark.rdd.EsRDDWriter$$anonfun$write$1.apply$mcV$sp(EsRDDWriter.scala:42)  
at org.apache.spark.TaskContextImpl$$anon$2.onTaskCompletion(TaskContextImpl.scala:68)  
at org.apache.spark.TaskContextImpl$$anonfun$markTaskCompleted$1.apply(TaskContextImpl.scala:79)  
at org.apache.spark.TaskContextImpl$$anonfun$markTaskCompleted$1.apply(TaskContextImpl.scala:77)  
at scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59)  
at scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:47)  
at org.apache.spark.TaskContextImpl.markTaskCompleted(TaskContextImpl.scala:77)  
at org.apache.spark.scheduler.Task.run(Task.scala:91)  
at org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:214)  
at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)  
at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)  
at java.lang.Thread.run(Thread.java:745)

Also sometimes, I have seen "Could not write all entries (maybe ES was overloaded?)."

We are trying to do POC which will essentially load 8 Billion Log entries(from about 180 days worth of Log Data). We are not being successful in loading even 7 days worth of Data.

Can someone help me point in right direction by pointing out what are we doing wrong? What are the best practices in terms of Sizing? I have gone through the ES Hadoop connector write performance section of documentation and have left all the settings to default.

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [May 31, 2016, 4:40pm UTC](https://discuss.elastic.co/t/elastic-search-hadoop-connector-spark-facing-issues-while-saving-to-es/50852/2 "2016-05-31T16:40:33Z")

</div>

What version of ES are you using?  
The error looks like a bug (raised [this issue](https://github.com/elastic/elasticsearch-hadoop/issues/776) and checking the code indicates this to be as such.  
Do you have any logs available (see the docs on how to enable them).

---

<div class="post-metadata">

**Author:** ![krishmah](https://avatars.discourse-cdn.com/v4/letter/k/d78d45/32.png) [@krishmah](https://discuss.elastic.co/u/krishmah)\
**Post date:** [May 31, 2016, 8:05pm UTC](https://discuss.elastic.co/t/elastic-search-hadoop-connector-spark-facing-issues-while-saving-to-es/50852/3 "2016-05-31T20:05:32Z")

</div>

Hi Costin,

Thanks for replying to my question. I was using Elastic Search v5.0 Alpha 2 release. Actually I do not get these errors while I am running smaller load. I get these errors when I try to push lot of volume and as a result I was thinking if I should add Trace or Debug , which will make it spend considerable amount of time logging . Do you suggest enabling them and run them again?

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [June 1, 2016, 11:24am UTC](https://discuss.elastic.co/t/elastic-search-hadoop-connector-spark-facing-issues-while-saving-to-es/50852/4 "2016-06-01T11:24:00Z")

</div>

Likely the message is caused by overload. I've fixed the bug and look into double checking the message structure again in ES as it might have changed between Alpha 1, 2 and 3.

I've pushed a fresh nightly build with the fix.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:24pm UTC](https://discuss.elastic.co/t/elastic-search-hadoop-connector-spark-facing-issues-while-saving-to-es/50852/5 "2017-07-06T13:24:34Z")

</div>


