# RDD saveToEs performance

**URL:** <https://discuss.elastic.co/t/rdd-savetoes-performance/42935>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [February 28, 2016, 9:05am UTC](https://discuss.elastic.co/t/rdd-savetoes-performance/42935 "2016-02-28T09:05:17Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![aaskey](https://avatars.discourse-cdn.com/v4/letter/a/5fc32e/32.png) [@aaskey](https://discuss.elastic.co/u/aaskey)\
**Post date:** [February 28, 2016, 9:05am UTC](https://discuss.elastic.co/t/rdd-savetoes-performance/42935/1 "2016-02-28T09:05:17Z")

</div>

Hi there,

With Es 2.2, Spark 1.6, Scala 2.10, SaveToEs performance is around 20 documents/second on MacBook Pro. Each document is less that 1KB. Is this expected? I am using the latest 2.2 elasticsearch-spark connector.

Adding logs of processing 10k messages:  
16/02/27 17:11:30 INFO SparkContext: Starting job: runJob at EsSpark.scala:67  
16/02/27 17:11:30 INFO DAGScheduler: Got job 0 (runJob at EsSpark.scala:67) with 1 output partitions  
16/02/27 17:11:30 INFO DAGScheduler: Final stage: ResultStage 0 (runJob at EsSpark.scala:67)  
16/02/27 17:14:28 INFO DAGScheduler: ResultStage 0 (runJob at EsSpark.scala:67) finished in 178.170 s  
16/02/27 17:14:28 INFO DAGScheduler: Job 0 finished: runJob at EsSpark.scala:67, took 178.347390 s  
16/02/27 17:14:29 INFO SparkContext: Starting job: runJob at EsSpark.scala:67  
16/02/27 17:14:29 INFO DAGScheduler: Got job 1 (runJob at EsSpark.scala:67) with 1 output partitions  
16/02/27 17:14:29 INFO DAGScheduler: Final stage: ResultStage 1 (runJob at EsSpark.scala:67)  
16/02/27 17:17:24 INFO DAGScheduler: ResultStage 1 (runJob at EsSpark.scala:67) finished in 175.553 s  
16/02/27 17:17:24 INFO DAGScheduler: Job 1 finished: runJob at EsSpark.scala:67, took 175.585507 s  
16/02/27 17:17:24 INFO SparkContext: Starting job: runJob at EsSpark.scala:67  
16/02/27 17:17:24 INFO DAGScheduler: Got job 2 (runJob at EsSpark.scala:67) with 1 output partitions  
16/02/27 17:17:24 INFO DAGScheduler: Final stage: ResultStage 2 (runJob at EsSpark.scala:67)  
16/02/27 17:20:18 INFO DAGScheduler: ResultStage 2 (runJob at EsSpark.scala:67) finished in 174.285 s  
16/02/27 17:20:18 INFO DAGScheduler: Job 2 finished: runJob at EsSpark.scala:67, took 174.286503 s

Thanks,

---

<div class="post-metadata">

**Author:** ![kucera.jan.cz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kucera.jan.cz/32/5869_2.png) [@kucera.jan.cz](https://discuss.elastic.co/u/kucera.jan.cz)\
**Post date:** [February 29, 2016, 7:59pm UTC](https://discuss.elastic.co/t/rdd-savetoes-performance/42935/2 "2016-02-29T19:59:36Z")

</div>

Hello aaskey,

20 docs/sec is definitely not expected throughput. I was recently prototyping some big doc size insertion and I was able to index 2k/sec on single-node.

Neverless I am wondering whether someone here can suggest best practices (except [official docs](https://www.elastic.co/guide/en/elasticsearch/hadoop/2.2/performance.html)) how to increase insertion rate to ES.

I have read interesting article about performance with spark-kafka which suggested [client caching](http://mkuthan.github.io/blog/2016/01/29/spark-kafka-integration2/) and wondering how much overhead `RestService.createWriter` causes for creating client for every call.

---

<div class="post-metadata">

**Author:** ![aaskey](https://avatars.discourse-cdn.com/v4/letter/a/5fc32e/32.png) [@aaskey](https://discuss.elastic.co/u/aaskey)\
**Post date:** [March 1, 2016, 4:13am UTC](https://discuss.elastic.co/t/rdd-savetoes-performance/42935/3 "2016-03-01T04:13:36Z")

</div>

Hi Kucera,

Is there any example code you can share or if you have done anything other than the official document suggested?

---

<div class="post-metadata">

**Author:** ![kucera.jan.cz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kucera.jan.cz/32/5869_2.png) [@kucera.jan.cz](https://discuss.elastic.co/u/kucera.jan.cz)\
**Post date:** [March 1, 2016, 5:23am UTC](https://discuss.elastic.co/t/rdd-savetoes-performance/42935/4 "2016-03-01T05:23:22Z")

</div>

On Spark side I didn't anything special. The elasticsearch itself was little tweaked from default configuration to support higher indexing throughput.

Regarding sharing scala code - when I was starting with spark-elastic this [github repo](https://github.com/holdenk/elasticsearchspark) was really useful.

-Jan

---

<div class="post-metadata">

**Author:** ![aaskey](https://avatars.discourse-cdn.com/v4/letter/a/5fc32e/32.png) [@aaskey](https://discuss.elastic.co/u/aaskey)\
**Post date:** [March 1, 2016, 8:14am UTC](https://discuss.elastic.co/t/rdd-savetoes-performance/42935/5 "2016-03-01T08:14:19Z")

</div>

Thanks Jan. I found out the the low performance as caused by other operations. Without those operations, I able to achieve the similar (2k/s) write performance to ES. Thanks again.

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [March 3, 2016, 9:53am UTC](https://discuss.elastic.co/t/rdd-savetoes-performance/42935/6 "2016-03-03T09:53:34Z")

</div>

@aaskey Adding some monitoring on ES side (such as [Marvel](https://www.elastic.co/products/marvel)) helps in figuring out whether ES itself is overloaded or whether the OS itself is under pressure.  
Always keep an eye on CPU/IO/Mem.

Glad to hear things are back to normal.

@kucera.jan.cz `RestService.createWriter` is used once _per_ client not once _for every call_. At least within ES-Hadoop itself - if the users keep creating jobs on every call then yes, the partition discovery and execution will be created quite often however at this point, likely Spark itself and the rest of the components will add a significant overhead already.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:26pm UTC](https://discuss.elastic.co/t/rdd-savetoes-performance/42935/7 "2017-07-06T13:26:00Z")

</div>


