# Throttling indexing to Elasticsearch in Spark

**URL:** <https://discuss.elastic.co/t/throttling-indexing-to-elasticsearch-in-spark/64203>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [October 27, 2016, 8:39pm UTC](https://discuss.elastic.co/t/throttling-indexing-to-elasticsearch-in-spark/64203 "2016-10-27T20:39:01Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![shakdoesspark](https://avatars.discourse-cdn.com/v4/letter/s/dec6dc/32.png) [@shakdoesspark](https://discuss.elastic.co/u/shakdoesspark)\
**Post date:** [October 27, 2016, 8:39pm UTC](https://discuss.elastic.co/t/throttling-indexing-to-elasticsearch-in-spark/64203/1 "2016-10-27T20:39:01Z")

</div>

Is there a way to throttle the number of tasks to index Elasticsearch but not throttle the number of tasks  
for the compute (map, flatmap) tasks in spark?

We're using the native ES-Hadoop plugin.

---

<div class="post-metadata">

**Author:** ![sukumar](https://avatars.discourse-cdn.com/v4/letter/s/b77776/32.png) [@sukumar](https://discuss.elastic.co/u/sukumar)\
**Post date:** [October 27, 2016, 8:45pm UTC](https://discuss.elastic.co/t/throttling-indexing-to-elasticsearch-in-spark/64203/2 "2016-10-27T20:45:00Z")

</div>

@shakdoesspark The number of task to index the ES is based on the shards in ES.

---

<div class="post-metadata">

**Author:** ![shakdoesspark](https://avatars.discourse-cdn.com/v4/letter/s/dec6dc/32.png) [@shakdoesspark](https://discuss.elastic.co/u/shakdoesspark)\
**Post date:** [October 27, 2016, 8:49pm UTC](https://discuss.elastic.co/t/throttling-indexing-to-elasticsearch-in-spark/64203/3 "2016-10-27T20:49:56Z")

</div>

So I have 5 shards on my index, 0 replicas, and I'm using the RDD.saveToEsWithMeta method, but I'm seeing 256 tasks being created. I have 8 nodes in my spark cluster, and I've set the --executor-cores to 4.

I'm seeing that it's running 32 Tasks( 8 nodes \* 4 cores per Executor ).

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [October 27, 2016, 9:55pm UTC](https://discuss.elastic.co/t/throttling-indexing-to-elasticsearch-in-spark/64203/4 "2016-10-27T21:55:16Z")

</div>

> [@sukumar](#):
>
> The number of task to index the ES is based on the shards in ES.

Not true: The number of tasks used to READ from ES is based on the number of shards it is reading from. You can write to ES using any number of tasks, there's no way for the library to control this as it is a user setting in both Hadoop and Spark. We just suggest using a number of partitions equal to the shards being written to as a starting point and for tuning to go from there.

@shakdoesspark I would advise using the `RDD.repartition(x: Int)` method to shrink the number of splits or to modify the original number of splits on the RDD to be a lower number.

---

<div class="post-metadata">

**Author:** ![jspooner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jspooner/32/12984_2.png) [@jspooner](https://discuss.elastic.co/u/jspooner)\
**Post date:** [January 17, 2017, 11:33pm UTC](https://discuss.elastic.co/t/throttling-indexing-to-elasticsearch-in-spark/64203/5 "2017-01-17T23:33:12Z")

</div>

From my experience the answer is no. I have been saving my processed data into HDFS or S3 then I have a separate job read and push to ES.

I've also been using Zeppelin notebooks so setting --num-executors and cores is a little more difficult. So I've been setting them directly on the config before calling saveToEs

```
spark.conf.set("spark.dynamicAllocation.enabled","false")
spark.conf.set("spark.executor.instances", "8") // --num-executors
spark.conf.set("spark.executor.cores", "4");
df.saveToEs(esConfig)

```

However when visiting the Executors tab in the SparkHistory UI the summary never matches up to what I set.

I find it difficult to get visibility into the task that are actually running. I have enabled the extra logging but those logs are created on each task node and I have not tried to use them to gage task usage yet.

---

<div class="post-metadata">

**Author:** ![shakdoesspark](https://avatars.discourse-cdn.com/v4/letter/s/dec6dc/32.png) [@shakdoesspark](https://discuss.elastic.co/u/shakdoesspark)\
**Post date:** [January 18, 2017, 2:51am UTC](https://discuss.elastic.co/t/throttling-indexing-to-elasticsearch-in-spark/64203/6 "2017-01-18T02:51:15Z")

</div>

RDD.repartition works great!  
@jspooner, you should give that a try!

---

<div class="post-metadata">

**Author:** ![no\_jihun](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/no_jihun/32/64028_2.png) [@no\_jihun](https://discuss.elastic.co/u/no_jihun)\
**Post date:** [January 28, 2017, 11:53pm UTC](https://discuss.elastic.co/t/throttling-indexing-to-elasticsearch-in-spark/64203/7 "2017-01-28T23:53:47Z")

</div>

You'd better use df.coalesce (...) than df.repartition (...) for spark effeciency.

---

<div class="post-metadata">

**Author:** ![shakdoesspark](https://avatars.discourse-cdn.com/v4/letter/s/dec6dc/32.png) [@shakdoesspark](https://discuss.elastic.co/u/shakdoesspark)\
**Post date:** [January 29, 2017, 12:21am UTC](https://discuss.elastic.co/t/throttling-indexing-to-elasticsearch-in-spark/64203/8 "2017-01-29T00:21:00Z")

</div>

After coalesce, why would you need to repartition ?

---

<div class="post-metadata">

**Author:** ![no\_jihun](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/no_jihun/32/64028_2.png) [@no\_jihun](https://discuss.elastic.co/u/no_jihun)\
**Post date:** [January 29, 2017, 6:44am UTC](https://discuss.elastic.co/t/throttling-indexing-to-elasticsearch-in-spark/64203/9 "2017-01-29T06:44:51Z")

</div>

no, I mean  
use coalesce instead of repartition.

---

<div class="post-metadata">

**Author:** ![shakdoesspark](https://avatars.discourse-cdn.com/v4/letter/s/dec6dc/32.png) [@shakdoesspark](https://discuss.elastic.co/u/shakdoesspark)\
**Post date:** [February 3, 2017, 4:52pm UTC](https://discuss.elastic.co/t/throttling-indexing-to-elasticsearch-in-spark/64203/10 "2017-02-03T16:52:32Z")

</div>

ah ok thanks for the clarification

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:22pm UTC](https://discuss.elastic.co/t/throttling-indexing-to-elasticsearch-in-spark/64203/11 "2017-07-06T13:22:32Z")

</div>


