# \[Hadoop\] - Difference between task creation for a write and read-update-write operation in ES

**URL:** <https://discuss.elastic.co/t/hadoop-difference-between-task-creation-for-a-write-and-read-update-write-operation-in-es/23573>\
**Category:** Elasticsearch\
**Created:** [May 6, 2015, 10:41am UTC](https://discuss.elastic.co/t/hadoop-difference-between-task-creation-for-a-write-and-read-update-write-operation-in-es/23573 "2015-05-06T10:41:39Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Piyush\_Goyal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/piyush_goyal/32/8025_2.png) [@Piyush\_Goyal](https://discuss.elastic.co/u/Piyush_Goyal)\
**Post date:** [May 6, 2015, 10:41am UTC](https://discuss.elastic.co/t/hadoop-difference-between-task-creation-for-a-write-and-read-update-write-operation-in-es/23573/1 "2015-05-06T10:41:39Z")

</div>

Hi Costin,

I saw a different behavior of creating task for write to ES operation while  
working on my project. The difference is as follows:

1.) Only write to ES - When I create an RDD of my own to insert data into  
ES, the task are created based on property "es.batch.size.bytes" and  
"es.batch.size.entries". Number of task created = Number of documents in  
RDD/the value of either of these properties. The request hits the node and  
node decides the shard to which document needs routed based on routing  
value(if specified).

2.) Read-Update-write to ES - Consider this case when I have to read data  
from ES, store it in RDD, do some updates in the documents in RDD and then  
index these documents back to ES. While reading, the number of tasks are  
created on basis of number of shards and I presume that each tasks fetch  
data from each Shard(not sure of how it works? - Task delagting request to  
node to serve data from a particular shard?). Now when I try to  
update/re-index data using same RDD and function saveToESWithMetadata, this  
time the number of task created is a number which is not based on point 1.  
If the data in each partition is less than property  
"es.batch.size.entries", it creates the same number of tasks as are the  
number of shards, else greater than it.

What's the reason behind this? Also like read operation where request is  
from particular shard, does write operation also write to a shard or all  
the task delegate their request to the node?

Thanks in advance  
Piyush Costin,

I saw a different behavior of creating task for write to ES operation while  
working on my project. The difference is as follows:

1.) Only write to ES - When I create an RDD of my own to insert data into  
ES, the task are created based on property "es.batch.size.bytes" and  
"es.batch.size.entries". Number of task created = Number of documents in  
RDD/the value of either of these properties. The request hits the node and  
node decides the shard to which document needs routed based on routing  
value(if specified).

2.) Read-Update-write to ES - Consider this case when I have to read data  
from ES, store it in RDD, do some updates in the documents in RDD and then  
index these documents back to ES. While reading, the number of tasks are  
created on basis of number of shards and I presume that each tasks fetch  
data from each Shard(not sure of how it works? - Task delagting request to  
node to serve data from a particular shard?). Now when I try to  
update/re-index data using same RDD and function saveToESWithMetadata, this  
time the number of task created is a number which is not based on point 1.  
If the data in each partition is less than property  
"es.batch.size.entries", it creates the same number of tasks as are the  
number of shards, else greater than it.

What's the reason behind this? Also like read operation where request is  
from particular shard, does write operation also write to a shard or all  
the task delegate their request to the node?

Thanks in advance  
Piyush

## -- Please update your bookmarks! We moved to [https://discuss.elastic.co/](https://discuss.elastic.co/)

You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/ec268e76-6220-430b-958a-884692283ca0%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ec268e76-6220-430b-958a-884692283ca0%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [May 14, 2015, 7:16am UTC](https://discuss.elastic.co/t/hadoop-difference-between-task-creation-for-a-write-and-read-update-write-operation-in-es/23573/2 "2015-05-14T07:16:30Z")

</div>

Hi,

First it would help to know what version of Elaticsearch, Elasticsearch Hadoop, JVM and Spark are you using.

On 5/6/15 1:41 PM, piyush goyal wrote:

> Hi Costin,
> 
> I saw a different behavior of creating task for write to ES operation while working on my project. The difference is  
> as follows:
> 
> 1.) Only write to ES - When I create an RDD of my own to insert data into ES, the task are created based on property  
> "es.batch.size.bytes" and "es.batch.size.entries". Number of task created = Number of documents in RDD/the value of  
> either of these properties. The request hits the node and node decides the shard to which document needs routed based  
> on routing value(if specified).

What makes you say that? In case of writing, the number of writers is determined by the number of shards of your target  
index. The more shards, the more concurrent writers. The behavior of all writers can be further tweaked through the  
properties mentioned however they do NOT affect the process parallelism.

> 2.) Read-Update-write to ES - Consider this case when I have to read data from ES, store it in RDD, do some updates in  
> the documents in RDD and then index these documents back to ES. While reading, the number of tasks are created on  
> basis of number of shards and I presume that each tasks fetch data from each Shard(not sure of how it works? - Task  
> delagting request to node to serve data from a particular shard?). Now when I try to update/re-index data using same  
> RDD and function saveToESWithMetadata, this time the number of task created is a number which is not based on point 1.  
> If the data in each partition is less than property "es.batch.size.entries", it creates the same number of tasks as  
> are the number of shards, else greater than it.

In case of reading, the number of tasks that can work in parallel is determine by the source parallelism - its number of  
partitions. So if you have an RDD with 1 partition likely it will result into one task that will write to another RDD  
down the line. Assuming that RDD is backed by Elastic - even if the index has 10 shards and thus can have a parallelism  
of 10, if the source has only one partition and there's only one task, there's nothing the connector can do to increase  
this number.

Again, remember that the connector is not an active component per se - rather it bridges two systems. It's spark and  
more importantly the RDDs structures that control the number of tasks/threads that can work in parallel at a given time.

Going forward, let's take this to Discuss. It's a discussion that I'm sure will benefit other users and hopefully it  
will be better archived there.

> What's the reason behind this? Also like read operation where request is from particular shard, does write operation  
> also write to a shard or all the task delegate their request to the node?
> 
> Thanks in advance  
> Piyush Costin,
> 
> I saw a different behavior of creating task for write to ES operation while working on my project. The difference is  
> as follows:
> 
> |1.) Only write to ES - When I create an RDD of my own to insert data into ES, the task are created based on property  
> "es.batch.size.bytes" and "es.batch.size.entries". Number of task created = Number of documents in RDD/the value of  
> either of these properties. The request hits the node and node decides the shard to which document needs routed based  
> on routing value(if specified).
> 
> 2.) Read-Update-write to ES - Consider this case when I have to read data from ES, store it in RDD, do some updates in  
> the documents in RDD and then index these documents back to ES. While reading, the number of tasks are created on  
> basis of number of shards and I presume that each tasks fetch data from each Shard(not sure of how it works? - Task  
> delagting request to node to serve data from a particular shard?). Now when I try to update/re-index data using same  
> RDD and function saveToESWithMetadata, this time the number of task created is a number which is not based on point 1.  
> If the data in each partition is less than property "es.batch.size.entries", it creates the same number of tasks as  
> are the number of shards, else greater than it.
> 
> What's the reason behind this? Also like read operation where request is from particular shard, does write operation  
> also write to a shard or all the task delegate their request to the node?
> 
> Thanks in advance
> 
> | Piyush |
> | --- |
> | Please update your bookmarks! We moved to [https://discuss.elastic.co/](https://discuss.elastic.co/) |
> 
> * * *
> 
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an email to  
> [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com) [mailto:elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/ec268e76-6220-430b-958a-884692283ca0%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ec268e76-6220-430b-958a-884692283ca0%40googlegroups.com)  
> [https://groups.google.com/d/msgid/elasticsearch/ec268e76-6220-430b-958a-884692283ca0%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/ec268e76-6220-430b-958a-884692283ca0%40googlegroups.com?utm_medium=email&utm_source=footer).  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
Costin

## -- Please update your bookmarks! We have moved to [https://discuss.elastic.co/](https://discuss.elastic.co/)

You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/55544BCE.6060201%40gmail.com](https://groups.google.com/d/msgid/elasticsearch/55544BCE.6060201%40gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:14am UTC](https://discuss.elastic.co/t/hadoop-difference-between-task-creation-for-a-write-and-read-update-write-operation-in-es/23573/3 "2017-07-06T00:14:07Z")

</div>


