# Is there a way to run parallel\_bulk with sleep time?

**URL:** <https://discuss.elastic.co/t/is-there-a-way-to-run-parallel-bulk-with-sleep-time/64584>\
**Category:** Elasticsearch\
**Created:** [November 1, 2016, 3:47pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-run-parallel-bulk-with-sleep-time/64584 "2016-11-01T15:47:36Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Moshe\_Sucaz](https://avatars.discourse-cdn.com/v4/letter/m/0ea827/32.png) [@Moshe\_Sucaz](https://discuss.elastic.co/u/Moshe_Sucaz)\
**Post date:** [November 1, 2016, 3:47pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-run-parallel-bulk-with-sleep-time/64584/1 "2016-11-01T15:47:36Z")

</div>

Hi,  
I am upgrading our elastic from 1.5.2 to 2.3.4. I wrote python re index script that using elasticsearch.helpers.parallel\_bulk, and also I keep indexing new incoming docs to 1.5.2 and 2.3.4 clusters. When I am running the script, the new incoming docs are not indexed and the queue get bigger and bigger. I guess its because parallel\_bulk is running without a break...  
Is there a way to run parallel\_bulk with sleep time once in a while, or something like that?  
I tried to put sleep in the for loop of the call to parallel\_bulk, but it didn't help

The function that I wrote to parallel\_bulk is:

> def parallel\_reindex(index\_name, doc\_type, chunk\_size=500, scroll='10m', scan\_kwargs={}, bulk\_kwargs={}):  
> target\_client = Elasticsearch(hosts = ['node01:9200', 'node02:9200', 'node03:9200'], retry\_on\_timeout = True, max\_retries = 10, timeout = 1000)  
> source\_client = Elasticsearch(hosts = ['node04:9200', 'node05:9200', 'node06:9200'], retry\_on\_timeout=True, max\_retries=10, timeout=1000)  
> query = {"query": {"match\_all": {}}}  
> docs = scan(source\_client,  
> query = query,  
> index = index\_name,  
> scroll = scroll,  
> doc\_type=doc\_type,  
> \*\* scan\_kwargs  
> )

> ```
> def _change_doc_params_to_elastic_2(hits, target_client):
> health = target_client.cluster.health()
> while (health.get("status") != "green"):
> print("waiting 10 min for cluster to be green")
> time.sleep(600)
> # reindex_log.info("cluster health is %s" % health)
> for h in hits:
> # changing dots to “_”
> if 'x.y.z' in h['_source']:
> h['_source']['x_y_z'] = h['_source']['x.y.z']
> del h['_source']['x.y.z']
> # removing _analyzer
> if '_analyzer' in h['_source']:
> del h['_source']['_analyzer']
> yield h
> kwargs = {
> 'stats_only': True,
> }
> kwargs.update(bulk_kwargs)
> for response in parallel_bulk(target_client, _change_doc_params_to_elastic_2(docs, target_client), thread_count=8, chunk_size=chunk_size, max_chunk_bytes=20 * 1014 * 1024):
> #time.sleep(1)  
> pass
> 
> ```

> ```
> print("Done parallel_reindex of doc_type %s in index %s" % (doc_type,index_name))
> 
> ```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:07pm UTC](https://discuss.elastic.co/t/is-there-a-way-to-run-parallel-bulk-with-sleep-time/64584/2 "2017-07-05T22:07:49Z")

</div>


