# Change threadpool queue size for batch process

**URL:** <https://discuss.elastic.co/t/change-threadpool-queue-size-for-batch-process/110788>\
**Category:** Elasticsearch\
**Created:** [December 8, 2017, 4:27am UTC](https://discuss.elastic.co/t/change-threadpool-queue-size-for-batch-process/110788 "2017-12-08T04:27:59Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Zhengcong\_Yin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/zhengcong_yin/32/62960_2.png) [@Zhengcong\_Yin](https://discuss.elastic.co/u/Zhengcong_Yin)\
**Post date:** [December 8, 2017, 4:27am UTC](https://discuss.elastic.co/t/change-threadpool-queue-size-for-batch-process/110788/1 "2017-12-08T04:27:59Z")

</div>

I wanted to do some batch process such as sequentially conducting 100 Million query. However, it will reject at some point. Can I just change the thread pool search queue size to -1(unbounded) ?  
If I can, do I have to restart the cluster and change the setting for each node I have?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 8, 2017, 9:18am UTC](https://discuss.elastic.co/t/change-threadpool-queue-size-for-batch-process/110788/2 "2017-12-08T09:18:55Z")

</div>

Why not just queue up queries at the application layer? If Elasticsearch is rejecting requests, it is generally for a good reason. Increasing the queue size in Elasticsearch will just result in increased memory usage and longer latencies as the size of the queue does not affect the query throughput (unless all the additional memory used actually slows it down).

---

<div class="post-metadata">

**Author:** ![Zhengcong\_Yin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/zhengcong_yin/32/62960_2.png) [@Zhengcong\_Yin](https://discuss.elastic.co/u/Zhengcong_Yin)\
**Post date:** [December 11, 2017, 3:34pm UTC](https://discuss.elastic.co/t/change-threadpool-queue-size-for-batch-process/110788/3 "2017-12-11T15:34:26Z")

</div>

Hi Christian,

Thanks for your reply. As you suggested, I queue up at my server side. However, when I set up the cluster, I found by adding more node, the time for batching query 10 thousand records also increased.  
These nodes settings are in default except for the role of the node and discovery IP. I was expected that the time will decrease linearly by adding more node to the cluster.

I checked the CPU utilization of each node, I found even when I conduct these queries, it almost remains the same.

Here is my set up: 43G data, 4 shards, no replica, 4 cores, 16G RAM. Virtual Machine

I am trying to use elasticsearch to do my thesis so I really appreciate your suggestions regarding this.

Zhengcong

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 11, 2017, 3:44pm UTC](https://discuss.elastic.co/t/change-threadpool-queue-size-for-batch-process/110788/4 "2017-12-11T15:44:22Z")

</div>

How many shards are you querying? How many nodes do you have in the cluster? How many parallel queries are you running?

---

<div class="post-metadata">

**Author:** ![Zhengcong\_Yin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/zhengcong_yin/32/62960_2.png) [@Zhengcong\_Yin](https://discuss.elastic.co/u/Zhengcong_Yin)\
**Post date:** [December 11, 2017, 3:52pm UTC](https://discuss.elastic.co/t/change-threadpool-queue-size-for-batch-process/110788/5 "2017-12-11T15:52:55Z")

</div>

My cluster always has 3 master node, and 1 client node, data node varies from one to four so that I could make the comparison.

I totally have 4 shards, right now I have 4 dedicated data note, so each data node has only one shard.

I just send these queries using FOR loop and track the 10000th one come back time using Callback.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 11, 2017, 4:02pm UTC](https://discuss.elastic.co/t/change-threadpool-queue-size-for-batch-process/110788/6 "2017-12-11T16:02:48Z")

</div>

How many of those shards are primary shards? Have you tried sending queries in parallel? As each shard is processed using a single thread per query, you will at most (assuming all shards are primary shards) use 4 cores (your number of shards) if you send all queries sequentially. It may therefore be that your setup is not able to benefit from the greater parallelism that a larger cluster can provide.

---

<div class="post-metadata">

**Author:** ![Zhengcong\_Yin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/zhengcong_yin/32/62960_2.png) [@Zhengcong\_Yin](https://discuss.elastic.co/u/Zhengcong_Yin)\
**Post date:** [December 11, 2017, 5:08pm UTC](https://discuss.elastic.co/t/change-threadpool-queue-size-for-batch-process/110788/7 "2017-12-11T17:08:48Z")

</div>

Thanks for your patience.

These four are all primary shards, no replica.

I checked the documentation, in my case, each node could handle 4 (cores) \*1.5 = 6 thread at the same time. So with 4 nodes, it should be 6 \*4 = 24, am I right?

Do I need to change anything regarding the client node in order to benefit from these 4 data nodes?

---

<div class="post-metadata">

**Author:** ![jasontedor](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jasontedor/32/66992_2.png) [@jasontedor](https://discuss.elastic.co/u/jasontedor)\
**Post date:** [December 11, 2017, 5:09pm UTC](https://discuss.elastic.co/t/change-threadpool-queue-size-for-batch-process/110788/8 "2017-12-11T17:09:37Z")

</div>

There is no such thing as an unbounded queue, instead you will eventually run out of heap space. That is, all queues are at least bounded by the heap space available to queue requests. Put differently: using an unbounded queue is dangerous and some day the places where "unbounded" queues are used within Elasticsearch will be [removed](https://github.com/elastic/elasticsearch/issues/18613) and the ability to set a queue as "unbounded" will be [removed](https://github.com/elastic/elasticsearch/issues/18613).

---

<div class="post-metadata">

**Author:** ![Zhengcong\_Yin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/zhengcong_yin/32/62960_2.png) [@Zhengcong\_Yin](https://discuss.elastic.co/u/Zhengcong_Yin)\
**Post date:** [December 11, 2017, 5:13pm UTC](https://discuss.elastic.co/t/change-threadpool-queue-size-for-batch-process/110788/9 "2017-12-11T17:13:56Z")

</div>

Here is my code for sending out these requests:  
I send out 5000 requests:

for (var i = 1; i \< 5000; i++){

```
 request.get(encodeURI(urls[i].split('\r').join('')), (error, response, body) => {
		 processedCount = processedCount +1;
		 console.log(processedCount)
		 let json = JSON.parse(body);
		console.log(json.hits.hits)			 
		 if(processedCount == 4999){
		 elapsed = new Date().getTime() / 1000 - timeStart
		 console.log(elapsed.toString())
	}		
 })

```

}

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 8, 2018, 5:14pm UTC](https://discuss.elastic.co/t/change-threadpool-queue-size-for-batch-process/110788/10 "2018-01-08T17:14:13Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
