# ElasticSearch and parallelization

**URL:** <https://discuss.elastic.co/t/elasticsearch-and-parallelization/10691>\
**Category:** Elasticsearch\
**Created:** [February 11, 2013, 12:39pm UTC](https://discuss.elastic.co/t/elasticsearch-and-parallelization/10691 "2013-02-11T12:39:04Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Dave\_O](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dave_o/32/69398_2.png) [@Dave\_O](https://discuss.elastic.co/u/Dave_O)\
**Post date:** [February 11, 2013, 12:39pm UTC](https://discuss.elastic.co/t/elasticsearch-and-parallelization/10691/1 "2013-02-11T12:39:04Z")

</div>

I have terabytes of data that I need to index. In my environment I have  
lots and lots of data but minimal concurrent queries. I'm curious how  
parallelization works with ES.

This is a hypothetical scenario. Lets say I have US state data that  
logically can be broken into 52 parts. For arguments sake lets say that  
the data counts are equivilent in all the states. Lets also assume I have  
52 Nodes/Servers

The 2 scenarios I'm looking at are

- Creating 1 large index with many shards. Lets say 52 in this case. (and  
putting on separate servers)
- Or I could create 52 separate indexes.

The questions I have.

- When searching on scenario 1 (52 shards) and my query spands all 52  
shards will this run in parallel. (ie, will I get 52 parallel threads)
- Similarly, If I query all 52 indexes with the same query such as  
index1,index2,index3...index52 will this kick off 52 parallel processes.

Is there a way to control the parallelization in either scenario?

Any opinions on which option be faster or a better configuration for both  
searches and indexing?

In either case I could direct single state queries to the appropriate  
Shard or Index by either using routing or just going directly to the  
appropriately index.

But routinely I will have queries that span some or all of the states.  
From everything I can tell the queries will be parallelized...but  
wondering to what extent.

Would the scenario be the same if I had all 52 Nodes on the same server.  
(Ie, would still get 52 parallel processes in both cases?)

Thanks for your guidance.  
_._[http://elasticsearch-users.115913.n3.nabble.com/Question-about-parallelization-td4029620.html#](http://elasticsearch-users.115913.n3.nabble.com/Question-about-parallelization-td4029620.html#)

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [February 11, 2013, 12:54pm UTC](https://discuss.elastic.co/t/elasticsearch-and-parallelization/10691/2 "2013-02-11T12:54:09Z")

</div>

Hi Mike

> This is a hypothetical scenario. Lets say I have US state data that  
> logically can be broken into 52 parts. For arguments sake lets say  
> that the data counts are equivilent in all the states. Lets also  
> assume I have 52 Nodes/Servers
> 
> The 2 scenarios I'm looking at are
> 
> - Creating 1 large index with many shards. Lets say 52 in this case.  
> (and putting on separate servers)
> - Or I could create 52 separate indexes.
> 
> The questions I have.
> 
> - When searching on scenario 1 (52 shards) and my query spands all 52  
> shards will this run in parallel. (ie, will I get 52 parallel  
> threads)
> - Similarly, If I query all 52 indexes with the same query such as  
> index1,index2,index3...index52 will this kick off 52 parallel  
> processes.

Searching 52 indices with 1 shard each is exactly equivalent to  
searching one index with 52 shards. The search happens at shard level.

All shards would be searched in parallel.

> Is there a way to control the parallelization in either scenario?

As in to limit the number of threads? You could look at configuring the  
thread pool:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

but I'm not sure if that would apply to the number of searches per  
shard, or for distributed search as you describe above.

> Any opinions on which option be faster or a better configuration for  
> both searches and indexing?

Exactly equivalent.

> In either case I could direct single state queries to the appropriate  
> Shard or Index by either using routing or just going directly to the  
> appropriately index.

Correct

> But routinely I will have queries that span some or all of the states.  
> From everything I can tell the queries will be parallelized...but  
> wondering to what extent.

The node that handles the request will "scatter" the request to all  
relevant shards in parallel, gather their results and return the  
collated results.

> Would the scenario be the same if I had all 52 Nodes on the same  
> server. (Ie, would still get 52 parallel processes in both cases?)

This I'm not sure about. I think this is where the threadpool config  
would kick in to limit the number of concurrent searches.

clint

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:52am UTC](https://discuss.elastic.co/t/elasticsearch-and-parallelization/10691/3 "2017-07-06T02:52:06Z")

</div>


