# How to do parallel filter search of 1M queries in minimal time?

**URL:** <https://discuss.elastic.co/t/how-to-do-parallel-filter-search-of-1m-queries-in-minimal-time/282023>\
**Category:** Elasticsearch\
**Created:** [August 20, 2021, 12:35am UTC](https://discuss.elastic.co/t/how-to-do-parallel-filter-search-of-1m-queries-in-minimal-time/282023 "2021-08-20T00:35:32Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![elastic\_yogi](https://avatars.discourse-cdn.com/v4/letter/e/71c47a/32.png) [@elastic\_yogi](https://discuss.elastic.co/u/elastic_yogi)\
**Post date:** [August 20, 2021, 12:35am UTC](https://discuss.elastic.co/t/how-to-do-parallel-filter-search-of-1m-queries-in-minimal-time/282023/1 "2021-08-20T00:35:32Z")

</div>

Hi, I would like to search 1M filter queries in parallel in order to reduce the search time. This means a boolean query with 1M filter values, an no ordering is necessary. My index size is about 1GB, and I've stored everything in one shard, one node at the moment. I'm currently using searchAfter on 10k filter queries per request. So that means I will run 100 searchAfter queries in sequence with each query taking in 10k values (of the total 1M). This currently takes a long time (more than a couple minutes). I would like to consider parallelizing this search such that each of the 100 queries can run in parallel (e.g. on 100 client threads), and be aggregated within my application. What's the best way to do this? Some thoughts:

- should I still use searchAfter if I want parallelism? or should i use multisearch? (I do not need ordering/sorting/pagination). Does multisearch also have a limit to max\_result\_window? The reason why I tried using searchAfter is to exceed max\_result\_window.
- should I create shards to introduce parallelism at my index size? Elasticsearch documentation recommends shard sizes between 10GB and 50GB however, and my index size is only 1GB.
- can I introduce hypothetically 100 replicas of the 1GB shard/node and create 100 search requests in parallel? Theoretically if I have a pool of 100 workers together serving a total of 100 simultaneous searches, with each search taking 10k queries, then each worker would only need to process one search operation.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [August 20, 2021, 4:36am UTC](https://discuss.elastic.co/t/how-to-do-parallel-filter-search-of-1m-queries-in-minimal-time/282023/2 "2021-08-20T04:36:23Z")

</div>

What does resource usage look like while you are querying? With a 1GB index/shard I would expect the full data set to be cached in memory assuming you have enough RAM available to the node? How much RAM and CPU cores do you have available? What does heap and CPU usage look like while you are querying?

> [@elastic\_yogi](#):
>
> should I create shards to introduce parallelism at my index size?

No. As you will have lots of queries running in parallel increasing the number of primary shards is likely to reduce performance rather than increase it.

> [@elastic\_yogi](#):
>
> can I introduce hypothetically 100 replicas of the 1GB shard/node and create 100 search requests in parallel?

You can run many searches in parallel against the single shard, either by using multiple client threads or multisearch. If system resources are maxed out you can add nodes and increase the number of replicas so all nodes jold a full copy of the index/shard.

This sound like an unusual set of requirements. What is the use case?

---

<div class="post-metadata">

**Author:** ![elastic\_yogi](https://avatars.discourse-cdn.com/v4/letter/e/71c47a/32.png) [@elastic\_yogi](https://discuss.elastic.co/u/elastic_yogi)\
**Post date:** [August 23, 2021, 9:40pm UTC](https://discuss.elastic.co/t/how-to-do-parallel-filter-search-of-1m-queries-in-minimal-time/282023/3 "2021-08-23T21:40:06Z")

</div>

@Christian_Dahlqvist I have variable number of RAM/CPU, and expanded it such that I am not saturating these resources. My use case is quite unusual. I have a large list of filter values, and I want to filter by them all as fast as possible, getting just the ID's of those documents. "VALUE1" OR "VALUE2" (where both values are exact terms, that can be filtered, and no scoring is necessary)

I have some questions before using multisearch.

Is multisearch bounded by max\_result\_window for each request, or all requests within the multisearch (the entire multisearch)? I want to return more than the max\_result\_window (which is why I switched to searchAfter originally) in total across all "subrequests" of the multisearch.

Second, is it possible to customize \_source field on search results with multisearch? This appears to be in the documentation for search and searchAfter, but not multisearch.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [August 24, 2021, 5:13am UTC](https://discuss.elastic.co/t/how-to-do-parallel-filter-search-of-1m-queries-in-minimal-time/282023/4 "2021-08-24T05:13:49Z")

</div>

How many documents do you have or plan to have? How many documents can each individual term match? Can a document match multiple terms?

---

<div class="post-metadata">

**Author:** ![elastic\_yogi](https://avatars.discourse-cdn.com/v4/letter/e/71c47a/32.png) [@elastic\_yogi](https://discuss.elastic.co/u/elastic_yogi)\
**Post date:** [August 24, 2021, 7:03pm UTC](https://discuss.elastic.co/t/how-to-do-parallel-filter-search-of-1m-queries-in-minimal-time/282023/5 "2021-08-24T19:03:28Z")

</div>

about 1M-10M documents per index. we only have 1 index right now.  
for this query/filter, each term can match a few (handful of) documents.  
a document will not match multiple terms, unless a term is repeated in the query.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [August 24, 2021, 8:43pm UTC](https://discuss.elastic.co/t/how-to-do-parallel-filter-search-of-1m-queries-in-minimal-time/282023/6 "2021-08-24T20:43:25Z")

</div>

This does not sound like something Elasticsearch is optimised for so I am not sure how much it is possible to speed it up. I am afraid I do not have any good suggestions.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 21, 2021, 8:43pm UTC](https://discuss.elastic.co/t/how-to-do-parallel-filter-search-of-1m-queries-in-minimal-time/282023/7 "2021-09-21T20:43:41Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
