# Wait-until-merges-finish-after-index operations

**URL:** <https://discuss.elastic.co/t/wait-until-merges-finish-after-index-operations/359286>\
**Category:** Elasticsearch\
**Tags:** rally\
**Created:** [May 10, 2024, 9:57pm UTC](https://discuss.elastic.co/t/wait-until-merges-finish-after-index-operations/359286 "2024-05-10T21:57:03Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![tommyd](https://avatars.discourse-cdn.com/v4/letter/t/f9ae1b/32.png) [@tommyd](https://discuss.elastic.co/u/tommyd)\
**Post date:** [May 10, 2024, 9:57pm UTC](https://discuss.elastic.co/t/wait-until-merges-finish-after-index-operations/359286/1 "2024-05-10T21:57:03Z")

</div>

Hello Rally Gurus,

I am testing my ES cluster using ESRally with openai\_vector track and I have a question about the "wait-until-merges-finish-after-index" operation. In the track README file, I see that there is this track parameter "parallel\_indexing\_time\_period (default: 1800)". Is this parameter the "wait-until-merges-finish-after-index" operation? If it is, is it wise to lower the wait time since that's 30 minutes of inactivity. Am I way off base? Please explain.

Best,

Tom

---

<div class="post-metadata">

**Author:** ![json](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/json/32/4125_2.png) [@json](https://discuss.elastic.co/u/json)\
**Post date:** [May 10, 2024, 10:23pm UTC](https://discuss.elastic.co/t/wait-until-merges-finish-after-index-operations/359286/2 "2024-05-10T22:23:19Z")

</div>

Hi Tom,

The default (and only) track challenge in OpenAI vector performs parallel index and search operations as its primary set of tasks. Before these parallel operations begin, there is a standalone initial indexing operation to pre-load the index, followed by a refresh (to commit any remaining indexing ops disk), then followed by the `wait-until-merges-finish-after-index` operation you mentioned. `wait-until-merges-finish-after-index` polls the cluster looking for any remaining in-flight segment merges that could pollute the benchmark and does not use the `parallel_indexing_time_period` track parameter. This task should not take long since it is only waiting for segment merges to finish.

`parallel_indexing_time_period` controls how long the indexing portion should run in the parallel search & indexing task execution.

Thank you,  
Jason

---

<div class="post-metadata">

**Author:** ![tommyd](https://avatars.discourse-cdn.com/v4/letter/t/f9ae1b/32.png) [@tommyd](https://discuss.elastic.co/u/tommyd)\
**Post date:** [May 10, 2024, 10:38pm UTC](https://discuss.elastic.co/t/wait-until-merges-finish-after-index-operations/359286/3 "2024-05-10T22:38:00Z")

</div>

Thank you, Jason, for a quick reply and explanation. Much appreciate it!

Best,

Tom

---

<div class="post-metadata">

**Author:** ![tommyd](https://avatars.discourse-cdn.com/v4/letter/t/f9ae1b/32.png) [@tommyd](https://discuss.elastic.co/u/tommyd)\
**Post date:** [May 12, 2024, 7:58pm UTC](https://discuss.elastic.co/t/wait-until-merges-finish-after-index-operations/359286/4 "2024-05-12T19:58:34Z")

</div>

Hi Jason,

Could I ask another question? How come I don't see either cohere\_vector or openai\_vector benchmarks in this [link](https://elasticsearch-benchmarks.elastic.co/)? Is it because the dataset is too big to run nightly tests?

Best,

Tom

---

<div class="post-metadata">

**Author:** ![json](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/json/32/4125_2.png) [@json](https://discuss.elastic.co/u/json)\
**Post date:** [May 13, 2024, 1:31pm UTC](https://discuss.elastic.co/t/wait-until-merges-finish-after-index-operations/359286/5 "2024-05-13T13:31:15Z")

</div>

Hi Tom,

There is no particular reason other than the developers of the cohere\_vector and openai\_vector tracks chose not to have them included in the nightly regression test benchmarks.

Thanks,  
Jason

---

<div class="post-metadata">

**Author:** ![tommyd](https://avatars.discourse-cdn.com/v4/letter/t/f9ae1b/32.png) [@tommyd](https://discuss.elastic.co/u/tommyd)\
**Post date:** [May 13, 2024, 1:44pm UTC](https://discuss.elastic.co/t/wait-until-merges-finish-after-index-operations/359286/6 "2024-05-13T13:44:20Z")

</div>

Thank you, sir!

Tom

---

<div class="post-metadata">

**Author:** ![tommyd](https://avatars.discourse-cdn.com/v4/letter/t/f9ae1b/32.png) [@tommyd](https://discuss.elastic.co/u/tommyd)\
**Post date:** [May 15, 2024, 7:14pm UTC](https://discuss.elastic.co/t/wait-until-merges-finish-after-index-operations/359286/7 "2024-05-15T19:14:41Z")

</div>

Hi Jason,

May I ask another question? I am not exactly sure what this operation does (standalone-search-knn-100-1000-multiple-clients) in the cohere\_vector track? I checked the default.json under the operations directory but still not clear what it does. Instead of guessing, could you help explain its operation?

Thanks,

Tom

---

<div class="post-metadata">

**Author:** ![json](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/json/32/4125_2.png) [@json](https://discuss.elastic.co/u/json)\
**Post date:** [May 15, 2024, 9:08pm UTC](https://discuss.elastic.co/t/wait-until-merges-finish-after-index-operations/359286/8 "2024-05-15T21:08:27Z")

</div>

> [@tommyd](#):
>
> I am not exactly sure what this operation does (standalone-search-knn-100-1000-multiple-clients) in the cohere\_vector track?

Hi Tom,

I will break it down from the top:

1. The `standalone-search-knn-100-1000-multiple-clients` task performs the `knn-search-100-1000` operation with some number of search clients (default `8`, configured with track parameter `standalone_search_clients`) and number of iterations (default `10000`, configured with track parameter `standalone_search_iterations`). Each client executes the same number of iterations.

2. In the `knn-search-100-1000` [operation](https://github.com/elastic/rally-tracks/blob/26603b08ab5e187665894413405ce62a89469c30/cohere_vector/operations/default.json#L35-L41):  
a. It is a search operation.  
b. The parameter source is `knn-param-source`  
c. Parameter `k` is set to `100`.  
d. Parameter `num-candidates` is `1000`.

What does it mean?

`knn-param-source` is [registered](https://github.com/elastic/rally-tracks/blob/26603b08ab5e187665894413405ce62a89469c30/cohere_vector/track.py#L7) in the track's `track.py` file from class `KnnParamSource`. Without getting too much into the specifics of `KnnParamSource`, parameters `k` and `num-candidates` are used to build a search request body in the form of a Knn query similar to those found at [Knn query | Elasticsearch Guide [8.13] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-knn-query.html). Each execution of the search (for the configured number of iterations) uses a query vector from the [`queries.json`](https://github.com/elastic/rally-tracks/blob/26603b08ab5e187665894413405ce62a89469c30/cohere_vector/queries.json), also found in the track.

The Knn query docs referenced above better describe Knn queries, composition, and functionality.

Thank you,  
Jason

---

<div class="post-metadata">

**Author:** ![tommyd](https://avatars.discourse-cdn.com/v4/letter/t/f9ae1b/32.png) [@tommyd](https://discuss.elastic.co/u/tommyd)\
**Post date:** [May 15, 2024, 9:33pm UTC](https://discuss.elastic.co/t/wait-until-merges-finish-after-index-operations/359286/9 "2024-05-15T21:33:52Z")

</div>

Thank you, Jason, for the detailed explanation! Very much appreciate it.

Best,

Tom
