# Force merge optimise in background process

**URL:** <https://discuss.elastic.co/t/force-merge-optimise-in-background-process/337768>\
**Category:** Elasticsearch\
**Created:** [July 6, 2023, 9:00am UTC](https://discuss.elastic.co/t/force-merge-optimise-in-background-process/337768 "2023-07-06T09:00:10Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![INS](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ins/32/92827_2.png) [@INS](https://discuss.elastic.co/u/INS)\
**Post date:** [July 6, 2023, 9:00am UTC](https://discuss.elastic.co/t/force-merge-optimise-in-background-process/337768/1 "2023-07-06T09:00:10Z")

</div>

Hi  
I need to change the force merge process in the background for count of segments.  
Also how I can steering/manipulating force merge. In my case I have the index which is really updating by data. So If this index is so tiered I can observe low performance on searching.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 6, 2023, 9:16am UTC](https://discuss.elastic.co/t/force-merge-optimise-in-background-process/337768/2 "2023-07-06T09:16:08Z")

</div>

If the index is still being updated or indexed to, why would you forcemerge?

Forcemerging can be I/O intensive. What type of storage are you using? Local SSDs?

---

<div class="post-metadata">

**Author:** ![INS](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ins/32/92827_2.png) [@INS](https://discuss.elastic.co/u/INS)\
**Post date:** [July 6, 2023, 10:35am UTC](https://discuss.elastic.co/t/force-merge-optimise-in-background-process/337768/3 "2023-07-06T10:35:58Z")

</div>

Hi We're using SSD disk with P30 tier on Azure,  
"Force merge" with less segments gives in performance search test, much more better results.

It allows to delete nested fields.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 6, 2023, 10:56am UTC](https://discuss.elastic.co/t/force-merge-optimise-in-background-process/337768/4 "2023-07-06T10:56:30Z")

</div>

Indexing and updating continously create new segments, so I do not see much point forcemerging an index actively indexed into. If you modified data only periodically it might make sense though.

Are you indexing/updating through bulk requests? What is your refresh interval set to?

---

<div class="post-metadata">

**Author:** ![INS](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ins/32/92827_2.png) [@INS](https://discuss.elastic.co/u/INS)\
**Post date:** [July 7, 2023, 1:18pm UTC](https://discuss.elastic.co/t/force-merge-optimise-in-background-process/337768/5 "2023-07-07T13:18:20Z")

</div>

Yes through bulk request, refresh interval was set to 60s

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 7, 2023, 1:30pm UTC](https://discuss.elastic.co/t/force-merge-optimise-in-background-process/337768/6 "2023-07-07T13:30:58Z")

</div>

How many inserts/updates are you performing per second? How many documents are there in the index?

Which version of Elasticsearch are you using?

> [@INS](#):
>
> "Force merge" with less segments gives in performance search test, much more better results.

How many segments did you forcemerge into? Did this test include concurrent indexing and updates?

---

<div class="post-metadata">

**Author:** ![INS](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ins/32/92827_2.png) [@INS](https://discuss.elastic.co/u/INS)\
**Post date:** [July 7, 2023, 1:45pm UTC](https://discuss.elastic.co/t/force-merge-optimise-in-background-process/337768/7 "2023-07-07T13:45:18Z")

</div>

Indexing rate is avr ~250/s so this index consists with 32mln of docs  
We're using 7.17 of ES

BTW. I can't find the topic but as far I as remember You've claimed that search performing on 1 shard and 2 replica (on 3 data nodes) should give much more better results then 3 shard with 1 replica. You was referring to using cache from this replica shard. But when I'checked it from .monitoring-es\* the param "index\_stats.total.query\_cache.hit\_count" was pointed out for the nodes with primary shards only.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 7, 2023, 1:53pm UTC](https://discuss.elastic.co/t/force-merge-optimise-in-background-process/337768/8 "2023-07-07T13:53:31Z")

</div>

The ideal number of primary and replica shards typically depend on whether you are optimising for latency or the number of concurrent queries that the cluster can support.

If a single primary shard give acceptable query latencies scaling out replicas will allow more nodes to handle queries. If a single primary shard is too large to support acceptable query latencies you may need more smaller primary shards to improve concurrency. This however often lead to fewer cobcurrent queries being possible on the same hardware.

---

<div class="post-metadata">

**Author:** ![INS](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ins/32/92827_2.png) [@INS](https://discuss.elastic.co/u/INS)\
**Post date:** [July 7, 2023, 2:33pm UTC](https://discuss.elastic.co/t/force-merge-optimise-in-background-process/337768/9 "2023-07-07T14:33:19Z")

</div>

The index as mentioned is 28GB in size, from the monitoring it can be seen that other nodes with replicas are used along with the query (I infer this from the mem and CPU usage) but I do not see that it uses cach. Does the ES mechanism only allow the use of cach from the primary shard?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 4, 2023, 2:34pm UTC](https://discuss.elastic.co/t/force-merge-optimise-in-background-process/337768/10 "2023-08-04T14:34:05Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
