# Optimizing for reads with very small and stable index

**URL:** <https://discuss.elastic.co/t/optimizing-for-reads-with-very-small-and-stable-index/328442>\
**Category:** Elasticsearch\
**Created:** [March 24, 2023, 9:00am UTC](https://discuss.elastic.co/t/optimizing-for-reads-with-very-small-and-stable-index/328442 "2023-03-24T09:00:50Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Meisseli](https://avatars.discourse-cdn.com/v4/letter/m/73ab20/32.png) [@Meisseli](https://discuss.elastic.co/u/Meisseli)\
**Post date:** [March 24, 2023, 9:00am UTC](https://discuss.elastic.co/t/optimizing-for-reads-with-very-small-and-stable-index/328442/1 "2023-03-24T09:00:50Z")

</div>

Hello!

I'm working with a very small index. It is only 65000 documents and is about 50MB in size. Index is also very "stable" since it is only written once a day and does not receive any writes meanwhile. Because of that, I do not have to worry about the performance for indexing.

My goal is to maximize performance for search: number of concurrent searches and the search latency. High availability is a nice bonus, but not the most important part.

I have read extensive number of guides, documentations and tutorials about this subject. I have also benchmarked several different setups. However, since my use case with very small index seems to be so uncommon, I do not know the suitable "basic setup" to start with. For example it is usually recommended to have 1 shard for 40G of data. But on the other hand there should be at least 1 shard per node...

I have now 3 node cluster with 1 shard and 3 read replicas. I have also experimented and benchmarked with other options. I will almost always end up with somehow unbalanced setup with only 2 nodes actually taking the load and 1 staying idle.

What would be your recommendations for basic setup for my use case from where I could begin? Number of nodes? Number of shards? Number of replicas? Anything else I should know? 😊 )

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [March 24, 2023, 9:11am UTC](https://discuss.elastic.co/t/optimizing-for-reads-with-very-small-and-stable-index/328442/2 "2023-03-24T09:11:22Z")

</div>

You probably want a 3 node cluster where all nodes have the same profile (master/data). As you have a single small index I would stick with 1 primary shard and 2 replica shards. Make sure your clients are set up to load balance requests across all three nodes in the cluster. You may also consider using ['\_local' preference](https://www.elastic.co/guide/en/elasticsearch/reference/8.6/search-search.html#search-preference) (do not think this is default).

This way all nodes hold a copy of the data and can serve it locally. As there is a single primary shard you optimize the number of concurrent searches the cluster can handle and the shard size should not cause any performance issues.

---

<div class="post-metadata">

**Author:** ![Meisseli](https://avatars.discourse-cdn.com/v4/letter/m/73ab20/32.png) [@Meisseli](https://discuss.elastic.co/u/Meisseli)\
**Post date:** [March 24, 2023, 11:40am UTC](https://discuss.elastic.co/t/optimizing-for-reads-with-very-small-and-stable-index/328442/3 "2023-03-24T11:40:43Z")

</div>

Thank you for your brief answer! 🙏

Couple of follow-ups, if you may:

- I have not defined node.roles at all at the moment, so I think that all nodes has now multiple roles (master, data, data\_content, data\_hot , ingest, ml etc....) as a default. Should I specifically define node.roles to [master, data] for all nodes instead of this default setup? Do I need to mark all nodes as master or just one node?

- I'm using REST API instead of "direct" client library. I think, that load balancing is set upped there out-of-the-box, am I right?

I'm excited to benchmark these new settings! It is great to get a decent starting point, so thank you very much for your input!

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [March 28, 2023, 9:41am UTC](https://discuss.elastic.co/t/optimizing-for-reads-with-very-small-and-stable-index/328442/4 "2023-03-28T09:41:57Z")

</div>

> [@Meisseli](#):
>
> I have not defined node.roles at all at the moment, so I think that all nodes has now multiple roles (master, data, data\_content, data\_hot , ingest, ml etc....) as a default. Should I specifically define node.roles to [master, data] for all nodes instead of this default setup? Do I need to mark all nodes as master or just one node?

You can leave all roles enabled for all nodes.

> [@Meisseli](#):
>
> I'm using REST API instead of "direct" client library. I think, that load balancing is set upped there out-of-the-box, am I right?

The client determines which node or nodes it connects to so this will depend on the client configuration and design.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 25, 2023, 9:42am UTC](https://discuss.elastic.co/t/optimizing-for-reads-with-very-small-and-stable-index/328442/5 "2023-04-25T09:42:21Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
