# Elasticsearch architecture and planing

**URL:** <https://discuss.elastic.co/t/elasticsearch-architecture-and-planing/125504>\
**Category:** Elasticsearch\
**Created:** [March 25, 2018, 7:13pm UTC](https://discuss.elastic.co/t/elasticsearch-architecture-and-planing/125504 "2018-03-25T19:13:11Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![learnhub17](https://avatars.discourse-cdn.com/v4/letter/l/91b2a8/32.png) [@learnhub17](https://discuss.elastic.co/u/learnhub17)\
**Post date:** [March 25, 2018, 7:13pm UTC](https://discuss.elastic.co/t/elasticsearch-architecture-and-planing/125504/1 "2018-03-25T19:13:11Z")

</div>

Hi, we are doing poc to use ES in production, I am trying to learn best practice how to use it in production.  
My concern point is:

1. When we have a heavy write environment then what is the best architecture we configure.
2. How to architect index and number of shards.
3. Configure index setting which only contains last 5 months data, and delete automatic rest of data.

My current setup is: Es 6.2.2  
3master (4GB, 2CPU), 2client(16GB RAM,4CPU), 4Data(64GB RAM,16CPU), all servers half of memory is configured for heap size.  
per day write data is: 20GB

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [March 25, 2018, 7:41pm UTC](https://discuss.elastic.co/t/elasticsearch-architecture-and-planing/125504/2 "2018-03-25T19:41:20Z")

</div>

May I suggest you look at the following resources about sizing:

[https://www.elastic.co/elasticon/conf/2016/sf/quantitative-cluster-sizing](https://www.elastic.co/elasticon/conf/2016/sf/quantitative-cluster-sizing)

> **[How many shards should I have in my Elasticsearch cluster?
	  	 | Elastic](https://www.elastic.co/blog/how-many-shards-should-i-have-in-my-elasticsearch-cluster)**
>
> Elasticsearch is a very versatile platform, that supports a variety of use cases, and provides great flexibility around data organisation and replication strategies. This flexibility can however somet...

> **[NetSecureDay: Managing your Black Friday Logs](https://speakerdeck.com/elastic/netsecureday-managing-your-black-friday-logs)**
>
> Surveiller une application complexe n’est pas une tâche aisée, mais avec les bons outils, ce n’est pas si sorcier. Néanmoins, des périodes fortes telles que les opérations de type « Black Friday » (Vendredi noir) ou période de Noël peuvent pousser...

---

<div class="post-metadata">

**Author:** ![rcowart](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rcowart/32/88091_2.png) [@rcowart](https://discuss.elastic.co/u/rcowart)\
**Post date:** [March 26, 2018, 8:12am UTC](https://discuss.elastic.co/t/elasticsearch-architecture-and-planing/125504/3 "2018-03-26T08:12:50Z")

</div>

First off 20GB of indexed data per day is not really very much data. Last week I setup a POC for a user, where 100GB per day is indexed and stored on a single node. On this node is Elasticsearch, Kibana and Logstash (with a very complex resource intensive pipeline). The hardware is 16 cores, 128GB RAM and data stored on 8 spinning disks. The disks are the real limiter here. They were already configured in RAID-5, which is clearly not ideal for write-biased workloads, and will be reconfigured.

My point is... your cluster is overkill for 20GB per day... in fact it is likely overkill for 100GB per day. I would recommend a more simple 3 node cluster, with these characteristics:

- Ideally these nodes would use SSD storage. If multiple drives these should be configured as JBOD. Index replicas will provide the necessary redundancy.

- 64GB is a good starting point for RAM. More can be even better as it provides more page cache space for the OS to cache disk IO.

- If you have a larger number of smaller indexes (\<5GB), start with 2 shards and 1 replica. For a small number of larger indicies, 3 shards and 1 replica will better spread the load. The rule is that increasing shards increases ingestion performance, while increasing replicas improves query performance.

Basically the resources saved reducing from the 9 nodes you mention, to the 3 nodes that you really need, can be invested in the best storage and more RAM for those 3 nodes.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 23, 2018, 8:12am UTC](https://discuss.elastic.co/t/elasticsearch-architecture-and-planing/125504/4 "2018-04-23T08:12:58Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
