Dear Team,
We are planning to implement an on-premises Elastic SIEM/EDR solution for an OT/IT security environment comprising approximately 60 endpoints and an estimated daily log ingestion volume of 50 GB/day.
Before finalizing the High-Level Design (HLD) and proceeding with implementation, we request your technical validation of the proposed architecture and VM resource allocation.
1. Proposed Architecture and VM Sizing
| Site | VM | Proposed Components | vCPU | RAM |
|---|---|---|---|---|
| DC | VM-1 | Vector + Kafka + Logstash (Ingestion Gateway) | 4 | 12 GB |
| DC | VM-2 | Elasticsearch (Hot/Master Role) + Kibana | 8 | 15 GB |
| DC | VM-3 | Elasticsearch (Cold Role) | 4 | 8 GB |
| DR | VM-4 | Vector + Kafka + Logstash (Ingestion Gateway) | 4 | 12 GB |
| DR | VM-5 | Elasticsearch (Hot/Master Role) + Kibana | 8 | 15 GB |
| DR | VM-6 | Elasticsearch (Cold Role) | 4 | 8 GB |
| Third/Neutral Site | VM-7 | Elasticsearch (Frozen/Master Role) and object storage connectivity | 4 | 8 GB |
Object storage: Approximately 10.1 TB proposed for archive/frozen data.
2. Validation Required
Kindly review and confirm the following:
-
Ingestion capacity: Are the proposed CPU and RAM resources sufficient to process 50 GB/day of OT/IT security logs, including peak ingestion rates and burst traffic?
-
Ingestion gateway sizing: Is it appropriate to deploy Vector, Kafka, and Logstash together on a single VM at each site with 4 vCPU and 12 GB RAM? Please validate Kafka throughput, queue retention, buffering, and back-pressure handling.
-
Elasticsearch sizing: Are the proposed resources for the Hot and Cold nodes sufficient for the expected ingestion volume, search performance, security analytics, alerting, dashboards, and concurrent SOC users?
-
Master-node architecture: Is the proposed combination of Elasticsearch Hot and Master roles on VM-2 and VM-5 recommended for production? Please advise on dedicated master nodes and the minimum supported topology for this deployment.
-
DC–DR architecture: Is the proposed architecture suitable for disaster recovery? Please clarify the recommended approach for cross-site replication, cluster separation, failover, data consistency, and recovery objectives (RPO/RTO).
-
Third/Neutral Site: Please validate the proposed Frozen/Master role on VM-7. Should the third-site VM act as a voting-only master/quorum node, a frozen-tier data node, or separate components? Please clarify the supported design and whether the proposed 4 vCPU and 8 GB RAM are sufficient.
-
Retention and storage: Please validate whether the proposed 10.1 TB object storage is adequate for the required retention period, considering 50 GB/day ingestion, Elasticsearch indexing overhead, compression, replicas, snapshots, and growth.
-
OT/IT security requirements: Please confirm whether the proposed sizing supports security monitoring and detection use cases for network devices, firewalls, servers, endpoints, and OT/ICS-related data sources.
-
EDR requirements: Please clarify whether the proposed sizing includes endpoint telemetry processing, EDR detection, and associated integrations, or whether additional Elastic Security/EDR components and resources are required.
-
Final sizing recommendation: Please provide the OEM-recommended production sizing for each VM, including vCPU, RAM, disk capacity, storage performance/IOPS, and network requirements, along with any architectural changes required.
3. Confirmation Requested
Please confirm whether the proposed architecture and VM resources can be approved for production implementation at the stated ingestion volume, or provide a revised, OEM-supported architecture and sizing recommendation..
Regards
Chandra kamal