# Handle 1 million inserts per second

**URL:** <https://discuss.elastic.co/t/handle-1-million-inserts-per-second/139387>\
**Category:** Elasticsearch\
**Created:** [July 10, 2018, 3:35pm UTC](https://discuss.elastic.co/t/handle-1-million-inserts-per-second/139387 "2018-07-10T15:35:54Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![hoey123](https://avatars.discourse-cdn.com/v4/letter/h/e480ec/32.png) [@hoey123](https://discuss.elastic.co/u/hoey123)\
**Post date:** [July 10, 2018, 3:35pm UTC](https://discuss.elastic.co/t/handle-1-million-inserts-per-second/139387/1 "2018-07-10T15:35:55Z")

</div>

we should implement really big data scenario with this features:

- input data is 1 million logs per second
- each log has almost 100 Bytes size
- we should retain data for 10 days
- system has 2,3 active users and they run maybe 1000 queries per day so our scenario is definitely heavy write

with this assumptions ( so many writes and small number of reads) , how much data each node in cluster should store ?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 10, 2018, 4:29pm UTC](https://discuss.elastic.co/t/handle-1-million-inserts-per-second/139387/2 "2018-07-10T16:29:06Z")

</div>

Is that 1 million events per second an average or a peak volume? What type of hardware are you planning to deploy this on? What type of data is being indexed?

---

<div class="post-metadata">

**Author:** ![hoey123](https://avatars.discourse-cdn.com/v4/letter/h/e480ec/32.png) [@hoey123](https://discuss.elastic.co/u/hoey123)\
**Post date:** [July 11, 2018, 8:00am UTC](https://discuss.elastic.co/t/handle-1-million-inserts-per-second/139387/3 "2018-07-11T08:00:27Z")

</div>

1 million is average volume  
data is structured and has 10-15 fields. must of columns are string and integer.  
My question is exactly that. what type of hardware we should use and how much resource we should consider for each node ( for maximum utilization of each server based on our heavy write scenario) ?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 11, 2018, 8:07am UTC](https://discuss.elastic.co/t/handle-1-million-inserts-per-second/139387/4 "2018-07-11T08:07:55Z")

</div>

If you want to keep up with the flow off data you probably need to size the cluster based on peak indexing rate rather than the average. The average will however tell you how much disk space you are likely to need in the cluster.

As it is an indexing heavy use case and indexing is very I/O intensive you will benefit from having data nodes with fast, locally attached SDDs. You probably also need a good amount of CPU and fast network.

When estimating the size of the cluster needed, you will need to run some benchmarks on realistic data and hardware. The following resources may help:

[https://www.elastic.co/elasticon/conf/2018/sf/the-seven-deadly-sins-of-elasticsearch-benchmarking](https://www.elastic.co/elasticon/conf/2018/sf/the-seven-deadly-sins-of-elasticsearch-benchmarking)

[https://www.elastic.co/elasticon/conf/2016/sf/quantitative-cluster-sizing](https://www.elastic.co/elasticon/conf/2016/sf/quantitative-cluster-sizing)

[https://www.elastic.co/webinars/using-rally-to-get-your-elasticsearch-cluster-size-right](https://www.elastic.co/webinars/using-rally-to-get-your-elasticsearch-cluster-size-right)

> **[Filebeat modules, access logs and Elasticsearch storage requirements
	  	 |...](https://www.elastic.co/blog/filebeat-modiles-access-logs-and-elasticsearch-storage-requirements)**
>
> Elastic recently introduced Filebeat Modules, which are designed to make it extremely easy to ingest and gain insights from common log formats. These follow the principle that

> **[How many shards should I have in my Elasticsearch cluster?
	  	 | Elastic](https://www.elastic.co/blog/how-many-shards-should-i-have-in-my-elasticsearch-cluster)**
>
> Elasticsearch is a very versatile platform, that supports a variety of use cases, and provides great flexibility around data organisation and replication strategies. This flexibility can however somet...

[https://www.elastic.co/guide/en/elasticsearch/reference/6.3/tune-for-indexing-speed.html](https://www.elastic.co/guide/en/elasticsearch/reference/6.3/tune-for-indexing-speed.html)

---

<div class="post-metadata">

**Author:** ![hoey123](https://avatars.discourse-cdn.com/v4/letter/h/e480ec/32.png) [@hoey123](https://discuss.elastic.co/u/hoey123)\
**Post date:** [July 11, 2018, 8:10am UTC](https://discuss.elastic.co/t/handle-1-million-inserts-per-second/139387/5 "2018-07-11T08:10:31Z")

</div>

is that any limitation for shard size or index size ?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 11, 2018, 8:13am UTC](https://discuss.elastic.co/t/handle-1-million-inserts-per-second/139387/6 "2018-07-11T08:13:06Z")

</div>

You need to find an optimal size through benchmarking as it depends on the data, queries and the use case. A shard can hold up to 2 billion documents, but you are likely to see query performance deteriorate before you get close to that. Each index can have many shards, so there is no strict limit there.

---

<div class="post-metadata">

**Author:** ![hoey123](https://avatars.discourse-cdn.com/v4/letter/h/e480ec/32.png) [@hoey123](https://discuss.elastic.co/u/hoey123)\
**Post date:** [July 11, 2018, 8:26am UTC](https://discuss.elastic.co/t/handle-1-million-inserts-per-second/139387/7 "2018-07-11T08:26:11Z")

</div>

ok i should get peak volume too.  
we should retain this data 10 days and I think SSD for this size is very expensive.  
how can I find detail of selecting server and resource recommendation for elastic ?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [July 11, 2018, 8:55am UTC](https://discuss.elastic.co/t/handle-1-million-inserts-per-second/139387/8 "2018-07-11T08:55:51Z")

</div>

Given that you have a relatively short retention period, the data will not sit idle on disk for long, so I think you will need SSDs at least for all the indices being indexed into. You may want to consider a [hot/warm architecture](https://www.elastic.co/blog/hot-warm-architecture-in-elasticsearch-5-x), but the retention period might be a bit short to properly warrant this.

Are you looking to deploy on bare-metal hardware or in the cloud?

---

<div class="post-metadata">

**Author:** ![hoey123](https://avatars.discourse-cdn.com/v4/letter/h/e480ec/32.png) [@hoey123](https://discuss.elastic.co/u/hoey123)\
**Post date:** [July 11, 2018, 11:07am UTC](https://discuss.elastic.co/t/handle-1-million-inserts-per-second/139387/9 "2018-07-11T11:07:06Z")

</div>

> [@Christian\_Dahlqvist](#):
>
> Are you looking to deploy on bare-metal hardware or in the cloud?

no I don't  
Ok thank you very much

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 8, 2018, 11:07am UTC](https://discuss.elastic.co/t/handle-1-million-inserts-per-second/139387/10 "2018-08-08T11:07:11Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
