# Massive scale elasticsearch

**URL:** <https://discuss.elastic.co/t/massive-scale-elasticsearch/9499>\
**Category:** Elasticsearch\
**Created:** [October 30, 2012, 10:54am UTC](https://discuss.elastic.co/t/massive-scale-elasticsearch/9499 "2012-10-30T10:54:16Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Robin\_Verlangen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/robin_verlangen/32/1542_2.png) [@Robin\_Verlangen](https://discuss.elastic.co/u/Robin_Verlangen)\
**Post date:** [October 30, 2012, 10:54am UTC](https://discuss.elastic.co/t/massive-scale-elasticsearch/9499/1 "2012-10-30T10:54:16Z")

</div>

Hi there,

I was wondering whether there are users that deploy ES on massive scale.  
For example 50TB of data. This is currently stored in HDFS and searching is  
possible with Map/Reduce jobs. However something like ES would be really  
sweet.

Is it possible? If so, what kind of cluster should you think of? Currently  
we run an eight-node cluster with 12x1TB disk per server, 16GB RAM and  
dual-quadcore.

Best regards,

Robin Verlangen  
_Software engineer_  
\*  
\*  
W [http://www.robinverlangen.nl](http://www.robinverlangen.nl)  
E robin@us2.nl

[http://goo.gl/Lt7BC](http://goo.gl/Lt7BC)

Disclaimer: The information contained in this message and attachments is  
intended solely for the attention and use of the named addressee and may be  
confidential. If you are not the intended recipient, you are reminded that  
the information remains the property of the sender. You must not use,  
disclose, distribute, copy, print or rely on this e-mail. If you have  
received this message in error, please contact the sender immediately and  
irrevocably delete this message and any copies.

--

---

<div class="post-metadata">

**Author:** ![radu\_gheorghe](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/radu_gheorghe/32/556_2.png) [@radu\_gheorghe](https://discuss.elastic.co/u/radu_gheorghe)\
**Post date:** [October 30, 2012, 11:49am UTC](https://discuss.elastic.co/t/massive-scale-elasticsearch/9499/2 "2012-10-30T11:49:41Z")

</div>

Hello Robin,

Yes, there are people using ES on massive scale. Take a look at this video  
for such an usecase:

> **[Elastic — The Search AI Company](https://www.elastic.co)**
>
> Power insights and outcomes with The Elastic Search AI Platform. See into your data and find answers that matter with enterprise solutions designed to help you accelerate time to insight. Try Elastic ...

Regarding hardware requirements, it depends on lots of stuff. How your data  
looks like, how your queries look like, if you want to search everywhere in  
your data or only on certain fields, etc.

I'd start with a fraction of those 50TB on one (or a few) test machines and  
do some performance testing to see how it goes. Then you will get a better  
estimation.

And if you want to monitor your cluster while testing, there are quite a  
few options. Of course I'd recommend ours 🙂

> **[Elasticsearch - Sematext Documentation](https://sematext.com/docs/integration/elasticsearch-integration/)**
>
> Collect and monitor key Elasticsearch metrics such as request latency, indexing rate, and segment merges with built-in anomaly detection, threshold, and heartbeat alerts. Send notifications to email and various chatops messaging services, correlate...

## Best regards, Radu

[http://sematext.com/](http://sematext.com/) -- Elasticsearch -- Solr -- Lucene

On Tue, Oct 30, 2012 at 12:54 PM, Robin Verlangen [robin@us2.nl](mailto:robin@us2.nl) wrote:

> Hi there,
> 
> I was wondering whether there are users that deploy ES on massive scale.  
> For example 50TB of data. This is currently stored in HDFS and searching is  
> possible with Map/Reduce jobs. However something like ES would be really  
> sweet.
> 
> Is it possible? If so, what kind of cluster should you think of? Currently  
> we run an eight-node cluster with 12x1TB disk per server, 16GB RAM and  
> dual-quadcore.
> 
> Best regards,
> 
> Robin Verlangen  
> _Software engineer_  
> \*  
> \*  
> W [http://www.robinverlangen.nl](http://www.robinverlangen.nl)  
> E robin@us2.nl

--

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [October 31, 2012, 1:36am UTC](https://discuss.elastic.co/t/massive-scale-elasticsearch/9499/3 "2012-10-31T01:36:18Z")

</div>

Hello Robin,

Some of Sematext's clients use ES on that sort of scale. 50TB of data is  
not precise enough if you are referring to raw data - who knows how much of  
that will be indexed, how much stored, etc. 16GB RAM sounds lowish, unless  
indices are not more than 50GB, say. But that is also not an accurate  
statement, because even a 100GB may be fine if you only ever hit it with a  
handful of queries. Or if you use routing well. Or if you are OK with  
high latency, of course 🙂

## Otis

Search Analytics - [Cloud Monitoring Tools & Services | Sematext](http://sematext.com/search-analytics/index.html)  
Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)

On Tuesday, October 30, 2012 6:54:19 AM UTC-4, Robin Verlangen wrote:

> Hi there,
> 
> I was wondering whether there are users that deploy ES on massive scale.  
> For example 50TB of data. This is currently stored in HDFS and searching is  
> possible with Map/Reduce jobs. However something like ES would be really  
> sweet.
> 
> Is it possible? If so, what kind of cluster should you think of? Currently  
> we run an eight-node cluster with 12x1TB disk per server, 16GB RAM and  
> dual-quadcore.
> 
> Best regards,
> 
> Robin Verlangen  
> _Software engineer_  
> \*  
> \*  
> W [http://www.robinverlangen.nl](http://www.robinverlangen.nl)  
> E ro...@us2.nl \<javascript:\>
> 
> [http://goo.gl/Lt7BC](http://goo.gl/Lt7BC)
> 
> Disclaimer: The information contained in this message and attachments is  
> intended solely for the attention and use of the named addressee and may be  
> confidential. If you are not the intended recipient, you are reminded that  
> the information remains the property of the sender. You must not use,  
> disclose, distribute, copy, print or rely on this e-mail. If you have  
> received this message in error, please contact the sender immediately and  
> irrevocably delete this message and any copies.

--

---

<div class="post-metadata">

**Author:** ![Chuck\_McKenzie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chuck_mckenzie/32/19660_2.png) [@Chuck\_McKenzie](https://discuss.elastic.co/u/Chuck_McKenzie)\
**Post date:** [November 3, 2012, 5:14am UTC](https://discuss.elastic.co/t/massive-scale-elasticsearch/9499/4 "2012-11-03T05:14:07Z")

</div>

Agreed on 50TB meaning a lot of different things. 50TB of .gz raw data is  
a lot different from 50TB of big fluffy JSON is a lot different from 50TB  
of raid-10, replicated indices. I've started telling our internal users  
that I can give them any number they want for "size", because it can vary  
by two orders of magnitude based on where it is in the pipe. If they want  
a number that actually means something reasonably consistent, I'll give  
them number of docs.

We have around 30 billion documents online in ES; for us, a billion docs is  
a good per-server limit. Each server uses about 1.5TB of disk for indices,  
double that for a replica, triple it if you want to be able to reindex from  
hdfs without taking existing data offline until you're done and switch  
aliases over. Double it again if you aren't using compression. (Use  
compression.) ES does very well at that sort of scale (beats the pants off  
SOLR, at any rate), although it will occasionally break in exotic ways, and  
it can be hard to figure out which query is problematic if you're throwing  
a lot of traffic at it. Our servers all have(and need) 96GB of ram, split  
evenly between java heap and OS cache, but your mileage will vary a lot on  
that based on your specific use case.

On Tuesday, October 30, 2012 8:36:18 PM UTC-5, Otis Gospodnetic wrote:

> Hello Robin,
> 
> Some of Sematext's clients use ES on that sort of scale. 50TB of data is  
> not precise enough if you are referring to raw data - who knows how much of  
> that will be indexed, how much stored, etc. 16GB RAM sounds lowish, unless  
> indices are not more than 50GB, say. But that is also not an accurate  
> statement, because even a 100GB may be fine if you only ever hit it with a  
> handful of queries. Or if you use routing well. Or if you are OK with  
> high latency, of course 🙂
> 
> ## Otis
> 
> Search Analytics - [Cloud Monitoring Tools & Services | Sematext](http://sematext.com/search-analytics/index.html)  
> Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> 
> On Tuesday, October 30, 2012 6:54:19 AM UTC-4, Robin Verlangen wrote:
> 
> > Hi there,
> > 
> > I was wondering whether there are users that deploy ES on massive scale.  
> > For example 50TB of data. This is currently stored in HDFS and searching is  
> > possible with Map/Reduce jobs. However something like ES would be really  
> > sweet.
> > 
> > Is it possible? If so, what kind of cluster should you think of?  
> > Currently we run an eight-node cluster with 12x1TB disk per server, 16GB  
> > RAM and dual-quadcore.
> > 
> > Best regards,
> > 
> > Robin Verlangen  
> > _Software engineer_  
> > \*  
> > \*  
> > W [http://www.robinverlangen.nl](http://www.robinverlangen.nl)  
> > E ro...@us2.nl
> > 
> > [http://goo.gl/Lt7BC](http://goo.gl/Lt7BC)
> > 
> > Disclaimer: The information contained in this message and attachments is  
> > intended solely for the attention and use of the named addressee and may be  
> > confidential. If you are not the intended recipient, you are reminded that  
> > the information remains the property of the sender. You must not use,  
> > disclose, distribute, copy, print or rely on this e-mail. If you have  
> > received this message in error, please contact the sender immediately and  
> > irrevocably delete this message and any copies.

--

---

<div class="post-metadata">

**Author:** ![Robin\_Verlangen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/robin_verlangen/32/1542_2.png) [@Robin\_Verlangen](https://discuss.elastic.co/u/Robin_Verlangen)\
**Post date:** [November 7, 2012, 11:26am UTC](https://discuss.elastic.co/t/massive-scale-elasticsearch/9499/5 "2012-11-07T11:26:14Z")

</div>

Thank you all for the information. The number I picked was an estimate  
based on a forecast of growth. We will be using daily indices and most of  
the time only using the last couple of days for searching. I think ES will  
be a good choice for this as it seems to scale pretty well!

Best regards,

Robin Verlangen  
_Software engineer_  
\*  
\*  
W [http://www.robinverlangen.nl](http://www.robinverlangen.nl)  
E robin@us2.nl

[http://goo.gl/Lt7BC](http://goo.gl/Lt7BC)

Disclaimer: The information contained in this message and attachments is  
intended solely for the attention and use of the named addressee and may be  
confidential. If you are not the intended recipient, you are reminded that  
the information remains the property of the sender. You must not use,  
disclose, distribute, copy, print or rely on this e-mail. If you have  
received this message in error, please contact the sender immediately and  
irrevocably delete this message and any copies.

On Sat, Nov 3, 2012 at 6:14 AM, Chuck McKenzie [redchuck@gmail.com](mailto:redchuck@gmail.com) wrote:

> Agreed on 50TB meaning a lot of different things. 50TB of .gz raw data is  
> a lot different from 50TB of big fluffy JSON is a lot different from 50TB  
> of raid-10, replicated indices. I've started telling our internal users  
> that I can give them any number they want for "size", because it can vary  
> by two orders of magnitude based on where it is in the pipe. If they want  
> a number that actually means something reasonably consistent, I'll give  
> them number of docs.
> 
> We have around 30 billion documents online in ES; for us, a billion docs  
> is a good per-server limit. Each server uses about 1.5TB of disk for  
> indices, double that for a replica, triple it if you want to be able to  
> reindex from hdfs without taking existing data offline until you're done  
> and switch aliases over. Double it again if you aren't using compression.  
> (Use compression.) ES does very well at that sort of scale (beats the  
> pants off SOLR, at any rate), although it will occasionally break in exotic  
> ways, and it can be hard to figure out which query is problematic if you're  
> throwing a lot of traffic at it. Our servers all have(and need) 96GB of  
> ram, split evenly between java heap and OS cache, but your mileage will  
> vary a lot on that based on your specific use case.
> 
> On Tuesday, October 30, 2012 8:36:18 PM UTC-5, Otis Gospodnetic wrote:
> 
> > Hello Robin,
> > 
> > Some of Sematext's clients use ES on that sort of scale. 50TB of data is  
> > not precise enough if you are referring to raw data - who knows how much of  
> > that will be indexed, how much stored, etc. 16GB RAM sounds lowish, unless  
> > indices are not more than 50GB, say. But that is also not an accurate  
> > statement, because even a 100GB may be fine if you only ever hit it with a  
> > handful of queries. Or if you use routing well. Or if you are OK with  
> > high latency, of course 🙂
> > 
> > ## Otis
> > 
> > Search Analytics - [http://sematext.com/search-\*\*analytics/index.html](http://sematext.com/search-**analytics/index.html)[http://sematext.com/search-analytics/index.html](http://sematext.com/search-analytics/index.html)  
> > Performance Monitoring - [http://sematext.com/spm/index.\*\*html](http://sematext.com/spm/index.**html)[http://sematext.com/spm/index.html](http://sematext.com/spm/index.html)
> > 
> > On Tuesday, October 30, 2012 6:54:19 AM UTC-4, Robin Verlangen wrote:
> > 
> > > Hi there,
> > > 
> > > I was wondering whether there are users that deploy ES on massive scale.  
> > > For example 50TB of data. This is currently stored in HDFS and searching is  
> > > possible with Map/Reduce jobs. However something like ES would be really  
> > > sweet.
> > > 
> > > Is it possible? If so, what kind of cluster should you think of?  
> > > Currently we run an eight-node cluster with 12x1TB disk per server, 16GB  
> > > RAM and dual-quadcore.
> > > 
> > > Best regards,
> > > 
> > > Robin Verlangen  
> > > _Software engineer_  
> > > \*  
> > > \*  
> > > W [http://www.robinverlangen.nl](http://www.robinverlangen.nl)  
> > > E ro...@us2.nl
> > > 
> > > [http://goo.gl/Lt7BC](http://goo.gl/Lt7BC)
> > > 
> > > Disclaimer: The information contained in this message and attachments is  
> > > intended solely for the attention and use of the named addressee and may be  
> > > confidential. If you are not the intended recipient, you are reminded that  
> > > the information remains the property of the sender. You must not use,  
> > > disclose, distribute, copy, print or rely on this e-mail. If you have  
> > > received this message in error, please contact the sender immediately and  
> > > irrevocably delete this message and any copies.
> > > 
> > > --

--

---

<div class="post-metadata">

**Author:** ![drewr](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/drewr/32/7803_2.png) [@drewr](https://discuss.elastic.co/u/drewr)\
**Post date:** [November 7, 2012, 2:26pm UTC](https://discuss.elastic.co/t/massive-scale-elasticsearch/9499/6 "2012-11-07T14:26:26Z")

</div>

Robin Verlangen wrote:

> > > > I was wondering whether there are users that deploy ES on  
> > > > massive scale. For example 50TB of data. This is currently  
> > > > stored in HDFS and searching is possible with Map/Reduce  
> > > > jobs. However something like ES would be really sweet.

[...]

> Thank you all for the information. The number I picked was an  
> estimate based on a forecast of growth. We will be using daily  
> indices and most of the time only using the last couple of days for  
> searching. I think ES will be a good choice for this as it seems to  
> scale pretty well!

Just to add a data point, we've comfortably run 60-100TiB clusters on  
15 m1.xlarges (16GB RAM) EC2 nodes with 8- or 6-stripe, 4-8TiB EBS  
volumes. That's with hundreds of 200GiB shards per node.

ES can easily handle the indexing and storage. Optimizing searching,  
as others have pointed out, is usually where you have to spend time  
getting it right.

-Drew

--

---

<div class="post-metadata">

**Author:** ![Robin\_Verlangen](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/robin_verlangen/32/1542_2.png) [@Robin\_Verlangen](https://discuss.elastic.co/u/Robin_Verlangen)\
**Post date:** [November 7, 2012, 6:45pm UTC](https://discuss.elastic.co/t/massive-scale-elasticsearch/9499/7 "2012-11-07T18:45:33Z")

</div>

Hi Drew,

Thank you for increasing the confidence I gained in ES. It seems exactly  
the right tool for what we need. If you want to search through such a  
volumes, it's obvious you'll need some hardware!

Best regards,

Robin Verlangen  
_Software engineer_  
\*  
\*  
W [http://www.robinverlangen.nl](http://www.robinverlangen.nl)  
E robin@us2.nl

[http://goo.gl/Lt7BC](http://goo.gl/Lt7BC)

Disclaimer: The information contained in this message and attachments is  
intended solely for the attention and use of the named addressee and may be  
confidential. If you are not the intended recipient, you are reminded that  
the information remains the property of the sender. You must not use,  
disclose, distribute, copy, print or rely on this e-mail. If you have  
received this message in error, please contact the sender immediately and  
irrevocably delete this message and any copies.

On Wed, Nov 7, 2012 at 3:26 PM, Drew Raines [aaraines@gmail.com](mailto:aaraines@gmail.com) wrote:

> Robin Verlangen wrote:
> 
> > > > > I was wondering whether there are users that deploy ES on  
> > > > > massive scale. For example 50TB of data. This is currently  
> > > > > stored in HDFS and searching is possible with Map/Reduce  
> > > > > jobs. However something like ES would be really sweet.
> 
> [...]
> 
> > Thank you all for the information. The number I picked was an  
> > estimate based on a forecast of growth. We will be using daily  
> > indices and most of the time only using the last couple of days for  
> > searching. I think ES will be a good choice for this as it seems to  
> > scale pretty well!
> 
> Just to add a data point, we've comfortably run 60-100TiB clusters on  
> 15 m1.xlarges (16GB RAM) EC2 nodes with 8- or 6-stripe, 4-8TiB EBS  
> volumes. That's with hundreds of 200GiB shards per node.
> 
> ES can easily handle the indexing and storage. Optimizing searching,  
> as others have pointed out, is usually where you have to spend time  
> getting it right.
> 
> -Drew
> 
> --

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:05am UTC](https://discuss.elastic.co/t/massive-scale-elasticsearch/9499/8 "2017-07-06T03:05:39Z")

</div>


