# Scaling Elasticsearch for 40GB of data

**URL:** <https://discuss.elastic.co/t/scaling-elasticsearch-for-40gb-of-data/9107>\
**Category:** Elasticsearch\
**Created:** [September 22, 2012, 4:55pm UTC](https://discuss.elastic.co/t/scaling-elasticsearch-for-40gb-of-data/9107 "2012-09-22T16:55:48Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jason\_Yankus](https://avatars.discourse-cdn.com/v4/letter/j/77aa72/32.png) [@Jason\_Yankus](https://discuss.elastic.co/u/Jason_Yankus)\
**Post date:** [September 22, 2012, 4:55pm UTC](https://discuss.elastic.co/t/scaling-elasticsearch-for-40gb-of-data/9107/1 "2012-09-22T16:55:48Z")

</div>

Greetings,

I am migrating a 40GB volume of mixed metadata and unstructured text into  
Elasticsearch (~500,000 documents with metadata + extracted text) and need  
some recommendations about setting up a cluster for high availability and  
high performance queries.

I have tried a few basic changes to the default configuration but have not  
had a lot of success avoiding:

1.) OOM/Heap Space Errors during indexing (I can only index about 150k docs  
during the ETL (Extract-Transform-Load) process before the system bails out  
after 3-4 hours)  
2.) Poor full text query performance (2+ seconds on an index size of  
roughly 30GB for 1 concurrent user)

We have two environments set up for test:

_Development_  
1x [8GB RAM 4x processor]

_Staging_  
2x [4GB RAM 2x processor]

My shard and replica settings are as follows:  
index.number\_of\_shards: 6  
index.number\_of\_replicas: 1

In each environment I have the ES\_HEAP\_SIZE value set to maximum physical  
memory - 1 GB (so 7192 on 8GB, 3072 on 4GB).

So, my questions are:

For a volume of data as I described, what is the expectation about the size  
of cluster (number of nodes, amount of ram) I will need to support high  
performance queries ( \<500ms per query with an expected average volume of  
25 concurrent users)

Second, after our initial ETL, we will only experience incremental indexing  
(as new docs come in or older docs are changed ~200/day). What sort of  
shard/replica settings should I use to facilitate our read-heavy behavior  
after launch. I understand the rough relationship between  
nodes/shards/replicas is that the more shards you have, the faster your  
expected index performance will be and a larger number of replicas you have  
distributed over your overall node count increases the possible query  
performance

Please forgive me if I've missed something obvious in the docs or the group  
list. I'm trying to plan and tune my setup without resorting to a lot of  
guess-and-test. I appreciate any pointers you can offer.

Thanks,

-jason

--

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [September 24, 2012, 6:09pm UTC](https://discuss.elastic.co/t/scaling-elasticsearch-for-40gb-of-data/9107/2 "2012-09-24T18:09:58Z")

</div>

Hi Jason,

Just a pair of initial suggestions:

You probably want to give ES less memory. Try half your RAM for a start  
and see if you still see OOMs.  
Use tools like SPM for Elasticsearch, or iostat, etc. to see if there is a  
lot of disk reading going on while querying or if the JMV GC is stealing  
the CPU.

## Otis

Search Analytics - [Cloud Monitoring Tools & Services | Sematext](http://sematext.com/search-analytics/index.html)  
Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)

On Saturday, September 22, 2012 12:55:48 PM UTC-4, Jason Yankus wrote:

> Greetings,
> 
> I am migrating a 40GB volume of mixed metadata and unstructured text into  
> Elasticsearch (~500,000 documents with metadata + extracted text) and need  
> some recommendations about setting up a cluster for high availability and  
> high performance queries.
> 
> I have tried a few basic changes to the default configuration but have not  
> had a lot of success avoiding:
> 
> 1.) OOM/Heap Space Errors during indexing (I can only index about 150k  
> docs during the ETL (Extract-Transform-Load) process before the system  
> bails out after 3-4 hours)  
> 2.) Poor full text query performance (2+ seconds on an index size of  
> roughly 30GB for 1 concurrent user)
> 
> We have two environments set up for test:
> 
> _Development_  
> 1x [8GB RAM 4x processor]
> 
> _Staging_  
> 2x [4GB RAM 2x processor]
> 
> My shard and replica settings are as follows:  
> index.number\_of\_shards: 6  
> index.number\_of\_replicas: 1
> 
> In each environment I have the ES\_HEAP\_SIZE value set to maximum physical  
> memory - 1 GB (so 7192 on 8GB, 3072 on 4GB).
> 
> So, my questions are:
> 
> For a volume of data as I described, what is the expectation about the  
> size of cluster (number of nodes, amount of ram) I will need to support  
> high performance queries ( \<500ms per query with an expected average volume  
> of 25 concurrent users)
> 
> Second, after our initial ETL, we will only experience incremental  
> indexing (as new docs come in or older docs are changed ~200/day). What  
> sort of shard/replica settings should I use to facilitate our read-heavy  
> behavior after launch. I understand the rough relationship between  
> nodes/shards/replicas is that the more shards you have, the faster your  
> expected index performance will be and a larger number of replicas you have  
> distributed over your overall node count increases the possible query  
> performance
> 
> Please forgive me if I've missed something obvious in the docs or the  
> group list. I'm trying to plan and tune my setup without resorting to a  
> lot of guess-and-test. I appreciate any pointers you can offer.
> 
> Thanks,
> 
> -jason

--

---

<div class="post-metadata">

**Author:** ![Jason\_Yankus](https://avatars.discourse-cdn.com/v4/letter/j/77aa72/32.png) [@Jason\_Yankus](https://discuss.elastic.co/u/Jason_Yankus)\
**Post date:** [September 25, 2012, 4:27pm UTC](https://discuss.elastic.co/t/scaling-elasticsearch-for-40gb-of-data/9107/3 "2012-09-25T16:27:15Z")

</div>

Thanks for the tips. I'll give them a try and let you know how it goes.

-jason

On Monday, September 24, 2012 2:09:59 PM UTC-4, Otis Gospodnetic wrote:

> Hi Jason,
> 
> Just a pair of initial suggestions:
> 
> You probably want to give ES less memory. Try half your RAM for a start  
> and see if you still see OOMs.  
> Use tools like SPM for Elasticsearch, or iostat, etc. to see if there is a  
> lot of disk reading going on while querying or if the JMV GC is stealing  
> the CPU.
> 
> ## Otis
> 
> Search Analytics - [Cloud Monitoring Tools & Services | Sematext](http://sematext.com/search-analytics/index.html)  
> Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> 
> On Saturday, September 22, 2012 12:55:48 PM UTC-4, Jason Yankus wrote:
> 
> > Greetings,
> > 
> > I am migrating a 40GB volume of mixed metadata and unstructured text into  
> > Elasticsearch (~500,000 documents with metadata + extracted text) and need  
> > some recommendations about setting up a cluster for high availability and  
> > high performance queries.
> > 
> > I have tried a few basic changes to the default configuration but have  
> > not had a lot of success avoiding:
> > 
> > 1.) OOM/Heap Space Errors during indexing (I can only index about 150k  
> > docs during the ETL (Extract-Transform-Load) process before the system  
> > bails out after 3-4 hours)  
> > 2.) Poor full text query performance (2+ seconds on an index size of  
> > roughly 30GB for 1 concurrent user)
> > 
> > We have two environments set up for test:
> > 
> > _Development_  
> > 1x [8GB RAM 4x processor]
> > 
> > _Staging_  
> > 2x [4GB RAM 2x processor]
> > 
> > My shard and replica settings are as follows:  
> > index.number\_of\_shards: 6  
> > index.number\_of\_replicas: 1
> > 
> > In each environment I have the ES\_HEAP\_SIZE value set to maximum physical  
> > memory - 1 GB (so 7192 on 8GB, 3072 on 4GB).
> > 
> > So, my questions are:
> > 
> > For a volume of data as I described, what is the expectation about the  
> > size of cluster (number of nodes, amount of ram) I will need to support  
> > high performance queries ( \<500ms per query with an expected average volume  
> > of 25 concurrent users)
> > 
> > Second, after our initial ETL, we will only experience incremental  
> > indexing (as new docs come in or older docs are changed ~200/day). What  
> > sort of shard/replica settings should I use to facilitate our read-heavy  
> > behavior after launch. I understand the rough relationship between  
> > nodes/shards/replicas is that the more shards you have, the faster your  
> > expected index performance will be and a larger number of replicas you have  
> > distributed over your overall node count increases the possible query  
> > performance
> > 
> > Please forgive me if I've missed something obvious in the docs or the  
> > group list. I'm trying to plan and tune my setup without resorting to a  
> > lot of guess-and-test. I appreciate any pointers you can offer.
> > 
> > Thanks,
> > 
> > -jason

--

---

<div class="post-metadata">

**Author:** ![Jason\_Yankus](https://avatars.discourse-cdn.com/v4/letter/j/77aa72/32.png) [@Jason\_Yankus](https://discuss.elastic.co/u/Jason_Yankus)\
**Post date:** [September 25, 2012, 4:28pm UTC](https://discuss.elastic.co/t/scaling-elasticsearch-for-40gb-of-data/9107/4 "2012-09-25T16:28:14Z")

</div>

Thanks for the tips. I'll give them a try and let you know how it goes.

-jason

--

---

<div class="post-metadata">

**Author:** ![anuj](https://avatars.discourse-cdn.com/v4/letter/a/dbc845/32.png) [@anuj](https://discuss.elastic.co/u/anuj)\
**Post date:** [September 26, 2012, 4:58am UTC](https://discuss.elastic.co/t/scaling-elasticsearch-for-40gb-of-data/9107/5 "2012-09-26T04:58:25Z")

</div>

Hi ,

I am creating index using ES with 8GB RAM ,1Node with 4 Shards. I am using bulk api to index my data.  
So i send 200 docs in one batch , my total batch is 2000.But i got performance issues ,I got java heap space exception when 20k docs get indexed.

Any one suggest me how i can solve this issue.

Thanks in advance

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:11am UTC](https://discuss.elastic.co/t/scaling-elasticsearch-for-40gb-of-data/9107/6 "2017-07-06T03:11:25Z")

</div>


