# Memory usage

**URL:** <https://discuss.elastic.co/t/memory-usage/2854>\
**Category:** Elasticsearch\
**Created:** [March 18, 2010, 2:29pm UTC](https://discuss.elastic.co/t/memory-usage/2854 "2010-03-18T14:29:51Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [March 18, 2010, 2:29pm UTC](https://discuss.elastic.co/t/memory-usage/2854/1 "2010-03-18T14:29:51Z")

</div>

Hiya

I've just tried loading a million documents into ElasticSearch, running  
on a small dev server, and memory usage grew until eventually there was  
none left, and it refused to accept any more docs.

I switched to using the file system rather than memory, and everything  
worked nicely (except a bit slower obviously)

However, I have another 4 million docs to load which will take up a LOT  
of memory.

Does the sharding mean that: if the memory usage of the node in a  
cluster with a single node is 4GB, then the memory usage on each node in  
a cluster with 4 nodes will be 1GB (approx)?

thanks

## clint

Web Announcements Limited is a company registered in England and Wales,  
with company number 05608868, with registered address at 10 Arvon Road,  
London, N5 1PR.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [March 18, 2010, 3:07pm UTC](https://discuss.elastic.co/t/memory-usage/2854/2 "2010-03-18T15:07:06Z")

</div>

On Thu, Mar 18, 2010 at 4:29 PM, Clinton Gormley [clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)wrote:

> Hiya
> 
> I've just tried loading a million documents into Elasticsearch, running  
> on a small dev server, and memory usage grew until eventually there was  
> none left, and it refused to accept any more docs.
> 
> I switched to using the file system rather than memory, and everything  
> worked nicely (except a bit slower obviously)

Did you use the native memory one (i.e. in 0.5 and above, type: memory in  
the store with no other argument)? In theory, you are then bounded by the  
physical memory, and not by how much memory you allocate to the JVM. In this  
case, by the way, I suggest using large bufferSize.

> However, I have another 4 million docs to load which will take up a LOT  
> of memory.
> 
> Does the sharding mean that: if the memory usage of the node in a  
> cluster with a single node is 4GB, then the memory usage on each node in  
> a cluster with 4 nodes will be 1GB (approx)?

Yep. Thats the idea. Don't forget the replicas though. If you have 5 shards  
with 1 replica each, then 1 node will take 4G, two nodes will each take 4G  
(because of the replicas).

> thanks
> 
> ## clint
> 
> Web Announcements Limited is a company registered in England and Wales,  
> with company number 05608868, with registered address at 10 Arvon Road,  
> London, N5 1PR.

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [March 18, 2010, 3:31pm UTC](https://discuss.elastic.co/t/memory-usage/2854/3 "2010-03-18T15:31:19Z")

</div>

> Did you use the native memory one (i.e. in 0.5 and above, type: memory  
> in the store with no other argument)?

yes

> In theory, you are then bounded by the physical memory, and not by how  
> much memory you allocate to the JVM.

yes - it was a small machine - only 2GB of memory. But that said,  
700,000 objects was using 1.4GB. I currently need to index 5 million  
objects, which will be a lot of memory 🙂

> In this case, by the way, I suggest using large bufferSize.

You mean when using index.storage.type = 'memory' ? Why a large  
bufferSize? And how big is considered large?

> Yep. Thats the idea. Don't forget the replicas though. If you have 5  
> shards with 1 replica each, then 1 node will take 4G, two nodes will  
> each take 4G (because of the replicas).

OK to understand this:

I have (eg) 5GB of data when running with one node, which has 5 shards.

If I start 5 nodes, with 5 shards and 2 replicas, then I would have:

- 2 nodes using 5GB
- 3 nodes using 1GB

Is this correct?

Still trying to get my head around how this all works 🙂

(and why aren't you on #elasticsearch on freenode 😉

ta

clint

> 

--  
Web Announcements Limited is a company registered in England and Wales,  
with company number 05608868, with registered address at 10 Arvon Road,  
London, N5 1PR.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [March 18, 2010, 8:54pm UTC](https://discuss.elastic.co/t/memory-usage/2854/4 "2010-03-18T20:54:49Z")

</div>

Answers below. But before that, let me point you to elasticsearch multi  
index support. It basically means that you can have one index in memory,  
which is smaller, and another index that is stored on the file system. For  
example, if you can break them based on types, it might make sense. Remember  
that you can search on several indices with elasticsearch.

On Thu, Mar 18, 2010 at 5:31 PM, Clinton Gormley [clinton@iannounce.co.uk](mailto:clinton@iannounce.co.uk)wrote:

> > Did you use the native memory one (i.e. in 0.5 and above, type: memory  
> > in the store with no other argument)?
> 
> yes
> 
> > In theory, you are then bounded by the physical memory, and not by how  
> > much memory you allocate to the JVM.
> 
> yes - it was a small machine - only 2GB of memory. But that said,  
> 700,000 objects was using 1.4GB. I currently need to index 5 million  
> objects, which will be a lot of memory 🙂

So, it means that if you have 5 simple machines with 4gb you would have 20g  
of memory :).

> > In this case, by the way, I suggest using large bufferSize.
> 
> You mean when using index.storage.type = 'memory' ? Why a large  
> bufferSize? And how big is considered large?

I would say 100k - 200k bufferSize is a good value.

> > Yep. Thats the idea. Don't forget the replicas though. If you have 5  
> > shards with 1 replica each, then 1 node will take 4G, two nodes will  
> > each take 4G (because of the replicas).
> 
> OK to understand this:
> 
> I have (eg) 5GB of data when running with one node, which has 5 shards.
> 
> If I start 5 nodes, with 5 shards and 2 replicas, then I would have:
> 
> - 2 nodes using 5GB
> - 3 nodes using 1GB
> 
> Is this correct?

Let me simplify the math. If you have 5 shards, each with 2 replicas, you  
have, in total, 5 \* (2 + 1) instances of shards running. In this case 15  
instances of shards (shards and their replicas).

If you have 5g with one node, that has 5 shards, then lets assume we have 1g  
per shard. This means that for 15 instances of shards you would need 15g.

> Still trying to get my head around how this all works 🙂
> 
> (and why aren't you on #elasticsearch on freenode 😉

Firewall..., home now, on now...

> ta
> 
> clint
> 
> > 
> 
> --  
> Web Announcements Limited is a company registered in England and Wales,  
> with company number 05608868, with registered address at 10 Arvon Road,  
> London, N5 1PR.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:25am UTC](https://discuss.elastic.co/t/memory-usage/2854/5 "2017-07-06T04:25:14Z")

</div>


