# Shard memory allocation and replicas

**URL:** <https://discuss.elastic.co/t/shard-memory-allocation-and-replicas/12327>\
**Category:** Elasticsearch\
**Created:** [June 8, 2013, 1:04pm UTC](https://discuss.elastic.co/t/shard-memory-allocation-and-replicas/12327 "2013-06-08T13:04:48Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ales\_Kafka\_2](https://avatars.discourse-cdn.com/v4/letter/a/77aa72/32.png) [@Ales\_Kafka\_2](https://discuss.elastic.co/u/Ales_Kafka_2)\
**Post date:** [June 8, 2013, 1:04pm UTC](https://discuss.elastic.co/t/shard-memory-allocation-and-replicas/12327/1 "2013-06-08T13:04:48Z")

</div>

Hi,

I have designed following scenario for ElasticSearch, but I need to answer  
some questions regarding memory allocation and replicas.

Scenario

I will have ES cluster composed of N nodes (will start with just 1, but  
could be up to 10 in future). Each node will have available sufficient  
amount of memory (like 10GB-30GB). I have two following indices.

_Index no. 1:_ Just few GB of data (probably no more than 3GB in any time,  
no more than million documents), few new documents per minute (can be  
manually regulated, doesn't depend on users behaviour). Lots of reads on  
this index (probably 80% and more percent of the estimated read-load on ES  
cluster).

_Index no. 2:_ User depended data (routed by user\_id into single shard).  
Order of magnitude more data than Index no. 1. Much more writes (can't  
estimate as now, but probably hundreds per minute, thousands in worst case  
scenario), but much less reads. Moreover, only several active users will  
access their data in short period of time. Every read will be routed to  
single user\_id. I can have arbitrary number of shards and replication  
factor can be even 0, as I can reindex missing data any time before next  
demand.

Questions

1. Is it good idea to have index no. 1 in just one shard and replicated to  
every node in cluster (maybe not all, but definitely those, who will handle  
HTTP requests - done via node tag && include in index setting).

2. Sometimes, I will need to do bulk update on index no. 1 (add integer  
value into array of single column in thousands of documents). Is it  
possible with so many replicas? Or are there gotchas?

3. When I read data via user\_id (from index no. 2) is whole shard accessed  
or only the required part?

4. If a shard of index no. 2 exceeds available memory and read is executed,  
does it knock out index n. 1 from memory or does ES keep it there?

5. Should I set index no. 1 as memory type (after I have more nodes in  
cluster) to have it all the time in memory?

Thanks for every single answer.  
Ales

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:32am UTC](https://discuss.elastic.co/t/shard-memory-allocation-and-replicas/12327/2 "2017-07-06T02:32:12Z")

</div>


