# Wouldn't this be cool?

**URL:** <https://discuss.elastic.co/t/wouldnt-this-be-cool/32307>\
**Category:** Elasticsearch\
**Created:** [October 15, 2015, 6:34pm UTC](https://discuss.elastic.co/t/wouldnt-this-be-cool/32307 "2015-10-15T18:34:00Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Chris\_Neal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chris_neal/32/3527_2.png) [@Chris\_Neal](https://discuss.elastic.co/u/Chris_Neal)\
**Post date:** [October 15, 2015, 6:34pm UTC](https://discuss.elastic.co/t/wouldnt-this-be-cool/32307/1 "2015-10-15T18:34:00Z")

</div>

Enticing subject I hope. 😄

Here's my thought. I've got servers with 256GB RAM. The optimal heap size for an ES JVM is 32GB (or just under that). How cool would it be if there was something like a "data-shared data node" (TM) that could run on the same server as my "regular" data node, but NOT need it's own copy of the shard data on disk? It would instead refer to the data already there from the "regular" data node.

This data-shared data node could use its entire heap for queries alone, and query off the data that is "owned" by the "regular" data node. I could run 1 regular and 2 shared JVMs per server, and dramatically increase my query potential without having to create new copies of the data!

That would make my day.

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [October 15, 2015, 6:40pm UTC](https://discuss.elastic.co/t/wouldnt-this-be-cool/32307/2 "2015-10-15T18:40:17Z")

</div>

> The optimal heap size for an ES JVM is 64GB (or just under that)

Almost—64 GB is a good _RAM size_ if you want to go with the recommendation of giving ~50% of the RAM to the JVM. The JVM heap should be kept below 30.5 GB to avoid uncompressed pointers.

---

<div class="post-metadata">

**Author:** ![Chris\_Neal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chris_neal/32/3527_2.png) [@Chris\_Neal](https://discuss.elastic.co/u/Chris_Neal)\
**Post date:** [October 15, 2015, 6:41pm UTC](https://discuss.elastic.co/t/wouldnt-this-be-cool/32307/3 "2015-10-15T18:41:46Z")

</div>

> [@magnusbaeck](#):
>
> Almost—64 GB is a good RAM size if you want to go with the recommendation of giving ~50% of the RAM to the JVM. The JVM heap should be kept below 30.5 GB to avoid uncompressed pointers.

Correct. I mistyped. Edited my post 🙂

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 15, 2015, 6:47pm UTC](https://discuss.elastic.co/t/wouldnt-this-be-cool/32307/4 "2015-10-15T18:47:45Z")

</div>

This sounds a lot like the [shadow replica](https://www.elastic.co/guide/en/elasticsearch/reference/current/indices-shadow-replicas.html) functionality, although that requires the use of a shared file system for the entire cluster, not just a single server.

---

<div class="post-metadata">

**Author:** ![Chris\_Neal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chris_neal/32/3527_2.png) [@Chris\_Neal](https://discuss.elastic.co/u/Chris_Neal)\
**Post date:** [October 15, 2015, 6:59pm UTC](https://discuss.elastic.co/t/wouldnt-this-be-cool/32307/5 "2015-10-15T18:59:08Z")

</div>

> [@Christian\_Dahlqvist](#):
>
> This sounds a lot like the shadow replica functionality, although that requires the use of a shared file system for the entire cluster, not just a single server.

Similar for sure. Seems like searching a shared cluster-wide file system would be a bit on the slow side. Multiple JVMs searching the same local disk data could be super fast.

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [October 15, 2015, 8:53pm UTC](https://discuss.elastic.co/t/wouldnt-this-be-cool/32307/6 "2015-10-15T20:53:53Z")

</div>

Considering the data stored on disk is not quite used as-is, the only real  
benefit you would see is lower disk space utilization, which is not the  
bottleneck.

Thanks to mmap/docvalues, more than just the JVM heap is used to store  
process data. Perhaps if you found a way to share doc value data, then you  
would have something.

Cheers,

Ivan

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:44pm UTC](https://discuss.elastic.co/t/wouldnt-this-be-cool/32307/7 "2017-07-05T23:44:31Z")

</div>


