# Elasticsearch HOT/WARM infrastructure with S3 mounted as a filesystem using S3FS problem

**URL:** <https://discuss.elastic.co/t/elasticsearch-hot-warm-infrastructure-with-s3-mounted-as-a-filesystem-using-s3fs-problem/159660>\
**Category:** Elasticsearch\
**Created:** [December 6, 2018, 6:37am UTC](https://discuss.elastic.co/t/elasticsearch-hot-warm-infrastructure-with-s3-mounted-as-a-filesystem-using-s3fs-problem/159660 "2018-12-06T06:37:42Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Zerobot](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/zerobot/32/48977_2.png) [@Zerobot](https://discuss.elastic.co/u/Zerobot)\
**Post date:** [December 6, 2018, 6:37am UTC](https://discuss.elastic.co/t/elasticsearch-hot-warm-infrastructure-with-s3-mounted-as-a-filesystem-using-s3fs-problem/159660/1 "2018-12-06T06:37:42Z")

</div>

Hello !

So I've got a cluster of 3xMASTER | 4x HOT | 4x WARM nodes. The goal of my project is to store indices on HOT nodes for about 1-2 weeks and then allocate them (with the help of Curator) on WARM nodes.  
We are talking about huge amounts of data and my company came with a bright idea of using ECS with S3 API for it. So we are mounting (we have to) S3 as a filesystem on RedHat, let's say on /data/, using S3FS and then setting Elasticsearch's **_path.data: /data/node\_1/_** or **_/node\_2/_** and so on. It's super slow of course but it doesn't blow up so that's something positive.

Cerebro / Shard statistic are saying that 100% of shard was sent to WARM but S3FS just cached this data locally on /app/, it is right now sending it to S3 mounted on /data/ as fast as it can but is slow as hell of course.

Problem is: when I tried to allocate a 100 GB index (4 x 25 GB shards) on WARM nodes it seems that Elasticsearch is hitting those 100% and after a few minutes (S3FS is still sending from it's cache od /app/ to S3 on /data) it starts to reallocate those shards all over again but in different shard-node setup.

For example I'm starting the allocation process and shards are aligned like this:  
**0**  
**1**  
**2**  
**3**  
Allocation hits 100%, shard are still green on HOT and purple on WARM, S3FS still has like 15GB of each shard to send to /data/ but Elasticsearch start's to reallocate shard all over again and now it looks like this:  
**2**  
**3**  
**1**  
**0**  
Whole process begins anew, S3FS can't send it fast enough so after hitting 100% it rerolls all over again until eternity.

Is there a possible way to force elasticsearch into "waiting" until those shard are completely sent to S3 by S3FS ? Maybe some timeout set to like 60 minutes or something? I tried to find something in Docs but sadly nothing helped me.

Also I have those settings:  
**_cluster.routing.allocation.enable: "all"_** --\> We're not using replicas right now so may as well set it to "primaries".  
**_cluster.routing.rebalance.enable: "none"_**

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [December 6, 2018, 6:39am UTC](https://discuss.elastic.co/t/elasticsearch-hot-warm-infrastructure-with-s3-mounted-as-a-filesystem-using-s3fs-problem/159660/2 "2018-12-06T06:39:14Z")

</div>

> [@Zerobot](#):
>
> It's super slow of course but it doesn't blow up so that's something positive.

It's a good idea, it won't work though.

> [@Zerobot](#):
>
> Problem is: when I tried to allocate a 100 GB index (4 x 25 GB shards) on WARM nodes it seems that Elasticsearch is hitting those 100% and after a few minutes (S3FS is still sending from it's cache od /app/ to S3 on /data) it starts to reallocate those shards all over again but in different shard-node setup.

Yep.

> [@Zerobot](#):
>
> Is there a possible way to force elasticsearch into "waiting" until those shard are completely sent to S3 by S3FS ?

Nope.

The problem is Elasticsearch is too fast, and expects things to be fast with it.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 6, 2018, 8:14am UTC](https://discuss.elastic.co/t/elasticsearch-hot-warm-infrastructure-with-s3-mounted-as-a-filesystem-using-s3fs-problem/159660/3 "2018-12-06T08:14:27Z")

</div>

This kind of set-up is a bad idea. I recommend you have a look at this thread:

> [@Elasticsearch indices stored on S3 mounted with S3FS](https://discuss.elastic.co/t/elasticsearch-indices-stored-on-s3-mounted-with-s3fs/157400):
>
> Hi, So I've a really specific infrastructure where I need to store my "Older than 30 days" indices on COLD/WARM nodes. Those nodes have a S3 bucket (1 bucket for all 4 nodes) mounted as a filesystem on each node in /data/ folder. Of course, /data/ is set as path for those nodes to store indices etc. Setup is : 4 Hot, 4 Cold/Warm, 15GB RAM each (7GB Heap) What I'd like to ask is: When we are talking about 100GB of data daily (right now) and something like 500GB of data daily in the future - do…

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 3, 2019, 8:14am UTC](https://discuss.elastic.co/t/elasticsearch-hot-warm-infrastructure-with-s3-mounted-as-a-filesystem-using-s3fs-problem/159660/4 "2019-01-03T08:14:31Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
