# Why is index not written to hdfs?

**URL:** <https://discuss.elastic.co/t/why-is-index-not-written-to-hdfs/7694>\
**Category:** Elasticsearch\
**Created:** [May 14, 2012, 10:36pm UTC](https://discuss.elastic.co/t/why-is-index-not-written-to-hdfs/7694 "2012-05-14T22:36:16Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mohit\_Anchlia](https://avatars.discourse-cdn.com/v4/letter/m/848f3c/32.png) [@Mohit\_Anchlia](https://discuss.elastic.co/u/Mohit_Anchlia)\
**Post date:** [May 14, 2012, 10:36pm UTC](https://discuss.elastic.co/t/why-is-index-not-written-to-hdfs/7694/1 "2012-05-14T22:36:16Z")

</div>

I have hadoop plugin with hdfs gateway but what I am seeing is that indexes  
are still being written locally. Can you please help me understand why it's  
being written locally?

# ls -ltr data/elasticsearch/nodes/0/indices/twitter/0/index/

total 12

-rw-r--r-- 1 root root 0 May 14 15:33 write.lock

-rw-r--r-- 1 root root 20 May 14 15:33 segments.gen

-rw-r--r-- 1 root root 58 May 14 15:33 segments\_1

-rw-r--r-- 1 root root 8 May 14 15:33 \_checksums-1337034789114

# hadoop fs -ls /elasticsearch

#returns nothing

config:

gateway.type: hdfs

gateway.hdfs.uri: hdfs://db1:54310

gateway.hdfs.path: elasticsearch

---

<div class="post-metadata">

**Author:** ![Mohit\_Anchlia](https://avatars.discourse-cdn.com/v4/letter/m/848f3c/32.png) [@Mohit\_Anchlia](https://discuss.elastic.co/u/Mohit_Anchlia)\
**Post date:** [May 14, 2012, 11:20pm UTC](https://discuss.elastic.co/t/why-is-index-not-written-to-hdfs/7694/2 "2012-05-14T23:20:57Z")

</div>

I finally got this working. Does anyone know when elasticsearch writes data  
to hdfs? Is it almost real time.

If I kill the node can I lose some data?

On Mon, May 14, 2012 at 3:36 PM, Mohit Anchlia [mohitanchlia@gmail.com](mailto:mohitanchlia@gmail.com)wrote:

> I have hadoop plugin with hdfs gateway but what I am seeing is that  
> indexes are still being written locally. Can you please help me understand  
> why it's being written locally?
> 
> # ls -ltr data/elasticsearch/nodes/0/indices/twitter/0/index/
> 
> total 12
> 
> -rw-r--r-- 1 root root 0 May 14 15:33 write.lock
> 
> -rw-r--r-- 1 root root 20 May 14 15:33 segments.gen
> 
> -rw-r--r-- 1 root root 58 May 14 15:33 segments\_1
> 
> -rw-r--r-- 1 root root 8 May 14 15:33 \_checksums-1337034789114
> 
> # hadoop fs -ls /elasticsearch
> 
> #returns nothing
> 
> config:
> 
> gateway.type: hdfs
> 
> gateway.hdfs.uri: hdfs://db1:54310
> 
> gateway.hdfs.path: elasticsearch

---

<div class="post-metadata">

**Author:** ![Berkay\_Mollamustafao](https://avatars.discourse-cdn.com/v4/letter/b/22d042/32.png) [@Berkay\_Mollamustafao](https://discuss.elastic.co/u/Berkay_Mollamustafao)\
**Post date:** [May 14, 2012, 11:23pm UTC](https://discuss.elastic.co/t/why-is-index-not-written-to-hdfs/7694/3 "2012-05-14T23:23:49Z")

</div>

There are two ways to configure ES for persistence (for the data to survive  
full cluster restart)

1. Local gateway, where the data persists on the servers
2. Shared or central gateway (S3, Hadoop, or shared file system) where data  
is stored elsewhere.

In either case, data is still stored locally. With the shared gateway, data  
is restored from that data store when a node restarts. For more  
information, highly recommend reading the docs thoroughly.

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

Regards,  
Berkay Mollamustafaoglu  
mberkay on yahoo, google and skype

On Mon, May 14, 2012 at 6:36 PM, Mohit Anchlia [mohitanchlia@gmail.com](mailto:mohitanchlia@gmail.com)wrote:

> I have hadoop plugin with hdfs gateway but what I am seeing is that  
> indexes are still being written locally. Can you please help me understand  
> why it's being written locally?
> 
> # ls -ltr data/elasticsearch/nodes/0/indices/twitter/0/index/
> 
> total 12
> 
> -rw-r--r-- 1 root root 0 May 14 15:33 write.lock
> 
> -rw-r--r-- 1 root root 20 May 14 15:33 segments.gen
> 
> -rw-r--r-- 1 root root 58 May 14 15:33 segments\_1
> 
> -rw-r--r-- 1 root root 8 May 14 15:33 \_checksums-1337034789114
> 
> # hadoop fs -ls /elasticsearch
> 
> #returns nothing
> 
> config:
> 
> gateway.type: hdfs
> 
> gateway.hdfs.uri: hdfs://db1:54310
> 
> gateway.hdfs.path: elasticsearch

---

<div class="post-metadata">

**Author:** ![Mohit\_Anchlia](https://avatars.discourse-cdn.com/v4/letter/m/848f3c/32.png) [@Mohit\_Anchlia](https://discuss.elastic.co/u/Mohit_Anchlia)\
**Post date:** [May 14, 2012, 11:40pm UTC](https://discuss.elastic.co/t/why-is-index-not-written-to-hdfs/7694/4 "2012-05-14T23:40:49Z")

</div>

Yes I read that and also have done recovery testing too, which seems to  
recover everything. My question was when does elasticsearch writes/commits  
data to Hadoop? Is it synchronously or async? Should I expect to lose any  
data that might be in elasticsearch memory? Just trying to understand the  
basics.

On Mon, May 14, 2012 at 4:23 PM, Berkay Mollamustafaoglu  
[mberkay@gmail.com](mailto:mberkay@gmail.com)wrote:

> There are two ways to configure ES for persistence (for the data to  
> survive full cluster restart)
> 
> 1. Local gateway, where the data persists on the servers
> 2. Shared or central gateway (S3, Hadoop, or shared file system) where  
> data is stored elsewhere.
> 
> In either case, data is still stored locally. With the shared gateway,  
> data is restored from that data store when a node restarts. For more  
> information, highly recommend reading the docs thoroughly.  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/modules/gateway/)
> 
> Regards,  
> Berkay Mollamustafaoglu  
> mberkay on yahoo, google and skype
> 
> On Mon, May 14, 2012 at 6:36 PM, Mohit Anchlia [mohitanchlia@gmail.com](mailto:mohitanchlia@gmail.com)wrote:
> 
> > I have hadoop plugin with hdfs gateway but what I am seeing is that  
> > indexes are still being written locally. Can you please help me understand  
> > why it's being written locally?
> > 
> > # ls -ltr data/elasticsearch/nodes/0/indices/twitter/0/index/
> > 
> > total 12
> > 
> > -rw-r--r-- 1 root root 0 May 14 15:33 write.lock
> > 
> > -rw-r--r-- 1 root root 20 May 14 15:33 segments.gen
> > 
> > -rw-r--r-- 1 root root 58 May 14 15:33 segments\_1
> > 
> > -rw-r--r-- 1 root root 8 May 14 15:33 \_checksums-1337034789114
> > 
> > # hadoop fs -ls /elasticsearch
> > 
> > #returns nothing
> > 
> > config:
> > 
> > gateway.type: hdfs
> > 
> > gateway.hdfs.uri: hdfs://db1:54310
> > 
> > gateway.hdfs.path: elasticsearch

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [May 17, 2012, 8:21pm UTC](https://discuss.elastic.co/t/why-is-index-not-written-to-hdfs/7694/5 "2012-05-17T20:21:06Z")

</div>

By default, elasticsearch will snapshot the data to HDFS every 10 seconds,  
I answered in another thread you posted regarding using the local gateway.

On Tue, May 15, 2012 at 2:40 AM, Mohit Anchlia [mohitanchlia@gmail.com](mailto:mohitanchlia@gmail.com)wrote:

> Yes I read that and also have done recovery testing too, which seems to  
> recover everything. My question was when does elasticsearch writes/commits  
> data to Hadoop? Is it synchronously or async? Should I expect to lose any  
> data that might be in elasticsearch memory? Just trying to understand the  
> basics.
> 
> On Mon, May 14, 2012 at 4:23 PM, Berkay Mollamustafaoglu \<  
> [mberkay@gmail.com](mailto:mberkay@gmail.com)\> wrote:
> 
> > There are two ways to configure ES for persistence (for the data to  
> > survive full cluster restart)
> > 
> > 1. Local gateway, where the data persists on the servers
> > 2. Shared or central gateway (S3, Hadoop, or shared file system) where  
> > data is stored elsewhere.
> > 
> > In either case, data is still stored locally. With the shared gateway,  
> > data is restored from that data store when a node restarts. For more  
> > information, highly recommend reading the docs thoroughly.  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/modules/gateway/)
> > 
> > Regards,  
> > Berkay Mollamustafaoglu  
> > mberkay on yahoo, google and skype
> > 
> > On Mon, May 14, 2012 at 6:36 PM, Mohit Anchlia [mohitanchlia@gmail.com](mailto:mohitanchlia@gmail.com)wrote:
> > 
> > > I have hadoop plugin with hdfs gateway but what I am seeing is that  
> > > indexes are still being written locally. Can you please help me understand  
> > > why it's being written locally?
> > > 
> > > # ls -ltr data/elasticsearch/nodes/0/indices/twitter/0/index/
> > > 
> > > total 12
> > > 
> > > -rw-r--r-- 1 root root 0 May 14 15:33 write.lock
> > > 
> > > -rw-r--r-- 1 root root 20 May 14 15:33 segments.gen
> > > 
> > > -rw-r--r-- 1 root root 58 May 14 15:33 segments\_1
> > > 
> > > -rw-r--r-- 1 root root 8 May 14 15:33 \_checksums-1337034789114
> > > 
> > > # hadoop fs -ls /elasticsearch
> > > 
> > > #returns nothing
> > > 
> > > config:
> > > 
> > > gateway.type: hdfs
> > > 
> > > gateway.hdfs.uri: hdfs://db1:54310
> > > 
> > > gateway.hdfs.path: elasticsearch

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:28am UTC](https://discuss.elastic.co/t/why-is-index-not-written-to-hdfs/7694/6 "2017-07-06T03:28:21Z")

</div>


