# Newbie help

**URL:** <https://discuss.elastic.co/t/newbie-help/3341>\
**Category:** Elasticsearch\
**Created:** [September 16, 2010, 4:24pm UTC](https://discuss.elastic.co/t/newbie-help/3341 "2010-09-16T16:24:29Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![thiago](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thiago/32/32096_2.png) [@thiago](https://discuss.elastic.co/u/thiago)\
**Post date:** [September 16, 2010, 4:24pm UTC](https://discuss.elastic.co/t/newbie-help/3341/1 "2010-09-16T16:24:29Z")

</div>

Hi there,

```
   I'm quite new to ES and I'm currently doing some experiments to

```

evaluate it.

```
   The experiment consisted of the following:

            1. Only one ES node up using a shared fs gateway dir

```

(the only config I did).  
2. A process doing live twitter indexing (using it's  
filtered streaming api).  
3. Used Win7 and only one hard disk.

```
   After ~6h indexing, I did a hard reboot of my computer and

```

restarted ES.

```
   In the end, I noticed the following:

           1. The gateway data dir had ~7000 files totaling ~9GB.
           2. It took ~2h for the ES node to become available.
           3. There were ~45k tweets indexed (not all tweets were

```

indexed due to applied filters).

```
   So, with so few documents indexed, why the cluster recovery

```

took so long? What configuration affects this behavior? And finally,  
why there were so many files in gateway dir? Any way to compact them?  
(maybe this slowed down recovery).

Thanks in advance,  
Thiago Souza

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [September 16, 2010, 6:58pm UTC](https://discuss.elastic.co/t/newbie-help/3341/2 "2010-09-16T18:58:07Z")

</div>

Hi,

The number of files in the gateway is very strange, it should not be that  
high. Where do you store the gateway?

-shay.banon

On Thu, Sep 16, 2010 at 6:24 PM, thiago [tcostasouza@gmail.com](mailto:tcostasouza@gmail.com) wrote:

> Hi there,
> 
> ```
> I'm quite new to ES and I'm currently doing some experiments to
> 
> ```
> 
> evaluate it.
> 
> ```
> The experiment consisted of the following:
> 
> 1. Only one ES node up using a shared fs gateway dir
> 
> ```
> 
> (the only config I did).  
> 2. A process doing live twitter indexing (using it's  
> filtered streaming api).  
> 3. Used Win7 and only one hard disk.
> 
> ```
> After ~6h indexing, I did a hard reboot of my computer and
> 
> ```
> 
> restarted ES.
> 
> ```
> In the end, I noticed the following:
> 
> 1. The gateway data dir had ~7000 files totaling ~9GB.
> 2. It took ~2h for the ES node to become available.
> 3. There were ~45k tweets indexed (not all tweets were
> 
> ```
> 
> indexed due to applied filters).
> 
> ```
> So, with so few documents indexed, why the cluster recovery
> 
> ```
> 
> took so long? What configuration affects this behavior? And finally,  
> why there were so many files in gateway dir? Any way to compact them?  
> (maybe this slowed down recovery).
> 
> Thanks in advance,  
> Thiago Souza

---

<div class="post-metadata">

**Author:** ![thiago](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thiago/32/32096_2.png) [@thiago](https://discuss.elastic.co/u/thiago)\
**Post date:** [September 16, 2010, 8:03pm UTC](https://discuss.elastic.co/t/newbie-help/3341/3 "2010-09-16T20:03:03Z")

</div>

Hi Shay,

```
  Thanks for you reply!

  I used fs (in a dir). This number is the sum of all data dir

```

subdirectorie (there were 5 of them). Each one had ~1500 files.

Regards,  
Thiago Souza

On Thu, Sep 16, 2010 at 15:58, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:

> Hi,
> 
> The number of files in the gateway is very strange, it should not be  
> that high. Where do you store the gateway?
> 
> -shay.banon
> 
> On Thu, Sep 16, 2010 at 6:24 PM, thiago [tcostasouza@gmail.com](mailto:tcostasouza@gmail.com) wrote:
> 
> > Hi there,
> > 
> > ```
> > I'm quite new to ES and I'm currently doing some experiments to
> > 
> > ```
> > 
> > evaluate it.
> > 
> > ```
> > The experiment consisted of the following:
> > 
> > 1. Only one ES node up using a shared fs gateway dir
> > 
> > ```
> > 
> > (the only config I did).  
> > 2. A process doing live twitter indexing (using it's  
> > filtered streaming api).  
> > 3. Used Win7 and only one hard disk.
> > 
> > ```
> > After ~6h indexing, I did a hard reboot of my computer and
> > 
> > ```
> > 
> > restarted ES.
> > 
> > ```
> > In the end, I noticed the following:
> > 
> > 1. The gateway data dir had ~7000 files totaling ~9GB.
> > 2. It took ~2h for the ES node to become available.
> > 3. There were ~45k tweets indexed (not all tweets were
> > 
> > ```
> > 
> > indexed due to applied filters).
> > 
> > ```
> > So, with so few documents indexed, why the cluster recovery
> > 
> > ```
> > 
> > took so long? What configuration affects this behavior? And finally,  
> > why there were so many files in gateway dir? Any way to compact them?  
> > (maybe this slowed down recovery).
> > 
> > Thanks in advance,  
> > Thiago Souza

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [September 16, 2010, 11:32pm UTC](https://discuss.elastic.co/t/newbie-help/3341/4 "2010-09-16T23:32:28Z")

</div>

That fs location is also on the local drive? Have you deleted the work dir  
in elasticsearch before restarting (you shouldn't for fast recovery)? Ohh,  
and one more thing, when you say it took 2 Hr for the nodes to become  
available, what is available mean for you? The fact that you can search on  
them, or a GREEN status in the cluster health API?

One more thing, if you can recreate it, then you can use the indices status  
API and it will give you details on how long the recovery took, what data  
was reused from the local work dir. It would be great if you can gist it.

-shay.banon

On Thu, Sep 16, 2010 at 10:03 PM, Thiago Souza [tcostasouza@gmail.com](mailto:tcostasouza@gmail.com)wrote:

> Hi Shay,
> 
> ```
> Thanks for you reply!
> 
> I used fs (in a dir). This number is the sum of all data dir
> 
> ```
> 
> subdirectorie (there were 5 of them). Each one had ~1500 files.
> 
> Regards,  
> Thiago Souza
> 
> On Thu, Sep 16, 2010 at 15:58, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:
> 
> > Hi,
> > 
> > The number of files in the gateway is very strange, it should not be  
> > that high. Where do you store the gateway?
> > 
> > -shay.banon
> > 
> > On Thu, Sep 16, 2010 at 6:24 PM, thiago [tcostasouza@gmail.com](mailto:tcostasouza@gmail.com) wrote:
> > 
> > > Hi there,
> > > 
> > > ```
> > > I'm quite new to ES and I'm currently doing some experiments to
> > > 
> > > ```
> > > 
> > > evaluate it.
> > > 
> > > ```
> > > The experiment consisted of the following:
> > > 
> > > 1. Only one ES node up using a shared fs gateway dir
> > > 
> > > ```
> > > 
> > > (the only config I did).  
> > > 2. A process doing live twitter indexing (using it's  
> > > filtered streaming api).  
> > > 3. Used Win7 and only one hard disk.
> > > 
> > > ```
> > > After ~6h indexing, I did a hard reboot of my computer and
> > > 
> > > ```
> > > 
> > > restarted ES.
> > > 
> > > ```
> > > In the end, I noticed the following:
> > > 
> > > 1. The gateway data dir had ~7000 files totaling ~9GB.
> > > 2. It took ~2h for the ES node to become available.
> > > 3. There were ~45k tweets indexed (not all tweets were
> > > 
> > > ```
> > > 
> > > indexed due to applied filters).
> > > 
> > > ```
> > > So, with so few documents indexed, why the cluster recovery
> > > 
> > > ```
> > > 
> > > took so long? What configuration affects this behavior? And finally,  
> > > why there were so many files in gateway dir? Any way to compact them?  
> > > (maybe this slowed down recovery).
> > > 
> > > Thanks in advance,  
> > > Thiago Souza

---

<div class="post-metadata">

**Author:** ![thiago](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thiago/32/32096_2.png) [@thiago](https://discuss.elastic.co/u/thiago)\
**Post date:** [September 17, 2010, 12:04am UTC](https://discuss.elastic.co/t/newbie-help/3341/5 "2010-09-17T00:04:45Z")

</div>

Hi Shay,

```
That fs location is also on the local drive?
Yes

Have you deleted the work dir in elasticsearch before restarting?
No

What is available mean for you?
I mean that I could not access

```

[http://localhost:9200/\_search?q=\*:](http://localhost:9200/_search?q=*:)\*(browser didn't respond)

```
I didn't know that the status API would give me these details, I'll

```

check this.

Cheers

On Thu, Sep 16, 2010 at 20:32, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:

> That fs location is also on the local drive? Have you deleted the work dir  
> in elasticsearch before restarting (you shouldn't for fast recovery)? Ohh,  
> and one more thing, when you say it took 2 Hr for the nodes to become  
> available, what is available mean for you? The fact that you can search on  
> them, or a GREEN status in the cluster health API?
> 
> One more thing, if you can recreate it, then you can use the indices status  
> API and it will give you details on how long the recovery took, what data  
> was reused from the local work dir. It would be great if you can gist it.
> 
> -shay.banon
> 
> On Thu, Sep 16, 2010 at 10:03 PM, Thiago Souza [tcostasouza@gmail.com](mailto:tcostasouza@gmail.com)wrote:
> 
> > Hi Shay,
> > 
> > ```
> > Thanks for you reply!
> > 
> > I used fs (in a dir). This number is the sum of all data dir
> > 
> > ```
> > 
> > subdirectorie (there were 5 of them). Each one had ~1500 files.
> > 
> > Regards,  
> > Thiago Souza
> > 
> > On Thu, Sep 16, 2010 at 15:58, Shay Banon [shay.banon@elasticsearch.com](mailto:shay.banon@elasticsearch.com)wrote:
> > 
> > > Hi,
> > > 
> > > The number of files in the gateway is very strange, it should not be  
> > > that high. Where do you store the gateway?
> > > 
> > > -shay.banon
> > > 
> > > On Thu, Sep 16, 2010 at 6:24 PM, thiago [tcostasouza@gmail.com](mailto:tcostasouza@gmail.com) wrote:
> > > 
> > > > Hi there,
> > > > 
> > > > ```
> > > > I'm quite new to ES and I'm currently doing some experiments to
> > > > 
> > > > ```
> > > > 
> > > > evaluate it.
> > > > 
> > > > ```
> > > > The experiment consisted of the following:
> > > > 
> > > > 1. Only one ES node up using a shared fs gateway dir
> > > > 
> > > > ```
> > > > 
> > > > (the only config I did).  
> > > > 2. A process doing live twitter indexing (using it's  
> > > > filtered streaming api).  
> > > > 3. Used Win7 and only one hard disk.
> > > > 
> > > > ```
> > > > After ~6h indexing, I did a hard reboot of my computer and
> > > > 
> > > > ```
> > > > 
> > > > restarted ES.
> > > > 
> > > > ```
> > > > In the end, I noticed the following:
> > > > 
> > > > 1. The gateway data dir had ~7000 files totaling ~9GB.
> > > > 2. It took ~2h for the ES node to become available.
> > > > 3. There were ~45k tweets indexed (not all tweets were
> > > > 
> > > > ```
> > > > 
> > > > indexed due to applied filters).
> > > > 
> > > > ```
> > > > So, with so few documents indexed, why the cluster recovery
> > > > 
> > > > ```
> > > > 
> > > > took so long? What configuration affects this behavior? And finally,  
> > > > why there were so many files in gateway dir? Any way to compact them?  
> > > > (maybe this slowed down recovery).
> > > > 
> > > > Thanks in advance,  
> > > > Thiago Souza

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:19am UTC](https://discuss.elastic.co/t/newbie-help/3341/6 "2017-07-06T04:19:14Z")

</div>


