# Index data in Mapreduce

**URL:** <https://discuss.elastic.co/t/index-data-in-mapreduce/7138>\
**Category:** Elasticsearch\
**Created:** [March 27, 2012, 2:50am UTC](https://discuss.elastic.co/t/index-data-in-mapreduce/7138 "2012-03-27T02:50:50Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![vreal](https://avatars.discourse-cdn.com/v4/letter/v/5f9b8f/32.png) [@vreal](https://discuss.elastic.co/u/vreal)\
**Post date:** [March 27, 2012, 2:50am UTC](https://discuss.elastic.co/t/index-data-in-mapreduce/7138/1 "2012-03-27T02:50:50Z")

</div>

Hi,

I'm currently writing a batch process which using amazon mapreduce and the results will be stored in database and indexed into elasticsearch cluster. The total amound of data will be less than 20k records. I've started 2 servers in the cluster with 2 cores and 3.75G memory (Amazon m1.medium instance) and I've also increate the max open file to 32000.

But during the process I've always get the "too many open files" error and the elastic search will stop serving for quite a while until I login to the server and restart the elasticsearch daemon.

Does elasticsearch support massive indexing operation? And if not what can I do to deal with the requirement?

Can anyone help me on this issue? Thanks.

Regards,  
Ye Zhou

---

<div class="post-metadata">

**Author:** ![vineeth\_mohan](https://avatars.discourse-cdn.com/v4/letter/v/bc79bd/32.png) [@vineeth\_mohan](https://discuss.elastic.co/u/vineeth_mohan)\
**Post date:** [March 27, 2012, 3:09am UTC](https://discuss.elastic.co/t/index-data-in-mapreduce/7138/2 "2012-03-27T03:09:21Z")

</div>

Increasing the number of open files at OS level is highly recommended for  
elasticSearch.  
So please go ahead and do that and then start elasticSearch.

Thanks  
Vineeth

On Tue, Mar 27, 2012 at 8:20 AM, Ye Zhou [zhouy.vreal@gmail.com](mailto:zhouy.vreal@gmail.com) wrote:

> Hi,
> 
> I'm currently writing a batch process which using amazon mapreduce and the  
> results will be stored in database and indexed into elasticsearch cluster.  
> The total amound of data will be less than 20k records. I've started 2  
> servers in the cluster with 2 cores and 3.75G memory (Amazon m1.medium  
> instance) and I've also increate the max open file to 32000.
> 
> But during the process I've always get the "too many open files" error and  
> the Elasticsearch will stop serving for quite a while until I login to the  
> server and restart the elasticsearch daemon.
> 
> Does elasticsearch support massive indexing operation? And if not what can  
> I do to deal with the requirement?
> 
> Can anyone help me on this issue? Thanks.
> 
> Regards,  
> Ye Zhou

---

<div class="post-metadata">

**Author:** ![vreal](https://avatars.discourse-cdn.com/v4/letter/v/5f9b8f/32.png) [@vreal](https://discuss.elastic.co/u/vreal)\
**Post date:** [March 27, 2012, 3:31am UTC](https://discuss.elastic.co/t/index-data-in-mapreduce/7138/3 "2012-03-27T03:31:02Z")

</div>

Hi,

The elasticsearch is running as root user.  
And my setting in /etc/security/limits.conf is:  
root soft nofile 102400  
root hard nofile 102400

Then I just reboot the instance.

Is that setting all right?

Regards,  
Ye Zhou

On 2012/03/27, at 11:9 , Vineeth Mohan wrote:

> Increasing the number of open files at OS level is highly recommended for elasticSearch.  
> So please go ahead and do that and then start elasticSearch.
> 
> Thanks  
> Vineeth
> 
> On Tue, Mar 27, 2012 at 8:20 AM, Ye Zhou [zhouy.vreal@gmail.com](mailto:zhouy.vreal@gmail.com) wrote:  
> Hi,
> 
> I'm currently writing a batch process which using amazon mapreduce and the results will be stored in database and indexed into elasticsearch cluster. The total amound of data will be less than 20k records. I've started 2 servers in the cluster with 2 cores and 3.75G memory (Amazon m1.medium instance) and I've also increate the max open file to 32000.
> 
> But during the process I've always get the "too many open files" error and the Elasticsearch will stop serving for quite a while until I login to the server and restart the elasticsearch daemon.
> 
> Does elasticsearch support massive indexing operation? And if not what can I do to deal with the requirement?
> 
> Can anyone help me on this issue? Thanks.
> 
> Regards,  
> Ye Zhou

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [March 27, 2012, 6:47pm UTC](https://discuss.elastic.co/t/index-data-in-mapreduce/7138/4 "2012-03-27T18:47:18Z")

</div>

Hi,

It's hard to tell if those particular numbers are good for you or not. Use  
`lsof' command as root (or with sudo) to check how many files ES is using.  
You may also want to run ES as a non-root user.  
ulimit -a will show you if your limits.conf numbers are really in effect.

## Otis

Hiring Elasticsearch Consultants --

> **[Jobs](https://sematext.com/jobs/)**
>
> We’re Hiring We are always looking for smart, passionate, motivated, and independent people regardless of where on the planet they may be. Learn more about the company Agent & Backend Engineer Full Stack Developer Backend Engineer Frontend...

On Tuesday, March 27, 2012 11:31:02 AM UTC+8, Ye Zhou wrote:

> Hi,
> 
> The elasticsearch is running as root user.  
> And my setting in /etc/security/limits.conf is:  
> root soft nofile 102400  
> root hard nofile 102400
> 
> Then I just reboot the instance.
> 
> Is that setting all right?
> 
> Regards,  
> Ye Zhou
> 
> On 2012/03/27, at 11:9 , Vineeth Mohan wrote:
> 
> Increasing the number of open files at OS level is highly recommended for  
> elasticSearch.  
> So please go ahead and do that and then start elasticSearch.
> 
> Thanks  
> Vineeth
> 
> On Tue, Mar 27, 2012 at 8:20 AM, Ye Zhou [zhouy.vreal@gmail.com](mailto:zhouy.vreal@gmail.com) wrote:
> 
> > Hi,
> > 
> > I'm currently writing a batch process which using amazon mapreduce and  
> > the results will be stored in database and indexed into elasticsearch  
> > cluster. The total amound of data will be less than 20k records. I've  
> > started 2 servers in the cluster with 2 cores and 3.75G memory (Amazon  
> > m1.medium instance) and I've also increate the max open file to 32000.
> > 
> > But during the process I've always get the "too many open files" error  
> > and the Elasticsearch will stop serving for quite a while until I login to  
> > the server and restart the elasticsearch daemon.
> > 
> > Does elasticsearch support massive indexing operation? And if not what  
> > can I do to deal with the requirement?
> > 
> > Can anyone help me on this issue? Thanks.
> > 
> > Regards,  
> > Ye Zhou

On Tuesday, March 27, 2012 11:31:02 AM UTC+8, Ye Zhou wrote:

> Hi,
> 
> The elasticsearch is running as root user.  
> And my setting in /etc/security/limits.conf is:  
> root soft nofile 102400  
> root hard nofile 102400
> 
> Then I just reboot the instance.
> 
> Is that setting all right?
> 
> Regards,  
> Ye Zhou
> 
> On 2012/03/27, at 11:9 , Vineeth Mohan wrote:
> 
> Increasing the number of open files at OS level is highly recommended for  
> elasticSearch.  
> So please go ahead and do that and then start elasticSearch.
> 
> Thanks  
> Vineeth
> 
> On Tue, Mar 27, 2012 at 8:20 AM, Ye Zhou [zhouy.vreal@gmail.com](mailto:zhouy.vreal@gmail.com) wrote:
> 
> > Hi,
> > 
> > I'm currently writing a batch process which using amazon mapreduce and  
> > the results will be stored in database and indexed into elasticsearch  
> > cluster. The total amound of data will be less than 20k records. I've  
> > started 2 servers in the cluster with 2 cores and 3.75G memory (Amazon  
> > m1.medium instance) and I've also increate the max open file to 32000.
> > 
> > But during the process I've always get the "too many open files" error  
> > and the Elasticsearch will stop serving for quite a while until I login to  
> > the server and restart the elasticsearch daemon.
> > 
> > Does elasticsearch support massive indexing operation? And if not what  
> > can I do to deal with the requirement?
> > 
> > Can anyone help me on this issue? Thanks.
> > 
> > Regards,  
> > Ye Zhou

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [March 28, 2012, 10:30am UTC](https://discuss.elastic.co/t/index-data-in-mapreduce/7138/5 "2012-03-28T10:30:02Z")

</div>

Can you make sure that the file limit configuration is actually applied?  
You can use teh nodes info API to see the max open file limit that the  
process actually runs with.

On Tue, Mar 27, 2012 at 5:31 AM, Ye Zhou [zhouy.vreal@gmail.com](mailto:zhouy.vreal@gmail.com) wrote:

> Hi,
> 
> The elasticsearch is running as root user.  
> And my setting in /etc/security/limits.conf is:  
> root soft nofile 102400  
> root hard nofile 102400
> 
> Then I just reboot the instance.
> 
> Is that setting all right?
> 
> Regards,  
> Ye Zhou
> 
> On 2012/03/27, at 11:9 , Vineeth Mohan wrote:
> 
> Increasing the number of open files at OS level is highly recommended for  
> elasticSearch.  
> So please go ahead and do that and then start elasticSearch.
> 
> Thanks  
> Vineeth
> 
> On Tue, Mar 27, 2012 at 8:20 AM, Ye Zhou [zhouy.vreal@gmail.com](mailto:zhouy.vreal@gmail.com) wrote:
> 
> > Hi,
> > 
> > I'm currently writing a batch process which using amazon mapreduce and  
> > the results will be stored in database and indexed into elasticsearch  
> > cluster. The total amound of data will be less than 20k records. I've  
> > started 2 servers in the cluster with 2 cores and 3.75G memory (Amazon  
> > m1.medium instance) and I've also increate the max open file to 32000.
> > 
> > But during the process I've always get the "too many open files" error  
> > and the Elasticsearch will stop serving for quite a while until I login to  
> > the server and restart the elasticsearch daemon.
> > 
> > Does elasticsearch support massive indexing operation? And if not what  
> > can I do to deal with the requirement?
> > 
> > Can anyone help me on this issue? Thanks.
> > 
> > Regards,  
> > Ye Zhou

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:34am UTC](https://discuss.elastic.co/t/index-data-in-mapreduce/7138/6 "2017-07-06T03:34:22Z")

</div>


