# Confused about tuning

**URL:** <https://discuss.elastic.co/t/confused-about-tuning/9279>\
**Category:** Elasticsearch\
**Created:** [October 8, 2012, 7:00pm UTC](https://discuss.elastic.co/t/confused-about-tuning/9279 "2012-10-08T19:00:00Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mohamed\_Lrhazi](https://avatars.discourse-cdn.com/v4/letter/m/74df32/32.png) [@Mohamed\_Lrhazi](https://discuss.elastic.co/u/Mohamed_Lrhazi)\
**Post date:** [October 8, 2012, 7:00pm UTC](https://discuss.elastic.co/t/confused-about-tuning/9279/1 "2012-10-08T19:00:00Z")

</div>

I indexed 20K documents using a 5 node ES setup, (RHEL 6.x)  
with everything in its default values. It took 15mins.

I then doubled the vCPUs on the VMs, from 4 to 8, and RAM from 4 to 8 GB.  
Rerun the indexing which took 16mins!

I then installed the service wrapper on all nodes, and added these lines at  
the top of the elasticsearch.conf:

set.default.ES\_HOME=/opt/elasticsearch-0.19.9  
set.default.ES\_HEAP\_SIZE=2048  
set.default.ES\_MIN\_MEM=4096  
set.default.ES\_MAX\_MEM=4096

Rerun my indexing and it took exactly 15mins again!!!

What am doing wrong? What is my bottleneck here?

Thanks a lot,  
Mohamed.

--

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [October 8, 2012, 7:23pm UTC](https://discuss.elastic.co/t/confused-about-tuning/9279/2 "2012-10-08T19:23:02Z")

</div>

Hey Mohammed,

Where are you loosing time? Is it when you get and build your docs or when you  
send it?  
How do you send it to ES? Are you using a bulk? Which size?

How does your documents look like?

It's best if you can provide more details about what you are doing. A curl  
recreation is perfect.

David.

Le 8 octobre 2012 à 21:00, Mohamed Lrhazi [ml623@georgetown.edu](mailto:ml623@georgetown.edu) a écrit :

> I indexed 20K documents using a 5 node ES setup, (RHEL 6.x) with everything in  
> its default values. It took 15mins.
> 
> I then doubled the vCPUs on the VMs, from 4 to 8, and RAM from 4 to 8 GB.  
> Rerun the indexing which took 16mins!
> 
> I then installed the service wrapper on all nodes, and added these lines at  
> the top of the elasticsearch.conf:
> 
> set.default.ES\_HOME=/opt/elasticsearch-0.19.9  
> set.default.ES\_HEAP\_SIZE=2048  
> set.default.ES\_MIN\_MEM=4096  
> set.default.ES\_MAX\_MEM=4096
> 
> Rerun my indexing and it took exactly 15mins again!!!
> 
> What am doing wrong? What is my bottleneck here?
> 
> Thanks a lot,  
> Mohamed.
> 
> --

--  
David Pilato  
[http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
[http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

--

---

<div class="post-metadata">

**Author:** ![Mohamed\_Lrhazi](https://avatars.discourse-cdn.com/v4/letter/m/74df32/32.png) [@Mohamed\_Lrhazi](https://discuss.elastic.co/u/Mohamed_Lrhazi)\
**Post date:** [October 8, 2012, 7:41pm UTC](https://discuss.elastic.co/t/confused-about-tuning/9279/3 "2012-10-08T19:41:33Z")

</div>

am using pyes... My script walks a dir tree looking for xml docs, for each  
file found:

- parses it using a python lib (lxml.objectify)
- index it a json dump fo the object.

I did run my script commenting out the indexing step... which means just  
walk the tree and parse the docs... it took 22 seconds!

I also notice, breaking my script after 1000 docs, that using 1, 2, 3, 4 or  
5 nodes, does not change the total time much!!

My documents have half a dozen attributes, one of which is a decent size  
HTML document.

I am using the default 5 shards and 1 replica.

am very very confused.

Thanks,  
Mohamed.

On Monday, October 8, 2012 3:23:15 PM UTC-4, David Pilato wrote:

> Hey Mohammed,
> 
> Where are you loosing time? Is it when you get and build your docs or  
> when you send it?  
> How do you send it to ES? Are you using a bulk? Which size?
> 
> How does your documents look like?
> 
> It's best if you can provide more details about what you are doing. A  
> curl recreation is perfect.
> 
> David.
> 
> Le 8 octobre 2012 à 21:00, Mohamed Lrhazi \<ml...@georgetown.edu\<javascript:\>\>  
> a écrit :
> 
> I indexed 20K documents using a 5 node ES setup, (RHEL 6.x)  
> with everything in its default values. It took 15mins.
> 
> I then doubled the vCPUs on the VMs, from 4 to 8, and RAM from 4 to 8 GB.  
> Rerun the indexing which took 16mins!
> 
> I then installed the service wrapper on all nodes, and added these lines  
> at the top of the elasticsearch.conf:
> 
> set.default.ES\_HOME=/opt/elasticsearch-0.19.9  
> set.default.ES\_HEAP\_SIZE=2048  
> set.default.ES\_MIN\_MEM=4096  
> set.default.ES\_MAX\_MEM=4096
> 
> Rerun my indexing and it took exactly 15mins again!!!
> 
> What am doing wrong? What is my bottleneck here?
> 
> Thanks a lot,  
> Mohamed.
> 
> --
> 
> --  
> David Pilato  
> [http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
> [http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
> Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

--

---

<div class="post-metadata">

**Author:** ![Mohamed\_Lrhazi](https://avatars.discourse-cdn.com/v4/letter/m/74df32/32.png) [@Mohamed\_Lrhazi](https://discuss.elastic.co/u/Mohamed_Lrhazi)\
**Post date:** [October 8, 2012, 7:42pm UTC](https://discuss.elastic.co/t/confused-about-tuning/9279/4 "2012-10-08T19:42:53Z")

</div>

On Monday, October 8, 2012 3:41:33 PM UTC-4, Mohamed Lrhazi wrote:

> - index it a json dump fo the object.

index a json.dump of the python dict object.

--

---

<div class="post-metadata">

**Author:** ![Mohamed\_Lrhazi](https://avatars.discourse-cdn.com/v4/letter/m/74df32/32.png) [@Mohamed\_Lrhazi](https://discuss.elastic.co/u/Mohamed_Lrhazi)\
**Post date:** [October 8, 2012, 7:50pm UTC](https://discuss.elastic.co/t/confused-about-tuning/9279/5 "2012-10-08T19:50:36Z")

</div>

OK. you mentioned "bulk", I was not using it... Using bulk I went from 15  
mins, to 35 seconds !!!

Thanks a lot,  
Mohamed.

On Monday, October 8, 2012 3:41:33 PM UTC-4, Mohamed Lrhazi wrote:

> am using pyes... My script walks a dir tree looking for xml docs, for each  
> file found:
> 
> - parses it using a python lib (lxml.objectify)
> - index it a json dump fo the object.
> 
> I did run my script commenting out the indexing step... which means just  
> walk the tree and parse the docs... it took 22 seconds!
> 
> I also notice, breaking my script after 1000 docs, that using 1, 2, 3, 4  
> or 5 nodes, does not change the total time much!!
> 
> My documents have half a dozen attributes, one of which is a decent size  
> HTML document.
> 
> I am using the default 5 shards and 1 replica.
> 
> am very very confused.
> 
> Thanks,  
> Mohamed.
> 
> On Monday, October 8, 2012 3:23:15 PM UTC-4, David Pilato wrote:
> 
> > Hey Mohammed,
> > 
> > Where are you loosing time? Is it when you get and build your docs or  
> > when you send it?  
> > How do you send it to ES? Are you using a bulk? Which size?
> > 
> > How does your documents look like?
> > 
> > It's best if you can provide more details about what you are doing. A  
> > curl recreation is perfect.
> > 
> > David.
> > 
> > Le 8 octobre 2012 à 21:00, Mohamed Lrhazi [ml...@georgetown.edu](mailto:ml...@georgetown.edu) a  
> > écrit :
> > 
> > I indexed 20K documents using a 5 node ES setup, (RHEL 6.x)  
> > with everything in its default values. It took 15mins.
> > 
> > I then doubled the vCPUs on the VMs, from 4 to 8, and RAM from 4 to 8  
> > GB. Rerun the indexing which took 16mins!
> > 
> > I then installed the service wrapper on all nodes, and added these lines  
> > at the top of the elasticsearch.conf:
> > 
> > set.default.ES\_HOME=/opt/elasticsearch-0.19.9  
> > set.default.ES\_HEAP\_SIZE=2048  
> > set.default.ES\_MIN\_MEM=4096  
> > set.default.ES\_MAX\_MEM=4096
> > 
> > Rerun my indexing and it took exactly 15mins again!!!
> > 
> > What am doing wrong? What is my bottleneck here?
> > 
> > Thanks a lot,  
> > Mohamed.
> > 
> > --
> > 
> > --  
> > David Pilato  
> > [http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
> > [http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
> > Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

--

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [October 8, 2012, 8:02pm UTC](https://discuss.elastic.co/t/confused-about-tuning/9279/6 "2012-10-08T20:02:40Z")

</div>

Check that everything is running quick from reading, parsing and generating  
JSon.  
Check your IO. If you are using VM, can it have a bad side effect on IO?

Let me say that I'm able to index on a "small" windows instance (1,5 Gb  
allocated to ES) to index about 300 docs per second.

Are you running 5 nodes on 5 boxes? If you are running 5 nodes on the same  
hardware, it's about the same than running 1 node.  
On one node, you have 5 lucene instances (5 shards) on one box. With 5 nodes on  
one nox, you have 1 shard per node, (1 Lucene instance per node), so you have 5  
Lucene intances on the same hardware!

That said, before growing the number of nodes, I think you should be able to get  
better results.  
You should be able to run it in less than 30 seconds (let's say 1 minute).

Le 8 octobre 2012 à 21:41, Mohamed Lrhazi [ml623@georgetown.edu](mailto:ml623@georgetown.edu) a écrit :

> am using pyes... My script walks a dir tree looking for xml docs, for each  
> file found:
> 
> - parses it using a python lib (lxml.objectify)
> - index it a json dump fo the object.
> 
> I did run my script commenting out the indexing step... which means just walk  
> the tree and parse the docs... it took 22 seconds!
> 
> I also notice, breaking my script after 1000 docs, that using 1, 2, 3, 4 or 5  
> nodes, does not change the total time much!!
> 
> My documents have half a dozen attributes, one of which is a decent size HTML  
> document.
> 
> I am using the default 5 shards and 1 replica.
> 
> am very very confused.
> 
> Thanks,  
> Mohamed.
> 
> On Monday, October 8, 2012 3:23:15 PM UTC-4, David Pilato wrote:
> 
> > > Hey Mohammed,
> > 
> > Where are you loosing time? Is it when you get and build your docs or  
> > when you send it?  
> > How do you send it to ES? Are you using a bulk? Which size?
> > 
> > How does your documents look like?
> > 
> > It's best if you can provide more details about what you are doing. A  
> > curl recreation is perfect.
> > 
> > David.
> > 
> > Le 8 octobre 2012 à 21:00, Mohamed Lrhazi \< ml...@georgetown.edu\> a écrit  
> > :
> > 
> > ```
> > > > > I indexed 20K documents using a 5 node ES setup, (RHEL 6.x) with
> > > > > everything in its default values. It took 15mins.
> > 
> > ```
> > 
> > > ```
> > > I then doubled the vCPUs on the VMs, from 4 to 8, and RAM from 4 to 8
> > > 
> > > ```
> > > 
> > > GB. Rerun the indexing which took 16mins!
> > > 
> > > ```
> > > I then installed the service wrapper on all nodes, and added these
> > > 
> > > ```
> > > 
> > > lines at the top of the elasticsearch.conf:
> > > 
> > > ```
> > > set.default.ES_HOME=/opt/elasticsearch-0.19.9
> > > set.default.ES_HEAP_SIZE=2048
> > > set.default.ES_MIN_MEM=4096
> > > set.default.ES_MAX_MEM=4096
> > > 
> > > Rerun my indexing and it took exactly 15mins again!!!
> > > 
> > > What am doing wrong? What is my bottleneck here?
> > > 
> > > Thanks a lot,
> > > Mohamed.
> > > 
> > > --
> > > 
> > > ```
> > > 
> > > > >
> > 
> > --  
> > David Pilato  
> > [http://www.scrutmydocs.org/](http://www.scrutmydocs.org/) [http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
> > [http://dev.david.pilato.fr/](http://dev.david.pilato.fr/) [http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
> > Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs
> > 
> > >
> 
> --

--  
David Pilato  
[http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
[http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

--

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [October 8, 2012, 8:03pm UTC](https://discuss.elastic.co/t/confused-about-tuning/9279/7 "2012-10-08T20:03:23Z")

</div>

Cool. That's a correct time now! 😉

Le 8 octobre 2012 à 21:50, Mohamed Lrhazi [ml623@georgetown.edu](mailto:ml623@georgetown.edu) a écrit :

> OK. you mentioned "bulk", I was not using it... Using bulk I went from 15  
> mins, to 35 seconds !!!
> 
> Thanks a lot,  
> Mohamed.
> 
> On Monday, October 8, 2012 3:41:33 PM UTC-4, Mohamed Lrhazi wrote:
> 
> > > am using pyes... My script walks a dir tree looking for xml docs, for  
> > > each file found:
> > 
> > - parses it using a python lib (lxml.objectify)
> > - index it a json dump fo the object.
> > 
> > I did run my script commenting out the indexing step... which means just  
> > walk the tree and parse the docs... it took 22 seconds!
> > 
> > I also notice, breaking my script after 1000 docs, that using 1, 2, 3, 4  
> > or 5 nodes, does not change the total time much!!
> > 
> > My documents have half a dozen attributes, one of which is a decent size  
> > HTML document.
> > 
> > I am using the default 5 shards and 1 replica.
> > 
> > am very very confused.
> > 
> > Thanks,  
> > Mohamed.
> > 
> > On Monday, October 8, 2012 3:23:15 PM UTC-4, David Pilato wrote:  
> > \> \> \> Hey Mohammed,
> > 
> > > ```
> > > Where are you loosing time? Is it when you get and build your docs or
> > > 
> > > ```
> > > 
> > > when you send it?  
> > > How do you send it to ES? Are you using a bulk? Which size?
> > > 
> > > ```
> > > How does your documents look like?
> > > 
> > > It's best if you can provide more details about what you are doing. A
> > > 
> > > ```
> > > 
> > > curl recreation is perfect.
> > > 
> > > ```
> > > David.
> > > 
> > > Le 8 octobre 2012 à 21:00, Mohamed Lrhazi < ml...@georgetown.edu> a
> > > 
> > > ```
> > > 
> > > écrit :
> > > 
> > > ```
> > > > > > > I indexed 20K documents using a 5 node ES setup, (RHEL 6.x)
> > > > > > > with everything in its default values. It took 15mins.
> > > 
> > > ```
> > > 
> > > > ```
> > > > I then doubled the vCPUs on the VMs, from 4 to 8, and RAM from 4
> > > > 
> > > > ```
> > > > 
> > > > to 8 GB. Rerun the indexing which took 16mins!
> > > > 
> > > > ```
> > > > I then installed the service wrapper on all nodes, and added these
> > > > 
> > > > ```
> > > > 
> > > > lines at the top of the elasticsearch.conf:
> > > > 
> > > > ```
> > > > set.default.ES_HOME=/opt/elasticsearch-0.19.9
> > > > set.default.ES_HEAP_SIZE=2048
> > > > set.default.ES_MIN_MEM=4096
> > > > set.default.ES_MAX_MEM=4096
> > > > 
> > > > Rerun my indexing and it took exactly 15mins again!!!
> > > > 
> > > > What am doing wrong? What is my bottleneck here?
> > > > 
> > > > Thanks a lot,
> > > > Mohamed.
> > > > 
> > > > --
> > > > 
> > > > > > > 
> > > > 
> > > > ```
> > > 
> > > ```
> > > --
> > > David Pilato
> > > http://www.scrutmydocs.org/ <http://www.scrutmydocs.org/>
> > > http://dev.david.pilato.fr/ <http://dev.david.pilato.fr/>
> > > Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs
> > > 
> > > ```
> > > 
> > > > > >
> 
> --

--  
David Pilato  
[http://www.scrutmydocs.org/](http://www.scrutmydocs.org/)  
[http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:09am UTC](https://discuss.elastic.co/t/confused-about-tuning/9279/8 "2017-07-06T03:09:41Z")

</div>


