# High volume Indexing of Documents

**URL:** <https://discuss.elastic.co/t/high-volume-indexing-of-documents/10113>\
**Category:** Elasticsearch\
**Created:** [December 19, 2012, 3:43am UTC](https://discuss.elastic.co/t/high-volume-indexing-of-documents/10113 "2012-12-19T03:43:12Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Meetu\_Maltiar](https://avatars.discourse-cdn.com/v4/letter/m/5f9b8f/32.png) [@Meetu\_Maltiar](https://discuss.elastic.co/u/Meetu_Maltiar)\
**Post date:** [December 19, 2012, 3:43am UTC](https://discuss.elastic.co/t/high-volume-indexing-of-documents/10113/1 "2012-12-19T03:43:12Z")

</div>

Hi,

We have an application that generates around 7000-10000 JSON messages  
per second. Each message size is around 2.6 KB. What are the best  
practices that needs to be followed at the java API level so that my  
application as well as Elastic-Search scales well.

Right now my application and ElasticSearch are residing on same box. I  
intend to use Java ElasticSearch client using a node of type client as  
suggested in documentation here [http://www.elasticsearch.org/guide/reference/java-api/client.html](http://www.elasticsearch.org/guide/reference/java-api/client.html).  
Since my application is multithreaded I will share client with them,  
is it ok?

For high data writes in ElasticSearch is using Bulk API better?

Please suggest any other best practices I can include in my  
implementation. I will like to scale to 13 nodes in a cluster soon.

Regards,  
Meetu Maltiar

--

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [December 19, 2012, 5:17am UTC](https://discuss.elastic.co/t/high-volume-indexing-of-documents/10113/2 "2012-12-19T05:17:01Z")

</div>

Hi,

Bulk is good indeed. -Xmx and JVM settings matter. If this is  
write-heavy, relatively speaking, any index merging params should be looked  
at. Refresh interval can/should be high unless you really need NRT.

May be best to wait until/if you hit issues and then you can provide  
concrete info about what you are doing and others can provide feedback.

## Otis

ELASTICSEARCH Performance Monitoring - [Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)

On Tuesday, December 18, 2012 10:43:12 PM UTC-5, Meetu Maltiar wrote:

> Hi,
> 
> We have an application that generates around 7000-10000 JSON messages  
> per second. Each message size is around 2.6 KB. What are the best  
> practices that needs to be followed at the java API level so that my  
> application as well as Elastic-Search scales well.
> 
> Right now my application and Elasticsearch are residing on same box. I  
> intend to use Java Elasticsearch client using a node of type client as  
> suggested in documentation here  
> [Elastic — The Search AI Company | Elastic](http://www.elasticsearch.org/guide/reference/java-api/client.html).  
> Since my application is multithreaded I will share client with them,  
> is it ok?
> 
> For high data writes in Elasticsearch is using Bulk API better?
> 
> Please suggest any other best practices I can include in my  
> implementation. I will like to scale to 13 nodes in a cluster soon.
> 
> Regards,  
> Meetu Maltiar

--

---

<div class="post-metadata">

**Author:** ![Meetu\_Maltiar](https://avatars.discourse-cdn.com/v4/letter/m/5f9b8f/32.png) [@Meetu\_Maltiar](https://discuss.elastic.co/u/Meetu_Maltiar)\
**Post date:** [December 19, 2012, 6:15am UTC](https://discuss.elastic.co/t/high-volume-indexing-of-documents/10113/3 "2012-12-19T06:15:41Z")

</div>

Thanks a lot Otis,

I am going with your suggestion of using bulk api. I will look at the  
a) JVM settings b) index merging patterns c) Refresh interval.

Right now I have a singleton node and have "node-client" that is  
shared by threads. Is this fine? I am trying to minimise client  
creation in each call otherwise I will have to create client for each  
document to be Indexed.

Regards,  
Meetu Maltiar

> **[Knoldus Blogs](https://blog.knoldus.com/)**
>
> Insights and Perspectives to keep you updated

On Dec 19, 10:17 am, Otis Gospodnetic [otis.gospodne...@gmail.com](mailto:otis.gospodne...@gmail.com)  
wrote:

> Hi,
> 
> Bulk is good indeed. -Xmx and JVM settings matter. If this is  
> write-heavy, relatively speaking, any index merging params should be looked  
> at. Refresh interval can/should be high unless you really need NRT.
> 
> May be best to wait until/if you hit issues and then you can provide  
> concrete info about what you are doing and others can provide feedback.
> 
> ## Otis
> 
> ELASTICSEARCH Performance Monitoring -[Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> 
> On Tuesday, December 18, 2012 10:43:12 PM UTC-5, Meetu Maltiar wrote:
> 
> > Hi,
> 
> > We have an application that generates around 7000-10000 JSON messages  
> > per second. Each message size is around 2.6 KB. What are the best  
> > practices that needs to be followed at the java API level so that my  
> > application as well as Elastic-Search scales well.
> 
> > Right now my application and Elasticsearch are residing on same box. I  
> > intend to use Java Elasticsearch client using a node of type client as  
> > suggested in documentation here  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/java-api/client.html).  
> > Since my application is multithreaded I will share client with them,  
> > is it ok?
> 
> > For high data writes in Elasticsearch is using Bulk API better?
> 
> > Please suggest any other best practices I can include in my  
> > implementation. I will like to scale to 13 nodes in a cluster soon.
> 
> > Regards,  
> > Meetu Maltiar

--

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 19, 2012, 6:32am UTC](https://discuss.elastic.co/t/high-volume-indexing-of-documents/10113/4 "2012-12-19T06:32:59Z")

</div>

Yes. Share the same client within all threads.  
BTW, If you are using Spring, you can have look at this: [GitHub - dadoonet/spring-elasticsearch: Spring factories for elasticsearch](https://github.com/dadoonet/spring-elasticsearch)

--  
David 😉  
Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

Le 19 déc. 2012 à 07:15, Meetu Maltiar [meetu@knoldus.com](mailto:meetu@knoldus.com) a écrit :

Thanks a lot Otis,

I am going with your suggestion of using bulk api. I will look at the  
a) JVM settings b) index merging patterns c) Refresh interval.

Right now I have a singleton node and have "node-client" that is  
shared by threads. Is this fine? I am trying to minimise client  
creation in each call otherwise I will have to create client for each  
document to be Indexed.

Regards,  
Meetu Maltiar

> **[Knoldus Blogs](https://blog.knoldus.com/)**
>
> Insights and Perspectives to keep you updated

On Dec 19, 10:17 am, Otis Gospodnetic [otis.gospodne...@gmail.com](mailto:otis.gospodne...@gmail.com)  
wrote:

> Hi,
> 
> Bulk is good indeed. -Xmx and JVM settings matter. If this is  
> write-heavy, relatively speaking, any index merging params should be looked  
> at. Refresh interval can/should be high unless you really need NRT.
> 
> May be best to wait until/if you hit issues and then you can provide  
> concrete info about what you are doing and others can provide feedback.
> 
> ## Otis
> 
> ELASTICSEARCH Performance Monitoring -[Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> 
> On Tuesday, December 18, 2012 10:43:12 PM UTC-5, Meetu Maltiar wrote:
> 
> > Hi,
> 
> > We have an application that generates around 7000-10000 JSON messages  
> > per second. Each message size is around 2.6 KB. What are the best  
> > practices that needs to be followed at the java API level so that my  
> > application as well as Elastic-Search scales well.
> 
> > Right now my application and Elasticsearch are residing on same box. I  
> > intend to use Java Elasticsearch client using a node of type client as  
> > suggested in documentation here  
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/java-api/client.html).  
> > Since my application is multithreaded I will share client with them,  
> > is it ok?
> 
> > For high data writes in Elasticsearch is using Bulk API better?
> 
> > Please suggest any other best practices I can include in my  
> > implementation. I will like to scale to 13 nodes in a cluster soon.
> 
> > Regards,  
> > Meetu Maltiar

--

--

---

<div class="post-metadata">

**Author:** ![Meetu\_Maltiar](https://avatars.discourse-cdn.com/v4/letter/m/5f9b8f/32.png) [@Meetu\_Maltiar](https://discuss.elastic.co/u/Meetu_Maltiar)\
**Post date:** [December 19, 2012, 7:37am UTC](https://discuss.elastic.co/t/high-volume-indexing-of-documents/10113/5 "2012-12-19T07:37:34Z")

</div>

Thanks David,

NIce will share client across. BTW, I am using Scala as a language,  
Elastic Search Java API and using Akka for parallelizing things.  
Though not using Spring at the moment, may do so after some time.  
Thanks for the github link it looks gr8 🙂 to use.

Meetu Maltiar  
Twitter: @meetumaltiar

> **[Knoldus Blogs](https://blog.knoldus.com/)**
>
> Insights and Perspectives to keep you updated

On Dec 19, 11:32 am, David Pilato [da...@pilato.fr](mailto:da...@pilato.fr) wrote:

> Yes. Share the same client within all threads.  
> BTW, If you are using Spring, you can have look at this:[GitHub - dadoonet/spring-elasticsearch: Spring factories for elasticsearch](https://github.com/dadoonet/spring-elasticsearch)
> 
> --  
> David 😉  
> Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs
> 
> Le 19 déc. 2012 à 07:15, Meetu Maltiar [me...@knoldus.com](mailto:me...@knoldus.com) a écrit :
> 
> Thanks a lot Otis,
> 
> I am going with your suggestion of using bulk api. I will look at the  
> a) JVM settings b) index merging patterns c) Refresh interval.
> 
> Right now I have a singleton node and have "node-client" that is  
> shared by threads. Is this fine? I am trying to minimise client  
> creation in each call otherwise I will have to create client for each  
> document to be Indexed.
> 
> Regards,  
> Meetu Maltiarhttp://blog.knoldus.com/
> 
> On Dec 19, 10:17 am, Otis Gospodnetic [otis.gospodne...@gmail.com](mailto:otis.gospodne...@gmail.com)  
> wrote:
> 
> > Hi,
> 
> > Bulk is good indeed. -Xmx and JVM settings matter. If this is  
> > write-heavy, relatively speaking, any index merging params should be looked  
> > at. Refresh interval can/should be high unless you really need NRT.
> 
> > May be best to wait until/if you hit issues and then you can provide  
> > concrete info about what you are doing and others can provide feedback.
> 
> > ## Otis
> > 
> > ELASTICSEARCH Performance Monitoring -[Sematext Monitoring | Infrastructure Monitoring Service](http://sematext.com/spm/index.html)
> 
> > On Tuesday, December 18, 2012 10:43:12 PM UTC-5, Meetu Maltiar wrote:
> 
> > > Hi,
> 
> > > We have an application that generates around 7000-10000 JSON messages  
> > > per second. Each message size is around 2.6 KB. What are the best  
> > > practices that needs to be followed at the java API level so that my  
> > > application as well as Elastic-Search scales well.
> 
> > > Right now my application and Elasticsearch are residing on same box. I  
> > > intend to use Java Elasticsearch client using a node of type client as  
> > > suggested in documentation here  
> > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/java-api/client.html).  
> > > Since my application is multithreaded I will share client with them,  
> > > is it ok?
> 
> > > For high data writes in Elasticsearch is using Bulk API better?
> 
> > > Please suggest any other best practices I can include in my  
> > > implementation. I will like to scale to 13 nodes in a cluster soon.
> 
> > > Regards,  
> > > Meetu Maltiar
> 
> --

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:59am UTC](https://discuss.elastic.co/t/high-volume-indexing-of-documents/10113/6 "2017-07-06T02:59:20Z")

</div>


