# Indexing large number of documents

**URL:** <https://discuss.elastic.co/t/indexing-large-number-of-documents/15838>\
**Category:** Elasticsearch\
**Created:** [February 17, 2014, 9:04am UTC](https://discuss.elastic.co/t/indexing-large-number-of-documents/15838 "2014-02-17T09:04:19Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Petr\_Jansky](https://avatars.discourse-cdn.com/v4/letter/p/ec9cab/32.png) [@Petr\_Jansky](https://discuss.elastic.co/u/Petr_Jansky)\
**Post date:** [February 17, 2014, 9:04am UTC](https://discuss.elastic.co/t/indexing-large-number-of-documents/15838/1 "2014-02-17T09:04:19Z")

</div>

Hello,

I'm trying to index \>300k docs using Java API.

_public class Fetcher {_

- public static String server = "localhost"; \*
- public static Integer port = 9300;\*
- public static String index = "default";\*
- public static String type = "default";\*
- public static String typeAttributename = null;\*
- static Client client = null;\*
- private static Fetcher inst;\*
- Settings settings = ImmutableSettings.settingsBuilder()\*
- .put("cluster.name", "elasticsearch")\*
- .put("node.name", "Killer")\*
- .build();\*
- public synchronized static Fetcher getInstace(){\*
- if(inst == null){\*
- inst = new Fetcher();\*
- }\*
- return inst;\*
- }\*
- public Fetcher() {\*
- client = new TransportClient(settings).addTransportAddress(new  
InetSocketTransportAddress(server, port));\*
- }\*
- public void index(DocumentVo document) {\*
- try {\*
- String type = Fetcher.type;\*
- if(typeAttributename != null && document.getData().get(typeAttributename)  
!= null){\*
- type = document.getData().get(typeAttributename).toString();\*
- type = type.toLowerCase();\*
- }\*
- IndexRequestBuilder rs =  
client.prepareIndex().setIndex(index).setType(type);\*
- rs.setTimeout(new TimeValue(10000));\*
- rs.setSource(document.getData());\*
- rs.execute().actionGet();\*
- } catch (Exception e) {\*
- e.printStackTrace();\*
- client.close();\*
- client = new TransportClient(settings).addTransportAddress(new  
InetSocketTransportAddress(server, port));\*
- index(document);\*
- } \*
- }\*
- public void close(){\*
- client.close();\*
- }\*  
_}_

in ~20 threads I run

_Fetcher.getInstace().index(document);_

I've created my own tokenizer filter that is quite slow so I'm getting

Feb 17, 2014 9:53:51 AM org.elasticsearch.client.transport  
INFO: [Killer] failed to get node info for  
[#transport#-1][inet[localhost/127.0.0.1:9300]], disconnecting...  
org.elasticsearch.transport.ReceiveTimeoutTransportException:  
[][inet[localhost/127.0.0.1:9300]][cluster/nodes/info] request\_id [2899]  
timed out after [5001ms]  
at  
org.elasticsearch.transport.TransportService$TimeoutHandler.run(TransportService.java:351)  
at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source)  
at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)  
at java.lang.Thread.run(Unknown Source)

org.elasticsearch.client.transport.NoNodeAvailableException: No node  
available  
at  
org.elasticsearch.client.transport.TransportClientNodesService$RetryListener.onFailure(TransportClientNodesService.java:249)  
at  
org.elasticsearch.action.TransportActionNodeProxy$1.handleException(TransportActionNodeProxy.java:84)  
at  
org.elasticsearch.transport.TransportService$Adapter$2$1.run(TransportService.java:311)  
at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source)  
at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)  
at java.lang.Thread.run(Unknown Source)

It seems that  
_rs.setTimeout(new TimeValue(10000));_  
in my index method doesn't work.

How can I setup timeout for indexing using API?

Is it correct to use one TransportCilent for multiple(10-60) threads?

Thanks  
Petr

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/f5f15b57-955c-4fcf-b225-3974e37e447b%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/f5f15b57-955c-4fcf-b225-3974e37e447b%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Jilles\_van\_Gurp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jilles_van_gurp/32/879_2.png) [@Jilles\_van\_Gurp](https://discuss.elastic.co/u/Jilles_van_Gurp)\
**Post date:** [February 17, 2014, 12:20pm UTC](https://discuss.elastic.co/t/indexing-large-number-of-documents/15838/2 "2014-02-17T12:20:21Z")

</div>

You'll want to use the batch API instead of indexing one document at the  
time. That scales a lot better. I've done tens of millions of documents  
like that in minutes. Basically, you can use mutlithreading with batch as  
well but you may want to not outnumber the number of cpus you can dedicate  
to indexing. Keep the batch sizes limited to a few hundred to a few  
thousand at most. Basically go for a size that es still handles in around a  
second or so.

If you really need to index one document at the time, you'll want to  
probably reduce the ambition level a bit with the number of threads. The  
error you are getting means that all nodes are busy with your previous  
requests. Increasing the timeout won't fix your problem; these requests  
normally should be in the range of a few ms and the fact that they are not,  
means you are hitting a bottleneck somewhere.

Jilles

On Monday, February 17, 2014 10:04:19 AM UTC+1, Petr Janský wrote:

> Hello,
> 
> I'm trying to index \>300k docs using Java API.
> 
> _public class Fetcher {_
> 
> - public static String server = "localhost"; \*
> - public static Integer port = 9300;\*
> - public static String index = "default";\*
> - public static String type = "default";\*
> - public static String typeAttributename = null;\*
> - static Client client = null;\*
> - private static Fetcher inst;\*
> - Settings settings = ImmutableSettings.settingsBuilder()\*
> - .put("cluster.name [http://cluster.name](http://cluster.name)", "elasticsearch")\*
> - .put("node.name [http://node.name](http://node.name)", "Killer")\*
> - .build();\*
> - public synchronized static Fetcher getInstace(){\*
> - if(inst == null){\*
> - inst = new Fetcher();\*
> - }\*
> - return inst;\*
> - }\*
> - public Fetcher() {\*
> - client = new TransportClient(settings).addTransportAddress(new  
> InetSocketTransportAddress(server, port));\*
> - }\*
> - public void index(DocumentVo document) {\*
> - try {\*
> - String type = Fetcher.type;\*
> - if(typeAttributename != null &&  
> document.getData().get(typeAttributename) != null){\*
> - type = document.getData().get(typeAttributename).toString();\*
> - type = type.toLowerCase();\*
> - }\*
> - IndexRequestBuilder rs =  
> client.prepareIndex().setIndex(index).setType(type);\*
> - rs.setTimeout(new TimeValue(10000));\*
> - rs.setSource(document.getData());\*
> - rs.execute().actionGet();\*
> - } catch (Exception e) {\*
> - e.printStackTrace();\*
> - client.close();\*
> - client = new TransportClient(settings).addTransportAddress(new  
> InetSocketTransportAddress(server, port));\*
> - index(document);\*
> - } \*
> - }\*
> - public void close(){\*
> - client.close();\*
> - }\*  
> _}_
> 
> in ~20 threads I run
> 
> _Fetcher.getInstace().index(document);_
> 
> I've created my own tokenizer filter that is quite slow so I'm getting
> 
> Feb 17, 2014 9:53:51 AM org.elasticsearch.client.transport  
> INFO: [Killer] failed to get node info for  
> [#transport#-1][inet[localhost/127.0.0.1:9300]], disconnecting...  
> org.elasticsearch.transport.ReceiveTimeoutTransportException:  
> [inet[localhost/127.0.0.1:9300]][cluster/nodes/info] request\_id [2899]  
> timed out after [5001ms]  
> at  
> org.elasticsearch.transport.TransportService$TimeoutHandler.run(TransportService.java:351)  
> at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source)  
> at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)  
> at java.lang.Thread.run(Unknown Source)
> 
> org.elasticsearch.client.transport.NoNodeAvailableException: No node  
> available  
> at  
> org.elasticsearch.client.transport.TransportClientNodesService$RetryListener.onFailure(TransportClientNodesService.java:249)  
> at  
> org.elasticsearch.action.TransportActionNodeProxy$1.handleException(TransportActionNodeProxy.java:84)  
> at  
> org.elasticsearch.transport.TransportService$Adapter$2$1.run(TransportService.java:311)  
> at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source)  
> at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)  
> at java.lang.Thread.run(Unknown Source)
> 
> It seems that  
> _rs.setTimeout(new TimeValue(10000));_  
> in my index method doesn't work.
> 
> How can I setup timeout for indexing using API?
> 
> Is it correct to use one TransportCilent for multiple(10-60) threads?
> 
> Thanks  
> Petr

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/88647d41-58fd-4e27-9e6a-80ee312fb439%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/88647d41-58fd-4e27-9e6a-80ee312fb439%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [February 17, 2014, 4:58pm UTC](https://discuss.elastic.co/t/indexing-large-number-of-documents/15838/3 "2014-02-17T16:58:30Z")

</div>

You are overwhelming the elasticsearch server. Instead of playing around  
with the timeout settings and the number of threads, consider using the  
Bulk API:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

The bulk processor class is extremely useful:  
[http://xbib.org/elasticsearch/1.0.0.Beta2-SNAPSHOT/apidocs/org/elasticsearch/action/bulk/BulkProcessor.html](http://xbib.org/elasticsearch/1.0.0.Beta2-SNAPSHOT/apidocs/org/elasticsearch/action/bulk/BulkProcessor.html)

--  
Ivan

On Mon, Feb 17, 2014 at 1:04 AM, Petr Janský [petr.jansky@6hats.cz](mailto:petr.jansky@6hats.cz) wrote:

> Hello,
> 
> I'm trying to index \>300k docs using Java API.
> 
> _public class Fetcher {_
> 
> - public static String server = "localhost"; \*
> - public static Integer port = 9300;\*
> - public static String index = "default";\*
> - public static String type = "default";\*
> - public static String typeAttributename = null;\*
> - static Client client = null;\*
> - private static Fetcher inst;\*
> - Settings settings = ImmutableSettings.settingsBuilder()\*
> - .put("cluster.name [http://cluster.name](http://cluster.name)", "elasticsearch")\*
> - .put("node.name [http://node.name](http://node.name)", "Killer")\*
> - .build();\*
> - public synchronized static Fetcher getInstace(){\*
> - if(inst == null){\*
> - inst = new Fetcher();\*
> - }\*
> - return inst;\*
> - }\*
> - public Fetcher() {\*
> - client = new TransportClient(settings).addTransportAddress(new  
> InetSocketTransportAddress(server, port));\*
> - }\*
> - public void index(DocumentVo document) {\*
> - try {\*
> - String type = Fetcher.type;\*
> - if(typeAttributename != null &&  
> document.getData().get(typeAttributename) != null){\*
> - type = document.getData().get(typeAttributename).toString();\*
> - type = type.toLowerCase();\*
> - }\*
> - IndexRequestBuilder rs =  
> client.prepareIndex().setIndex(index).setType(type);\*
> - rs.setTimeout(new TimeValue(10000));\*
> - rs.setSource(document.getData());\*
> - rs.execute().actionGet();\*
> - } catch (Exception e) {\*
> - e.printStackTrace();\*
> - client.close();\*
> - client = new TransportClient(settings).addTransportAddress(new  
> InetSocketTransportAddress(server, port));\*
> - index(document);\*
> - } \*
> - }\*
> - public void close(){\*
> - client.close();\*
> - }\*  
> _}_
> 
> in ~20 threads I run
> 
> _Fetcher.getInstace().index(document);_
> 
> I've created my own tokenizer filter that is quite slow so I'm getting
> 
> Feb 17, 2014 9:53:51 AM org.elasticsearch.client.transport  
> INFO: [Killer] failed to get node info for  
> [#transport#-1][inet[localhost/127.0.0.1:9300]], disconnecting...  
> org.elasticsearch.transport.ReceiveTimeoutTransportException:  
> [inet[localhost/127.0.0.1:9300]][cluster/nodes/info] request\_id [2899]  
> timed out after [5001ms]  
> at  
> org.elasticsearch.transport.TransportService$TimeoutHandler.run(TransportService.java:351)  
> at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source)  
> at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)  
> at java.lang.Thread.run(Unknown Source)
> 
> org.elasticsearch.client.transport.NoNodeAvailableException: No node  
> available  
> at  
> org.elasticsearch.client.transport.TransportClientNodesService$RetryListener.onFailure(TransportClientNodesService.java:249)  
> at  
> org.elasticsearch.action.TransportActionNodeProxy$1.handleException(TransportActionNodeProxy.java:84)  
> at  
> org.elasticsearch.transport.TransportService$Adapter$2$1.run(TransportService.java:311)  
> at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source)  
> at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)  
> at java.lang.Thread.run(Unknown Source)
> 
> It seems that  
> _rs.setTimeout(new TimeValue(10000));_  
> in my index method doesn't work.
> 
> How can I setup timeout for indexing using API?
> 
> Is it correct to use one TransportCilent for multiple(10-60) threads?
> 
> Thanks  
> Petr
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/f5f15b57-955c-4fcf-b225-3974e37e447b%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/f5f15b57-955c-4fcf-b225-3974e37e447b%40googlegroups.com)  
> .  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQCea8t7Mr3fFgpJQhWANxVttbA2KGh4qtGiRAq9TUimXw%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQCea8t7Mr3fFgpJQhWANxVttbA2KGh4qtGiRAq9TUimXw%40mail.gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [February 17, 2014, 5:17pm UTC](https://discuss.elastic.co/t/indexing-large-number-of-documents/15838/4 "2014-02-17T17:17:47Z")

</div>

Yes, the BulkProcessor is useful - the official link to the source is

[https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/action/bulk/BulkProcessor.java](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/action/bulk/BulkProcessor.java)

Thanks Ivan for pointing to my Javadoc but I think it is better to  
reference the source 😉

Petr, what ES cluster is this, how many nodes, how much heap?

You should carefully design your cluster and your indexing process before  
putting some indexing load on it - or you must travel down the bumpy road  
and learn for yourself and fix all kinds of issues to get it run smoothly.

It is ok to use TranspotrClient in a single instance but not like you do in  
the catch clause.

Also it is bad practice to drop the IndexResponse by ...  
execute().actionGet(). Because of the async nature of the API, you are  
sending far too many requests one after another. Please evaluate the  
responses, and continue only if limits are not exceeded and there is no  
error in a response - there are no exceptions thrown.

Jörg

On Mon, Feb 17, 2014 at 5:58 PM, Ivan Brusic [ivan@brusic.com](mailto:ivan@brusic.com) wrote:

> You are overwhelming the elasticsearch server. Instead of playing around  
> with the timeout settings and the number of threads, consider using the  
> Bulk API:  
> [http://www.elasticsearch.org/guide/en/elasticsearch/client/java-api/current/bulk.html](http://www.elasticsearch.org/guide/en/elasticsearch/client/java-api/current/bulk.html)
> 
> The bulk processor class is extremely useful:  
> [http://xbib.org/elasticsearch/1.0.0.Beta2-SNAPSHOT/apidocs/org/elasticsearch/action/bulk/BulkProcessor.html](http://xbib.org/elasticsearch/1.0.0.Beta2-SNAPSHOT/apidocs/org/elasticsearch/action/bulk/BulkProcessor.html)
> 
> --  
> Ivan
> 
> On Mon, Feb 17, 2014 at 1:04 AM, Petr Janský [petr.jansky@6hats.cz](mailto:petr.jansky@6hats.cz) wrote:
> 
> > Hello,
> > 
> > I'm trying to index \>300k docs using Java API.
> > 
> > _public class Fetcher {_
> > 
> > - public static String server = "localhost"; \*
> > - public static Integer port = 9300;\*
> > - public static String index = "default";\*
> > - public static String type = "default";\*
> > - public static String typeAttributename = null;\*
> > - static Client client = null;\*
> > - private static Fetcher inst;\*
> > - Settings settings = ImmutableSettings.settingsBuilder()\*
> > - .put("cluster.name [http://cluster.name](http://cluster.name)", "elasticsearch")\*
> > - .put("node.name [http://node.name](http://node.name)", "Killer")\*
> > - .build();\*
> > - public synchronized static Fetcher getInstace(){\*
> > - if(inst == null){\*
> > - inst = new Fetcher();\*
> > - }\*
> > - return inst;\*
> > - }\*
> > - public Fetcher() {\*
> > - client = new TransportClient(settings).addTransportAddress(new  
> > InetSocketTransportAddress(server, port));\*
> > - }\*
> > - public void index(DocumentVo document) {\*
> > - try {\*
> > - String type = Fetcher.type;\*
> > - if(typeAttributename != null &&  
> > document.getData().get(typeAttributename) != null){\*
> > - type = document.getData().get(typeAttributename).toString();\*
> > - type = type.toLowerCase();\*
> > - }\*
> > - IndexRequestBuilder rs =  
> > client.prepareIndex().setIndex(index).setType(type);\*
> > - rs.setTimeout(new TimeValue(10000));\*
> > - rs.setSource(document.getData());\*
> > - rs.execute().actionGet();\*
> > - } catch (Exception e) {\*
> > - e.printStackTrace();\*
> > - client.close();\*
> > - client = new TransportClient(settings).addTransportAddress(new  
> > InetSocketTransportAddress(server, port));\*
> > - index(document);\*
> > - } \*
> > - }\*
> > - public void close(){\*
> > - client.close();\*
> > - }\*  
> > _}_
> > 
> > in ~20 threads I run
> > 
> > _Fetcher.getInstace().index(document);_
> > 
> > I've created my own tokenizer filter that is quite slow so I'm getting
> > 
> > Feb 17, 2014 9:53:51 AM org.elasticsearch.client.transport  
> > INFO: [Killer] failed to get node info for  
> > [#transport#-1][inet[localhost/127.0.0.1:9300]], disconnecting...  
> > org.elasticsearch.transport.ReceiveTimeoutTransportException:  
> > [inet[localhost/127.0.0.1:9300]][cluster/nodes/info] request\_id [2899]  
> > timed out after [5001ms]  
> > at  
> > org.elasticsearch.transport.TransportService$TimeoutHandler.run(TransportService.java:351)  
> > at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source)  
> > at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)  
> > at java.lang.Thread.run(Unknown Source)
> > 
> > org.elasticsearch.client.transport.NoNodeAvailableException: No node  
> > available  
> > at  
> > org.elasticsearch.client.transport.TransportClientNodesService$RetryListener.onFailure(TransportClientNodesService.java:249)  
> > at  
> > org.elasticsearch.action.TransportActionNodeProxy$1.handleException(TransportActionNodeProxy.java:84)  
> > at  
> > org.elasticsearch.transport.TransportService$Adapter$2$1.run(TransportService.java:311)  
> > at java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source)  
> > at java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)  
> > at java.lang.Thread.run(Unknown Source)
> > 
> > It seems that  
> > _rs.setTimeout(new TimeValue(10000));_  
> > in my index method doesn't work.
> > 
> > How can I setup timeout for indexing using API?
> > 
> > Is it correct to use one TransportCilent for multiple(10-60) threads?
> > 
> > Thanks  
> > Petr
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/f5f15b57-955c-4fcf-b225-3974e37e447b%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/f5f15b57-955c-4fcf-b225-3974e37e447b%40googlegroups.com)  
> > .  
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQCea8t7Mr3fFgpJQhWANxVttbA2KGh4qtGiRAq9TUimXw%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQCea8t7Mr3fFgpJQhWANxVttbA2KGh4qtGiRAq9TUimXw%40mail.gmail.com)  
> .
> 
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoEsVSU09jeJCtxTwagAAf3V%2BWAj-EjuOugkyZDd8AOADA%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoEsVSU09jeJCtxTwagAAf3V%2BWAj-EjuOugkyZDd8AOADA%40mail.gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [February 17, 2014, 5:32pm UTC](https://discuss.elastic.co/t/indexing-large-number-of-documents/15838/5 "2014-02-17T17:32:15Z")

</div>

I respectively disagree. 🙂

In object oriented programming, you code to the interface, not the  
implementation. Then again, most people should be using code aware IDEs,  
which makes code lookup even easier.

Judging by his settings, I am assuming he is using a single local instance.  
Before getting into further design, using bulk indexing should be his  
baseline to measure by.

Cheers,

Ivan

On Mon, Feb 17, 2014 at 9:17 AM, [joergprante@gmail.com](mailto:joergprante@gmail.com) \<  
[joergprante@gmail.com](mailto:joergprante@gmail.com)\> wrote:

> Yes, the BulkProcessor is useful - the official link to the source is
> 
> [https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/action/bulk/BulkProcessor.java](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/action/bulk/BulkProcessor.java)
> 
> Thanks Ivan for pointing to my Javadoc but I think it is better to  
> reference the source 😉

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQA3KndkCaxpvnQ%3DZKCsO-UM%2BjAY-OGA%2BW7Ct0zo7TS2OA%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CALY%3DcQA3KndkCaxpvnQ%3DZKCsO-UM%2BjAY-OGA%2BW7Ct0zo7TS2OA%40mail.gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:49am UTC](https://discuss.elastic.co/t/indexing-large-number-of-documents/15838/6 "2017-07-06T01:49:36Z")

</div>


