# What is the best way for huge bulk file indexing?

**URL:** <https://discuss.elastic.co/t/what-is-the-best-way-for-huge-bulk-file-indexing/6074>\
**Category:** Elasticsearch\
**Created:** [December 6, 2011, 1:12pm UTC](https://discuss.elastic.co/t/what-is-the-best-way-for-huge-bulk-file-indexing/6074 "2011-12-06T13:12:29Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![ko526so](https://avatars.discourse-cdn.com/v4/letter/k/82dd89/32.png) [@ko526so](https://discuss.elastic.co/u/ko526so)\
**Post date:** [December 6, 2011, 1:12pm UTC](https://discuss.elastic.co/t/what-is-the-best-way-for-huge-bulk-file-indexing/6074/1 "2011-12-06T13:12:29Z")

</div>

I have to index huge volume of data frequently for research purpose.  
60,000,000 docs are one of my recent task for indexing. Fortunately, the  
size of docs is very small, so the total size of bulk index file for 60 M  
docs is only 11 G.

I used the following command for Solr to prevent memory error and high  
performance. And it was good.

curl [http://localhost:8080/example/update](http://localhost:8080/example/update) -F stream.file=/tmp/artists.xml

Is there any similar command with ES like the above?

Thanks always.

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [December 6, 2011, 8:46pm UTC](https://discuss.elastic.co/t/what-is-the-best-way-for-huge-bulk-file-indexing/6074/2 "2011-12-06T20:46:03Z")

</div>

You need to chunk it yourself into bulk indexing requests.

On Tue, Dec 6, 2011 at 3:12 PM, ko526so [kono.kim@gmail.com](mailto:kono.kim@gmail.com) wrote:

> I have to index huge volume of data frequently for research purpose.  
> 60,000,000 docs are one of my recent task for indexing. Fortunately, the  
> size of docs is very small, so the total size of bulk index file for 60 M  
> docs is only 11 G.
> 
> I used the following command for Solr to prevent memory error and high  
> performance. And it was good.
> 
> curl [http://localhost:8080/example/update](http://localhost:8080/example/update) -F stream.file=/tmp/artists.xml
> 
> Is there any similar command with ES like the above?
> 
> Thanks always.

---

<div class="post-metadata">

**Author:** ![colinsurprenant](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/colinsurprenant/32/14776_2.png) [@colinsurprenant](https://discuss.elastic.co/u/colinsurprenant)\
**Post date:** [December 6, 2011, 10:01pm UTC](https://discuss.elastic.co/t/what-is-the-best-way-for-huge-bulk-file-indexing/6074/3 "2011-12-06T22:01:09Z")

</div>

For this I wrote a multithreaded writer which reads a file, bundle n  
(usually 500) documents, queue the chunks which are picked up by the  
writer threads which bulk index over http in round robin over all my  
cluster nodes.

Now, there's a lot of tweeking that can be done to optimize  
performance, see this thread for some guidelines:  
[https://groups.google.com/a/elasticsearch.com/group/users/msg/06d62ea3ceb4db30](https://groups.google.com/a/elasticsearch.com/group/users/msg/06d62ea3ceb4db30)

Colin

On Tue, Dec 6, 2011 at 8:12 AM, ko526so [kono.kim@gmail.com](mailto:kono.kim@gmail.com) wrote:

> I have to index huge volume of data frequently for research purpose.  
> 60,000,000 docs are one of my recent task for indexing. Fortunately, the  
> size of docs is very small, so the total size of bulk index file for 60 M  
> docs is only 11 G.
> 
> I used the following command for Solr to prevent memory error and high  
> performance. And it was good.
> 
> curl [http://localhost:8080/example/update](http://localhost:8080/example/update) -F stream.file=/tmp/artists.xml
> 
> Is there any similar command with ES like the above?
> 
> Thanks always.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 6, 2011, 10:20pm UTC](https://discuss.elastic.co/t/what-is-the-best-way-for-huge-bulk-file-indexing/6074/4 "2011-12-06T22:20:03Z")

</div>

> For this I wrote a multithreaded writer which reads a file, bundle n  
> (usually 500) documents, queue the chunks which are picked up by the  
> writer threads which bulk index over http in round robin over all my  
> cluster nodes.  
> Is it opensourced somewhere ?

Thanks,  
David.

--  
David Pilato  
[http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
Twitter : @dadoonet

---

<div class="post-metadata">

**Author:** ![colinsurprenant](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/colinsurprenant/32/14776_2.png) [@colinsurprenant](https://discuss.elastic.co/u/colinsurprenant)\
**Post date:** [December 6, 2011, 10:46pm UTC](https://discuss.elastic.co/t/what-is-the-best-way-for-huge-bulk-file-indexing/6074/5 "2011-12-06T22:46:58Z")

</div>

No its not, sorry... this code is just a part of another project. It  
wouldn't be a bad idea to make this piece generic and opensource it.  
It's in Ruby. If you still have interest, I'll see what I can do.

Colin

On Tue, Dec 6, 2011 at 5:20 PM, [david@pilato.fr](mailto:david@pilato.fr) [david@pilato.fr](mailto:david@pilato.fr) wrote:

> > For this I wrote a multithreaded writer which reads a file, bundle n
> 
> > (usually 500) documents, queue the chunks which are picked up by the
> 
> > writer threads which bulk index over http in round robin over all my
> 
> > cluster nodes.
> 
> Is it opensourced somewhere ?
> 
> Thanks,
> 
> David.
> 
> --  
> David Pilato  
> [http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
> Twitter : @dadoonet

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 6, 2011, 10:53pm UTC](https://discuss.elastic.co/t/what-is-the-best-way-for-huge-bulk-file-indexing/6074/6 "2011-12-06T22:53:53Z")

</div>

Thanks Colin. I thought it was in Java. I don't know Ruby at this time ☹

I don't need it by now. I was just curious on the way you implemented it.

Cheers  
David

Le 6 déc. 2011 à 23:46, Colin Surprenant [colin.surprenant@gmail.com](mailto:colin.surprenant@gmail.com) a écrit :

> No its not, sorry... this code is just a part of another project. It  
> wouldn't be a bad idea to make this piece generic and opensource it.  
> It's in Ruby. If you still have interest, I'll see what I can do.
> 
> Colin
> 
> On Tue, Dec 6, 2011 at 5:20 PM, [david@pilato.fr](mailto:david@pilato.fr) [david@pilato.fr](mailto:david@pilato.fr) wrote:
> 
> > > For this I wrote a multithreaded writer which reads a file, bundle n
> > 
> > > (usually 500) documents, queue the chunks which are picked up by the
> > 
> > > writer threads which bulk index over http in round robin over all my
> > 
> > > cluster nodes.
> > 
> > Is it opensourced somewhere ?
> > 
> > Thanks,
> > 
> > David.
> > 
> > --  
> > David Pilato  
> > [http://dev.david.pilato.fr/](http://dev.david.pilato.fr/)  
> > Twitter : @dadoonet

---

<div class="post-metadata">

**Author:** ![Karussell1](https://avatars.discourse-cdn.com/v4/letter/k/50afbb/32.png) [@Karussell1](https://discuss.elastic.co/u/Karussell1)\
**Post date:** [December 8, 2011, 12:00am UTC](https://discuss.elastic.co/t/what-is-the-best-way-for-huge-bulk-file-indexing/6074/7 "2011-12-08T00:00:23Z")

</div>

> Thanks Colin. I thought it was in Java. I don't know Ruby at this time ☹

Not complicated at all

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

Regards,  
Peter.

---

<div class="post-metadata">

**Author:** ![Karussell1](https://avatars.discourse-cdn.com/v4/letter/k/50afbb/32.png) [@Karussell1](https://discuss.elastic.co/u/Karussell1)\
**Post date:** [December 8, 2011, 12:01am UTC](https://discuss.elastic.co/t/what-is-the-best-way-for-huge-bulk-file-indexing/6074/8 "2011-12-08T00:01:23Z")

</div>

ups, ok. its the multithreaded reader which is interesting 🙂 ...  
sorry.

On 8 Dez., 01:00, Karussell [tableyourt...@googlemail.com](mailto:tableyourt...@googlemail.com) wrote:

> > Thanks Colin. I thought it was in Java. I don't know Ruby at this time ☹
> 
> Not complicated at all
> 
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/reference/java-api/bulk.html)
> 
> Regards,  
> Peter.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:46am UTC](https://discuss.elastic.co/t/what-is-the-best-way-for-huge-bulk-file-indexing/6074/9 "2017-07-06T03:46:04Z")

</div>


