# Indexing 570 millions rows

**URL:** <https://discuss.elastic.co/t/indexing-570-millions-rows/63646>\
**Category:** Elasticsearch\
**Created:** [October 21, 2016, 6:51pm UTC](https://discuss.elastic.co/t/indexing-570-millions-rows/63646 "2016-10-21T18:51:09Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![jgenari](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jgenari/32/12646_2.png) [@jgenari](https://discuss.elastic.co/u/jgenari)\
**Post date:** [October 21, 2016, 6:51pm UTC](https://discuss.elastic.co/t/indexing-570-millions-rows/63646/1 "2016-10-21T18:51:09Z")

</div>

Hello my name is João Sakai and I'm in the middle of the greatest challenge of my life;

**"Using a logstash I have to index one CSV with 46 columns and 570 millions rows on elasticsearch as soon as possible"**

Environment Configurations:

**Logstash Configuration**

* * *

**_Amazon Instance Type_:** _m3.medium_

**_Logstash config_:**

![](https://us1.discourse-cdn.com/elastic/original/2X/9/98e4e54a1edeba7563dcb94ddd0a529c70b0beb9.png)

**_Index Template Config_:**

 ![](https://us1.discourse-cdn.com/elastic/original/2X/e/e9895c1e9c826d5625cef9181c76941446d00114.png)

* * *

**Note:**  
As you can see I already do the optimization for _" number of replicas: 0 "_ and _" refresh interval: -1 "_;

* * *
\*\*Elasticsearch Configuration\*\*
* * *

**_Amazon Instance Type_:** _m3.2xlarge_

_ **Cluster Config:** _  
 ![](https://us1.discourse-cdn.com/elastic/original/2X/9/93672bbcb40ffca45b6715d28ea7b372e03e76c7.png)

* * *

**Note:**  
I'm following the Indexing performance guide: [https://www.elastic.co/guide/en/elasticsearch/guide/current/indexing-performance.html](https://www.elastic.co/guide/en/elasticsearch/guide/current/indexing-performance.html)

* * *

My results were really disappointing! 😢

Executing the index process in 1 hour the total of indexed documents was only **2 millions** which leads me to think that there is something wrong about logstash configuration, elasticsearch configuration or anything else.

Is there some configuration wrong? Have I change the ec2 instances configurations?

Someone could give me some insights about how to index a large bulk of data?

---

<div class="post-metadata">

**Author:** ![eperry](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/eperry/32/551_2.png) [@eperry](https://discuss.elastic.co/u/eperry)\
**Post date:** [October 21, 2016, 7:35pm UTC](https://discuss.elastic.co/t/indexing-570-millions-rows/63646/2 "2016-10-21T19:35:25Z")

</div>

how many data nodes do you have spawned.

A quick test for logstash would be to send the data output to /dev/null just to see the CSV get read, parsed and outputed without disk speeds  
replace your elasticsearch section with

output{  
file{  
path =\> "/dev/null"  
}  
}

for example I index about 10K docuements per second on 9 data nodes, each with 24 cpu's and 30GB heap, and EMC SAN. but my documents are also pretty complex.

---

<div class="post-metadata">

**Author:** ![eperry](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/eperry/32/551_2.png) [@eperry](https://discuss.elastic.co/u/eperry)\
**Post date:** [October 21, 2016, 7:36pm UTC](https://discuss.elastic.co/t/indexing-570-millions-rows/63646/3 "2016-10-21T19:36:16Z")

</div>

Oh and I would install marvel to see your cluster's performance see where it is bottle necking

---

<div class="post-metadata">

**Author:** ![anhlqn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anhlqn/32/5454_2.png) [@anhlqn](https://discuss.elastic.co/u/anhlqn)\
**Post date:** [October 21, 2016, 11:20pm UTC](https://discuss.elastic.co/t/indexing-570-millions-rows/63646/4 "2016-10-21T23:20:51Z")

</div>

You should remove the `stdout { codec => json }`. It should be enabled only when you are debugging. Having `stdout` slows processing down a lot.

In addition to @eperry suggestion, you can use the logstash metrics plugin [https://www.elastic.co/guide/en/logstash/current/plugins-filters-metrics.html](https://www.elastic.co/guide/en/logstash/current/plugins-filters-metrics.html) to measure messages through LS.

I think that the `flush_size => 100` is a bit low. The default one is already 125.

The size of your csv file also affects read speed.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:10pm UTC](https://discuss.elastic.co/t/indexing-570-millions-rows/63646/5 "2017-07-05T22:10:22Z")

</div>


