# Best method - Importing 50x10gb CSV files into Elasticsearch on GCE

**URL:** <https://discuss.elastic.co/t/best-method-importing-50x10gb-csv-files-into-elasticsearch-on-gce/1758>\
**Category:** Elasticsearch\
**Created:** [June 2, 2015, 4:40pm UTC](https://discuss.elastic.co/t/best-method-importing-50x10gb-csv-files-into-elasticsearch-on-gce/1758 "2015-06-02T16:40:12Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jordan\_Schwartz](https://avatars.discourse-cdn.com/v4/letter/j/46a35a/32.png) [@Jordan\_Schwartz](https://discuss.elastic.co/u/Jordan_Schwartz)\
**Post date:** [June 2, 2015, 4:40pm UTC](https://discuss.elastic.co/t/best-method-importing-50x10gb-csv-files-into-elasticsearch-on-gce/1758/1 "2015-06-02T16:40:12Z")

</div>

I am looking to use ES to index about 400 million records broken into 50 files with about 360 columns in each file.  
Once indexed, the data will remain static. I am just looking for the best approach to load up this data initially.

The data is in CSV format. I signed up for Google Compute Engine and spun up 3 ES instances.

I attempted to use logstash locally on my mac-book and send the files to the remote ES server but I am only getting about 400 documents per second.

There has to be a better approach at loading this big data.

Any suggestions?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [June 2, 2015, 10:10pm UTC](https://discuss.elastic.co/t/best-method-importing-50x10gb-csv-files-into-elasticsearch-on-gce/1758/2 "2015-06-02T22:10:18Z")

</div>

Logstash will do a lot more than that, try running it closer to the ES server.

---

<div class="post-metadata">

**Author:** ![javadevmtl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javadevmtl/32/45613_2.png) [@javadevmtl](https://discuss.elastic.co/u/javadevmtl)\
**Post date:** [June 3, 2015, 5:36am UTC](https://discuss.elastic.co/t/best-method-importing-50x10gb-csv-files-into-elasticsearch-on-gce/1758/3 "2015-06-03T05:36:06Z")

</div>

360 column is sugesting "large" documents what's the average size of doc?

Are you using bulk?

- You can probably disable replicas of the index until indexing is done.
- Maybe set the refresh interval to higher then default for the index for the duration of the bulk inserts.

The rest could depend on your hardware possibly. I have to check some of the setting I did on mine.

---

<div class="post-metadata">

**Author:** ![Jordan\_Schwartz](https://avatars.discourse-cdn.com/v4/letter/j/46a35a/32.png) [@Jordan\_Schwartz](https://discuss.elastic.co/u/Jordan_Schwartz)\
**Post date:** [June 5, 2015, 2:37pm UTC](https://discuss.elastic.co/t/best-method-importing-50x10gb-csv-files-into-elasticsearch-on-gce/1758/4 "2015-06-05T14:37:16Z")

</div>

I did try running it straight from the GCE VM. I was actually getting slower indexing rates (~300-400/sec).  
I am using out of the box configurations. I'm assuming I should still see better results.

@javadevmtl I like the idea of disabling replicas. I'm going to try that. Here is my logstash config. See anything wrong?

input {  
file {  
path =\> "/elasticsearch/data.csv"  
start\_position =\> "beginning"  
type =\> "data"  
}  
}  
filter {  
csv{  
separator =\> "|"  
}  
}  
output {

```
 elasticsearch {
     action => "index"
     host => "localhost"
     port => "9200"
     index => "indextest-data1" 
     workers => 2
     #cluster => "elasticsearch-cluster"
    protocol => "http"
cluster => "elasticsearch-cluster"
 }

 #stdout { codec => json }

```

}

Thanks for the help so far.

---

<div class="post-metadata">

**Author:** ![javadevmtl](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/javadevmtl/32/45613_2.png) [@javadevmtl](https://discuss.elastic.co/u/javadevmtl)\
**Post date:** [June 5, 2015, 2:51pm UTC](https://discuss.elastic.co/t/best-method-importing-50x10gb-csv-files-into-elasticsearch-on-gce/1758/5 "2015-06-05T14:51:53Z")

</div>

Do you only have one ES node? I would have said try to load balance your log stash request to ES node(s) you may have. The JAVA client will do that and I think you can set logstash config to use Native API instead of HTTP.

Also on ES cluster try to set...  
"indices.store.throttle.max\_bytes\_per\_sec": "200mb" For my hardware 200mb is good (Defaults to 20mb).

And on your index settings try setting...  
"index.refresh\_interval": "30s" (Defaults to 1s)  
"index.translog.flush\_threshold\_size": "1000mb" \<-- This again depends on your hardware (Defaults to 512mb).

---

<div class="post-metadata">

**Author:** ![Jordan\_Schwartz](https://avatars.discourse-cdn.com/v4/letter/j/46a35a/32.png) [@Jordan\_Schwartz](https://discuss.elastic.co/u/Jordan_Schwartz)\
**Post date:** [June 5, 2015, 8:43pm UTC](https://discuss.elastic.co/t/best-method-importing-50x10gb-csv-files-into-elasticsearch-on-gce/1758/6 "2015-06-05T20:43:29Z")

</div>

I switched to Amazon EC2  
I have 2 ES nodes in a load balanced environment.  
Still getting a slow rate but I will try the config settings now

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:09am UTC](https://discuss.elastic.co/t/best-method-importing-50x10gb-csv-files-into-elasticsearch-on-gce/1758/7 "2017-07-06T00:09:20Z")

</div>


