# Indexing 150 G of data

**URL:** <https://discuss.elastic.co/t/indexing-150-g-of-data/13915>\
**Category:** Elasticsearch\
**Created:** [October 10, 2013, 11:31pm UTC](https://discuss.elastic.co/t/indexing-150-g-of-data/13915 "2013-10-10T23:31:30Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![gautam\_singhania](https://avatars.discourse-cdn.com/v4/letter/g/a183cd/32.png) [@gautam\_singhania](https://discuss.elastic.co/u/gautam_singhania)\
**Post date:** [October 10, 2013, 11:31pm UTC](https://discuss.elastic.co/t/indexing-150-g-of-data/13915/1 "2013-10-10T23:31:30Z")

</div>

Hi

I am planning to index 150 Gb of jsob objects (200 million rows of  
index+json object) using elastic search.

The box i have is 250 Gb with 8 Gb of ram  
i am planning to use the bulk entry point

also planning to increase RAm by doing

set JAVA\_OPTS=-Xmx2g -Xms2g

then split the files using unix split into chunks and do the following

#!/usr/bin/perl  
use strict;  
use warnings;  
my @files = glob("input/\*");  
foreach my $f (@files) {  
my $cmd = "curl -s -XPOST '[http://brigho.com:9200/\_bulk](http://brigho.com:9200/_bulk)' --data-binary  
@".$f;

print `$cmd`;  
#print $cmd."\n";  
}

has anyone tried doing this? will this work? are the system configs  
enough? any pointers would be helpful

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Hendrik](https://avatars.discourse-cdn.com/v4/letter/h/839c29/32.png) [@Hendrik](https://discuss.elastic.co/u/Hendrik)\
**Post date:** [October 19, 2013, 5:49pm UTC](https://discuss.elastic.co/t/indexing-150-g-of-data/13915/2 "2013-10-19T17:49:29Z")

</div>

depends

first: es is a distributed search engine, so maybe its a good idea to setup  
more than one node/machine if performance is not sufficient  
then: 8gb ram is not bad but make sure youre on a 64bit operation system  
and 64bit java so that the full meory can be addressed by the operation  
system

To use your 8 GB memory do (leave 2 gig to the operation system)  
set JAVA\_OPTS=%JAVA\_OPTS% -Xmx6g -Xms6g

To get your data into ea you maybe want consider the UDP Bulk endpoint (

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

)  
or use your perl scripts (iam not into perl, so i cannot say anything about  
it). For perl there is also a ea client available:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

Am Freitag, 11. Oktober 2013 01:31:30 UTC+2 schrieb gautam singhania:

> Hi
> 
> I am planning to index 150 Gb of jsob objects (200 million rows of  
> index+json object) using Elasticsearch.
> 
> The box i have is 250 Gb with 8 Gb of ram  
> i am planning to use the bulk entry point
> 
> also planning to increase RAm by doing
> 
> set JAVA\_OPTS=-Xmx2g -Xms2g
> 
> then split the files using unix split into chunks and do the following
> 
> #!/usr/bin/perl  
> use strict;  
> use warnings;  
> my @files = glob("input/\*");  
> foreach my $f (@files) {  
> my $cmd = "curl -s -XPOST '[http://brigho.com:9200/\_bulk](http://brigho.com:9200/_bulk)' --data-binary  
> @".$f;
> 
> print `$cmd`;  
> #print $cmd."\n";  
> }
> 
> has anyone tried doing this? will this work? are the system configs  
> enough? any pointers would be helpful

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:11am UTC](https://discuss.elastic.co/t/indexing-150-g-of-data/13915/3 "2017-07-06T02:11:37Z")

</div>


