# Perl Script to insert data into ES ~13M to ~17M documents a day

**URL:** <https://discuss.elastic.co/t/perl-script-to-insert-data-into-es-13m-to-17m-documents-a-day/9709>\
**Category:** Elasticsearch\
**Created:** [November 14, 2012, 4:32am UTC](https://discuss.elastic.co/t/perl-script-to-insert-data-into-es-13m-to-17m-documents-a-day/9709 "2012-11-14T04:32:01Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Richard\_Pardue](https://avatars.discourse-cdn.com/v4/letter/r/ee59a6/32.png) [@Richard\_Pardue](https://discuss.elastic.co/u/Richard_Pardue)\
**Post date:** [November 14, 2012, 4:32am UTC](https://discuss.elastic.co/t/perl-script-to-insert-data-into-es-13m-to-17m-documents-a-day/9709/1 "2012-11-14T04:32:01Z")

</div>

Fast background on System setup:  
5 Nodes (JVM defaults setting)  
10 shards; 0 replicas per index (indexes based on date)  
Cluster Templates for new index base on git logstash - Works great

=== Template Start ===  
{  
"template": "al-logstash-_",  
"settings" : {  
"number\_of\_shards" : 10,  
"number\_of\_replicas" : 0,  
"index" : {  
"query" : { "default\_field" : "@message" },  
"store" : { "compress" : { "stored" : true, "tv": true } }  
}  
},  
"mappings": {  
"logs": {  
"\_all": { "enabled": false },  
"\_source": { "compress": true },  
"dynamic\_templates": [  
{  
"string\_template" : {  
"match" : "_",  
"mapping": { "type": "string", "index": "analyzed"  
},  
"match\_mapping\_type" : "string"  
}  
}  
],  
"properties" : {  
"@type" : { "type" : "string", "index" : "analyzed" },  
"@message" : { "type" : "string", "index" : "analyzed" },  
"@timestamp" : { "type" : "date", "index" : "analyzed" },  
"@ip" : { "type" : "string", "index" : "analyzed" },  
"@ident" : { "type" : "string", "index" : "analyzed" },  
"@authuser" : { "type" : "string", "index" : "analyzed" },  
"@protocol" : { "type" : "string", "index" : "analyzed" },  
"@method" : { "type" : "string", "index" : "analyzed" },  
"@request" : { "type" : "string", "index" : "analyzed" },  
"@cs\_referer" : { "type" : "string", "index" : "analyzed" },  
"@user\_agent" : { "type" : "string", "index" : "analyzed" },  
"@bytes" : { "type" : "integer", "index" : "analyzed" },  
"@response\_code" : { "type" : "integer", "index" :  
"analyzed" }  
}  
}  
}  
}  
=== Template End ===

Perl Script:  
=== Perl Script End ===

#!/usr/bin/perl

### 

# Parser for Logs

# Version 2.1

# Date: 2012-11-09

# Richard Pardue

### 

# Notes:

# 2012-11-06

# Removed the index and replaced with index\_bulk

# Changed index to bulk\_index

# Started adding error handling -\> beta

# 

# 2012-11-08

# Move the client connection object outside the loop

# 

# 2012-11-09

# Remove the server refresh after insert and the 1sec system pause

# 

### 

# Set perl to use ElasticSeach

use ElasticSearch;

# Set app elasticseach client

my $es = ElasticSearch-\>new(  
servers =\> ['127.0.0.1:9200',  
'127.0.0.1:9201',  
'127.0.0.1:9202',  
'127.0.0.1:9203',  
'127.0.0.1:9204'], # default  
'127.0.0.1:9200'  
transport =\> 'http', # default 'http'  
timeout =\> 30,  
#max\_requests =\> 10\_000, # default 10\_000

```
     #trace_calls => 'log_file.log',                                 
     #no_refresh => 0 | 1,

```

);

# Starts loop from STDIN

while (\<\>)  
{  
# Data that comes in from the pipe STDIN  
my $data = $\_;

```
    # Create 1st set of array data
    my @values = split('"', $data);                                     

```

# Main Data Set fields

```
    my @values0 = split(' ', @values[0]); # 

```

IP/DTY plus?  
my @values7 = split(' ', @values[7]); # ??  
my @values2 = split(' ', @values[2]); #  
Response Code and BW  
my @values1 = split(' ', @values[1]); #  
Method , request , protocol

```
    # Fields 
    my $field0 = @values0[0];                                         

```

# Client IP Address

```
    my $field1 = @values0[1];                                         

```

# -

```
    my $field2 = @values0[2];                                         

```

# -

```
    my $field3 = substr(substr("@values0[3] @values0[4]",1),0,-1); # 

```

DTS  
#my $field4 = @values7[7]; #  
CARTCOOKIEUUID  
#my $field5 = @values7[8]; #  
ASP.NET\_SessionId  
my $field6 = @values[5];

# User Agent

```
    my $field7 = @values2[0]; # 

```

ResponseCode  
my $field8 = @values2[1]; #  
bytes  
my $field9 = @values1[0]; #  
Method  
my $field10 = @values1[1]; #  
Request  
my $field11 = @values1[2]; #  
Protocol  
my $field12 = @values0[1]; #  
ident  
my $field13 = @values0[2]; #  
authuser  
my $field14 = @values[3]; #  
cs(Reerer)

```
    # Charge val for date changes
    my %mo = (
            'Jan'=>'01',
            'Feb'=>'02',
            'Mar'=>'03',
            'Apr'=>'04',
            'May'=>'05',
            'Jun'=>'06',                                               
  
            'Jul'=>'07',                                               
  
            'Aug'=>'08',
            'Sep'=>'09',
            'Oct'=>'10',
            'Nov'=>'11',
            'Dec'=>'12'
    );
    
    # Formats the log date/time from apache to ISOFormat
    my $tmptime = substr($field3,7,4) . '-' . $mo{substr($field3,3,3)} 

```

. '-' . substr($field3,0,2) . 'T' . substr($field3,12);

```
    # Format the Date for creating logstash index
    my $mylogdts = substr($field3,7,4) . '.' . $mo{substr($field3,3,3)} 

```

. '.' . substr($field3,0,2);

```
    # Set new log date/time to field
    $field3 = substr($tmptime,0,-6);
    
    # Set the logstash index name string
    $mylogdts = "al-logstash-$mylogdts";
                                                                       
# Forces a lookup of live nodes
#$results = $es->refresh_servers();
  
    # Uses the elasticseach client to bluk insert data into ElasticSeach
    $results = $es->bulk_index(
                                index => $mylogdts,               
            
                                type => 'logs',
                               refresh => 1,
                                #on_conflict => 'IGNORE',
                                #on_error => 'IGNORE',
                                on_error => sub { myError },
                                docs => [
                                            {
                                            data => {
                                                    '@type' => 'logs',
                                                    '@message' => 

```

$data,  
'@timestamp' =\>  
$field3,  
'@ip' =\> $field0,

```
                                                    '@ident' => 

```

$field12,  
'@authuser' =\>  
$field13,  
'@protocol' =\>  
$field11,  
'@method' =\>  
$field9,  
'@request' =\>  
$field10,  
'@cs\_referer' =\>  
$field14,  
'@user\_agent' =\>  
$field6,  
'@bytes' =\>  
$field8,  
'@response\_code' =\>  
$field7,  
},  
},  
]  
);  
}

### My Error

sub myError {  
print "\*\*\* ERROR: \*\*\*";  
print $results;  
$results = $es-\>refresh\_servers();  
print $results;  
}

exit 0;

=== Perl Script End ===

All log files are compressed

Then using zcat \* | perl my script.pl (~24 to 28 files)

The scripts then takes of running... inserting ~1000+ documents per sec  
base on the 'head' plug-in when the refresh button is clocked... very fast  
but then over a period of time ~ less then 30 mins the inserting drops to  
~60 to 300 inserts per click of refresh button and the inserts take over  
day to 2 days to finish.

Connect to each node: (Based on BigDesk plugin)  
Node 1 = ~8 to 10  
Node 2..5 = ~2 to 4

The System CPU running wide open... working great... still running query  
I have also ran multi zcat and scripts at the same time.

Is there a better way to kept the inserts running as over 1000+ per sec?  
Is there a better way to change the script and stop the CPU from running  
wide open?

--

---

<div class="post-metadata">

**Author:** ![chenryn](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/chenryn/32/44917_2.png) [@chenryn](https://discuss.elastic.co/u/chenryn)\
**Post date:** [November 14, 2012, 9:19am UTC](https://discuss.elastic.co/t/perl-script-to-insert-data-into-es-13m-to-17m-documents-a-day/9709/2 "2012-11-14T09:19:22Z")

</div>

I use perl script bulk\_index too. Suggest post your own map template  
with '"index"  
: "not\_analyzed"', and use httplite transport. I got ~5000+ at 2 nodes.

2012/11/14 Richard Pardue [richard.pardue@gmail.com](mailto:richard.pardue@gmail.com)

> Fast background on System setup:  
> 5 Nodes (JVM defaults setting)  
> 10 shards; 0 replicas per index (indexes based on date)  
> Cluster Templates for new index base on git logstash - Works great
> 
> === Template Start ===  
> {  
> "template": "al-logstash-_",  
> "settings" : {  
> "number\_of\_shards" : 10,  
> "number\_of\_replicas" : 0,  
> "index" : {  
> "query" : { "default\_field" : "@message" },  
> "store" : { "compress" : { "stored" : true, "tv": true } }  
> }  
> },  
> "mappings": {  
> "logs": {  
> "\_all": { "enabled": false },  
> "\_source": { "compress": true },  
> "dynamic\_templates": [  
> {  
> "string\_template" : {  
> "match" : "_",  
> "mapping": { "type": "string", "index": "analyzed"  
> },  
> "match\_mapping\_type" : "string"  
> }  
> }  
> ],  
> "properties" : {  
> "@type" : { "type" : "string", "index" : "analyzed" },  
> "@message" : { "type" : "string", "index" : "analyzed" },  
> "@timestamp" : { "type" : "date", "index" : "analyzed" },  
> "@ip" : { "type" : "string", "index" : "analyzed" },  
> "@ident" : { "type" : "string", "index" : "analyzed" },  
> "@authuser" : { "type" : "string", "index" : "analyzed" },  
> "@protocol" : { "type" : "string", "index" : "analyzed" },  
> "@method" : { "type" : "string", "index" : "analyzed" },  
> "@request" : { "type" : "string", "index" : "analyzed" },  
> "@cs\_referer" : { "type" : "string", "index" : "analyzed"  
> },  
> "@user\_agent" : { "type" : "string", "index" : "analyzed"  
> },  
> "@bytes" : { "type" : "integer", "index" : "analyzed" },  
> "@response\_code" : { "type" : "integer", "index" :  
> "analyzed" }  
> }  
> }  
> }  
> }  
> === Template End ===
> 
> Perl Script:  
> === Perl Script End ===
> 
> #!/usr/bin/perl
> 
> ### 
> 
> # Parser for Logs
> 
> # Version 2.1
> 
> # Date: 2012-11-09
> 
> # Richard Pardue
> 
> ### 
> 
> # Notes:
> 
> # 2012-11-06
> 
> # Removed the index and replaced with index\_bulk
> 
> # Changed index to bulk\_index
> 
> # Started adding error handling -\> beta
> 
> # 
> 
> # 2012-11-08
> 
> # Move the client connection object outside the loop
> 
> # 
> 
> # 2012-11-09
> 
> # Remove the server refresh after insert and the 1sec system pause
> 
> # 
> 
> ### 
> 
> # Set perl to use ElasticSeach
> 
> use Elasticsearch;
> 
> # Set app elasticseach client
> 
> my $es = Elasticsearch-\>new(  
> servers =\> ['127.0.0.1:9200',  
> '127.0.0.1:9201',  
> '127.0.0.1:9202',  
> '127.0.0.1:9203',  
> '127.0.0.1:9204'], # default '  
> 127.0.0.1:9200'  
> transport =\> 'http', # default 'http'  
> timeout =\> 30,  
> #max\_requests =\> 10\_000, # default 10\_000
> 
> ```
> #trace_calls => 'log_file.log',
> #no_refresh => 0 | 1,
> 
> ```
> 
> );
> 
> # Starts loop from STDIN
> 
> while (\<\>)  
> {  
> # Data that comes in from the pipe STDIN  
> my $data = $\_;
> 
> ```
> # Create 1st set of array data
> my @values = split('"', $data);
> 
> ```
> 
> # Main Data Set fields
> 
> ```
> my @values0 = split(' ', @values[0]); #
> 
> ```
> 
> IP/DTY plus?  
> my @values7 = split(' ', @values[7]); #  
> ??  
> my @values2 = split(' ', @values[2]); #  
> Response Code and BW  
> my @values1 = split(' ', @values[1]); #  
> Method , request , protocol
> 
> ```
> # Fields
> my $field0 = @values0[0];
> 
> ```
> 
> # Client IP Address
> 
> ```
> my $field1 = @values0[1];
> 
> ```
> 
> # -
> 
> ```
> my $field2 = @values0[2];
> 
> ```
> 
> # -
> 
> ```
> my $field3 = substr(substr("@values0[3] @values0[4]",1),0,-1); #
> 
> ```
> 
> DTS  
> #my $field4 = @values7[7];
> 
> # CARTCOOKIEUUID
> 
> ```
> #my $field5 = @values7[8];
> 
> ```
> 
> # ASP.NET\_SessionId
> 
> ```
> my $field6 = @values[5];
> 
> ```
> 
> # User Agent
> 
> ```
> my $field7 = @values2[0];
> 
> ```
> 
> # ResponseCode
> 
> ```
> my $field8 = @values2[1];
> 
> ```
> 
> # bytes
> 
> ```
> my $field9 = @values1[0];
> 
> ```
> 
> # Method
> 
> ```
> my $field10 = @values1[1]; #
> 
> ```
> 
> Request  
> my $field11 = @values1[2]; #  
> Protocol  
> my $field12 = @values0[1]; #  
> ident  
> my $field13 = @values0[2]; #  
> authuser  
> my $field14 = @values[3]; #  
> cs(Reerer)
> 
> ```
> # Charge val for date changes
> my %mo = (
> 'Jan'=>'01',
> 'Feb'=>'02',
> 'Mar'=>'03',
> 'Apr'=>'04',
> 'May'=>'05',
> 'Jun'=>'06',
> 
> 'Jul'=>'07',
> 
> 'Aug'=>'08',
> 'Sep'=>'09',
> 'Oct'=>'10',
> 'Nov'=>'11',
> 'Dec'=>'12'
> );
> 
> # Formats the log date/time from apache to ISOFormat
> my $tmptime = substr($field3,7,4) . '-' . $mo{substr($field3,3,3)}
> 
> ```
> 
> . '-' . substr($field3,0,2) . 'T' . substr($field3,12);
> 
> ```
> # Format the Date for creating logstash index
> my $mylogdts = substr($field3,7,4) . '.' .
> 
> ```
> 
> $mo{substr($field3,3,3)} . '.' . substr($field3,0,2);
> 
> ```
> # Set new log date/time to field
> $field3 = substr($tmptime,0,-6);
> 
> # Set the logstash index name string
> $mylogdts = "al-logstash-$mylogdts";
> 
> # Forces a lookup of live nodes
> #$results = $es->refresh_servers();
> 
> # Uses the elasticseach client to bluk insert data into
> 
> ```
> 
> ElasticSeach  
> $results = $es-\>bulk\_index(  
> index =\> $mylogdts,
> 
> ```
> type => 'logs',
> refresh => 1,
> #on_conflict => 'IGNORE',
> #on_error => 'IGNORE',
> on_error => sub { myError },
> docs => [
> {
> data => {
> '@type' => 'logs',
> '@message' =>
> 
> ```
> 
> $data,  
> '@timestamp' =\>  
> $field3,  
> '@ip' =\> $field0,
> 
> ```
> '@ident' =>
> 
> ```
> 
> $field12,  
> '@authuser' =\>  
> $field13,  
> '@protocol' =\>  
> $field11,  
> '@method' =\>  
> $field9,  
> '@request' =\>  
> $field10,  
> '@cs\_referer' =\>  
> $field14,  
> '@user\_agent' =\>  
> $field6,  
> '@bytes' =\>  
> $field8,  
> '@response\_code'  
> =\> $field7,  
> },  
> },  
> ]  
> );  
> }
> 
> ### My Error
> 
> sub myError {  
> print "\*\*\* ERROR: \*\*\*";  
> print $results;  
> $results = $es-\>refresh\_servers();  
> print $results;  
> }
> 
> exit 0;
> 
> === Perl Script End ===
> 
> All log files are compressed
> 
> Then using zcat \* | perl my script.pl (~24 to 28 files)
> 
> The scripts then takes of running... inserting ~1000+ documents per sec  
> base on the 'head' plug-in when the refresh button is clocked... very fast  
> but then over a period of time ~ less then 30 mins the inserting drops to  
> ~60 to 300 inserts per click of refresh button and the inserts take over  
> day to 2 days to finish.
> 
> Connect to each node: (Based on BigDesk plugin)  
> Node 1 = ~8 to 10  
> Node 2..5 = ~2 to 4
> 
> The System CPU running wide open... working great... still running query  
> I have also ran multi zcat and scripts at the same time.
> 
> Is there a better way to kept the inserts running as over 1000+ per sec?  
> Is there a better way to change the script and stop the CPU from running  
> wide open?
> 
> --

--

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [November 14, 2012, 12:26pm UTC](https://discuss.elastic.co/t/perl-script-to-insert-data-into-es-13m-to-17m-documents-a-day/9709/3 "2012-11-14T12:26:34Z")

</div>

Hiya

Couple of notes on this:

> my $es = Elasticsearch-\>new(
> 
> ```
> servers => ['127.0.0.1:9200',
> '127.0.0.1:9201',
> '127.0.0.1:9202',
> '127.0.0.1:9203',
> '127.0.0.1:9204'], # default
> 
> ```
> 
> '127.0.0.1:9200'

Why are you running several nodes on the same machine? You won't get any  
performance benefit out of this, and in fact you're probably hurting  
performance.

> ```
> transport => 'http', # default
> 
> ```
> 
> 'http'

for fastest performance, use the 'curl' backend.

> ```
> # Forces a lookup of live nodes
> #$results = $es->refresh_servers();
> 
> ```

No need to refresh - this is automatic

> ```
> $results = $es->bulk_index(
> index => $mylogdts,
>                   
> type => 'logs',
> refresh => 1,
> #on_conflict => 'IGNORE',
> #on_error => 'IGNORE',
> on_error => sub { myError },
> docs => [
> {
> data => {
> 
> ```

You're using bulk, but only indexing one document at a time. Accumulate  
eg 1,000 documents in @docs then pass all of them to bulk at the same  
time.

clint

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:04am UTC](https://discuss.elastic.co/t/perl-script-to-insert-data-into-es-13m-to-17m-documents-a-day/9709/4 "2017-07-06T03:04:22Z")

</div>


