# One time batch processing big number of files

**URL:** <https://discuss.elastic.co/t/one-time-batch-processing-big-number-of-files/46461>\
**Category:** Logstash\
**Created:** [April 5, 2016, 10:37pm UTC](https://discuss.elastic.co/t/one-time-batch-processing-big-number-of-files/46461 "2016-04-05T22:37:55Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![Alexander\_Popov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alexander_popov/32/7154_2.png) [@Alexander\_Popov](https://discuss.elastic.co/u/Alexander_Popov)\
**Post date:** [April 5, 2016, 10:37pm UTC](https://discuss.elastic.co/t/one-time-batch-processing-big-number-of-files/46461/1 "2016-04-05T22:37:55Z")

</div>

Have ~ 50K files in single directory( ~200Gb ) of logs.  
trying to process to parse and add them to elasticsearch  
my config:

```
input {
    file {
    path => "/mnt/storage/*.txt"
    sincedb_path => "/dev/null"
    codec => "json_lines"
    start_position => "beginning"
    ignore_older => 123123123131
      }
}

filter {
    useragent {
      source => "[req][headers][user-agent]"
      target => "user-agent"
    }
    mutate {
       rename => { "msg" => "message" }
       convert => { # unify some field types
        "level" => "string"
        "page" => "string"
         }
       remove_field => ["[f1][f1nested]", "[f2][f2-nested][f2nested-nested]", "metadata" ] #removes obsolete fileds, which should not goes even to s3 or elasticsearch
    }

 if [level] == "10" {
          mutate {
              update => { "level" => "trace" }
                }
          }

    if [level] == "20" {
          mutate {
              update => { "level" => "debug" }
                }
          }
    if [level] == "30" {
          mutate {
              update => { "level" => "info" }
                }
          }
    if [level] == "40" {
          mutate {
              update => { "level" => "warn" }
                }
          }
    if [level] == "50" {
          mutate {
              update => { "level" => "error" }
                }
          }
    if [level] == "60" {
          mutate {
              update => { "level" => "fatal" }
                }
          }

  clone { #clone to put original in s3
    clones => ["details"]
  }

    if [type] != "details" { # simplify record for elasticsearch

      if [some][some][some-name] == "some" {
        drop { }
          }
      if [some][some][some-name] == "some" {
        drop { }
          }
  
    ruby {        
        code => "event['fileId'] = event['version'] if event['version'].is_a?(String)"
    }

      mutate {
      #de-neste some fields like:
       rename => { 
          "[command][command]" => "myCommand"
          "[command][channel_id]" => "myChannelId"
          "[command][channel_name]" => "myChannelName"
          "[email][address]" => "email"
             ...... 
          }
       remove_field => ["command"......, "version", ....]
      }

    ruby {        
        code => "event['email'] = '' if event['email'] and not event['email'].is_a?(String)"
    }

    }
}

output {
  stdout { codec => dots }
  if [type] != "details" {
  # for debug
  # file {
  # path => "_out/elastic.log"   
  # }
   elasticsearch{
      hosts => "localhost:9200"
    }
  }
 # if [type] == "details" {
 # s3 {
 # ......
 # }  
 # }
}

```

running as

LS\_HEAP\_SIZE="15g" /opt/logstash/bin/logstash -f logstash.conf

after some time it is crashed with error:

Settings: Default pipeline workers: 8  
Pipeline main started  
java.lang.OutOfMemoryError: Java heap space  
Dumping heap to /opt/logstash/heapdump.hprof ...  
Unable to create /opt/logstash/heapdump.hprof: File exists  
Pipeline main has been shutdown  
stopping pipeline {:id=\>"main"}  
Error: Your application used more memory than the safety cap of 15G.

and Logs not even started pushed in elasticsearch

with single file all works fine

Where I'm wrong? does it any other way to parse big log files?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [April 5, 2016, 10:41pm UTC](https://discuss.elastic.co/t/one-time-batch-processing-big-number-of-files/46461/2 "2016-04-05T22:41:17Z")

</div>

At a guess it's probably not handling the number of files, can you break things up a bit more on the input?

---

<div class="post-metadata">

**Author:** ![Alexander\_Popov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alexander_popov/32/7154_2.png) [@Alexander\_Popov](https://discuss.elastic.co/u/Alexander_Popov)\
**Post date:** [April 5, 2016, 10:59pm UTC](https://discuss.elastic.co/t/one-time-batch-processing-big-number-of-files/46461/3 "2016-04-05T22:59:37Z")

</div>

for test purpose - yes,  
but big folders with huge amount of logs is usual for us

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [April 6, 2016, 6:25am UTC](https://discuss.elastic.co/t/one-time-batch-processing-big-number-of-files/46461/4 "2016-04-06T06:25:55Z")

</div>

Which version of LS and the File input are you using? The most recent version of the file input and filewatch are better at handling these scenarios. Please bear in mind though, the file input was designed for tailing files - it tries its best to work for the read files case.

---

<div class="post-metadata">

**Author:** ![Alexander\_Popov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alexander_popov/32/7154_2.png) [@Alexander\_Popov](https://discuss.elastic.co/u/Alexander_Popov)\
**Post date:** [April 6, 2016, 6:48am UTC](https://discuss.elastic.co/t/one-time-batch-processing-big-number-of-files/46461/5 "2016-04-06T06:48:33Z")

</div>

logstash 2.3.0, just recent installed.  
Since I need file input as initial seed, i don't need to watch for files changes. my files is immutable

Does it any other ways to seed big data in LS? I tried get same data from S3 input, but after 2.5 hr  
it still reads files list

```
S3 input: Found key {:key=>"full-log/ls.s3.ip-10-0-0-46.2016-01-14T23.07.part16152.txt", :level=>:debug, :file=>"logstash/inputs/s3.rb", :line=>"111", :method=>"list_new_files"}
S3 input: Adding to objects[] {:key=>"full-log/ls.s3.ip-10-0-0-46.2016-01-14T23.07.part16152.txt", :level=>:debug, :file=>"logstash/inputs/s3.rb", :line=>"116", :method=>"list_new_files"}
```

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [April 6, 2016, 6:55am UTC](https://discuss.elastic.co/t/one-time-batch-processing-big-number-of-files/46461/6 "2016-04-06T06:55:43Z")

</div>

You should try the json codec? The file input is already line oriented. I think the JsonLines line buffer is filling up with concatenated "lines" because it is looking for a "\n" but the file input has taken them out already.

---

<div class="post-metadata">

**Author:** ![Alexander\_Popov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alexander_popov/32/7154_2.png) [@Alexander\_Popov](https://discuss.elastic.co/u/Alexander_Popov)\
**Post date:** [April 6, 2016, 12:24pm UTC](https://discuss.elastic.co/t/one-time-batch-processing-big-number-of-files/46461/7 "2016-04-06T12:24:20Z")

</div>

Wow!  
Looks better now. Good to know about this behavior.

I also reduce close\_older value and max\_open\_files to faster cycle among files  
now config:

```
input {
    file {
        path => "/mnt/storage/*.txt"
        sincedb_path => "/tmp/sincedb"
        codec => "json"
        start_position => "beginning"
        ignore_older => 864000000
        close_older => 2
        max_open_files => 10
   }
}
output {
        stdout { codec => dots }
}
```

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [April 6, 2016, 1:00pm UTC](https://discuss.elastic.co/t/one-time-batch-processing-big-number-of-files/46461/8 "2016-04-06T13:00:59Z")

</div>

Perfect. I am glad to see you using close\_older and max\_open\_files. Did you read my blog post about these changes? [the evolving story of the file input](https://www.elastic.co/blog/the-evolving-story-about-the-logstash-file-input)

So some background on the codec mismatch, in LS there are three types of sources in the inputs 1) provides bytes 2) provides lines 3) provides protocol string; and there are various codecs that accept only one of these three. Unfortunately, we don't have a mechanism to establish when the input source to codec is mismatched. If we did, we could warn when file is used with json-lines.

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [April 6, 2016, 1:08pm UTC](https://discuss.elastic.co/t/one-time-batch-processing-big-number-of-files/46461/9 "2016-04-06T13:08:21Z")

</div>

Out of curiosity, how are your lines of JSON being generated? Nginx or Apache log formatting perhaps? If so, be aware that this technique can generate invalid JSON - some user supplied data can be incorrectly escaped (0xHH instead of u00HH).

---

<div class="post-metadata">

**Author:** ![Alexander\_Popov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alexander_popov/32/7154_2.png) [@Alexander\_Popov](https://discuss.elastic.co/u/Alexander_Popov)\
**Post date:** [April 6, 2016, 1:08pm UTC](https://discuss.elastic.co/t/one-time-batch-processing-big-number-of-files/46461/10 "2016-04-06T13:08:55Z")

</div>

Not seen your blog post yet, decide to use it myself.  
Will read it now.

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [April 6, 2016, 1:10pm UTC](https://discuss.elastic.co/t/one-time-batch-processing-big-number-of-files/46461/11 "2016-04-06T13:10:01Z")

</div>

Glad to help, happy stashing 😃

---

<div class="post-metadata">

**Author:** ![Alexander\_Popov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alexander_popov/32/7154_2.png) [@Alexander\_Popov](https://discuss.elastic.co/u/Alexander_Popov)\
**Post date:** [April 6, 2016, 1:37pm UTC](https://discuss.elastic.co/t/one-time-batch-processing-big-number-of-files/46461/12 "2016-04-06T13:37:03Z")

</div>

My logs was generated with logstash s3 output with codec =\> "json\_lines"  
It was a "hardcopy" of all logs which was tried to push into elasticsearch.

Since our log records is very heterogeneous, many entries was failed while parsing in elasticsearch  
To not lose anything We stored it also in S3.

Last days I added many normalization filters and tried to reparse S3 logs back 🙂

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [April 6, 2016, 1:37pm UTC](https://discuss.elastic.co/t/one-time-batch-processing-big-number-of-files/46461/13 "2016-04-06T13:37:52Z")

</div>

Ahhhhh OK.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 5:03am UTC](https://discuss.elastic.co/t/one-time-batch-processing-big-number-of-files/46461/14 "2017-07-06T05:03:29Z")

</div>


