# Processing large amoung of data with Logstash

**URL:** <https://discuss.elastic.co/t/processing-large-amoung-of-data-with-logstash/242020>\
**Category:** Logstash\
**Created:** [July 21, 2020, 11:56am UTC](https://discuss.elastic.co/t/processing-large-amoung-of-data-with-logstash/242020 "2020-07-21T11:56:59Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![evannobre](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/evannobre/32/72502_2.png) [@evannobre](https://discuss.elastic.co/u/evannobre)\
**Post date:** [July 21, 2020, 11:56am UTC](https://discuss.elastic.co/t/processing-large-amoung-of-data-with-logstash/242020/1 "2020-07-21T11:56:59Z")

</div>

Hi,

Sorry for the text, but I tried to detail my problem as much as possible to make it clear.

I am part of a medical devices team. We run hundreds of tests over the devices to assure the quality and we save the results in a .txt file.

In that file, each line represent a campaign of tests where we have the informations about the campaign, the tests, the results of each test and the comments of each test as well.

To have an idea, is something like (each ";" represent a new field):  
\* **Line 1 :** Name\_campaign; Timestamp; PC\_host; OS; IP; PixDyn\_version; Test1; Result1; Comment1; Test2; Result2; Comment2  
\* **Line 2 :** Name\_campaign; Timestamp; PC\_host; OS; IP; PixDyn\_version; Test1; Result1; Comment1; Test2; Result2; Comment2; Test3; Result3; Comment3; Test4; Result4; Comment4

As you can notice, the number os tests can vary from line to line. Expecting to make my life easier, I created a .py to modify the file. The .py get the line with the maximum number of fields (X) and complete the other lines with an empty string ('') until X. Like that, I could use just one grok filter to match each occurrence instead of have one filter per line.

My grok filter is a bizarre thing that looks like that :

```auto
grok {
		patterns_dir => ["/patterns"]
		match => { "message" => ["^%{NUMBER:VersionTableauStat};(?<SessionName>%{YEAR}\_%{MONTHNUM}\_%{MONTHDAY}\__%{HOUR}\_%{MINUTE}\_%{SECOND});%{HOSTNAME:PCName};%{CISCO_REASON:OS};%{IP:Host_DLL};%{IP:NIOS};%{IP:FPGA1};%{IP:FPGA2};%{IP:PULL_DLL};%{WORD:SN_PU};%{WORD:SN_Détecteur};%{WORD:SignOn};(?<PULB_PN>%{AD_TYPE:AD};%{WORD:Test_name};%{RESULT_TEST:Result};%{COMMENT_TEST:Comment})"] }
}

```

That is the short version, because, for example, in one project I can have 150 tests, i.e., the patterns {WORD:Test\_name};%{RESULT\_TEST:Result};%{COMMENT\_TEST:Comment} will be replicated 150 times. Before someone ask, I created a regex to the last two patterns and they work (not the problem).

**With that grok I was expecting to have in the variables "Test\_name", "Result" and "Comment" the name of all the tests I ran in one campaign, its results and comments respectively. Like that I could use Kibana or Grafana to visualize which tests failed in some campaign and to monitore in real time.**

And the grok works if I do not have a large amount of tests. When I run over a campaign that has 20 tests, for example, the grok matches. But when I have a huge amount of tests, like 150, it shows "groktimeout". To avoid that, I put "timeout\_milis =\> 0" and the file is running for 2:30 hours and no results yet.

**My question is : there is a way to make the process faster/to optimize the filter/an easier solution ?**

Thank you for read it. 😚.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 21, 2020, 1:00pm UTC](https://discuss.elastic.co/t/processing-large-amoung-of-data-with-logstash/242020/2 "2020-07-21T13:00:26Z")

</div>

How are the RESULT\_TEST and COMMENT\_TEST patterns defined?

---

<div class="post-metadata">

**Author:** ![evannobre](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/evannobre/32/72502_2.png) [@evannobre](https://discuss.elastic.co/u/evannobre)\
**Post date:** [July 21, 2020, 1:14pm UTC](https://discuss.elastic.co/t/processing-large-amoung-of-data-with-logstash/242020/3 "2020-07-21T13:14:35Z")

</div>

They are defined as follow :

```auto
RESULT_TEST \s*[0-9]*
COMMENT_TEST \s*[-\a-zA-Z0-9àâäéèëîïôùûüÿ\(\)\_]*

```

With RESULT\_TEST I am expecting to read and empty string or numbers and with COMMENT\_TEST I am expecting to read an empty string or the special characters from French and some special characters such as (, ), - and \_

---

<div class="post-metadata">

**Author:** ![Jenni](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jenni/32/29684_2.png) [@Jenni](https://discuss.elastic.co/u/Jenni)\
**Post date:** [July 21, 2020, 1:40pm UTC](https://discuss.elastic.co/t/processing-large-amoung-of-data-with-logstash/242020/4 "2020-07-21T13:40:09Z")

</div>

I haven't tested the performance, but for simplicity I'd ignore your patterns, have trust in the semicolon as a delimiter and loop over the repeating fields with Ruby.

```auto
dissect {
  mapping => { "message" => "%{Name_campaign}; %{Timestamp}; %{PC_host}; %{OS}; %{IP}; %{PixDyn_version}; %{tests_string}" }
}
ruby {
  code => '
    counter = 0
    tests = []
    event.get("tests_string").split("; ").each_with_index { |v,k|
      i = k%3
      tests[counter] = {} if i == 0
      fieldname = i==0?"name":i==1?"comment":"result"
      tests[counter][fieldname] = v
      counter = counter+1 if i==2
    }
    event.set("tests", tests)
  '
}
```

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 21, 2020, 2:05pm UTC](https://discuss.elastic.co/t/processing-large-amoung-of-data-with-logstash/242020/5 "2020-07-21T14:05:14Z")

</div>

I agree with Jenni. Use dissect to parse the first part of the line then use ruby. You could use the .scan method of the String class

```
    ruby {
        code => '
            s = event.get("tests_string")
            if s
                event.set("matches", s.scan(/\s*([^;]+); ([^;]+); ([^;]+)(;|$)/))
            end
        '
    }

```

which will result in a variable length array such as

```
       "matches" => [
    [0] [
        [0] "Test1",
        [1] "Result1",
        [2] "Comment1",
        [3] ";"
    ],
    [1] [
        [0] "Test2",
        [1] "Result2",
        [2] "Comment2",
        [3] ""
    ]
],

```

You will likely want to iterate over the array and reformat the data.

---

<div class="post-metadata">

**Author:** ![evannobre](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/evannobre/32/72502_2.png) [@evannobre](https://discuss.elastic.co/u/evannobre)\
**Post date:** [July 23, 2020, 7:28am UTC](https://discuss.elastic.co/t/processing-large-amoung-of-data-with-logstash/242020/6 "2020-07-23T07:28:21Z")

</div>

Thank you, Badger and Jenni.

Both solutions worked and they are much faster than my idea (less than 7 minutes to send all the document).

🙂

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 20, 2020, 7:28am UTC](https://discuss.elastic.co/t/processing-large-amoung-of-data-with-logstash/242020/7 "2020-08-20T07:28:30Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
