# How to extract things from beginning and end of a long single line log not parsing or extracting the thing in the middle of log

**URL:** <https://discuss.elastic.co/t/how-to-extract-things-from-beginning-and-end-of-a-long-single-line-log-not-parsing-or-extracting-the-thing-in-the-middle-of-log/136770>\
**Category:** Logstash\
**Created:** [June 20, 2018, 10:49pm UTC](https://discuss.elastic.co/t/how-to-extract-things-from-beginning-and-end-of-a-long-single-line-log-not-parsing-or-extracting-the-thing-in-the-middle-of-log/136770 "2018-06-20T22:49:36Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![lserranov](https://avatars.discourse-cdn.com/v4/letter/l/3bc359/32.png) [@lserranov](https://discuss.elastic.co/u/lserranov)\
**Post date:** [June 20, 2018, 10:49pm UTC](https://discuss.elastic.co/t/how-to-extract-things-from-beginning-and-end-of-a-long-single-line-log-not-parsing-or-extracting-the-thing-in-the-middle-of-log/136770/1 "2018-06-20T22:49:36Z")

</div>

Hi,

Suppose that I have a single line log. This log has like 500 words. Suppose that I'm interested in extracting the initial 5 words and the last 5 words and I'm not interested on parsing the middle.

Is there a way to do this?

I'm trying to save the compute power that i would be spending by filtering too many words which are in the middle but which I'm not interested on parsing.

Thanks  
Luis

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [June 20, 2018, 11:57pm UTC](https://discuss.elastic.co/t/how-to-extract-things-from-beginning-and-end-of-a-long-single-line-log-not-parsing-or-extracting-the-thing-in-the-middle-of-log/136770/2 "2018-06-20T23:57:29Z")

</div>

I would use a pair of grok patterns, one anchored with ^ (start of line) and one with $ (end of line).

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [June 21, 2018, 4:25am UTC](https://discuss.elastic.co/t/how-to-extract-things-from-beginning-and-end-of-a-long-single-line-log-not-parsing-or-extracting-the-thing-in-the-middle-of-log/136770/3 "2018-06-21T04:25:48Z")

</div>

> I would use a pair of grok patterns, one anchored with ^ (start of line) and one with $ (end of line).

Is that more efficient than having a single expression with `.*` in the middle?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [June 21, 2018, 3:47pm UTC](https://discuss.elastic.co/t/how-to-extract-things-from-beginning-and-end-of-a-long-single-line-log-not-parsing-or-extracting-the-thing-in-the-middle-of-log/136770/5 "2018-06-21T15:47:31Z")

</div>

> [@magnusbaeck](#):
>
> Is that more efficient than having a single expression with `.*` in the middle?

Definitely not. But if I think I am logically doing two different things (grabbing words from the beginning and grabbing words from the end) then I would still be tempted to use two patterns.

If I have a kafka topic that sends 20,000 identical lines each of which contains 500 words, then I can read it in two pipelines, one of which does

```
grok {
    match => { "message" => ["^%{WORD:a} %{WORD:b} %{WORD:c} %{WORD:d} %{WORD:e}.*%{WORD:v} %{WORD:w} %{WORD:x} %{WORD:y} %{WORD:z}$"] }
    id => "grok with 1 pattern"
}

```

and another that does

```
grok {
    match => { "message" => ["^%{WORD:a} %{WORD:b} %{WORD:c} %{WORD:d} %{WORD:e}", "%{WORD:v} %{WORD:w} %{WORD:x} %{WORD:y} %{WORD:z}$"] }
    break_on_match => false
    id => "grok with 2 patterns"
}

```

Then I can look at the pipeline stats and check duration\_in\_millis. I see 17088 when I use 1 pattern vs. 56211 for 2 patterns.

---

<div class="post-metadata">

**Author:** ![lserranov](https://avatars.discourse-cdn.com/v4/letter/l/3bc359/32.png) [@lserranov](https://discuss.elastic.co/u/lserranov)\
**Post date:** [June 21, 2018, 9:49pm UTC](https://discuss.elastic.co/t/how-to-extract-things-from-beginning-and-end-of-a-long-single-line-log-not-parsing-or-extracting-the-thing-in-the-middle-of-log/136770/6 "2018-06-21T21:49:21Z")

</div>

Thanks for the pointers!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 19, 2018, 9:49pm UTC](https://discuss.elastic.co/t/how-to-extract-things-from-beginning-and-end-of-a-long-single-line-log-not-parsing-or-extracting-the-thing-in-the-middle-of-log/136770/7 "2018-07-19T21:49:27Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
