# Logstash filtering. Extract data between two strings

**URL:** <https://discuss.elastic.co/t/logstash-filtering-extract-data-between-two-strings/251276>\
**Category:** Logstash\
**Created:** [October 7, 2020, 1:26pm UTC](https://discuss.elastic.co/t/logstash-filtering-extract-data-between-two-strings/251276 "2020-10-07T13:26:15Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Igor\_Olikh](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_olikh/32/46434_2.png) [@Igor\_Olikh](https://discuss.elastic.co/u/Igor_Olikh)\
**Post date:** [October 7, 2020, 1:26pm UTC](https://discuss.elastic.co/t/logstash-filtering-extract-data-between-two-strings/251276/1 "2020-10-07T13:26:15Z")

</div>

I have a UNIX log looks like:

ACTION **started** Lorem Ipsum is simply dummy text of the printing and typesetting industry.  
Lorem Ipsum has been the industry's standard dummy text ever since the 1500s **finished**  
when an unknown printer took a galley of type **started** and scrambled it to make a type specimen book. **finished** It has survived not only five centuries **started** but also the leap into electronic typesetting, remaining essentially unchanged **started** but also the leap into electronic typesetting, remaining essentially unchanged **finished**

I have to extract all the data between:

1. started and finished words (ex. and scrambled it to make a type specimen book)
2. if no finished for the started then between started and the next started (ex. but also the leap into electronic typesetting, remaining essentially unchanged)

Thanks.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [October 7, 2020, 3:00pm UTC](https://discuss.elastic.co/t/logstash-filtering-extract-data-between-two-strings/251276/2 "2020-10-07T15:00:40Z")

</div>

Maybe

```
 grok { match => { "message" => "started%{DATA:someText}(finished|started)" } }
```

---

<div class="post-metadata">

**Author:** ![Igor\_Olikh](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_olikh/32/46434_2.png) [@Igor\_Olikh](https://discuss.elastic.co/u/Igor_Olikh)\
**Post date:** [October 7, 2020, 3:53pm UTC](https://discuss.elastic.co/t/logstash-filtering-extract-data-between-two-strings/251276/3 "2020-10-07T15:53:16Z")

</div>

Thank you Badger, I'll check it later.

---

<div class="post-metadata">

**Author:** ![Igor\_Olikh](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_olikh/32/46434_2.png) [@Igor\_Olikh](https://discuss.elastic.co/u/Igor_Olikh)\
**Post date:** [October 8, 2020, 5:57am UTC](https://discuss.elastic.co/t/logstash-filtering-extract-data-between-two-strings/251276/4 "2020-10-08T05:57:17Z")

</div>

Hi Badger,  
What is the meaning of "someText" in this case? Could you please explain?

---

<div class="post-metadata">

**Author:** ![Igor\_Olikh](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_olikh/32/46434_2.png) [@Igor\_Olikh](https://discuss.elastic.co/u/Igor_Olikh)\
**Post date:** [October 8, 2020, 1:47pm UTC](https://discuss.elastic.co/t/logstash-filtering-extract-data-between-two-strings/251276/5 "2020-10-08T13:47:10Z")

</div>

I tried the grok. It seems working but if there are multiple rows between started and ended words, Logstash creates for each row a different document in the index.  
I need the text between these two words to be in the one document.  
Maybe the setting in input plugin are incorrect?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [October 8, 2020, 2:34pm UTC](https://discuss.elastic.co/t/logstash-filtering-extract-data-between-two-strings/251276/6 "2020-10-08T14:34:44Z")

</div>

> [@Igor\_Olikh](#):
>
> It seems working but if there are multiple rows between started and ended words

Then you would have to use a multiline codec on the input or the multiline option in filebeat (if applicable) to combine them.

---

<div class="post-metadata">

**Author:** ![Igor\_Olikh](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/igor_olikh/32/46434_2.png) [@Igor\_Olikh](https://discuss.elastic.co/u/Igor_Olikh)\
**Post date:** [October 8, 2020, 3:04pm UTC](https://discuss.elastic.co/t/logstash-filtering-extract-data-between-two-strings/251276/7 "2020-10-08T15:04:33Z")

</div>

Thank you!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 5, 2020, 3:04pm UTC](https://discuss.elastic.co/t/logstash-filtering-extract-data-between-two-strings/251276/8 "2020-11-05T15:04:44Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
