# How to prevent empty lines

**URL:** <https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490>\
**Category:** Logstash\
**Created:** [May 28, 2018, 8:38am UTC](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490 "2018-05-28T08:38:24Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![npontes](https://avatars.discourse-cdn.com/v4/letter/n/ba8739/32.png) [@npontes](https://discuss.elastic.co/u/npontes)\
**Post date:** [May 28, 2018, 8:38am UTC](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490/1 "2018-05-28T08:38:24Z")

</div>

Hello,

I'm currently using Logstash 2.1.1 to retrieve data from Elasticsearch 1.4.4 and put it on a CSV to be imported to a database. My problem is that there are lines where I only have commas (Like this: ",,,,,,,,,,").  
Obviously when trying to import to a PostgreSQL it retrieves an error. is there any way for logstash to prevent this from happening?

My configuration file is the following:

```
input {
	elasticsearch {
		hosts => "xxx"
		index => "xxx"
		scroll => "1m"
		
	}
}

output {
	csv {
		path => "output_Int_1thread_nofilter.csv"
		fields => ["xxx", "xxx","xxx","xxx", "xxx","xxx", "xxx", "xxx"]
	}
}
```

---

<div class="post-metadata">

**Author:** ![atira](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/atira/32/28699_2.png) [@atira](https://discuss.elastic.co/u/atira)\
**Post date:** [May 28, 2018, 7:56pm UTC](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490/2 "2018-05-28T19:56:15Z")

</div>

Drop all messages before the output that only contain commas?

```auto
filter {
    if [message] =~ '/^,+$/' {
        drop { }
    }
}

```

---

<div class="post-metadata">

**Author:** ![npontes](https://avatars.discourse-cdn.com/v4/letter/n/ba8739/32.png) [@npontes](https://discuss.elastic.co/u/npontes)\
**Post date:** [May 29, 2018, 7:11am UTC](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490/3 "2018-05-29T07:11:06Z")

</div>

Hello,

I've tried to use that, but I had an error:

```
SyntaxError: (eval):54: syntax error, unexpected ','
              if (((event["[message]"] =~ //^,+$//))) # if [message] =~ '/^,+$/'
```

---

<div class="post-metadata">

**Author:** ![atira](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/atira/32/28699_2.png) [@atira](https://discuss.elastic.co/u/atira)\
**Post date:** [May 29, 2018, 7:28am UTC](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490/4 "2018-05-29T07:28:46Z")

</div>

Unexpected comma? Try to escape it.

`if [message] =~ '/^\,+$/'`

---

<div class="post-metadata">

**Author:** ![npontes](https://avatars.discourse-cdn.com/v4/letter/n/ba8739/32.png) [@npontes](https://discuss.elastic.co/u/npontes)\
**Post date:** [May 29, 2018, 7:31am UTC](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490/5 "2018-05-29T07:31:54Z")

</div>

Different error:

```
SyntaxError: (eval):54: syntax error, unexpected null
              if (((event["[message]"] =~ //^\,+$//))) # if [message] =~ '/^\,+$/'
```

---

<div class="post-metadata">

**Author:** ![atira](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/atira/32/28699_2.png) [@atira](https://discuss.elastic.co/u/atira)\
**Post date:** [May 29, 2018, 7:51am UTC](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490/6 "2018-05-29T07:51:00Z")

</div>

You just gotta love how regex syntax differs from app to app.

I can't test it myself, so two other variants to try:

`if [message] =~ /^,+$/`  
`if [message] =~ '^,+$'`

---

<div class="post-metadata">

**Author:** ![rcowart](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rcowart/32/88091_2.png) [@rcowart](https://discuss.elastic.co/u/rcowart)\
**Post date:** [May 29, 2018, 7:55am UTC](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490/7 "2018-05-29T07:55:46Z")

</div>

You don't need the single quotes. I use this method all the time... `if [message] =~ /^,+$/`

For example...

```auto
if [log][message] =~ /^DHCPOFFER .*$/ {

```

---

<div class="post-metadata">

**Author:** ![npontes](https://avatars.discourse-cdn.com/v4/letter/n/ba8739/32.png) [@npontes](https://discuss.elastic.co/u/npontes)\
**Post date:** [May 29, 2018, 7:57am UTC](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490/8 "2018-05-29T07:57:25Z")

</div>

I'm currently using the first suggestion @atira made without quotes like you said. It didn't return any error. I'll let it run and when it finishes I'll let you know if this configuration worked.

---

<div class="post-metadata">

**Author:** ![npontes](https://avatars.discourse-cdn.com/v4/letter/n/ba8739/32.png) [@npontes](https://discuss.elastic.co/u/npontes)\
**Post date:** [May 29, 2018, 9:42am UTC](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490/9 "2018-05-29T09:42:45Z")

</div>

After the execution of the configuration I went to check the file and I still have some lines with nothing but commas. Do you have other suggestions?

---

<div class="post-metadata">

**Author:** ![npontes](https://avatars.discourse-cdn.com/v4/letter/n/ba8739/32.png) [@npontes](https://discuss.elastic.co/u/npontes)\
**Post date:** [May 29, 2018, 12:28pm UTC](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490/10 "2018-05-29T12:28:47Z")

</div>

Just to clarify, right now my configuration is like this:

```
input {
	elasticsearch {
		hosts => "xxx"
		index => "xxx"
		scroll => "1m"
		
	}
}

filter {
    if [message] =~ /^,+$/ {
        drop { }
    }
}

output {
	csv {
		path => "output_Int_1thread_nofilter.csv"
		fields => ["xxx", "xxx","xxx","xxx", "xxx","xxx", "xxx", "xxx"]
	}
}

```

Despite this configuration I still have lines with only commas.

---

<div class="post-metadata">

**Author:** ![atira](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/atira/32/28699_2.png) [@atira](https://discuss.elastic.co/u/atira)\
**Post date:** [May 29, 2018, 3:20pm UTC](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490/11 "2018-05-29T15:20:15Z")

</div>

Ah okay. If your message field contains more than just commas, or no commas at all, then the conditional won't work.  
All we know that the output produces lines with only commas. So the input produces an empty line? Though I wonder how.

If that's the case, modify the filter to this:

```
filter {
    if [message] =~ /^$/ {
        drop { }
    }
}

```

This should delete all empty messages that come from the input.

---

<div class="post-metadata">

**Author:** ![npontes](https://avatars.discourse-cdn.com/v4/letter/n/ba8739/32.png) [@npontes](https://discuss.elastic.co/u/npontes)\
**Post date:** [May 30, 2018, 8:40am UTC](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490/12 "2018-05-30T08:40:21Z")

</div>

Hello @atira the solution you proposed didn't work ☹

In a universe of more than 800k lines I still hava 34 that are just ,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,,

---

<div class="post-metadata">

**Author:** ![atira](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/atira/32/28699_2.png) [@atira](https://discuss.elastic.co/u/atira)\
**Post date:** [May 30, 2018, 9:31am UTC](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490/13 "2018-05-30T09:31:35Z")

</div>

Uncharted territory for me, maybe someone else will look here too.

Brainstorming mode.  
It'd be good to know what events the input filter creates exactly.  
Can you configure eg. a file output? I presume that would simply write out the events to a file without any transformation. Then we would know what the csv filter receives.

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [May 30, 2018, 10:50am UTC](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490/14 "2018-05-30T10:50:23Z")

</div>

This is what is happening under the covers...

```auto
bin/logstash -i irb
Sending Logstash's logs to /Users/guy/tmp/logstash-6.2.4/logs which is now configured via log4j2.properties
irb(main):001:0> event = LogStash::Event.new
=> #<LogStash::Event:0x6ed0a4aa>
irb(main):002:0> event.get("[foo]")
=> nil
irb(main):003:0> line = 5.times.map{ event.get("[foo]") }
=> [nil, nil, nil, nil, nil]
irb(main):004:0> require 'csv'
=> true
irb(main):005:0> line.to_csv
=> ",,,,\n"
irb(main):006:0>

```

You have some documents coming from Elasticsearch that are a different schema and do not have any of the fields that the CSV output needs.

---

<div class="post-metadata">

**Author:** ![npontes](https://avatars.discourse-cdn.com/v4/letter/n/ba8739/32.png) [@npontes](https://discuss.elastic.co/u/npontes)\
**Post date:** [May 30, 2018, 12:24pm UTC](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490/15 "2018-05-30T12:24:47Z")

</div>

Right now I'm extracting everything I have on Index and then I'll try to find the positions where I have problems. Once I have any news I'll let you guys know

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [May 30, 2018, 2:40pm UTC](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490/16 "2018-05-30T14:40:30Z")

</div>

You can use a conditional to check for the existence of important fields and drop in the else branch.

```auto
  if [field1] and [field2] and [fieldN] {
    # do transforms on the "good" docs
  } else {
    drop {}
  }

```

---

<div class="post-metadata">

**Author:** ![npontes](https://avatars.discourse-cdn.com/v4/letter/n/ba8739/32.png) [@npontes](https://discuss.elastic.co/u/npontes)\
**Post date:** [May 30, 2018, 3:50pm UTC](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490/17 "2018-05-30T15:50:41Z")

</div>

I'll give that a go. But now I'm trying to analyse a 5GB text file to try to spot what might be causing this problem.

The solution that @atira proposed with :

```
if [message] =~ /^$/ {
        drop { }
}

```

made sense to me, and it didn't work.

Regarding what you proposed what transforms should I do?

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [May 30, 2018, 3:59pm UTC](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490/18 "2018-05-30T15:59:14Z")

</div>

Any transforms you need can go there - its just a bit harder to do a negating if condition.

```auto
if !( [field1] and [field2] and [fieldN] ) {
  drop {}
}

```

The reason why Attila's suggestion does not work in this case, is because the event or document sourced from Elasticsearch already has all the fields.

Do you understand the "whats is happening under the covers" code block?

---

<div class="post-metadata">

**Author:** ![npontes](https://avatars.discourse-cdn.com/v4/letter/n/ba8739/32.png) [@npontes](https://discuss.elastic.co/u/npontes)\
**Post date:** [May 31, 2018, 9:46am UTC](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490/19 "2018-05-31T09:46:39Z")

</div>

Yes, negating would work for me either because this index has more than 300 fields and I just want 40.

I think I understood. I'm just a little bit confused regarding the syntax you suggested.

It shouldbe like this:

```
if [xxx] and [xxx] and [xxx] {
    #can I do nothing here?
  } else {
    drop {}
  }
```

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [May 31, 2018, 10:20am UTC](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490/20 "2018-05-31T10:20:01Z")

</div>

The syntax of the negated conditional?

[Next page](https://discuss.elastic.co/t/how-to-prevent-empty-lines/133490.md?page=2)
