# Csv filter output splitting a row in two because of some special char

**URL:** <https://discuss.elastic.co/t/csv-filter-output-splitting-a-row-in-two-because-of-some-special-char/288710>\
**Category:** Logstash\
**Created:** [November 9, 2021, 5:54am UTC](https://discuss.elastic.co/t/csv-filter-output-splitting-a-row-in-two-because-of-some-special-char/288710 "2021-11-09T05:54:34Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![stillfreem](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stillfreem/32/85628_2.png) [@stillfreem](https://discuss.elastic.co/u/stillfreem)\
**Post date:** [November 9, 2021, 5:54am UTC](https://discuss.elastic.co/t/csv-filter-output-splitting-a-row-in-two-because-of-some-special-char/288710/1 "2021-11-09T05:54:34Z")

</div>

I have a `csv` file with **7543** entries that I'd like to get to cloud SIEM.

My config file is as follows

```auto
input {
                        file
                                {
                                        path => "/home/<user>/myfile.csv"
                                        start_position => beginning
                                        #sincedb_path => ""
                                }
                }

filter {
                        csv
                                {
                                        autodetect_column_names => true
                                        separator => ","
                                        skip_empty_columns => false
                                }
                }

output {
       microsoft-logstash-output-azure-loganalytics {
            workspace_id => "MyID"
            workspace_key => "Mykey"
            custom_log_table_name => "<Mytablename>"
            key_names => ['name'. 'of', 'my', 'columns']
                                                        }
        stdout{}
        }

```

The configuration is working perfectly but I have a couple of problems.

1. Not all entries were sent to the SIEM. I run it several times, renaming the file and the first time 4000 something were sent, second time 6000 something was sent, but not all 7543.

2. Second problem is that I had some errors popping out on the standard output and it seems that the Logstash split several entries on two because of what I believe is a char that needs to be escaped or something like that. Check the below entry and pay closer attention on the bolt part

The console output

 ![logstash](https://us1.discourse-cdn.com/elastic/original/3X/2/f/2f1addbac1c2553b3a7c7103ce6cf587381bc9e7.png)

1. The file that Logstash reads can sometimes change some of the values in the same entries, meaning it's not adding new rows but just updating some of the values in the old ones. I did a test, changing one value and Logstash didn't recognize this. The way sincedb works is that its just waiting for new rows but how about old entries with changed values, is there a way I tell Logstash to watch for this too?

Thank you all in advance 🙂

---

<div class="post-metadata">

**Author:** ![stillfreem](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stillfreem/32/85628_2.png) [@stillfreem](https://discuss.elastic.co/u/stillfreem)\
**Post date:** [November 9, 2021, 6:27am UTC](https://discuss.elastic.co/t/csv-filter-output-splitting-a-row-in-two-because-of-some-special-char/288710/2 "2021-11-09T06:27:48Z")

</div>

I found what the problem is regarding question 2.  
There was a new line as a value in the Comment column

```auto
SLES 11_x000d_
Adabas test server

```

Once I remove the new line it worked out for me 🙂  
`SLES 11_x000d_Adabas test server`  
 ![image](https://us1.discourse-cdn.com/elastic/original/3X/3/2/32c4b7e0693cd9a5b3f2b903f878c3de7138ea21.png)

Then I wrote the following piece of code and it worked. All entries were ingested into my SIEM without csv parse errors 🙂

```auto
codec => multiline {
                                                                pattern => "^[0-9]"
                                                                negate => true
                                                                what => "previous"
                                                           }

```

Only Q3 remains unanswered for now 🙂

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [November 9, 2021, 5:18pm UTC](https://discuss.elastic.co/t/csv-filter-output-splitting-a-row-in-two-because-of-some-special-char/288710/3 "2021-11-09T17:18:04Z")

</div>

> [@stillfreem](#):
>
> The way sincedb works is that its just waiting for new rows but how about old entries with changed values, is there a way I tell Logstash to watch for this too?

No, in "tail" mode the file input assumes all new data is appended to the end of the file. It will not re-read data it has already read.

---

<div class="post-metadata">

**Author:** ![stillfreem](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stillfreem/32/85628_2.png) [@stillfreem](https://discuss.elastic.co/u/stillfreem)\
**Post date:** [November 9, 2021, 6:00pm UTC](https://discuss.elastic.co/t/csv-filter-output-splitting-a-row-in-two-because-of-some-special-char/288710/4 "2021-11-09T18:00:33Z")

</div>

Thank you @Badger, so is it possible at all, then to Tell Logstash to watch for mods on already read data?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [November 9, 2021, 6:04pm UTC](https://discuss.elastic.co/t/csv-filter-output-splitting-a-row-in-two-because-of-some-special-char/288710/5 "2021-11-09T18:04:20Z")

</div>

I cannot think of a way to do that.

---

<div class="post-metadata">

**Author:** ![stillfreem](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stillfreem/32/85628_2.png) [@stillfreem](https://discuss.elastic.co/u/stillfreem)\
**Post date:** [November 9, 2021, 6:06pm UTC](https://discuss.elastic.co/t/csv-filter-output-splitting-a-row-in-two-because-of-some-special-char/288710/6 "2021-11-09T18:06:43Z")

</div>

Thank you anyway @Badger

---

<div class="post-metadata">

**Author:** ![stillfreem](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stillfreem/32/85628_2.png) [@stillfreem](https://discuss.elastic.co/u/stillfreem)\
**Post date:** [November 10, 2021, 1:01pm UTC](https://discuss.elastic.co/t/csv-filter-output-splitting-a-row-in-two-because-of-some-special-char/288710/7 "2021-11-10T13:01:04Z")

</div>

Hi again @Badger and everyone,  
Some of my entries contain German characters and this is breaking my parser.

Check this out

```auto
[WARN] 2021-11-10 13:42:57.718 [[main]<file] plain - Received an event that has a different character encoding than you configured. {:text=>"xxx,Windows,ACTIVE,10.21.36.164,xxx,Gr\\xE4der, XXX,\\r", :expected_charset=>"UTF-8"}

```

For instance in the above entry the name is is not `Gr\\xE4der` and this char `\\xE4` seems to be in German. How can I tell Logstash to look for UTF-8 and German chars?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [November 10, 2021, 2:29pm UTC](https://discuss.elastic.co/t/csv-filter-output-splitting-a-row-in-two-because-of-some-special-char/288710/8 "2021-11-10T14:29:17Z")

</div>

> [@stillfreem](#):
>
> Gr\xE4der

Specify the charset on the input...

```
file {
    codec => plain { charset => "someValue" }
    ....

```

If your text contains `Gr\\xE4der` for Gräder then it is not UTF-8. It could be ISO 8859, CP-1252 or even some other encoding.

---

<div class="post-metadata">

**Author:** ![stillfreem](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stillfreem/32/85628_2.png) [@stillfreem](https://discuss.elastic.co/u/stillfreem)\
**Post date:** [November 10, 2021, 2:59pm UTC](https://discuss.elastic.co/t/csv-filter-output-splitting-a-row-in-two-because-of-some-special-char/288710/9 "2021-11-10T14:59:24Z")

</div>

Yeap ISO 8859 did the trick, many thanks 🙂

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 8, 2021, 3:00pm UTC](https://discuss.elastic.co/t/csv-filter-output-splitting-a-row-in-two-because-of-some-special-char/288710/10 "2021-12-08T15:00:02Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
