# How to import CSV files, where couple of fields have multiline content

**URL:** <https://discuss.elastic.co/t/how-to-import-csv-files-where-couple-of-fields-have-multiline-content/280453>\
**Category:** Logstash\
**Created:** [August 4, 2021, 3:02pm UTC](https://discuss.elastic.co/t/how-to-import-csv-files-where-couple-of-fields-have-multiline-content/280453 "2021-08-04T15:02:11Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![itokai](https://avatars.discourse-cdn.com/v4/letter/i/db5fbb/32.png) [@itokai](https://discuss.elastic.co/u/itokai)\
**Post date:** [August 4, 2021, 3:02pm UTC](https://discuss.elastic.co/t/how-to-import-csv-files-where-couple-of-fields-have-multiline-content/280453/1 "2021-08-04T15:02:11Z")

</div>

Hello,

how to accurately import CSV where lines contain fields with multiline content? Default separator is comma, but multiline content is surrounded with double-quotes. For example:

```auto
Summary,Issue key,Issue id,Parent id,Issue Type,Status,Project key
Content1,Content2,"Multiline content 3 line1
Multiline content line2

Multiline content line3
",Content4,"Multiline content 5 line1
Multiline content 5 line2
Multiline content 5 line3
",Content6,Content7

```

Input file is actually Jira issues export jira.issueviews:searchrequest-csv-all-fields.  
What's the optimal way to load Jira issues into ELK? If it turns out that Jira can be imported without medior CSV step, new topic should be created I guess.

Regards

---

<div class="post-metadata">

**Author:** ![itokai](https://avatars.discourse-cdn.com/v4/letter/i/db5fbb/32.png) [@itokai](https://discuss.elastic.co/u/itokai)\
**Post date:** [August 5, 2021, 8:29am UTC](https://discuss.elastic.co/t/how-to-import-csv-files-where-couple-of-fields-have-multiline-content/280453/2 "2021-08-05T08:29:27Z")

</div>

There is a [Multiline codec plugin](https://www.elastic.co/guide/en/logstash/current/plugins-codecs-multiline.html)  
but what pattern/configuration to use to distinguish

1. simple fields (just comma separated)
2. multiline fields (double-quotes+comma separated)  
?

Regards

---

<div class="post-metadata">

**Author:** ![itokai](https://avatars.discourse-cdn.com/v4/letter/i/db5fbb/32.png) [@itokai](https://discuss.elastic.co/u/itokai)\
**Post date:** [August 6, 2021, 8:26am UTC](https://discuss.elastic.co/t/how-to-import-csv-files-where-couple-of-fields-have-multiline-content/280453/3 "2021-08-06T08:26:15Z")

</div>

All fields are now quoted for sake of consistency.

```auto
        codec => multiline {        
            pattern => "^\""
            negate => true
            what => "previous"
        }  

```

works for most of the line-breaks, except when the line starts with closing double-quote, followed by comma (",). Pseudo example above covers these cases.  
So, how to write proper regexp: _every new document starts with double-quote and that double-quote is not followed by comma_?

Regards

---

<div class="post-metadata">

**Author:** ![itokai](https://avatars.discourse-cdn.com/v4/letter/i/db5fbb/32.png) [@itokai](https://discuss.elastic.co/u/itokai)\
**Post date:** [August 6, 2021, 12:28pm UTC](https://discuss.elastic.co/t/how-to-import-csv-files-where-couple-of-fields-have-multiline-content/280453/4 "2021-08-06T12:28:29Z")

</div>

Have also tried

```auto
        codec => multiline {        
            pattern => "^\"[^,]"
            negate => true
            what => "previous"
        } 

```

but getting all sorts of troubles like _:exception=\>#\<CSV::MalformedCSVError: Missing or stray quote in line 1\>_

Finally, after pre-processing (like removing empty lines), file is imported. However, how to automate it?  
_Multiline (rich content), quoted fields_ are common use-case. Is there a configuration and import example? Might want to include it in _Multiline codec plugin_ documentation?

Regards

---

<div class="post-metadata">

**Author:** ![xeraa](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xeraa/32/48181_2.png) [@xeraa](https://discuss.elastic.co/u/xeraa)\
**Post date:** [August 10, 2021, 11:13pm UTC](https://discuss.elastic.co/t/how-to-import-csv-files-where-couple-of-fields-have-multiline-content/280453/5 "2021-08-10T23:13:51Z")

</div>

Hm, I think your regular expression should have two groups to capture this — one for the blank line (`\n`, assuming Unix line endings) and one for `",`. Maybe something like this?

```
codec => multiline {        
    pattern => "^(\",|\n)"
    negate => true
    what => "previous"
}
```

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [August 11, 2021, 2:00am UTC](https://discuss.elastic.co/t/how-to-import-csv-files-where-couple-of-fields-have-multiline-content/280453/6 "2021-08-11T02:00:43Z")

</div>

I do not believe a multiline codec maintains enough state to solve this in the general case. Consider

field1,field2  
"field1, ya know",field2  
"line1 of field1  
line2  
",field2

A full solution would need a codec that consumes a character at a time, not a line at a time. There are all sorts of corner cases where this mismatch breaks things. It has been discussed a lot (you can tell that from [this thread](https://github.com/logstash-plugins/logstash-input-file/issues/210) between Colin and Guy).

Another variant is a line oriented input consuming a line that ends in \n and feeding to a codec that reads character pairs in UTF16. Because the input does not consume the second byte of the UTF16 character that starts with \n, the endianess of the rest of the file is flipped, and the text turned into gibberish.

It will (IMHO) never get fixed. logstash functionality is getting pulled back into beat processors, or pushed forward into elasticsearch processing pipelines. I very much doubt Elastic have an appetite to re-architect the logstash input design that almost always works.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 8, 2021, 2:01am UTC](https://discuss.elastic.co/t/how-to-import-csv-files-where-couple-of-fields-have-multiline-content/280453/7 "2021-09-08T02:01:06Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
