# Logstash: Ingesting CSV, filtered by CSV columns

**URL:** <https://discuss.elastic.co/t/logstash-ingesting-csv-filtered-by-csv-columns/272831>\
**Category:** Logstash\
**Created:** [May 12, 2021, 4:37pm UTC](https://discuss.elastic.co/t/logstash-ingesting-csv-filtered-by-csv-columns/272831 "2021-05-12T16:37:58Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![hsalim](https://avatars.discourse-cdn.com/v4/letter/h/c67d28/32.png) [@hsalim](https://discuss.elastic.co/u/hsalim)\
**Post date:** [May 12, 2021, 4:37pm UTC](https://discuss.elastic.co/t/logstash-ingesting-csv-filtered-by-csv-columns/272831/1 "2021-05-12T16:37:58Z")

</div>

Is it possible to ingest data via logstash by applying a filter on data in the CSV file being ingested?

For example:

```auto
input {
  file {
path => "blah/blah/*.gz"
max_open_files => 16000
mode => "read"
start_position => "beginning"
sincedb_path => "/dev/null"
exit_after_read => true
   }
}

*# Ingest only rows where [field1]=1000*
filter{
csv {
  separator => "|"
  columns => [
  "time",
  "field 1",
  "field 2"
  ]
}
}
output {
blah...blah...
 }

```

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [May 12, 2021, 5:19pm UTC](https://discuss.elastic.co/t/logstash-ingesting-csv-filtered-by-csv-columns/272831/2 "2021-05-12T17:19:11Z")

</div>

You could add

```
if [field1] != "1000" { drop {} }

```

after the csv filter.

---

<div class="post-metadata">

**Author:** ![hsalim](https://avatars.discourse-cdn.com/v4/letter/h/c67d28/32.png) [@hsalim](https://discuss.elastic.co/u/hsalim)\
**Post date:** [May 12, 2021, 5:55pm UTC](https://discuss.elastic.co/t/logstash-ingesting-csv-filtered-by-csv-columns/272831/3 "2021-05-12T17:55:51Z")

</div>

So a follow up question, because I neglected to explain my first question.

I have multiple CSV files ( **field1 is same within a file** ), assume each has three fields and each file is uniquely identifiable by filed1, and field1 value will repeat for entirety of that particular file, for example:

CSV1 with field1 = 1000:  
field1|field2|field3

CSV2 with field1 = 2000:  
field1|field2|field3

CSV3 with field1 = 3000:  
field1|field2|field3

by adding if [field1] != "1000" { drop {} }, it will will read through **all** entries for **all** CSVs? correct? Reason i ask is because each CSV will have millions of rows, and reading a file that is not relevant can have performance impact. in that case i would need to think of another solution.

Thank you!

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [May 12, 2021, 7:08pm UTC](https://discuss.elastic.co/t/logstash-ingesting-csv-filtered-by-csv-columns/272831/4 "2021-05-12T19:08:41Z")

</div>

> [@hsalim](#):
>
> if [field1] != "1000" { drop {} }

That will drop any event where [field1] is not "1000", so it will drop every event from the second and third CSVs. But if you want to do that then why read them in the first place?

---

<div class="post-metadata">

**Author:** ![hsalim](https://avatars.discourse-cdn.com/v4/letter/h/c67d28/32.png) [@hsalim](https://discuss.elastic.co/u/hsalim)\
**Post date:** [May 12, 2021, 7:30pm UTC](https://discuss.elastic.co/t/logstash-ingesting-csv-filtered-by-csv-columns/272831/5 "2021-05-12T19:30:20Z")

</div>

sure, because all files arrive in a shared directory and we want correct pipeline to ingest the relevant file. for example, we don't want pipeline for field1=1000 ingesting data from CSV with field1=2000. So, at least I am not aware of any way to check before ingesting, the value of field1.

and the way source system is setup we cannot have separate directories.

Your suggestion worked in my test and i did not see any performance issue.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 9, 2021, 7:30pm UTC](https://discuss.elastic.co/t/logstash-ingesting-csv-filtered-by-csv-columns/272831/6 "2021-06-09T19:30:28Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
