# CSV Output Plugin - Duplicate Entries

**URL:** <https://discuss.elastic.co/t/csv-output-plugin-duplicate-entries/233368>\
**Category:** Logstash\
**Created:** [May 19, 2020, 3:55pm UTC](https://discuss.elastic.co/t/csv-output-plugin-duplicate-entries/233368 "2020-05-19T15:55:51Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![muraliv](https://avatars.discourse-cdn.com/v4/letter/m/f05b48/32.png) [@muraliv](https://discuss.elastic.co/u/muraliv)\
**Post date:** [May 19, 2020, 3:55pm UTC](https://discuss.elastic.co/t/csv-output-plugin-duplicate-entries/233368/1 "2020-05-19T15:55:51Z")

</div>

Hi All,

I am using logstash to write to a csv file and elasticsearch. Elasticsearch output is fine but csv output plugin generates entries each time it runs. I have set the following based on the csv filter:

- skip\_headers: true
- defined column names
- defined pipeline workers as 1
- defined java\_execution as false

> documentation: skip\_header: If skip\_header is set without autodetect\_column\_names being set then columns should be set which will result in the skipping of any row that exactly matches the specified column values. Logstash pipeline workers must be set to 1 for this option to work.

Here is the configuration file

```auto
input
{
  http_poller
  {
    urls =>
    {
      mispevents =>
      {
        method => post
        url => "https://hostname.domain.com/events/csv/download"
        headers =>
        {
          Authorization => "${MISP_TOKEN}"
          "Content-Type" => "application/json"
        }
        body => '{"ignore": "True", "tags": ["con=high"], "type": ["type1"], "last": "400d"}'
      }
    }
  cacert => "/u/elasticStack/cacerts/cacerts.pem"
  schedule => { every => "1m" }
  codec => "line"
  }
}

filter
{

  csv
  {
    skip_header => "true"
    columns => ["uuid","event_id","category","type","value","comment","to_ids","date","object_uuid","object_name","object_meta_category"]
    add_field => { "priority" => "6"}
  }

  if [type] == "type1" {
     mutate {
        copy => { "value" => "type1" }
     }
     mutate {
        add_field => { "misp_key" => "%{value}" "misp_value" => " %{category},%{comment},%{priority}" }
     }
  }

  mutate {
     copy => { "category" => "name" "event_id" => "eventid" "comment" => "description" }
  }

  mutate {
    remove_field => ["message", "object_meta_category", "object_name", "@version", "@timestamp", "object_uuid", "date", "to_ids", "event_id"]
  }
}

output
{

  stdout { codec => rubydebug }

  csv {
     path => "/u/elasticStack/data/misp.csv"
     csv_options => { "col_sep" => ":" }
     fields => ["misp_key","misp_value"]
  }

}

```

Please let me know if there is anything I am missing

Thanks  
Murali

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [May 19, 2020, 8:03pm UTC](https://discuss.elastic.co/t/csv-output-plugin-duplicate-entries/233368/2 "2020-05-19T20:03:10Z")

</div>

> [@muraliv](#):
>
> csv output plugin generates entries each time it runs

How does that differ to the behaviour that you want?

---

<div class="post-metadata">

**Author:** ![muraliv](https://avatars.discourse-cdn.com/v4/letter/m/f05b48/32.png) [@muraliv](https://discuss.elastic.co/u/muraliv)\
**Post date:** [May 19, 2020, 8:26pm UTC](https://discuss.elastic.co/t/csv-output-plugin-duplicate-entries/233368/3 "2020-05-19T20:26:43Z")

</div>

Hi Badger,

I would prefer the csv file to contain unique entries each time it runs.

Thanks  
Murali

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [May 19, 2020, 10:53pm UTC](https://discuss.elastic.co/t/csv-output-plugin-duplicate-entries/233368/4 "2020-05-19T22:53:58Z")

</div>

> [@muraliv](#):
>
> I would prefer the csv file to contain unique entries each time it runs.

That's not how it works. Whatever the URL you are polling every minute is fed to logstash as events. If you want to check whether you have seen the events before you would need a database containing them.

---

<div class="post-metadata">

**Author:** ![muraliv](https://avatars.discourse-cdn.com/v4/letter/m/f05b48/32.png) [@muraliv](https://discuss.elastic.co/u/muraliv)\
**Post date:** [May 19, 2020, 11:57pm UTC](https://discuss.elastic.co/t/csv-output-plugin-duplicate-entries/233368/5 "2020-05-19T23:57:14Z")

</div>

> [@muraliv](#):
>
> If skip\_header is set without autodetect\_column\_names being set then columns should be set which will result in the skipping of any row that exactly matches the specified column values. Logstash pipeline workers must be set to 1 for this option to work.

Hi Badger,

Thank you

Murali

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 16, 2020, 11:57pm UTC](https://discuss.elastic.co/t/csv-output-plugin-duplicate-entries/233368/6 "2020-06-16T23:57:18Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
