# Logstash ingest and export to elasticsearch files twice

**URL:** <https://discuss.elastic.co/t/logstash-ingest-and-export-to-elasticsearch-files-twice/297247>\
**Category:** Logstash\
**Created:** [February 15, 2022, 1:24pm UTC](https://discuss.elastic.co/t/logstash-ingest-and-export-to-elasticsearch-files-twice/297247 "2022-02-15T13:24:48Z")\
**Posts on this page:** 17\
**Page:** 1

<div class="post-metadata">

**Author:** ![Greninja\_San](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/greninja_san/32/96119_2.png) [@Greninja\_San](https://discuss.elastic.co/u/Greninja_San)\
**Post date:** [February 15, 2022, 1:24pm UTC](https://discuss.elastic.co/t/logstash-ingest-and-export-to-elasticsearch-files-twice/297247/1 "2022-02-15T13:24:48Z")

</div>

Hello!

I have a logstash config that gather lines from CSVs and then send them to Elasticsearch. However, for some unknown reason, lines are duplicated in Elasticsearch, taking twice the storage space, and making statistics wrong.

Any ideas of the cause?

Thanks

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [February 15, 2022, 1:36pm UTC](https://discuss.elastic.co/t/logstash-ingest-and-export-to-elasticsearch-files-twice/297247/2 "2022-02-15T13:36:55Z")

</div>

You need to share your configuration, use the preformatted button in the forum `</>`, and paste your configuration.

---

<div class="post-metadata">

**Author:** ![Greninja\_San](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/greninja_san/32/96119_2.png) [@Greninja\_San](https://discuss.elastic.co/u/Greninja_San)\
**Post date:** [February 15, 2022, 1:51pm UTC](https://discuss.elastic.co/t/logstash-ingest-and-export-to-elasticsearch-files-twice/297247/3 "2022-02-15T13:51:55Z")

</div>

Woops sorry

```auto
input {
  file {
    path => "/srv/edm/ftp/bcp/*.csv"
    start_position => "beginning"
    ignore_older => 17280000
    sincedb_path => "/dev/null"
    codec => plain {
      charset => "ANSI_X3.4-1968"
    }
  }
}

filter {
     csv {
        separator => ";"
        columns => ["sender", "receiver", "flow_type", "end_traitement", "start_traitement", "size_end_f", "size_start_f", "format_end", "format_start", "platform_end", "platform_start", "transport_end", "transport_start", "prod"]
        skip_header => true
     }
    ruby {
        code => 'event.set("date_start_formated", event.get("start_traitement").ljust(21, "0"))'
    }
    ruby {
        code => 'event.set("date_end_formated", event.get("end_traitement").ljust(21, "0"))'
    }
    ruby {
        code => 'event.set("date_start_formated_cut", ((event.get("date_start_formated"))[0,21]))'
    }
    ruby {
        code => 'event.set("date_end_formated_cut", ((event.get("date_end_formated"))[0,21]))'
    }
   ruby {
        code => 'event.set("date_start", (event.get("date_start_formated_cut").gsub(",",".")))'
    }
    ruby {
        code => 'event.set("date_end", (event.get("date_end_formated_cut").gsub(",",".")))'
    }
    ruby {
    code => 'event.set("size_start", event.get("size_start_f").to_i)'
    }

    ruby {
    code => 'event.set("size_end", event.get("size_end_f").to_i)'
    }
    date {
    match => ["date_start", "dd/MM/yy HH:mm:ss.SSS"]
    target => "start_traitement_true"
    }
    date {
    match => ["date_end", "dd/MM/yy HH:mm:ss.SSS"]
    target => "end_traitement_true"
    }
    ruby {
    code => 'event.set("date_epoch_end", event.get("end_traitement_true").to_i)'
    }
    ruby {
    code => 'event.set("date_epoch_start", event.get("start_traitement_true").to_i)'
    }
    ruby {
    code => 'event.set("time_between", ((event.get("date_epoch_end"))-(event.get("date_epoch_start"))).abs)'
    }
    ruby {
    code => 'event.set("fixedProd", (event.get("prod")).tr("\r", ""))'
    }

   ruby {
    code => 'event.set("cat", "BCP")'
    }

    mutate {
      convert => {
         "sender" => "string"
         "receiver" => "string"
         "flow_type" => "string"
         "size_start" => "integer"
         "size_end" => "integer"
         "format_start" => "string"
         "format_end" => "string"
         "platform_start" => "string"
         "platform_end" => "string"
         "date_epoch_end" => "float"
         "date_epoch_start" => "float"
         "transport_start" => "string"
         "transport_end" => "string"
         "prod" => "string"
         }
    }

    prune{
      blacklist_names => ["message","date_end_formated_cut","date_start_formated_cut","date_start_formated","date_end_formated","date_start","date_end"]
    }

}

output {
  stdout{}
  elasticsearch {
    hosts => "http://XXXXXXX:9200"
    index => "index_bcp"
    user => "XXXXXXX"
    password => "XXXXXXX"
    ssl_certificate_verification => false
  }
}

```

---

<div class="post-metadata">

**Author:** ![Tomo\_M](https://avatars.discourse-cdn.com/v4/letter/t/848f3c/32.png) [@Tomo\_M](https://discuss.elastic.co/u/Tomo_M)\
**Post date:** [February 15, 2022, 5:54pm UTC](https://discuss.elastic.co/t/logstash-ingest-and-export-to-elasticsearch-files-twice/297247/4 "2022-02-15T17:54:53Z")

</div>

Have you run the script just once?

---

<div class="post-metadata">

**Author:** ![Greninja\_San](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/greninja_san/32/96119_2.png) [@Greninja\_San](https://discuss.elastic.co/u/Greninja_San)\
**Post date:** [February 15, 2022, 7:14pm UTC](https://discuss.elastic.co/t/logstash-ingest-and-export-to-elasticsearch-files-twice/297247/5 "2022-02-15T19:14:25Z")

</div>

It is launched using systemctl, and just once, yes.

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [February 15, 2022, 8:05pm UTC](https://discuss.elastic.co/t/logstash-ingest-and-export-to-elasticsearch-files-twice/297247/6 "2022-02-15T20:05:00Z")

</div>

Any reason to use `sincedb_path` as `/dev/null` ? This makes Logstash reads the file again if your service is restarted for some reason.

There is nothing in your configuration that could duplicate the documents unless your logstash is being restarted.

Do you have automatic reload enabled? I think that if you have automatic reload enabled and you change the pipeline configuration file, this will trigger a pipeline reload that could reread the files.

What does your `pipelines.yml` looks like? Do you have multiple pipelines? Are pointing to a directory with multiple configuration files?

---

<div class="post-metadata">

**Author:** ![Greninja\_San](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/greninja_san/32/96119_2.png) [@Greninja\_San](https://discuss.elastic.co/u/Greninja_San)\
**Post date:** [February 16, 2022, 9:24am UTC](https://discuss.elastic.co/t/logstash-ingest-and-export-to-elasticsearch-files-twice/297247/7 "2022-02-16T09:24:00Z")

</div>

If not /dev/null, I could set any folder?

---

<div class="post-metadata">

**Author:** ![Greninja\_San](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/greninja_san/32/96119_2.png) [@Greninja\_San](https://discuss.elastic.co/u/Greninja_San)\
**Post date:** [February 16, 2022, 9:38am UTC](https://discuss.elastic.co/t/logstash-ingest-and-export-to-elasticsearch-files-twice/297247/8 "2022-02-16T09:38:49Z")

</div>

Also I don't think it is due to a restart, since on duplicate documents, @timestamp is the exact same.

---

<div class="post-metadata">

**Author:** ![Greninja\_San](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/greninja_san/32/96119_2.png) [@Greninja\_San](https://discuss.elastic.co/u/Greninja_San)\
**Post date:** [February 16, 2022, 12:59pm UTC](https://discuss.elastic.co/t/logstash-ingest-and-export-to-elasticsearch-files-twice/297247/9 "2022-02-16T12:59:43Z")

</div>

I just verified, I have just one pipeline, leading to a folder with 2 different config files. And no automatic reload as well.

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [February 16, 2022, 1:00pm UTC](https://discuss.elastic.co/t/logstash-ingest-and-export-to-elasticsearch-files-twice/297247/10 "2022-02-16T13:00:16Z")

</div>

The `sincedb_path` needs to point to a [file](https://www.elastic.co/guide/en/logstash/current/plugins-inputs-file.html#plugins-inputs-file-sincedb_path) where logstash will keep track of the position already read from the file, you can remove this option of your config and let logstash set the sincedb path itself.

Using `/dev/null` means that logstash does not keep track of the position in the file already read and it can read the file again on a restart.

If the `@timestamp` is the same, then you are right, it may not be related to a restart as this would change the `@timestamp` when logstash reread the file.

You didn't share your `pipelines.yml`, are you using multiple pipelines or just one `main` pipeline pointing to `/etc/logstash/conf.d/*.conf`? Share your `pipelines.yml`.

The only other way that I can think of that could create duplicate documents is if you are pointing to a folder with multiple configurations and have another Elasticsearch output indexing to the same index name.

---

<div class="post-metadata">

**Author:** ![Greninja\_San](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/greninja_san/32/96119_2.png) [@Greninja\_San](https://discuss.elastic.co/u/Greninja_San)\
**Post date:** [February 16, 2022, 1:01pm UTC](https://discuss.elastic.co/t/logstash-ingest-and-export-to-elasticsearch-files-twice/297247/11 "2022-02-16T13:01:38Z")

</div>

![image](https://us1.discourse-cdn.com/elastic/original/3X/4/6/46cd0829f7943bfc7990b48daeef573e3bd84b3d.png)

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [February 16, 2022, 1:04pm UTC](https://discuss.elastic.co/t/logstash-ingest-and-export-to-elasticsearch-files-twice/297247/12 "2022-02-16T13:04:33Z")

</div>

Ok, and do you have any other `*.conf` files in the `/etc/logstash/conf.d` directory?

If there is other files, please share them as well.

---

<div class="post-metadata">

**Author:** ![Greninja\_San](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/greninja_san/32/96119_2.png) [@Greninja\_San](https://discuss.elastic.co/u/Greninja_San)\
**Post date:** [February 16, 2022, 1:08pm UTC](https://discuss.elastic.co/t/logstash-ingest-and-export-to-elasticsearch-files-twice/297247/13 "2022-02-16T13:08:37Z")

</div>

These 2 configs :

```auto
input {
  file {
    path => "/srv/edm/ftp/not_reversed/*.csv"
    start_position => "beginning"
    ignore_older => 17280000
    sincedb_path => "/dev/null"
    codec => plain {
      charset => "ANSI_X3.4-1968"
    }
  }
}

filter {
     csv {
        separator => ";"
        columns => ["sender", "receiver", "flow_type", "start_traitement", "end_traitement", "size_start_f", "size_end_f", "format_start", "format_end", "platform_start", "platform_end", "transport_start", "transport_end", "prod"]
        skip_header => true
     }
    ruby {
        code => 'event.set("date_start_formated", event.get("start_traitement").ljust(21, "0"))'
    }
    ruby {
        code => 'event.set("date_end_formated", event.get("end_traitement").ljust(21, "0"))'
    }
    ruby {
        code => 'event.set("date_start_formated_cut", ((event.get("date_start_formated"))[0,21]))'
    }
    ruby {
        code => 'event.set("date_end_formated_cut", ((event.get("date_end_formated"))[0,21]))'
    }
   ruby {
        code => 'event.set("date_start", (event.get("date_start_formated_cut").gsub(",",".")))'
    }
    ruby {
        code => 'event.set("date_end", (event.get("date_end_formated_cut").gsub(",",".")))'
    }
    ruby {
    code => 'event.set("size_start", event.get("size_start_f").to_i)'
    }

    ruby {
    code => 'event.set("size_end", event.get("size_end_f").to_i)'
    }

    date {
    match => ["date_start", "dd/MM/yy HH:mm:ss.SSS"]
    target => "start_traitement_true"
    }
    date {
    match => ["date_end", "dd/MM/yy HH:mm:ss.SSS"]
    target => "end_traitement_true"
    }
    ruby {
    code => 'event.set("date_epoch_end", event.get("end_traitement_true").to_i)'
    }
    ruby {
    code => 'event.set("date_epoch_start", event.get("start_traitement_true").to_i)'
    }
    ruby {
    code => 'event.set("time_between", ((event.get("date_epoch_end"))-(event.get("date_epoch_start"))).abs)'
    }
    ruby {
    code => 'event.set("fixedProd", (event.get("prod")).tr("\r", ""))'
    }

   ruby {
    code => 'event.set("cat", "BCP")'
    }

    mutate {
      convert => {
         "sender" => "string"
         "receiver" => "string"
         "flow_type" => "string"
         "size_start" => "integer"
         "size_end" => "integer"
         "format_start" => "string"
         "format_end" => "string"
         "platform_start" => "string"
         "platform_end" => "string"
         "date_epoch_end" => "float"
         "date_epoch_start" => "float"
         "transport_start" => "string"
         "transport_end" => "string"
         "prod" => "string"
         }
    }

    prune{
      blacklist_names => ["message","date_end_formated_cut","date_start_formated_cut","date_start_formated","date_end_formated","date_start","date_end"]
    }

}

output {
  stdout{}
  elasticsearch {
    hosts => "http://XXXXXXX:9200"
    index => "index_bcp"
    user => "XXXXX"
    password => "XXXXXXX"
    ssl_certificate_verification => false
  }
}

```

AND

```auto
input {
  file {
    path => "/srv/edm/ftp/bcp/*.csv"
    start_position => "beginning"
    ignore_older => 17280000
    sincedb_path => "/dev/null"
    codec => plain {
      charset => "ANSI_X3.4-1968"
    }
  }
}

filter {
     csv {
        separator => ";"
        columns => ["sender", "receiver", "flow_type", "end_traitement", "start_traitement", "size_end_f", "size_start_f", "format_end", "format_start", "platform_end", "platform_start", "transport_end", "transport_start", "prod"]
        skip_header => true
     }
    ruby {
        code => 'event.set("date_start_formated", event.get("start_traitement").ljust(21, "0"))'
    }
    ruby {
        code => 'event.set("date_end_formated", event.get("end_traitement").ljust(21, "0"))'
    }
    ruby {
        code => 'event.set("date_start_formated_cut", ((event.get("date_start_formated"))[0,21]))'
    }
    ruby {
        code => 'event.set("date_end_formated_cut", ((event.get("date_end_formated"))[0,21]))'
    }
   ruby {
        code => 'event.set("date_start", (event.get("date_start_formated_cut").gsub(",",".")))'
    }
    ruby {
        code => 'event.set("date_end", (event.get("date_end_formated_cut").gsub(",",".")))'
    }
    ruby {
    code => 'event.set("size_start", event.get("size_start_f").to_i)'
    }

    ruby {
    code => 'event.set("size_end", event.get("size_end_f").to_i)'
    }
    date {
    match => ["date_start", "dd/MM/yy HH:mm:ss.SSS"]
    target => "start_traitement_true"
    }
    date {
    match => ["date_end", "dd/MM/yy HH:mm:ss.SSS"]
    target => "end_traitement_true"
    }
    ruby {
    code => 'event.set("date_epoch_end", event.get("end_traitement_true").to_i)'
    }
    ruby {
    code => 'event.set("date_epoch_start", event.get("start_traitement_true").to_i)'
    }
    ruby {
    code => 'event.set("time_between", ((event.get("date_epoch_end"))-(event.get("date_epoch_start"))).abs)'
    }
    ruby {
    code => 'event.set("fixedProd", (event.get("prod")).tr("\r", ""))'
    }

   ruby {
    code => 'event.set("cat", "BCP")'
    }

    mutate {
      convert => {
         "sender" => "string"
         "receiver" => "string"
         "flow_type" => "string"
         "size_start" => "integer"
         "size_end" => "integer"
         "format_start" => "string"
         "format_end" => "string"
         "platform_start" => "string"
         "platform_end" => "string"
         "date_epoch_end" => "float"
         "date_epoch_start" => "float"
         "transport_start" => "string"
         "transport_end" => "string"
         "prod" => "string"
         }
    }

    prune{
      blacklist_names => ["message","date_end_formated_cut","date_start_formated_cut","date_start_formated","date_end_formated","date_start","date_end"]
    }

}

output {
  stdout{}
  elasticsearch {
    hosts => "http://XXXXXX:9200"
    index => "index_bcp"
    user => "XXXXXXX"
    password => "XXXXXXX"
    ssl_certificate_verification => false
  }
}

```

Both hosts, user and password are the same.

---

<div class="post-metadata">

**Author:** ![Greninja\_San](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/greninja_san/32/96119_2.png) [@Greninja\_San](https://discuss.elastic.co/u/Greninja_San)\
**Post date:** [February 16, 2022, 1:09pm UTC](https://discuss.elastic.co/t/logstash-ingest-and-export-to-elasticsearch-files-twice/297247/14 "2022-02-16T13:09:41Z")

</div>

The 2nd one is the one I already gave.

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [February 16, 2022, 1:25pm UTC](https://discuss.elastic.co/t/logstash-ingest-and-export-to-elasticsearch-files-twice/297247/15 "2022-02-16T13:25:22Z")

</div>

That's the issue.

Your `pipeline.yml` has only one pipeline pointing to `/etc/logstash/conf.d/*.conf`, so when Logstash starts all the files in this directory will be merged and the events will pass through every filter and output.

In reality what you have is something like this:

```auto
input {
  file {
    path => "/srv/edm/ftp/not_reversed/*.csv"
    start_position => "beginning"
    ignore_older => 17280000
    sincedb_path => "/dev/null"
    codec => plain {
      charset => "ANSI_X3.4-1968"
    }
  }
  file {
    path => "/srv/edm/ftp/bcp/*.csv"
    start_position => "beginning"
    ignore_older => 17280000
    sincedb_path => "/dev/null"
    codec => plain {
      charset => "ANSI_X3.4-1968"
    }
  }
}
filter {
    all your filters
}
output {
  stdout{}
  elasticsearch {
    hosts => "http://XXXXXXX:9200"
    index => "index_bcp"
    user => "XXXXX"
    password => "XXXXXXX"
    ssl_certificate_verification => false
  }
  stdout{}
  elasticsearch {
    hosts => "http://XXXXXX:9200"
    index => "index_bcp"
    user => "XXXXXXX"
    password => "XXXXXXX"
    ssl_certificate_verification => false
  }
}

```

Since you use the same index, you have your output twice, this is what get your documents duplicated.

You should isolate your pipelines changing your `pipelines.yml` to something similar as this:

```auto
- pipeline.id: "pipeline-one"
  path.config: "/etc/logstash/conf.d/file-one.conf"

- pipeline.id: "pipeline-two"
  path.config: "/etc/logstash/conf.d/file-two.conf

```

This will isolate your pipelines and you will have just one output per pipeline.

---

<div class="post-metadata">

**Author:** ![Greninja\_San](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/greninja_san/32/96119_2.png) [@Greninja\_San](https://discuss.elastic.co/u/Greninja_San)\
**Post date:** [February 16, 2022, 2:06pm UTC](https://discuss.elastic.co/t/logstash-ingest-and-export-to-elasticsearch-files-twice/297247/16 "2022-02-16T14:06:26Z")

</div>

That was the solution! Thank you very much \<3

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 16, 2022, 2:06pm UTC](https://discuss.elastic.co/t/logstash-ingest-and-export-to-elasticsearch-files-twice/297247/17 "2022-03-16T14:06:57Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
