# Avoid reloading (duplicate) csv records into same index

**URL:** <https://discuss.elastic.co/t/avoid-reloading-duplicate-csv-records-into-same-index/215734>\
**Category:** Logstash\
**Created:** [January 20, 2020, 12:20pm UTC](https://discuss.elastic.co/t/avoid-reloading-duplicate-csv-records-into-same-index/215734 "2020-01-20T12:20:39Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![bharat1](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@bharat1](https://discuss.elastic.co/u/bharat1)\
**Post date:** [January 20, 2020, 12:20pm UTC](https://discuss.elastic.co/t/avoid-reloading-duplicate-csv-records-into-same-index/215734/1 "2020-01-20T12:20:40Z")

</div>

Hi All,  
I have a working config for loading csv records from logstash to elasticsearch. However when I try to restart the logstash service, the same csv file records are reloaded again into same index and creating duplicate record entries. I want to avoid this happening. Can someone point me whats wrong with this conf or anything missing.

```auto
input {
  file {
    path => "/etc/logstash/1.csv"
    start_position => "beginning"
    sincedb_path => "/etc/logstash/sincedb_sample1csv"
  }
  file {
    path => "/etc/logstash/2.csv"
    start_position => "beginning"
    sincedb_path => "/etc/logstash/sincedb_sample2csv"
  }
}

filter {
  if [path] == "/etc/logstash/1.csv"
  {
    csv {
       separator => ","
       columns => ["column1","column2"]
        }
    }
  if [path] == "/etc/logstash/2.csv"
  {
    csv {
       separator => ","
       columns => ["column1","column2"]
        }
    }
}
output {
  if [path] == "/etc/logstash/1.csv"
    {
     elasticsearch {
       action => "index"
       hosts => ["http://192.168.1.1:9200"]
       index => "sample1"
     }
    }
  if [path] == "/etc/logstash/2.csv"
    {
     elasticsearch {
       action => "index"
       hosts => ["http://192.168.1.1:9200"]
       index => "sample2"
     }
   }

```

---

<div class="post-metadata">

**Author:** ![Fabio-sama](https://avatars.discourse-cdn.com/v4/letter/f/b9e5f3/32.png) [@Fabio-sama](https://discuss.elastic.co/u/Fabio-sama)\
**Post date:** [January 20, 2020, 1:45pm UTC](https://discuss.elastic.co/t/avoid-reloading-duplicate-csv-records-into-same-index/215734/2 "2020-01-20T13:45:22Z")

</div>

Hi there,

here you are not forcing a document\_id, so a random one is assigned to your documents. It means that even if all the fields of your documents are identical, the `_id` is different, hence Elasticsearch will treat is as a new document.

Check out the `fingerprint` filter [https://www.elastic.co/guide/en/logstash/current/plugins-filters-fingerprint.html](https://www.elastic.co/guide/en/logstash/current/plugins-filters-fingerprint.html) 😉

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 17, 2020, 1:45pm UTC](https://discuss.elastic.co/t/avoid-reloading-duplicate-csv-records-into-same-index/215734/3 "2020-02-17T13:45:23Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
