# Updating data in CSV logstash is pushing the entire CSV file again with updated data which duplicates my records in index. But I just want to sync the data

**URL:** <https://discuss.elastic.co/t/updating-data-in-csv-logstash-is-pushing-the-entire-csv-file-again-with-updated-data-which-duplicates-my-records-in-index-but-i-just-want-to-sync-the-data/175354>\
**Category:** Logstash\
**Created:** [April 4, 2019, 8:36am UTC](https://discuss.elastic.co/t/updating-data-in-csv-logstash-is-pushing-the-entire-csv-file-again-with-updated-data-which-duplicates-my-records-in-index-but-i-just-want-to-sync-the-data/175354 "2019-04-04T08:36:00Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![Sumit\_Kumar1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sumit_kumar1/32/43471_2.png) [@Sumit\_Kumar1](https://discuss.elastic.co/u/Sumit_Kumar1)\
**Post date:** [April 4, 2019, 8:36am UTC](https://discuss.elastic.co/t/updating-data-in-csv-logstash-is-pushing-the-entire-csv-file-again-with-updated-data-which-duplicates-my-records-in-index-but-i-just-want-to-sync-the-data/175354/1 "2019-04-04T08:36:01Z")

</div>

```
input{

file {
path => "/home/elastic/elk/logstash-6.4.3/csv_data/attrition_dump_v3.csv"
start_position => "beginning"
sincedb_path => "/dev/null"
}

file {
path => "/home/elastic/elk/logstash-6.4.3/csv_data/workday_dump_v3.csv"
start_position => "beginning"
sincedb_path => "/dev/null"
}

```

}

filter{

```
if [path] == "/home/elastic/elk/logstash-6.4.3/csv_data/attrition_dump_v3.csv"
{
    csv {
    
    columns => ["Emp id","Full Name","Original Hire date","Hire Date","Last Day of Work","Quarter","Tenure","Tenure in years","Termination Reason","Reason","Vol/Invol",
    "Position Title","Grade","Cost Center - Name","Mgr Name","Goal","Competency","Leader","Kumar -1","BU","Team"] 
    separator => ","
    }

    mutate {
    add_field => { "doc_type" => "attrition" }
    add_field => {"id" => ""}
    copy => {"Emp id" => "id" }
    remove_field => ["message"]
    }
}

else if [path] == "/home/elastic/elk/logstash-6.4.3/csv_data/workday_dump_v3.csv"
{
    csv {
    columns => ["Employee ID","Employee","Last,First Name","Email - Primary Work","Hire Date","Original Hire Date","Is Rehire","Years of Service","Company Service Date",
    "Continuous Service Date","Seniority Date","Time in Job Profile","Time in Job Profile Start Date","Time in Position","Position","Job Title","Job Profile",
    "Grade Profile ID","Grade","Grade Effective Date","Employment Status","Leave Type","Employee Type","Worker Type","Full/Part","Reg/Temp","Worker SubType","Exempt/Non-Exempt","Pay Rate Type","Scheduled Std Hours - Calculated FTE","Location Std Hours","Default Weekly Hours","FTE","Cost Center - ID","Cost Center - Name","HFM-Code","HFM-Function","HFM-SubFunction","Profit Center","Product Code","IES/Novella","Project ID","Project Description","Tech/Non-Tech","Client Facing Y/N","Worker's Business Unit","HR BU","SBU","Finance BU","Department","FM Entity","Custom 1","Custom 2","HR Category","Company","Location Code","Location","City","State","Country Name","Mature/Emerging","Geo Region","No. of Directs","Manager ID","Manager Name","Tier 1","Tier 2","Tier 3","Tier 4","Tier 5","Tier 6","Tier 7","Last Base Pay Increase - Date","Last Base Pay Increase Reason","Total Pay - Amount","Total Base Pay - Amount","Total Base Pay (Base or Basis) Local","Hourly Rate - Amount","Total Base Pay - Frequency","Total Base Pay - Currency","Total Base Pay in USD","Total Base Pay (Base or Basis) USD","Pay Range - Minimum","Pay Range - Midpoint","Pay Range - Maximum","Compa Ratio (Base or Basis)","Compa Ratio Bucket","VC Plan ID","VC Plan Name","Target Bonus - Percent","Target Bonus - Amount","Target Bonus - Currency","Target Bonus Amount in USD","CS Summary Role","Billable Stat","Job Family","Job Family Group","Competency Rating (2017/18)","Goal Rating (2017)","Competency Rating (2016/17)","Goal Rating (2016)","Competency Rating (2015/16)","Goal Rating (2015)","Legacy Organization"]
    separator => ","
    }

    mutate {
    add_field => { "doc_type" => "workday" }
    add_field => {"id" => ""}
    copy => {"Employee ID" => "id" }
    remove_field => ["message"]
    }

}
uuid {
target => "uuid"

```

}  
}

output {  
elasticsearch  
{  
index =\> "xg\_hr\_details-000001"  
action =\> "update"  
document\_id =\> "%{[Employee ID]}"  
doc\_as\_upsert =\> "true"  
hosts =\> "[http://caruelsatic01p:9200/](http://caruelsatic01p:9200/)"  
}  
}

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [April 4, 2019, 12:15pm UTC](https://discuss.elastic.co/t/updating-data-in-csv-logstash-is-pushing-the-entire-csv-file-again-with-updated-data-which-duplicates-my-records-in-index-but-i-just-want-to-sync-the-data/175354/2 "2019-04-04T12:15:17Z")

</div>

Two things. Firstly, you are referencing "Employee ID" as the document\_id, but that does not exist for your attrition records.

Secondly...

```
mutate {
    [...]
    add_field => {"id" => ""}
    copy => {"Emp id" => "id" }
    [...]
}

```

A mutate filter performs operations in a fixed order, and add\_field comes after copy. So this will copy "Emp id" to id, then it will add the string "" to id, resulting in an array. Just remove the add\_field.

Then you probably want to reference id in the document\_id option on the output.

---

<div class="post-metadata">

**Author:** ![Sumit\_Kumar1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sumit_kumar1/32/43471_2.png) [@Sumit\_Kumar1](https://discuss.elastic.co/u/Sumit_Kumar1)\
**Post date:** [April 4, 2019, 12:28pm UTC](https://discuss.elastic.co/t/updating-data-in-csv-logstash-is-pushing-the-entire-csv-file-again-with-updated-data-which-duplicates-my-records-in-index-but-i-just-want-to-sync-the-data/175354/3 "2019-04-04T12:28:54Z")

</div>

Hi Badger ,

Thanks , for the suggestions i have implemented the changes suggested by you please , find my latest logstash config . As while updating i am getting duplicate data . Whenever i make changes in my CSV the data is getting pushed thrice...  
input{

```
file {
path => "/home/elastic/elk/logstash-6.4.3/csv_data/attrition_dump_v3.csv"
start_position => "beginning"
sincedb_path => "/dev/null"
}

file {
path => "/home/elastic/elk/logstash-6.4.3/csv_data/workday_dump_v3.csv"
start_position => "beginning"
sincedb_path => "/dev/null"
}

```

}

filter{

```
if [path] == "/home/elastic/elk/logstash-6.4.3/csv_data/attrition_dump_v3.csv"
{
    csv {
    
    columns => ["Emp id","Full Name","Original Hire date","Hire Date","Last Day of Work","Quarter","Tenure","Tenure in years","Termination Reason","Reason","Vol/Invol",
    "Position Title","Grade","Cost Center - Name","Mgr Name","Goal","Competency","Leader","Kumar -1","BU","Team"] 
    separator => ","
    }

    mutate {
    add_field => { "doc_type" => "attrition" }
    copy => {"Emp id" => "id"}
    add_field => {"id" => ""}
    split => ["id", ","]
    add_field => { "ID" => "%{id[0]}"}
    remove_field => ["message"]
    }
}

else if [path] == "/home/elastic/elk/logstash-6.4.3/csv_data/workday_dump_v3.csv"
{
    csv {
    columns => ["Employee ID","Employee","Last,First Name","Email - Primary Work","Hire Date","Original Hire Date","Is Rehire","Years of Service","Company Service Date",
    "Continuous Service Date","Seniority Date","Time in Job Profile","Time in Job Profile Start Date","Time in Position","Position","Job Title","Job Profile",
    "Grade Profile ID","Grade","Grade Effective Date","Employment Status","Leave Type","Employee Type","Worker Type","Full/Part","Reg/Temp","Worker SubType","Exempt/Non-Exempt","Pay Rate Type","Scheduled Std Hours - Calculated FTE","Location Std Hours","Default Weekly Hours","FTE","Cost Center - ID","Cost Center - Name","HFM-Code","HFM-Function","HFM-SubFunction","Profit Center","Product Code","IES/Novella","Project ID","Project Description","Tech/Non-Tech","Client Facing Y/N","Worker's Business Unit","HR BU","SBU","Finance BU","Department","FM Entity","Custom 1","Custom 2","HR Category","Company","Location Code","Location","City","State","Country Name","Mature/Emerging","Geo Region","No. of Directs","Manager ID","Manager Name","Tier 1","Tier 2","Tier 3","Tier 4","Tier 5","Tier 6","Tier 7","Last Base Pay Increase - Date","Last Base Pay Increase Reason","Total Pay - Amount","Total Base Pay - Amount","Total Base Pay (Base or Basis) Local","Hourly Rate - Amount","Total Base Pay - Frequency","Total Base Pay - Currency","Total Base Pay in USD","Total Base Pay (Base or Basis) USD","Pay Range - Minimum","Pay Range - Midpoint","Pay Range - Maximum","Compa Ratio (Base or Basis)","Compa Ratio Bucket","VC Plan ID","VC Plan Name","Target Bonus - Percent","Target Bonus - Amount","Target Bonus - Currency","Target Bonus Amount in USD","CS Summary Role","Billable Stat","Job Family","Job Family Group","Competency Rating (2017/18)","Goal Rating (2017)","Competency Rating (2016/17)","Goal Rating (2016)","Competency Rating (2015/16)","Goal Rating (2015)","Legacy Organization"]
    separator => ","
    }

    mutate {
    add_field => { "doc_type" => "workday" }
    copy => {"Employee ID" => "id"}
    add_field => {"id" => ""}
    split => ["id", ","]
    add_field => { "[ID]" => "%{id[0]}"}
    remove_field => ["message"]
    }

}
uuid {
target => "uuid"

```

}  
}

output {

if [document\_id] {  
elasticsearch  
{  
index =\> "xg\_hr\_details-000001"  
action =\> "update"  
document\_id =\> "%{[ID]}"  
doc\_as\_upsert =\> "true"  
hosts =\> "[http://caruelsatic01p:9200/](http://caruelsatic01p:9200/)"  
}  
}  
else {  
elasticsearch  
{  
index =\> "xg\_hr\_details-000001"  
hosts =\> "[http://caruelsatic01p:9200/](http://caruelsatic01p:9200/)"  
}  
}  
}

Please , find the kibana discover image to see duplicate logs. Circled RED document is original document , i have made change in Name . Then it pushed thrice....  
Please, provide suggestion for avoiding duplicates.

 ![duplicate%20error](https://us1.discourse-cdn.com/elastic/original/3X/4/e/4efe4642aef4d87385fc28011cf762fc729eb5e0.png)

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [April 4, 2019, 12:41pm UTC](https://discuss.elastic.co/t/updating-data-in-csv-logstash-is-pushing-the-entire-csv-file-again-with-updated-data-which-duplicates-my-records-in-index-but-i-just-want-to-sync-the-data/175354/4 "2019-04-04T12:41:33Z")

</div>

In your output section you are testing for the existence of the [document\_id] field, which does not exist, so it will always go through the else section, which unconditionally creates a new document.

---

<div class="post-metadata">

**Author:** ![Sumit\_Kumar1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sumit_kumar1/32/43471_2.png) [@Sumit\_Kumar1](https://discuss.elastic.co/u/Sumit_Kumar1)\
**Post date:** [April 4, 2019, 12:43pm UTC](https://discuss.elastic.co/t/updating-data-in-csv-logstash-is-pushing-the-entire-csv-file-again-with-updated-data-which-duplicates-my-records-in-index-but-i-just-want-to-sync-the-data/175354/5 "2019-04-04T12:43:05Z")

</div>

Can you suggest my output filter for that ..... ?

---

<div class="post-metadata">

**Author:** ![Sumit\_Kumar1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sumit_kumar1/32/43471_2.png) [@Sumit\_Kumar1](https://discuss.elastic.co/u/Sumit_Kumar1)\
**Post date:** [April 4, 2019, 12:46pm UTC](https://discuss.elastic.co/t/updating-data-in-csv-logstash-is-pushing-the-entire-csv-file-again-with-updated-data-which-duplicates-my-records-in-index-but-i-just-want-to-sync-the-data/175354/6 "2019-04-04T12:46:19Z")

</div>

I have used ID but, still duplicates are getting inserted

output {

if [ID] {  
elasticsearch  
{  
index =\> "xg\_hr\_details-000001"  
action =\> "update"  
document\_id =\> "%{[ID]}"  
doc\_as\_upsert =\> "true"  
hosts =\> "[http://caruelsatic01p:9200/](http://caruelsatic01p:9200/)"  
}  
}  
else {  
elasticsearch  
{  
index =\> "xg\_hr\_details-000001"  
hosts =\> "[http://caruelsatic01p:9200/](http://caruelsatic01p:9200/)"  
}  
}  
}

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [April 4, 2019, 12:52pm UTC](https://discuss.elastic.co/t/updating-data-in-csv-logstash-is-pushing-the-entire-csv-file-again-with-updated-data-which-duplicates-my-records-in-index-but-i-just-want-to-sync-the-data/175354/7 "2019-04-04T12:52:37Z")

</div>

If you start over with a fresh index and run logstash twice, what does a single document in the index look like. Copy it from the JSON tab in Kibana.

---

<div class="post-metadata">

**Author:** ![Sumit\_Kumar1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sumit_kumar1/32/43471_2.png) [@Sumit\_Kumar1](https://discuss.elastic.co/u/Sumit_Kumar1)\
**Post date:** [April 4, 2019, 1:04pm UTC](https://discuss.elastic.co/t/updating-data-in-csv-logstash-is-pushing-the-entire-csv-file-again-with-updated-data-which-duplicates-my-records-in-index-but-i-just-want-to-sync-the-data/175354/8 "2019-04-04T13:04:34Z")

</div>

I have the index setting - refresh\_interval to 1 sec. Logstash i am running it as a service. Whenever i am saving my CSV then it's getting pushed again.

I am having duplicate logs . I changed the name to Sumit test

{  
"\_index": "xg\_hr\_details-000001",  
"\_type": "doc",  
"\_id": "52202",  
"\_version": 8,  
"\_score": null,  
"\_source": {  
"Last Day of Work": null,  
"path": "/home/elastic/elk/logstash-6.4.3/csv\_data/attrition\_dump\_v3.csv",  
"ID": "52202",  
"host": "carulogbdf01p",  
"Competency": null,  
"uuid": "8968604e-8b1b-4582-8a83-ca3654b34666",  
"Mgr Name": null,  
"id": [  
"52202",  
""  
],  
"Kumar -1": null,  
"@version": "1",  
"@timestamp": "2019-04-04T12:52:30.944Z",  
"Vol/Invol": null,  
"Original Hire date": "26-Apr-18",  
"Tenure in years": null,  
"Position Title": null,  
"Reason": null,  
"doc\_type": "attrition",  
"Cost Center - Name": null,  
"Tenure": null,  
"Grade": null,  
"Team": null,  
"Termination Reason": null,  
"Leader": null,  
"Emp id": "52202",  
"Hire Date": "26-Apr-18",  
"Goal": null,  
"BU": null,  
"Quarter": null,  
"Full Name": "Sumit test"  
},  
"fields": {  
"avg\_HC": [  
787  
],  
"vol\_attr\_count": [  
0  
],  
"@timestamp": [  
"2019-04-04T12:52:30.944Z"  
]  
},  
"highlight": {  
"Full Name": [  
"@kibana-highlighted-field@Sumit@/kibana-highlighted-field@ test"  
],  
"doc\_type": [  
"@kibana-highlighted-field@attrition@/kibana-highlighted-field@"  
]  
},  
"sort": [  
1554382350944  
]  
}

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [April 4, 2019, 1:19pm UTC](https://discuss.elastic.co/t/updating-data-in-csv-logstash-is-pushing-the-entire-csv-file-again-with-updated-data-which-duplicates-my-records-in-index-but-i-just-want-to-sync-the-data/175354/9 "2019-04-04T13:19:37Z")

</div>

> [@Sumit\_Kumar1](#):
>
> "\_id": "52202"

You are setting the document id, so you should not be getting duplicates. You will get a new version of the record every time you restart logstash.

---

<div class="post-metadata">

**Author:** ![Sumit\_Kumar1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sumit_kumar1/32/43471_2.png) [@Sumit\_Kumar1](https://discuss.elastic.co/u/Sumit_Kumar1)\
**Post date:** [April 4, 2019, 1:30pm UTC](https://discuss.elastic.co/t/updating-data-in-csv-logstash-is-pushing-the-entire-csv-file-again-with-updated-data-which-duplicates-my-records-in-index-but-i-just-want-to-sync-the-data/175354/10 "2019-04-04T13:30:30Z")

</div>

Yes, I know .. it should happen but it's not happening....

---

<div class="post-metadata">

**Author:** ![Sumit\_Kumar1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sumit_kumar1/32/43471_2.png) [@Sumit\_Kumar1](https://discuss.elastic.co/u/Sumit_Kumar1)\
**Post date:** [April 5, 2019, 5:16am UTC](https://discuss.elastic.co/t/updating-data-in-csv-logstash-is-pushing-the-entire-csv-file-again-with-updated-data-which-duplicates-my-records-in-index-but-i-just-want-to-sync-the-data/175354/11 "2019-04-05T05:16:08Z")

</div>

Logstash is running as a service , so whenever i make any changes or simply save it whole data of csv is getting pushed...

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [April 5, 2019, 12:34pm UTC](https://discuss.elastic.co/t/updating-data-in-csv-logstash-is-pushing-the-entire-csv-file-again-with-updated-data-which-duplicates-my-records-in-index-but-i-just-want-to-sync-the-data/175354/12 "2019-04-05T12:34:37Z")

</div>

> [@Sumit\_Kumar1](#):
>
> start\_position =\> "beginning" sincedb\_path =\> "/dev/null"

You have told logsash not to preserve state across restarts, so that is expected.

---

<div class="post-metadata">

**Author:** ![Sumit\_Kumar1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sumit_kumar1/32/43471_2.png) [@Sumit\_Kumar1](https://discuss.elastic.co/u/Sumit_Kumar1)\
**Post date:** [April 8, 2019, 5:22am UTC](https://discuss.elastic.co/t/updating-data-in-csv-logstash-is-pushing-the-entire-csv-file-again-with-updated-data-which-duplicates-my-records-in-index-but-i-just-want-to-sync-the-data/175354/13 "2019-04-08T05:22:02Z")

</div>

> [@Badger](#):
>
> You have told logsash not to preserve state across restarts, so that is expected.

After removing - start\_position =\> "beginning" sincedb\_path =\> "/dev/null"  
also whole CSV data is getting pushed .... can you please provide any sample config for updating and syncing of data for CSV.

---

<div class="post-metadata">

**Author:** ![Sumit\_Kumar1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sumit_kumar1/32/43471_2.png) [@Sumit\_Kumar1](https://discuss.elastic.co/u/Sumit_Kumar1)\
**Post date:** [April 11, 2019, 9:57am UTC](https://discuss.elastic.co/t/updating-data-in-csv-logstash-is-pushing-the-entire-csv-file-again-with-updated-data-which-duplicates-my-records-in-index-but-i-just-want-to-sync-the-data/175354/14 "2019-04-11T09:57:02Z")

</div>

Hi Badger,

Different issue that i came across is i am reading a file from windows folder

file {  
path =\> "C:\Users\Public\Documents\Elastic\logstash-6.3.0\csv\_data\attrition\_dump\*.csv"  
}

My CSV file will keep on changing so for that i used \* but, data of that csv is not getting pushed but when i give full name everything works fine.

Please, suggest the appropriate pattern for reading the file from windows directory

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 9, 2019, 9:57am UTC](https://discuss.elastic.co/t/updating-data-in-csv-logstash-is-pushing-the-entire-csv-file-again-with-updated-data-which-duplicates-my-records-in-index-but-i-just-want-to-sync-the-data/175354/15 "2019-05-09T09:57:07Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
