# Logstash - update CSV content changes into ES

**URL:** <https://discuss.elastic.co/t/logstash-update-csv-content-changes-into-es/175984>\
**Category:** Logstash\
**Created:** [April 9, 2019, 8:49am UTC](https://discuss.elastic.co/t/logstash-update-csv-content-changes-into-es/175984 "2019-04-09T08:49:54Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![cheriemilk](https://avatars.discourse-cdn.com/v4/letter/c/c37758/32.png) [@cheriemilk](https://discuss.elastic.co/u/cheriemilk)\
**Post date:** [April 9, 2019, 8:49am UTC](https://discuss.elastic.co/t/logstash-update-csv-content-changes-into-es/175984/1 "2019-04-09T08:49:54Z")

</div>

Hi Team.

I have a user scenario that to maintain automation test cases in elk.

Step1(No Issue) - my colleges will send me first CSV file which containing below columns. And I use CSV filter to output them in to ES.  
1. Test Case ID  
2. Author  
3. Feature  
4. Function  
5. Verification

Step2(has issue here) - after 1 moth, my colleages will send me second CSV file which contains the newly created test cases or existing test cases updated(for example, Verification and Function are updated.)

I want logstash to judge the Test Case ID existed in ES or not.

If test Case ID already existed in ES, then to overwrite the existing event .  
If test case ID not found in ES, then to newly created one Event in ES.

How should I write the judet by IF condition in logstahs?

```
filter { 
          csv { columns => [ "Test Case ID",
                              "Author",
                              "Feature",
                              "Function",
                              "Verification"]
               separator => ","
			   skip_header => "true"
			   }   
	   
                   
        If Test Case ID in the CSV
        { overwrite the existing event in ES}
        else
        {Create a new event in ES}
 }
```

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [April 9, 2019, 10:55pm UTC](https://discuss.elastic.co/t/logstash-update-csv-content-changes-into-es/175984/2 "2019-04-09T22:55:56Z")

</div>

> [@cheriemilk](#):
>
> If test Case ID already existed in ES, then to overwrite the existing event .  
> If test case ID not found in ES, then to newly created one Event in ES.

Don't do this using if-else in the filter section. Use an elasticsearch output, set document\_id to your test case id (or a hash of it using a fingerprint filter) and set the [doc\_as\_upsert](https://www.elastic.co/guide/en/logstash/current/plugins-outputs-elasticsearch.html#plugins-outputs-elasticsearch-doc_as_upsert) option on the output

---

<div class="post-metadata">

**Author:** ![cheriemilk](https://avatars.discourse-cdn.com/v4/letter/c/c37758/32.png) [@cheriemilk](https://discuss.elastic.co/u/cheriemilk)\
**Post date:** [April 10, 2019, 5:16am UTC](https://discuss.elastic.co/t/logstash-update-csv-content-changes-into-es/175984/3 "2019-04-10T05:16:36Z")

</div>

Hi Badger,

Per the official user guide, it says that type of "document\_id" is a string. So I add this configuration document\_id =\> " TID". But in Kibana, the result is not expected.  
1. The \_id field value is TID, instead of 1 or 2.  
2. only 2nd record is indexed to ES. Where is first 1 record?

```
  **Data:** 
  TID,Author,Module,Feature,Function,Verification,Creation Date
  1,Cherie Zhou,SCM,SOC,Nomination,successfully,2019-04-10
  2,Cherie Zhou,CAL,Analyzer,Analyzer,successfully,2019-04-09

    **configuration**              
     output {
        elasticsearch {
    	   action => "index"
    	   hosts => "localhost:9200"
    	   index => "testcase"
         manage_template => true
         template => "C:/elkstack/elasticsearch-6.5.1/indextemp/dtemplate.json"
         template_name=> "dtemplate"
         template_overwrite => true
         document_id => "Test Case ID"
         doc_as_upsert => true }              
    	stdout { codec => rubydebug {metadata => true}}
    }
```

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [April 10, 2019, 11:53am UTC](https://discuss.elastic.co/t/logstash-update-csv-content-changes-into-es/175984/4 "2019-04-10T11:53:46Z")

</div>

> [@cheriemilk](#):
>
> document\_id =\> "Test Case ID"

```
document_id => "%{[Test Case ID]}"

```

---

<div class="post-metadata">

**Author:** ![cheriemilk](https://avatars.discourse-cdn.com/v4/letter/c/c37758/32.png) [@cheriemilk](https://discuss.elastic.co/u/cheriemilk)\
**Post date:** [April 10, 2019, 2:10pm UTC](https://discuss.elastic.co/t/logstash-update-csv-content-changes-into-es/175984/5 "2019-04-10T14:10:00Z")

</div>

Figure it out. It should be document\_id =\> "%{TID}"

---

<div class="post-metadata">

**Author:** ![cheriemilk](https://avatars.discourse-cdn.com/v4/letter/c/c37758/32.png) [@cheriemilk](https://discuss.elastic.co/u/cheriemilk)\
**Post date:** [April 10, 2019, 2:11pm UTC](https://discuss.elastic.co/t/logstash-update-csv-content-changes-into-es/175984/6 "2019-04-10T14:11:27Z")

</div>

why there is "[" "]"? it could work as well without the square brackets.

Does it mean hash? If it's hash scenario, multiple fields can be putted in it. For example [TID, UID, EID]. How does it know from which field the value of \_id should come from?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [April 10, 2019, 2:32pm UTC](https://discuss.elastic.co/t/logstash-update-csv-content-changes-into-es/175984/7 "2019-04-10T14:32:28Z")

</div>

The square brackets are optional in your case. If you are referencing a field inside an object they are not optional, so if your event had a beat field that contains a hostname field, a sprintf reference to it would be %{[beat][hostname]}

---

<div class="post-metadata">

**Author:** ![cheriemilk](https://avatars.discourse-cdn.com/v4/letter/c/c37758/32.png) [@cheriemilk](https://discuss.elastic.co/u/cheriemilk)\
**Post date:** [April 10, 2019, 9:59pm UTC](https://discuss.elastic.co/t/logstash-update-csv-content-changes-into-es/175984/8 "2019-04-10T21:59:39Z")

</div>

Ok. Thank you. it’s the case of referencing to a nested field.

One more question. If I want the value of \_id comes from the combination of TID and Author. Is the syntax like this?? document\_id =\> “%{TID}+%{Author}”

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [April 10, 2019, 10:18pm UTC](https://discuss.elastic.co/t/logstash-update-csv-content-changes-into-es/175984/9 "2019-04-10T22:18:04Z")

</div>

> [@cheriemilk](#):
>
> document\_id =\> “%{TID}+%{Author}”

Yes, that will work.

If you need to do that with any more than those 2 switch to fingerprint...

```
fingerprint { source => ["TID", "Author"] target => "[@metadata][docid]" method => "MURMUR3" }

```

Then you can use document\_id =\> %{[@metadata][docid]}

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 8, 2019, 10:18pm UTC](https://discuss.elastic.co/t/logstash-update-csv-content-changes-into-es/175984/10 "2019-05-08T22:18:07Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
