# How to update a index based on fields other than the document\_id using logstash?

**URL:** <https://discuss.elastic.co/t/how-to-update-a-index-based-on-fields-other-than-the-document-id-using-logstash/133481>\
**Category:** Logstash\
**Created:** [May 28, 2018, 7:35am UTC](https://discuss.elastic.co/t/how-to-update-a-index-based-on-fields-other-than-the-document-id-using-logstash/133481 "2018-05-28T07:35:33Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Rajkumar\_Selvakumar](https://avatars.discourse-cdn.com/v4/letter/r/a88e4f/32.png) [@Rajkumar\_Selvakumar](https://discuss.elastic.co/u/Rajkumar_Selvakumar)\
**Post date:** [May 28, 2018, 7:35am UTC](https://discuss.elastic.co/t/how-to-update-a-index-based-on-fields-other-than-the-document-id-using-logstash/133481/1 "2018-05-28T07:35:33Z")

</div>

Here is my use case,

I have created a elastic index ( name : order\_index) and I able to load the index with the inputs from my database through logstash jdbc plugin ( all is fine ). the index looks like this seq\_id (database sequence), order id, sender , receiver,proessed time.

doucment\_id is seq\_id ..  
order id + sender + receiver is unique ( once the order is received the receivers are determined based on a routing logic), so there are more than one record for each order\_id.

Now I have a CSV file, which has order ID and request Time.

I need to read the CSV ( i use logstash csv plugin ) and upsert in to the order\_index.

problem:

CSV file has only order Id ( no sender or receiver) and request date.

I need to add the request date in order\_index based on order id.

order\_index cannot have order\_id as document\_id because order\_id is unique.

> Blockquote

input {  
beats {  
port =\> "5044"  
}  
}

filter {

csv {  
separator =\> ","  
columns =\> ["order\_id","request\_time"]  
}

}  
output {

elasticsearch {  
hosts =\> ["localhost:9200"]  
index =\> ["order\_index"]  
}

> Blockquote

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [May 30, 2018, 9:53am UTC](https://discuss.elastic.co/t/how-to-update-a-index-based-on-fields-other-than-the-document-id-using-logstash/133481/2 "2018-05-30T09:53:52Z")

</div>

Where does the CSV file come from? A database?  
If not, how often is new data added to the file?

I think you could use a SQLite DB, import the CSV and use the jdbc\_streaming filter to add the request date to all documents with the same order\_id.

What analysis are you intending to do in Kibana?

---

<div class="post-metadata">

**Author:** ![Rajkumar\_Selvakumar](https://avatars.discourse-cdn.com/v4/letter/r/a88e4f/32.png) [@Rajkumar\_Selvakumar](https://discuss.elastic.co/u/Rajkumar_Selvakumar)\
**Post date:** [May 30, 2018, 10:16am UTC](https://discuss.elastic.co/t/how-to-update-a-index-based-on-fields-other-than-the-document-id-using-logstash/133481/3 "2018-05-30T10:16:42Z")

</div>

Thanks for your response.

The CSV is generated from jmeter ( lod run tests).

The CSV file will be generated after the end of test run. So, its not frequently updated.

importing the CSV to the database and then reading it through jdbc input plugin would defenitely work. thats an option which I have.

However, I am trying to find a solution where in which I am trying to by-pass the database,  
If logstash can provide an option to update the existing documents in the index based on certain fields ( similar to the search query), then I can use that, instead of load the CSV into database and then stream it to Elastic.

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [May 30, 2018, 10:30am UTC](https://discuss.elastic.co/t/how-to-update-a-index-based-on-fields-other-than-the-document-id-using-logstash/133481/4 "2018-05-30T10:30:46Z")

</div>

Maybe scripted updates in the elasticsearch output might work for you?

> [@Using script params - elasticsearch output plugin](https://discuss.elastic.co/t/using-script-params-elasticsearch-output-plugin/71693):
>
> Hello, I was trying to use [script params in scripted updates](https://www.elastic.co/guide/en/elasticsearch/reference/current/docs-update.html#_scripted_updates) in the elasticsearch output plugin of logstash but I was unsure of how to do it. Currently, we two types of scripted updates: script\_lang =\> "painless" script\_type =\> "inline" script =\> ' ctx.\_source.name\_servers = "%{name\_servers}"; ctx.\_source.past\_name\_servers.add("%{past\_name\_servers}"); ' Where my event has f…

This would be in a second LS pipeline that reads the CSV file. Be aware that the document you are trying to update must be in ES already. You may need to use the elasticsearch filter to add the sequence\_id of the first doc matching against order\_id. Alternatively you could try to add an array of sequence\_id of all matching docs by order\_id and then split the array into multiple docs.

---

<div class="post-metadata">

**Author:** ![Rajkumar\_Selvakumar](https://avatars.discourse-cdn.com/v4/letter/r/a88e4f/32.png) [@Rajkumar\_Selvakumar](https://discuss.elastic.co/u/Rajkumar_Selvakumar)\
**Post date:** [June 4, 2018, 2:14am UTC](https://discuss.elastic.co/t/how-to-update-a-index-based-on-fields-other-than-the-document-id-using-logstash/133481/5 "2018-06-04T02:14:50Z")

</div>

Hi ,

Have been busy on other stuff. I have started learning scripts in elastic. Let me try to create scripts and will test whether it helps to resolve my problem. I will share the updated logstash config once done.

Thanks,  
Raj

---

<div class="post-metadata">

**Author:** ![Rajkumar\_Selvakumar](https://avatars.discourse-cdn.com/v4/letter/r/a88e4f/32.png) [@Rajkumar\_Selvakumar](https://discuss.elastic.co/u/Rajkumar_Selvakumar)\
**Post date:** [June 5, 2018, 9:08am UTC](https://discuss.elastic.co/t/how-to-update-a-index-based-on-fields-other-than-the-document-id-using-logstash/133481/6 "2018-06-05T09:08:28Z")

</div>

Have found a solution but fews issues have to be sorted out.Since I can read the CSV file , i extract the order id , sender & receiver.  
Use ruby flter plugin to query the index based on order id , sender & receiver to get the document id.Once the document ID is available , i can use the regular upsert function as below.

I am sharing my configuration , please review and assist

> output {  
> elasticsearch {
> 
> hosts =\> ["host1:9200", "host2:9200","host3:9200"]  
> document\_id =\> "%{my\_doc\_id}"  
> index =\> "My\_Index"  
> doc\_as\_upsert =\> true  
> action =\> "update"  
> manage\_template =\> true  
> }  
> }

My ruby filter plugin:

> input {  
> beats {  
> port =\> "5044"  
> }  
> }
> 
> filter {  
> csv {  
> separator =\> ","  
> columns =\> ["publish\_time","order\_id","sender","receiver"]  
> }  
> mutate {  
> remove\_field =\> ["source", "message" , "beat", "host"]  
> }  
> ruby {  
> code =\> "  
> require 'elasticsearch'  
> client = Elasticsearch::Client.new hosts: ["host1:9200", "host2:9200","host3:9200"]  
> response = client.search index: 'My\_Index', body: { query: { match\_all:  
> { 'order\_id': event.get('correlation\_id'),  
> 'sender': event.get('sender'),  
> 'receiver': event.get('receiver')  
> }  
> }  
> }  
> event.set('my\_doc\_id', response ['hits']['hits'][0]['\_source']['\_id'])
> 
> ```
> "
> 
> ```
> 
> }  
> }

problem: the query returns numerous documents instead of one because , order\_id , sender and receiver are treated as text. In kibana, i can get unique results if I use .keyword ( example order\_id.keyword ) , but when I use it inside ruby plugin , i encounter the below error.

SyntaxError: (ruby filter code):4: syntax error, unexpected ':'  
response = client.search index: 'perf\_report\_by\_audit\_and\_jmeter', body: { query: { match: { 'order\_id.keyword': event.get(order\_id') } } }  
^  
eval at org/jruby/RubyKernel.java:1079  
register at /bpms/ELK/logstash-5.6.3/vendor/bundle/jruby/1.9/gems/logstash-filter-ruby-3.0.4/lib/logstash/filters/ruby.rb:38  
register at /bpms/ELK/logstash-5.6.3/vendor/jruby/lib/ruby/1.9/forwardable.rb:201  
register\_plugin at /bpms/ELK/logstash-5.6.3/logstash-core/lib/logstash/pipeline.rb:290  
register\_plugins at /bpms/ELK/logstash-5.6.3/logstash-core/lib/logstash/pipeline.rb:301  
each at org/jruby/RubyArray.java:1613  
register\_plugins at /bpms/ELK/logstash-5.6.3/logstash-core/lib/logstash/pipeline.rb:301  
start\_workers at /bpms/ELK/logstash-5.6.3/logstash-core/lib/logstash/pipeline.rb:311  
run at /bpms/ELK/logstash-5.6.3/logstash-core/lib/logstash/pipeline.rb:235  
start\_pipeline at /bpms/ELK/logstash-5.6.3/logstash-core/lib/logstash/agent.rb:398

---

<div class="post-metadata">

**Author:** ![Rajkumar\_Selvakumar](https://avatars.discourse-cdn.com/v4/letter/r/a88e4f/32.png) [@Rajkumar\_Selvakumar](https://discuss.elastic.co/u/Rajkumar_Selvakumar)\
**Post date:** [June 22, 2018, 6:05am UTC](https://discuss.elastic.co/t/how-to-update-a-index-based-on-fields-other-than-the-document-id-using-logstash/133481/7 "2018-06-22T06:05:09Z")

</div>

I have resolved the issue.

Used Ruby code with ElasticSearch gem to address the issue.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 20, 2018, 6:05am UTC](https://discuss.elastic.co/t/how-to-update-a-index-based-on-fields-other-than-the-document-id-using-logstash/133481/8 "2018-07-20T06:05:11Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
