# Ruby filter - How to access some elastic index data from a ruby filter

**URL:** <https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701>\
**Category:** Logstash\
**Created:** [June 20, 2018, 1:25pm UTC](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701 "2018-06-20T13:25:52Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![akapit](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/akapit/32/32490_2.png) [@akapit](https://discuss.elastic.co/u/akapit)\
**Post date:** [June 20, 2018, 1:25pm UTC](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701/1 "2018-06-20T13:25:52Z")

</div>

Hi,  
I'm trying to let logstash add a field based in the calculation of some incoming event and data that I have in another index or external parameter.

How can I access to data from any elasticsearch index in the ruby code?

This is what I'm trying to do:

```
ruby {
    code => "event.set('new_field', event.get('score').to_i * another_index.config.rate"
}

```

Where 'another\_index.config.rate' should be\*\* data from a elastic index from outside the scope of the logstash import.

I actually tried by performing a GET http request to elasticsearch from the ruby code and it seems to work, but I feel this way is not right if it's actually making a GET request for every one of the millions records that logstash is importing... this is the code i'm using:

```
uri = URI.parse('http://localhost:9200/test/config/1')
           response = Net::HTTP.get_response(uri)
           if response.code == '200'
             result = JSON.parse(response.body)
             rate = result['_source']['rate']
             event.set('new_field', event.get('Installs').to_i * rate.to_i)

           else
             event.set('new_field', '0')
           end

```

What's the right way to achieve this?

Thanks in advance

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 20, 2018, 3:07pm UTC](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701/2 "2018-06-20T15:07:35Z")

</div>

This is what the [elasticsearch filter plugin](https://www.elastic.co/guide/en/logstash/current/plugins-filters-elasticsearch.html) does, but this adds a lot of overhead and reduces throughput as you correctly point out. That will be the case even if you do it through Ruby code.

If you have a limited set of data you are looking up that does not change frequently, you could put it in a file and use the [translate filter plugin](https://www.elastic.co/guide/en/logstash/current/plugins-filters-translate.html) which stores the data in memory and therefore have a significantly smaller effect on throughput.

---

<div class="post-metadata">

**Author:** ![akapit](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/akapit/32/32490_2.png) [@akapit](https://discuss.elastic.co/u/akapit)\
**Post date:** [June 21, 2018, 7:28am UTC](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701/3 "2018-06-21T07:28:54Z")

</div>

Might the jdbc\_streaming filter help someway to reuse the connection and make it better?  
It must be an efficient way of doing this as It looks to me as a very common scenario.

Thanks a lot

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 21, 2018, 7:32am UTC](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701/4 "2018-06-21T07:32:05Z")

</div>

If you have the data in a database, the jdbc streaming plugin is certainly an option.

---

<div class="post-metadata">

**Author:** ![akapit](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/akapit/32/32490_2.png) [@akapit](https://discuss.elastic.co/u/akapit)\
**Post date:** [June 21, 2018, 7:33am UTC](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701/5 "2018-06-21T07:33:15Z")

</div>

About the translate filter recommendation, could work but I need actually data that's coming from database.. I could eventually consume that data from mongodb instead of elastic if that helps somehow.

And last... in the case that there is no way rather than my code above... the 2nd question is, what happens when logstash bombs elastic so much so many times? how does elastic react to that?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 21, 2018, 7:38am UTC](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701/6 "2018-06-21T07:38:00Z")

</div>

Issuing a lot of queries will put extra load on the cluster, so I would avoid this for streams with high throughput rates.

For this to be efficient you need a filter that can cache client-side, and unfortunately I do not think the Elasticsearch filter does that yet. There is [an open issue](https://github.com/logstash-plugins/logstash-filter-elasticsearch/issues/9), which had some comments not too far back. Maybe @guyboertje can provide some further details?

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [June 21, 2018, 7:57am UTC](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701/7 "2018-06-21T07:57:21Z")

</div>

You really don't want to make a http call to elasticsearch in the ruby filter. You will need to handle failures etc.

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [June 21, 2018, 7:59am UTC](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701/8 "2018-06-21T07:59:38Z")

</div>

If your data is (or can be) in a JDBC accessible database then JDBC Streaming is your best bet.

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [June 21, 2018, 8:04am UTC](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701/9 "2018-06-21T08:04:14Z")

</div>

Another option is to run a second LS pipeline that takes the data out of ES with an ES input plugin and writes it to a KV or JSON file that is used by the translate filter in the first LS pipeline.

---

<div class="post-metadata">

**Author:** ![akapit](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/akapit/32/32490_2.png) [@akapit](https://discuss.elastic.co/u/akapit)\
**Post date:** [June 21, 2018, 8:18am UTC](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701/10 "2018-06-21T08:18:58Z")

</div>

This is what u mean in ur last comment?

[https://www.elastic.co/guide/en/logstash/current/pipeline-to-pipeline.html](https://www.elastic.co/guide/en/logstash/current/pipeline-to-pipeline.html)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [June 21, 2018, 8:19am UTC](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701/11 "2018-06-21T08:19:11Z")

</div>

> [@guyboertje](#):
>
> Another option is to run a second LS pipeline that takes the data out of ES with an ES input plugin and writes it to a KV or JSON file that is used by the translate filter in the first LS pipeline.

I think this could be very error prone. As it is not possible to control when the translate plugin will refresh, it could end up refreshing halfway through the file being written. I would probably prefer having a script prepare the file and then just replace it when complete.

---

<div class="post-metadata">

**Author:** ![akapit](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/akapit/32/32490_2.png) [@akapit](https://discuss.elastic.co/u/akapit)\
**Post date:** [June 21, 2018, 8:20am UTC](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701/12 "2018-06-21T08:20:07Z")

</div>

I hear u,.. so it looks to me so far that the JDBC Streaming option is the safer one, isn't it?

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [June 21, 2018, 8:23am UTC](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701/13 "2018-06-21T08:23:43Z")

</div>

Actually, I don't think the second pipeline idea works as the file will not be written once at the end of the ES input data collection run.

---

<div class="post-metadata">

**Author:** ![akapit](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/akapit/32/32490_2.png) [@akapit](https://discuss.elastic.co/u/akapit)\
**Post date:** [June 21, 2018, 8:59am UTC](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701/14 "2018-06-21T08:59:21Z")

</div>

Wait, i'm confused, so the Elasticsearch filter plugin isn't a good fit for this?

([https://www.elastic.co/guide/en/logstash/current/plugins-filters-elasticsearch.html](https://www.elastic.co/guide/en/logstash/current/plugins-filters-elasticsearch.html))

Does it make a request for every record imported?

---

<div class="post-metadata">

**Author:** ![akapit](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/akapit/32/32490_2.png) [@akapit](https://discuss.elastic.co/u/akapit)\
**Post date:** [June 21, 2018, 10:15am UTC](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701/15 "2018-06-21T10:15:23Z")

</div>

I'm actually having this issue:

"Error: java::mongodb.jdbc.MongoDriver not loaded"

> <https://github.com/logstash-plugins/logstash-input-jdbc/issues/215>

And couldn't yet find any solution.

This is my config:

```
    jdbc_streaming {
    jdbc_driver_library => "/Applications/UnityJDBC/mongodb_unityjdbc_full.jar"
    jdbc_driver_class => "java::mongodb.jdbc.MongoDriver"
    jdbc_connection_string => "jdbc:mongodb://localhost:27018/mydb"
    statement => "select value from app_values where setting = 'rate'"
    target => "rate"
    add_field => {"yeah" => "This is my rate: %{rate}"}
}

```

Any direction?

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [June 21, 2018, 10:50am UTC](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701/16 "2018-06-21T10:50:13Z")

</div>

`jdbc_driver_class => "java::mongodb.jdbc.MongoDriver"` is wrong.

Have a look at [http://www.unityjdbc.com/mongojdbc/mongo\_jdbc.php](http://www.unityjdbc.com/mongojdbc/mongo_jdbc.php)

---

<div class="post-metadata">

**Author:** ![akapit](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/akapit/32/32490_2.png) [@akapit](https://discuss.elastic.co/u/akapit)\
**Post date:** [June 21, 2018, 11:10am UTC](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701/17 "2018-06-21T11:10:18Z")

</div>

You're right, however i've tried also what they wrote in theirs documentation.  
Also

> jdbc\_connection\_string =\> "jdbc:mongodb://localhost:27018/mydb"

Wasn't right, but also tried with theirs "version" and I get the same error:

```
Pipeline aborted due to error {:pipeline_id=>"main", :exception=>#<Sequel::AdapterNotFound: 
java::mongodb.jdbc.MongoDriver not loaded>, :backtrace=> 
["/Users/akapit/workspace/outflink/outflink_platform/logstash-6.2.4/vendor/bundle/jruby/2.3.0/gems/sequel-5.7.1/lib/sequel/adapters/jdbc.rb:44:in `load_driver'", 

```

@Anybody here made this work?

Thanks

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [June 21, 2018, 4:31pm UTC](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701/18 "2018-06-21T16:31:26Z")

</div>

Well, it is true that we don't have an adapter for MongoDB. AFAIR, when a specific adapter is not found a generic one is used. You may have trouble with the generic adapter because it does vanilla SQL only.

Is this what you are trying?

```auto
    jdbc_driver_class => "mongodb.jdbc.MongoDriver"
    jdbc_connection_string => "jdbc:mongo://localhost:27018/mydb"

```

The above is from the link I gave.

---

<div class="post-metadata">

**Author:** ![akapit](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/akapit/32/32490_2.png) [@akapit](https://discuss.elastic.co/u/akapit)\
**Post date:** [June 26, 2018, 11:34am UTC](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701/19 "2018-06-26T11:34:48Z")

</div>

Yes, that's what i'm trying

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [June 26, 2018, 11:36am UTC](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701/20 "2018-06-26T11:36:27Z")

</div>

Post your config please

[Next page](https://discuss.elastic.co/t/ruby-filter-how-to-access-some-elastic-index-data-from-a-ruby-filter/136701.md?page=2)
