# Optimize logstash filter plugin with million lines of dictionary look up

**URL:** <https://discuss.elastic.co/t/optimize-logstash-filter-plugin-with-million-lines-of-dictionary-look-up/164723>\
**Category:** Logstash\
**Created:** [January 18, 2019, 2:44am UTC](https://discuss.elastic.co/t/optimize-logstash-filter-plugin-with-million-lines-of-dictionary-look-up/164723 "2019-01-18T02:44:29Z")\
**Posts on this page:** 17\
**Page:** 1

<div class="post-metadata">

**Author:** ![ronchav](https://avatars.discourse-cdn.com/v4/letter/r/35a633/32.png) [@ronchav](https://discuss.elastic.co/u/ronchav)\
**Post date:** [January 18, 2019, 2:44am UTC](https://discuss.elastic.co/t/optimize-logstash-filter-plugin-with-million-lines-of-dictionary-look-up/164723/1 "2019-01-18T02:44:29Z")

</div>

Hi,

I am implementing data masking which is based in a dictionary lookup. Currently there are four dictionary files (total of ~1.2 million lines) reference to translate my greedy message.

When transformation runs using translate plugin (four individual translate plugin mapped to each dictionary), execution and transformation of each line of the log file is taking ~25-30 secs. each, which is too high.

Can you please advise how to optimize the data transformation? I don't want to reinvent the wheel but in case you have any idea or somebody also faced this performance issue. Thank you in advance.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [January 18, 2019, 3:01am UTC](https://discuss.elastic.co/t/optimize-logstash-filter-plugin-with-million-lines-of-dictionary-look-up/164723/2 "2019-01-18T03:01:00Z")

</div>

What about putting the data into an index and then using the Elasticsearch filter to query it?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 18, 2019, 6:07am UTC](https://discuss.elastic.co/t/optimize-logstash-filter-plugin-with-million-lines-of-dictionary-look-up/164723/3 "2019-01-18T06:07:40Z")

</div>

What does your data and config look like? Can you show some sample dictionary records?

---

<div class="post-metadata">

**Author:** ![ronchav](https://avatars.discourse-cdn.com/v4/letter/r/35a633/32.png) [@ronchav](https://discuss.elastic.co/u/ronchav)\
**Post date:** [January 18, 2019, 8:57am UTC](https://discuss.elastic.co/t/optimize-logstash-filter-plugin-with-million-lines-of-dictionary-look-up/164723/4 "2019-01-18T08:57:31Z")

</div>

@warkolm: Technically possible but we cannot do that, sensitive data shouldn't reach elastic engine. At the ETL data integration layer it should be masked before forwarding to elastic. Thanks

---

<div class="post-metadata">

**Author:** ![ronchav](https://avatars.discourse-cdn.com/v4/letter/r/35a633/32.png) [@ronchav](https://discuss.elastic.co/u/ronchav)\
**Post date:** [January 18, 2019, 9:04am UTC](https://discuss.elastic.co/t/optimize-logstash-filter-plugin-with-million-lines-of-dictionary-look-up/164723/5 "2019-01-18T09:04:27Z")

</div>

My config looks like this:

if [message] =~ /\d/ {  
mutate {  
gsub =\> ["message", "\d", "#"]  
add\_tag =\> "Masked"  
}  
}

translate {  
field =\> "message"  
destination =\> "message"  
override =\> true  
exact =\> false  
dictionary\_path =\> "dictionary1.csv"  
fallback =\> "NoMatch: %{message}"  
}

translate {  
field =\> "message"  
destination =\> "message"  
override =\> true  
exact =\> false  
dictionary\_path =\> "dictionary2.csv"  
fallback =\> "NoMatch: %{message}"  
}

translate {  
field =\> "message"  
destination =\> "message"  
override =\> true  
exact =\> false  
dictionary\_path =\> "dictionary3.csv"  
fallback =\> "NoMatch: %{message}"  
}

translate {  
field =\> "message"  
destination =\> "message"  
override =\> true  
exact =\> false  
dictionary\_path =\> "dictionary4.csv"  
fallback =\> "NoMatch: %{message}"  
}

if "NoMatch" in [message] {  
mutate { gsub =\> ["message", "NoMatch: ", ""] }  
} else {  
if !("Masked" in [tags]) {  
mutate {  
add\_tag =\> "Masked"  
}  
}  
}

Sample dictionary:  
dictionary1.csv  
Christian, C#######n  
Jayson, J####n  
.  
.  
.  
etc... to thousand of lines

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 18, 2019, 10:56am UTC](https://discuss.elastic.co/t/optimize-logstash-filter-plugin-with-million-lines-of-dictionary-look-up/164723/6 "2019-01-18T10:56:55Z")

</div>

That sounds like a very expensive brute force way to address the problem. Is there any pattern to the words/phrases you are masking? If not I suspect it would be more efficient to create a custom plugin that loads the full dictionary and the processes the message word by word comparing it to the dictionary and generating a new, updated message based on the matched data.

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [January 18, 2019, 12:54pm UTC](https://discuss.elastic.co/t/optimize-logstash-filter-plugin-with-million-lines-of-dictionary-look-up/164723/7 "2019-01-18T12:54:00Z")

</div>

We are releasing the memcached filter as part of Logstash 6.6.0 in 3 days time. You can install it separately on earlier LS versions though.

`bin/logstash-plugin install logstash-filter-memcached`

You will have to run a memcached daemon though and pre-load it with the KV data - make sure you understand the expiry side of things.

> <https://github.com/RedisLabs/memcache_populator/blob/master/mcpopulator.py>

> **[jorisroovers/memclient](https://github.com/jorisroovers/memclient)**
>
> Simple memcached commandline client written in Go. Contribute to jorisroovers/memclient development by creating an account on GitHub.

> <https://stackoverflow.com/questions/46540380/storing-million-key-value-in-memcached-good-or-bad-idea>

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [January 18, 2019, 11:16pm UTC](https://discuss.elastic.co/t/optimize-logstash-filter-plugin-with-million-lines-of-dictionary-look-up/164723/9 "2019-01-18T23:16:14Z")

</div>

It'd be better if you create a new topic for this question 🙂

---

<div class="post-metadata">

**Author:** ![ronchav](https://avatars.discourse-cdn.com/v4/letter/r/35a633/32.png) [@ronchav](https://discuss.elastic.co/u/ronchav)\
**Post date:** [January 21, 2019, 9:24am UTC](https://discuss.elastic.co/t/optimize-logstash-filter-plugin-with-million-lines-of-dictionary-look-up/164723/11 "2019-01-21T09:24:10Z")

</div>

HI Christian, there is no such pattern as this masking will happen in the greedy message. I wrote a Java program to perform the masking and it is doing it well in terms of performance. I only need to call it from Ruby as custom filter plugin. Will this be fine? Any thoughts on this

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [January 21, 2019, 9:32am UTC](https://discuss.elastic.co/t/optimize-logstash-filter-plugin-with-million-lines-of-dictionary-look-up/164723/12 "2019-01-21T09:32:03Z")

</div>

I believe there is a new Java API that can be used to create plugins, so you might be able to convert your code into a filter plugin. Not sure how well documented this is or whether it has been finalised. Maybe @guyboertje or someone else from the Logstash team knows?

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [January 21, 2019, 9:36am UTC](https://discuss.elastic.co/t/optimize-logstash-filter-plugin-with-million-lines-of-dictionary-look-up/164723/13 "2019-01-21T09:36:55Z")

</div>

We are releasing experimental support for plugins written in Java in 6.6.0 in the not too distant future. There will be a specific blog post by Dan Hermann explaining about the Java Plugin API.

---

<div class="post-metadata">

**Author:** ![ronchav](https://avatars.discourse-cdn.com/v4/letter/r/35a633/32.png) [@ronchav](https://discuss.elastic.co/u/ronchav)\
**Post date:** [January 22, 2019, 6:41am UTC](https://discuss.elastic.co/t/optimize-logstash-filter-plugin-with-million-lines-of-dictionary-look-up/164723/14 "2019-01-22T06:41:49Z")

</div>

Hi @warkolm, another approach I can think of is if we let the data flow in to elastic and then restrict the field to the user and recreate a new masked field using ingest pipeline processors. But I am not sure if in the ingest processors that elastic have is supporting lookup from external dictionary reference. Any idea? Thanks

---

<div class="post-metadata">

**Author:** ![ronchav](https://avatars.discourse-cdn.com/v4/letter/r/35a633/32.png) [@ronchav](https://discuss.elastic.co/u/ronchav)\
**Post date:** [January 22, 2019, 6:42am UTC](https://discuss.elastic.co/t/optimize-logstash-filter-plugin-with-million-lines-of-dictionary-look-up/164723/15 "2019-01-22T06:42:32Z")

</div>

Thanks, we can check on this in the future but we decided to go to production with 6.3 version.

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [January 29, 2019, 4:28pm UTC](https://discuss.elastic.co/t/optimize-logstash-filter-plugin-with-million-lines-of-dictionary-look-up/164723/16 "2019-01-29T16:28:49Z")

</div>

I have a similar use case, I'm trying to use a dictionary with ~ 19 millions lines and 700 MB, but I'm still looking on how to implement it, since loading it on memory as a `.yml` file does not seem as a good idea.

The memcached filter is already available? I saw that 6.6.0 was launched today, but didn't find any information about a memcached filter.

---

<div class="post-metadata">

**Author:** ![guyboertje](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guyboertje/32/31592_2.png) [@guyboertje](https://discuss.elastic.co/u/guyboertje)\
**Post date:** [January 29, 2019, 5:54pm UTC](https://discuss.elastic.co/t/optimize-logstash-filter-plugin-with-million-lines-of-dictionary-look-up/164723/17 "2019-01-29T17:54:49Z")

</div>

We are still in the process of releasing 6.6.0. Download artifacts are up but docs and blog posts are in progress.

In the meantime...

[https://www.elastic.co/guide/en/logstash-versioned-plugins/current/v0.1.1-plugins-filters-memcached.html](https://www.elastic.co/guide/en/logstash-versioned-plugins/current/v0.1.1-plugins-filters-memcached.html)

The memcached filter plugin is installable now on 6.5.4 etc.

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [January 29, 2019, 6:05pm UTC](https://discuss.elastic.co/t/optimize-logstash-filter-plugin-with-million-lines-of-dictionary-look-up/164723/18 "2019-01-29T18:05:51Z")

</div>

Thanks! I will try that!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 26, 2019, 6:08pm UTC](https://discuss.elastic.co/t/optimize-logstash-filter-plugin-with-million-lines-of-dictionary-look-up/164723/19 "2019-02-26T18:08:53Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
