# Translate filter with +2.5 million dictionary entries

**URL:** <https://discuss.elastic.co/t/translate-filter-with-2-5-million-dictionary-entries/217390>\
**Category:** Logstash\
**Created:** [January 31, 2020, 2:01pm UTC](https://discuss.elastic.co/t/translate-filter-with-2-5-million-dictionary-entries/217390 "2020-01-31T14:01:52Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Anabella\_Cristaldi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/anabella_cristaldi/32/23612_2.png) [@Anabella\_Cristaldi](https://discuss.elastic.co/u/Anabella_Cristaldi)\
**Post date:** [January 31, 2020, 2:01pm UTC](https://discuss.elastic.co/t/translate-filter-with-2-5-million-dictionary-entries/217390/1 "2020-01-31T14:01:52Z")

</div>

Hi,  
I have a case when I need to use a translate filter with a very large YAML dictionary (around 2.5 million of entries).  
I was reading this post

> [@Optimize logstash filter plugin with million lines of dictionary look up](https://discuss.elastic.co/t/optimize-logstash-filter-plugin-with-million-lines-of-dictionary-look-up/164723):
>
> Hi, I am implementing data masking which is based in a dictionary lookup. Currently there are four dictionary files (total of ~1.2 million lines) reference to translate my greedy message. When transformation runs using translate plugin (four individual translate plugin mapped to each dictionary), execution and transformation of each line of the log file is taking ~25-30 secs. each, which is too high. Can you please advise how to optimize the data transformation? I don't want to reinvent the w…

where it recommends to use memcached in order to speed up the translation.  
But my case is I use the regex =\> true

```
                                    translate{
                                            field => "a_key"
                                            destination => "data_A_Billing"
                                            regex => true
                                            exact => true
                                            dictionary_path => '/data/sbc/tables/rates_a_billing.yaml'
                                            fallback => "NF|NF|-1|Destination Not Found|Origin Not Found|NF"
                                    }

```

where a\_key is a combination of the destination prefix and an originating number in a phone call  
For example:

**`34655#376688785`**

That will match the following line in the yaml

**'^34655#376[0-9]\*$':**"34655|376|0,083200|Spain -Mob ORANGE|ZONE 3|A"

The dictionary file is sorted from more specific to more general regex.

Originally the dictionary was of 400.000 entries, but now is 7 times bigger and latency in processing events are now a drawback.

Is there a better way to implement this?  
Any feddback will be appreciated

Thank you!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 28, 2020, 2:01pm UTC](https://discuss.elastic.co/t/translate-filter-with-2-5-million-dictionary-entries/217390/2 "2020-02-28T14:01:56Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
