# Slow performance of Logstash elasticsearch filter plugin

**URL:** <https://discuss.elastic.co/t/slow-performance-of-logstash-elasticsearch-filter-plugin/201609>\
**Category:** Logstash\
**Created:** [September 30, 2019, 9:47am UTC](https://discuss.elastic.co/t/slow-performance-of-logstash-elasticsearch-filter-plugin/201609 "2019-09-30T09:47:41Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![msk\_76](https://avatars.discourse-cdn.com/v4/letter/m/dbc845/32.png) [@msk\_76](https://discuss.elastic.co/u/msk_76)\
**Post date:** [September 30, 2019, 9:47am UTC](https://discuss.elastic.co/t/slow-performance-of-logstash-elasticsearch-filter-plugin/201609/1 "2019-09-30T09:47:41Z")

</div>

I have around 4.5Million records in my input data of logstash to which I am doing a lookup of an existing index in ES using following ES filter plugin. This is just like adding department information to a user\_name field.

elasticsearch {  
hosts =\> ["[http://10.129.212.45:9200](http://10.129.212.45:9200)"]  
index =\> "sys\_username\_mapping"  
query =\> "user\_name:%{[user\_name]}"  
fields =\> { "email" =\> "email" "site" =\> "site" "group" =\> "group" "division" =\> "division" "cad\_cc" =\> "cad\_cc" "ldap\_cc" =\> "ldap\_cc" }  
}

After this lookup, I am doing indexing of this complete data in a new index in ES.

If I comment es filter plugin ( i.e without department information) it takes about 5-6 minutes to load all input data in elasticsearch and with having this filter plugin, it 's not even completing in 40 minutes.

Does translate filter can be an alternative to this? Will it perform better than ES filter plugin if I translate ( lookup ) to a text file than an already indexed data?

This user\_name to department kind of lookup is important for me.

Please suggest

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [September 30, 2019, 1:25pm UTC](https://discuss.elastic.co/t/slow-performance-of-logstash-elasticsearch-filter-plugin/201609/2 "2019-09-30T13:25:19Z")

</div>

An elasticsearch output makes one API call to elasticsearch for each batch of events. By default the batch size is 125. An elasticsearch filter makes one API call to elasticsearch for each event, so it is making 125 times as many calls. Thus it is not surprising to me that it would take more than 10 times as long.

I would expect a translate filter to be very much faster.

---

<div class="post-metadata">

**Author:** ![msk\_76](https://avatars.discourse-cdn.com/v4/letter/m/dbc845/32.png) [@msk\_76](https://discuss.elastic.co/u/msk_76)\
**Post date:** [October 1, 2019, 4:26am UTC](https://discuss.elastic.co/t/slow-performance-of-logstash-elasticsearch-filter-plugin/201609/3 "2019-10-01T04:26:31Z")

</div>

Thanks Badger for this explanation.

I am surprised when you said that a translate filter where the lookup file is stored on local disk will work faster than a ES query response in case of elasticsearch filter plugin. Because the disk IO throughput for ES cluster( due to parallelism) is multiple time higher than local disk of logstash server. May be I am wrong here.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 1, 2019, 5:12am UTC](https://discuss.elastic.co/t/slow-performance-of-logstash-elasticsearch-filter-plugin/201609/4 "2019-10-01T05:12:19Z")

</div>

The file the translate filter uses is read into memory so might require a larger heap if it is big but does not result in a lot of disk I/O.

---

<div class="post-metadata">

**Author:** ![msk\_76](https://avatars.discourse-cdn.com/v4/letter/m/dbc845/32.png) [@msk\_76](https://discuss.elastic.co/u/msk_76)\
**Post date:** [October 1, 2019, 6:30am UTC](https://discuss.elastic.co/t/slow-performance-of-logstash-elasticsearch-filter-plugin/201609/5 "2019-10-01T06:30:54Z")

</div>

Thanks Christian,

Yes, I wasn’t aware of this in memory read of translate filter. However could you also tell how to workaround the lookup file rollover because overwriting/updating  
the lookup on disk could cause issue during the time when file/inode is getting updated? Is there any parameter in logstash with translate filter which keeps the last in-memory read of file and refresh it in memory only on our command.

Just for example steps:

1. 

Previous lookup file loaded in memory of logstash.

1. 

Lookup file updated or replaced or overwritten

1. 

Refresh the new file in-memory of logstash by some schedule.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 29, 2019, 6:30am UTC](https://discuss.elastic.co/t/slow-performance-of-logstash-elasticsearch-filter-plugin/201609/6 "2019-10-29T06:30:58Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
