# Understanding regexp query better to avoid query failures and OOMs

**URL:** https://discuss.elastic.co/t/understanding-regexp-query-better-to-avoid-query-failures-and-ooms/20473
**Category:** Elasticsearch
**Created:** [October 29, 2014, 3:42am UTC](https://discuss.elastic.co/t/understanding-regexp-query-better-to-avoid-query-failures-and-ooms/20473 "2014-10-29T03:42:40Z")
**Posts on this page:** 2
**Page:** 1

<div class="post-metadata">

### Author: ![vaidik](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vaidik/32/1134_2.png) [@vaidik](https://discuss.elastic.co/u/vaidik)
#### Post date: [October 29, 2014, 3:42am UTC](https://discuss.elastic.co/t/understanding-regexp-query-better-to-avoid-query-failures-and-ooms/20473/1 "2014-10-29T03:42:40Z")

</div>

Hi Guys,

I have been trying to get my head around how Regexp Query works in  
Elasticsearch. To my knowledge, it uses Lucene's Regex Engine, which is  
limited. A problem with running regexp query on a particular field can be  
expensive depending upon the number of unique terms in the index for that  
field. So if a field has a value "brown sugar cake", and if the standard  
default tokenizer is in-use, then any regex expression provided in the  
regexp query for the field holding the mentioned value will run against  
brown, sugar and cake and not on the entire string. For this reason, regex  
in Elasticsearch (and Lucene) becomes expensive. Am I correct?

Assuming that I am, I have a further question. If the performance of Regex  
really depends on the number of unique terms in a field, then reducing the  
number of unique tokens should significantly boost up the performance. So  
running regexp queries on not\_analyzed fields should help. But that's not  
the case really and regexp is still extremely slow. In my case, the field  
is called URL and it holds URL with the query parameters. The field is  
not\_analyzed. In most of the cases, a simple regex is fast enough but if  
the regex gets slightly complicated, I never get a response from the  
server. I also noticed on a local ES server, that the memory starts  
increasing and eventually I get an OOM exception.

Another thing that is beyond my understanding is the variables on which  
performance of a regexp query works. Just to test that, I created a new  
index with just 1 document. The document looks something like this:

{  
"url": "[https://abc.com/launchingsoon?product=imgburn&](https://abc.com/launchingsoon?product=imgburn&)",  
"ts": 123456679,  
"os": "Linux",  
...  
}

Remember there is just 1 document in the index. I ran the following regex  
query:

GET /INDEX/\_search  
{  
"query": {  
"filtered": {  
"query": {  
"bool": {  
"must": [  
{  
"regexp": {  
"url":  
"._(cacaoweb|youtube-to-mp3-converter|google-chrome|itunes|adwcleaner|msn-messenger-skype|skype|adobe-flash-player-ie|firefox|jpeg-to-pdf|avira-antivir-personal---free-antivirus|irfanview|mp3-converter|realplayer|adobe-reader|youtube-download--convert|internet-explorer-8|windows-live-mail|windows-live-movie-maker-2011|ccleaner|zune-software|vanbascos-karaoke-player|amule|karaoke|imgburn|google-earth|internet-explorer-9|mp3jam|media-downloader|avg-anti-virus-free-edition|k-lite-codec-pack-full|vwo|windows-media-player|opera|kmplayer|sopcast|drweb-cureit|vwo)._"  
}  
}  
]  
}  
}  
}  
}  
}

This query ran but it took about 400 ms on my local machine. Then I ran the  
following query which has the same regular expression but a very  
unoptimized regular expression:

GET /INDEX/\_search  
{  
"query": {  
"filtered": {  
"query": {  
"bool": {  
"must": [  
{  
"regexp": {  
"url.not\_analyzed":  
"._cacaoweb._|._youtube-to-mp3-converter._|._google-chrome._|._itunes._|._adwcleaner._|._msn-messenger-skype._|._skype._|._adobe-flash-player-ie._|._firefox._|._jpeg-to-pdf._|._avira-antivir-personal---free-antivirus._|._irfanview._|._mp3-converter._|._realplayer._|._adobe-reader._|._youtube-download--convert._|._internet-explorer-8._|._windows-live-mail._|._windows-live-movie-maker-2011._|._ccleaner._|._zune-software._|._vanbascos-karaoke-player._|._amule._|._karaoke._|._imgburn._|._google-earth._|._internet-explorer-9._|._mp3jam._|._media-downloader._|._avg-anti-virus-free-edition._|._k-lite-codec-pack-full._|._photoscape._|._windows-media-player._|._opera._|._kmplayer._|._sopcast._|._drweb-cureit._"  
}  
}  
]  
}  
}  
}  
}  
}

This query took a lot of time. Logs were showing that the GC would kicking  
in after every 3-5 seconds. And finally the query fails with an OOM  
exception. I have been trying to understand what's the reason for this  
query to make OOM happen. After OOM, the ES node just becomes unresponsive  
until the GC is actually able to clear up some m.emory. This is the exact  
exception I get in the logs: [http://pastebin.mozilla.org/6975835](http://pastebin.mozilla.org/6975835).

In the above case, I understand the regex is not optimized for  
Elasticsearch's (or rather Lucene's) regex engine. But an unoptimized regex  
requires a lot of memory? I don't quite understand that.

I don't know what's causing this and I really need to understand how Regexp  
Queries work in Elasticsearch and how they work in Lucene.

Vaidik Kapoor  
vaidikkapoor.info

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CACWtv5nqQr-RqLeSp4t1KBaojByff8\_nnpi38V-zhSodB3b%3D8g%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CACWtv5nqQr-RqLeSp4t1KBaojByff8_nnpi38V-zhSodB3b%3D8g%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 6, 2017, 12:53am UTC](https://discuss.elastic.co/t/understanding-regexp-query-better-to-avoid-query-failures-and-ooms/20473/2 "2017-07-06T00:53:16Z")

</div>


