# Minimise dataset for regexp query

**URL:** https://discuss.elastic.co/t/minimise-dataset-for-regexp-query/37558
**Category:** Elasticsearch
**Created:** [December 18, 2015, 9:49am UTC](https://discuss.elastic.co/t/minimise-dataset-for-regexp-query/37558 "2015-12-18T09:49:26Z")
**Posts on this page:** 4
**Page:** 1

<div class="post-metadata">

### Author: ![hrathore](https://avatars.discourse-cdn.com/v4/letter/h/ed655f/32.png) [@hrathore](https://discuss.elastic.co/u/hrathore)
#### Post date: [December 18, 2015, 9:49am UTC](https://discuss.elastic.co/t/minimise-dataset-for-regexp-query/37558/1 "2015-12-18T09:49:26Z")

</div>

I am running a regexp query that starts from (._) and it runs on filetext field for a specific employeeId.  
The searched text can be anywhere inside the token, so I am wrapping the regex with ._

Index Size: 100GB  
Mapping for Index is: employeeId, filetext  
Search Type: Scan  
Scroll=5m

```auto
{
  "query": {
    "filtered" : {
      "query" : { "regexp" : { "filetext" : ".*1234.*" }},
      "filter" : {
        "bool" : {
          "must": { "term" : { "employeeId" : 5725 }},
        }
      }
    }
  } 
} 
```

This query scans the full index instead of searching data specific to that employeeId (confirmed as the full disk is read, as per IO stats), and as it is a regex query, it's very slow.

1. How do I make sure that query runs only on specific employeeId, and not on complete dataset?
2. How do I speed this up?

---

<div class="post-metadata">

### Author: ![cbuescher](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cbuescher/32/60402_2.png) [@cbuescher](https://discuss.elastic.co/u/cbuescher)
#### Post date: [December 18, 2015, 11:21am UTC](https://discuss.elastic.co/t/minimise-dataset-for-regexp-query/37558/2 "2015-12-18T11:21:57Z")

</div>

Hi,

I assume this is Elasticsearch 1.7, since the filtered query has been deprecated in 2.0. Have you tried moving the regex query after the term query in the must section of the bool-filter like this? I haven't tested this with a large number of documents, but the must clauses should be [executed in order](https://www.elastic.co/guide/en/elasticsearch/guide/current/_filter_order.html), so the term-filter should reduce the number of documents the regexp-query runs on.

```auto
"query" : {
    "filtered" : {
      "filter" : {
        "bool" : {
          "must": [
            { "term" : { "employeeId" : 5725 }},
            { "regexp" : { "filetext" : ".*1234.*" }}
        ]}
      }
    }
  } 

```

---

<div class="post-metadata">

### Author: ![hrathore](https://avatars.discourse-cdn.com/v4/letter/h/ed655f/32.png) [@hrathore](https://discuss.elastic.co/u/hrathore)
#### Post date: [December 22, 2015, 5:35am UTC](https://discuss.elastic.co/t/minimise-dataset-for-regexp-query/37558/4 "2015-12-22T05:35:17Z")

</div>

@Christoph  
I am using version 1.5.2.  
I have tried the query you suggested, but it didn't make any difference. The new query took the same time as earlier, and searched the whole index instead of the employeeId filter.

Does it change in version 2.0 ?

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 5, 2017, 11:29pm UTC](https://discuss.elastic.co/t/minimise-dataset-for-regexp-query/37558/5 "2017-07-05T23:29:26Z")

</div>


