# How to access \_id from Painless in Query context?

**URL:** <https://discuss.elastic.co/t/how-to-access-id-from-painless-in-query-context/155211>\
**Category:** Elasticsearch\
**Created:** [November 2, 2018, 5:02pm UTC](https://discuss.elastic.co/t/how-to-access-id-from-painless-in-query-context/155211 "2018-11-02T17:02:29Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![m9aertner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/m9aertner/32/39588_2.png) [@m9aertner](https://discuss.elastic.co/u/m9aertner)\
**Post date:** [November 2, 2018, 5:02pm UTC](https://discuss.elastic.co/t/how-to-access-id-from-painless-in-query-context/155211/1 "2018-11-02T17:02:29Z")

</div>

I would like to sanity-check indexed data.

Is it possible to access the \_id field in a Painless script condition, in query context?

That is, something like (ES 5.3.x):

```
{
  "query": {
     "bool": {
        "must": [
           { "term": { ... other conditions ... } },
           {
              "script" : {
                 "script" : {
                    "inline": "doc['_id'].value.length()>10",
                    "lang": "painless"
                 }
              }
           }
        ]
     }
  }
}

```

This yields

```
    "type": "illegal_argument_exception",
    "reason": "Fielddata is not supported on field [_id] of type [_id]"
```

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [November 6, 2018, 9:12am UTC](https://discuss.elastic.co/t/how-to-access-id-from-painless-in-query-context/155211/2 "2018-11-06T09:12:59Z")

</div>

Hey,

I am not sure this works in ES 5.x without mapping changes, but it does with ES 6.4

```auto
GET foo/_search
{
  "query": {
    "bool": {
      "must": [
        {
          "script": {
            "script": "doc['_id'][0].length() > 1"
          }
        }
      ]
    }
  }
}

```

Still, I think the better approach here would be to use an ingest processor and store the length of the field on indexing (I only tested this on 6.x as well)

```auto
POST _ingest/pipeline/_simulate
{
  "pipeline" :
  {
    "description": "_description",
    "processors": [
      {
        "script" : {
          "source" : "ctx.len = ctx._id.length()"
        }
      }
    ]
  },
  "docs": [
    {
      "_index": "index",
      "_type": "doc",
      "_id": "1",
      "_source": { "foo": "bar" }
    },
    {
      "_index": "index",
      "_type": "_doc",
      "_id": "second",
      "_source": { "foo": "rab" }
    }
  ]
}

# returns
{
  "docs": [
    {
      "doc": {
        "_index": "index",
        "_type": "doc",
        "_id": "1",
        "_source": {
          "len": 1,
          "foo": "bar"
        },
        "_ingest": {
          "timestamp": "2018-11-06T09:12:12.310133Z"
        }
      }
    },
    {
      "doc": {
        "_index": "index",
        "_type": "_doc",
        "_id": "second",
        "_source": {
          "len": 6,
          "foo": "rab"
        },
        "_ingest": {
          "timestamp": "2018-11-06T09:12:12.310166Z"
        }
      }
    }
  ]
}

```

The advantage of this would be, that you will have really fast queries, as you do not need to invoke a script for each hit.

Hope this helps!

--Alex

---

<div class="post-metadata">

**Author:** ![m9aertner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/m9aertner/32/39588_2.png) [@m9aertner](https://discuss.elastic.co/u/m9aertner)\
**Post date:** [November 6, 2018, 10:45am UTC](https://discuss.elastic.co/t/how-to-access-id-from-painless-in-query-context/155211/3 "2018-11-06T10:45:49Z")

</div>

Hello Alexander, thanks for checking this, and your detailed reply!

Unfortunately, for ES 5.3.x the `doc['_id']` bit already produces the error `"Fielddata is not supported on field [_id] of type [_id]"`, no matter what's written after this. (One more reason to update...)

The idea with storing the length right away is nice. Alas, for _checking_ multiple million documents one would have to "query-by-update" or "\_reindex", which both require running a script again once per document, even if it's very simple.

We'll solve the issue by looking at the _upstream_ data from which the `_id` is generated.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 4, 2018, 10:45am UTC](https://discuss.elastic.co/t/how-to-access-id-from-painless-in-query-context/155211/4 "2018-12-04T10:45:58Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
