# Accessing Unique Token or Term ID

**URL:** <https://discuss.elastic.co/t/accessing-unique-token-or-term-id/35947>\
**Category:** Elasticsearch\
**Created:** [November 30, 2015, 8:12pm UTC](https://discuss.elastic.co/t/accessing-unique-token-or-term-id/35947 "2015-11-30T20:12:56Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![neal](https://avatars.discourse-cdn.com/v4/letter/n/49beb7/32.png) [@neal](https://discuss.elastic.co/u/neal)\
**Post date:** [November 30, 2015, 8:12pm UTC](https://discuss.elastic.co/t/accessing-unique-token-or-term-id/35947/1 "2015-11-30T20:12:56Z")

</div>

Does ES assign a token id to each token / term? I imagine that it does.

Is there any way to get the token ID? I've tried through the term vector API, but it seems that IDs are not a valid field.

So, I'd like to do something like:

$ curl [http://localhost:9200/twitter/\_term?t='test'](http://localhost:9200/twitter/_term?t='test')

The response back would be something like

'{ 'term':'test', '\_id':12938}

Any ideas?

Thanks!

---

<div class="post-metadata">

**Author:** ![vtst2412](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vtst2412/32/6228_2.png) [@vtst2412](https://discuss.elastic.co/u/vtst2412)\
**Post date:** [November 30, 2015, 8:40pm UTC](https://discuss.elastic.co/t/accessing-unique-token-or-term-id/35947/2 "2015-11-30T20:40:43Z")

</div>

I have never seen the `_term` function before. Why aren't you using `_search`?

How about:

`GET /twitter/_search?q=term:test`

FYI: `_id` is an absolutely valid field to search on.

---

<div class="post-metadata">

**Author:** ![neal](https://avatars.discourse-cdn.com/v4/letter/n/49beb7/32.png) [@neal](https://discuss.elastic.co/u/neal)\
**Post date:** [November 30, 2015, 9:19pm UTC](https://discuss.elastic.co/t/accessing-unique-token-or-term-id/35947/3 "2015-11-30T21:19:13Z")

</div>

> [@vtst2412](#):
>
> twitter/\_search?q=term:test

I just contrived of the \_term function just now as an example of what I'm looking for. I don't think that it exists.

I'm not looking for a _document id_. I'm looking for a _token id_. There must be a token - id dictionary somewhere, and I'd like to get the id of a token.

---

<div class="post-metadata">

**Author:** ![colings86](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/colings86/32/44960_2.png) [@colings86](https://discuss.elastic.co/u/colings86)\
**Post date:** [December 1, 2015, 9:11am UTC](https://discuss.elastic.co/t/accessing-unique-token-or-term-id/35947/4 "2015-12-01T09:11:29Z")

</div>

AFAIK there are no term IDs in the index. Terms can be uniquely identified by there fieldname and value. What problem are you wanting to solve? What are you wanting to use the 'token ID' for? Maybe there is another way to achieve what you are trying to do.

---

<div class="post-metadata">

**Author:** ![neal](https://avatars.discourse-cdn.com/v4/letter/n/49beb7/32.png) [@neal](https://discuss.elastic.co/u/neal)\
**Post date:** [December 1, 2015, 2:15pm UTC](https://discuss.elastic.co/t/accessing-unique-token-or-term-id/35947/5 "2015-12-01T14:15:26Z")

</div>

Thank you for your reply!

Well, I'm sure that the underlying Lucene Index converts terms and tokens to a uniq id over all of the documents. Either via hashing or just counting - there must be an token ID symbol table somewhere.

We're using Elasticsearch as an document store / index for alot of our machine learning / NLP analytics that we run via Spark. Instead of building and maintaining our own dictionary ID store (or worse yet - re assign ids for each subset of corpora that we work with ), we like to keep all the docs in ES and have a single source for tokenization / token ids.

While the TermVectors api is useful for us, it's not enough yet. We also would need tokenized docs. Is there a way to access those outside of term vectors? Something similar to UIMA CAS?

Keeping all of our docs / index / tokenization in a single place greatly simplifies our analytic processing. So, maybe ES just isn't the way to go, or we just build a separate token ID symbol table somewhere else.

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [December 1, 2015, 5:30pm UTC](https://discuss.elastic.co/t/accessing-unique-token-or-term-id/35947/6 "2015-12-01T17:30:26Z")

</div>

Lucene and Elasticsearch do not assign numeric ids to tokens, they just use the term bytes to identify terms. For instance if you index an analyzed string, it will be broken into tokens that are essentially a char[] and then Lucene stores them in the index using their utf-8 representation.

---

<div class="post-metadata">

**Author:** ![neal](https://avatars.discourse-cdn.com/v4/letter/n/49beb7/32.png) [@neal](https://discuss.elastic.co/u/neal)\
**Post date:** [December 1, 2015, 6:04pm UTC](https://discuss.elastic.co/t/accessing-unique-token-or-term-id/35947/7 "2015-12-01T18:04:37Z")

</div>

I see!!!

Thank you for letting me know that! I was unaware of that approach before.  
This is really helpful for me to know.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:34pm UTC](https://discuss.elastic.co/t/accessing-unique-token-or-term-id/35947/8 "2017-07-05T23:34:22Z")

</div>


