# Performance of doc\_values field vs analysed field

**URL:** <https://discuss.elastic.co/t/performance-of-doc-values-field-vs-analysed-field/99445>\
**Category:** Elasticsearch\
**Created:** [September 5, 2017, 2:10pm UTC](https://discuss.elastic.co/t/performance-of-doc-values-field-vs-analysed-field/99445 "2017-09-05T14:10:26Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![ndtreviv](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ndtreviv/32/22494_2.png) [@ndtreviv](https://discuss.elastic.co/u/ndtreviv)\
**Post date:** [September 5, 2017, 2:10pm UTC](https://discuss.elastic.co/t/performance-of-doc-values-field-vs-analysed-field/99445/1 "2017-09-05T14:10:26Z")

</div>

Hi!

We have an elasticsearch that contains over half a billion documents that each have a `url` field that stores a URL.

The `url` field mapping currently has the settings:

```auto
{
    index: not_analyzed
    doc_values: true
    ...
}

```

We want our users to be able to search URLs, or portions of URLs without having to use wildcards.  
For example, taking the URL: `https://www.domain.com/part1/user@site/part2/part3.ext`

They should be able to bring back a matching document by searching:

- `part3.ext`
- `user@site`
- `part1`
- `part2/part3.ext`

The way I see it, we have two options:

1. Implement an analysed version of this field (which can no longer have `doc_values: true`) and do match querying instead of wildcards. This would also require using a custom analyser to leverage the `pattern` tokeniser to make the extracted terms correct (the standard tokeniser would split `user@site` into `user` and `site`).
2. Go through our database and for each document create a new field that is a list of URL parts. This field could have `doc_values: true` still so would be stored off-heap, and we could do term querying on exact field values instead of wildcards.

**My question is this:**  
Which is better for performance: having a list of variable lengths that has `doc_values` on, or having an analysed field? (ie: option 1 or option 2) OR is there an option 3 that would be even better yet?!

Thanks for your help!

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [September 6, 2017, 5:09pm UTC](https://discuss.elastic.co/t/performance-of-doc-values-field-vs-analysed-field/99445/2 "2017-09-06T17:09:25Z")

</div>

Will you really need to aggregate over parts of URLs? I suspect not. You could just index it twice: once with type:text and once with `type:keyword` and `index:false`. The first field will be useful to search over url parts and the second field will be be useful to aggregate entire urls.

---

<div class="post-metadata">

**Author:** ![ndtreviv](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ndtreviv/32/22494_2.png) [@ndtreviv](https://discuss.elastic.co/u/ndtreviv)\
**Post date:** [September 19, 2017, 2:18pm UTC](https://discuss.elastic.co/t/performance-of-doc-values-field-vs-analysed-field/99445/3 "2017-09-19T14:18:23Z")

</div>

So I tried it myself against 10 million documents.

I did one test with a `filename` field that was created by `copy_to` from two other URL-type fields (also set `store: false`), and used a custom analyser with a pattern tokeniser. This one I called `Filename Analysed`.

I did another test where I ran an enrichment over all 10 million documents to create a list field that contained all permutations of segments for both URL-type fields that I wanted to include. This one I called `Filename Array`.

Then I ran gatling tests against each index to see how they compared, whilst comparing netdata ([https://github.com/firehol/netdata](https://github.com/firehol/netdata)) CPU/Memory stats.

The results were interesting.

**Compare the centiles:**

 ![28](https://us1.discourse-cdn.com/elastic/original/3X/5/a/5a95e08c02d6e29cb1e2d77802620f5e537335b6.png)

In terms of netdata, the Filename Array index performed much better - it hardly registered usage on the CPU at all in comparison to the Filename Analysed index.

Both solutions solve the problem above, but Filename Array also lets me do exact substring matching (without the need for wildcards) whereas Filename Analysed does not.

There is a problem with Filename Array, though in that it nearly doubles the index size (bloats it by about 80%).

So now I'm wondering if I can write a custom tokenizer script that creates the same tokens that I was creating in my Filename Array enrichment?

ie: Taking the url above: `https://www.domain.com/part1/user@site/part2/part3.ext` I want the following tokens:

- `www.domain.com`
- `part1`
- `user@site`
- `part2`
- `part3.ext`
- `www.domain.com/part1`
- `www.domain.com/part1/user@site`
- `www.domain.com/part1/user@site/part2`
- `www.domain.com/part1/user@site/part2/part3.ext`
- `part1/user@site/part2/part3.ext`
- `user@site/part2/part3.ext`
- `part2/part3.ext`
- `part1/user@site/part2`
- `part1/user@site`
- `user@site/part2`

My regex capabilities don't live up to this requirement! Is it possible to write a java plugin that can be used as a custom tokenizer? Or a groovy script? Or can someone suggest a regular expression that might work?!

Thanks for any help!

---

<div class="post-metadata">

**Author:** ![ndtreviv](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ndtreviv/32/22494_2.png) [@ndtreviv](https://discuss.elastic.co/u/ndtreviv)\
**Post date:** [September 20, 2017, 9:31am UTC](https://discuss.elastic.co/t/performance-of-doc-values-field-vs-analysed-field/99445/4 "2017-09-20T09:31:56Z")

</div>

OK, I've discovered that it should be possible to write a plugin that contains a custom tokenizer.  
It's not a process that's documented anywhere, and I'm hitting problems: [Building a custom tokenizer: "Could not find suitable constructor"](https://discuss.elastic.co/t/building-a-custom-tokenizer-could-not-find-suitable-constructor/101138)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 18, 2017, 9:32am UTC](https://discuss.elastic.co/t/performance-of-doc-values-field-vs-analysed-field/99445/5 "2017-10-18T09:32:07Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
