# Huge fields - Mapping best practice

**URL:** <https://discuss.elastic.co/t/huge-fields-mapping-best-practice/307261>\
**Category:** Elasticsearch\
**Created:** [June 15, 2022, 11:45am UTC](https://discuss.elastic.co/t/huge-fields-mapping-best-practice/307261 "2022-06-15T11:45:25Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![stevesimpson](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stevesimpson/32/41437_2.png) [@stevesimpson](https://discuss.elastic.co/u/stevesimpson)\
**Post date:** [June 15, 2022, 11:45am UTC](https://discuss.elastic.co/t/huge-fields-mapping-best-practice/307261/1 "2022-06-15T11:45:25Z")

</div>

Hey all 👋

I'm wondering what the best practice and recommendation would be for handling huge fields at scale? The situation I have is that I need to include an XML document (which is almost always massive) in the Kibana discover app when users submit searches. This XML document does not need to be searchable or indexed or anything, just needs to be viewable and ideally included in reports.

I’ve tried setting the mapping to `xml_field: {enabled: false, type: object}` which does not analyse and index the field but query performance is awful and will not scale. I am assuming that this is because the field is still included in `_source`?

If I add a `source filter` in `Index Management` then query performance is great again but I cannot view the field 🙃

Feels like I'm a bit stuck between a rock and a hard place here, having the XML response viewable as part of the search results is vital for our debugging needs and it's not possible to parse the XML contents before indexing using Logstash etc as the format is pretty unstructured.

Would love to get some advice and discussion around what steps I might be able to take to resolve this, I've asked this in the slack group but thought this might be a better forum.

---

<div class="post-metadata">

**Author:** ![intrepid1](https://avatars.discourse-cdn.com/v4/letter/i/dbc845/32.png) [@intrepid1](https://discuss.elastic.co/u/intrepid1)\
**Post date:** [June 15, 2022, 11:50am UTC](https://discuss.elastic.co/t/huge-fields-mapping-best-practice/307261/2 "2022-06-15T11:50:25Z")

</div>

For our massive fields we use the wildcard type.

You might find this useful to read.

[Find strings within strings faster with the Elasticsearch wildcard field | Elastic Blog](https://www.elastic.co/blog/find-strings-within-strings-faster-with-the-new-elasticsearch-wildcard-field)

---

<div class="post-metadata">

**Author:** ![stevesimpson](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stevesimpson/32/41437_2.png) [@stevesimpson](https://discuss.elastic.co/u/stevesimpson)\
**Post date:** [June 15, 2022, 12:19pm UTC](https://discuss.elastic.co/t/huge-fields-mapping-best-practice/307261/3 "2022-06-15T12:19:22Z")

</div>

Thanks for the reply @intrepid1

Wouldn't a wildcard type make the problem worse? The field is currently set to `enabled: false`

> The `enabled` setting, which can be applied only to the top-level mapping definition and to [`object`](https://www.elastic.co/guide/en/elasticsearch/reference/7.9/object.html) fields, causes Elasticsearch to skip parsing of the contents of the field entirely. The JSON can still be retrieved from the [`_source`](https://www.elastic.co/guide/en/elasticsearch/reference/7.9/mapping-source-field.html) field, but it is not searchable or stored in any other way

---

<div class="post-metadata">

**Author:** ![RabBit\_BR](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rabbit_br/32/82261_2.png) [@RabBit\_BR](https://discuss.elastic.co/u/RabBit_BR)\
**Post date:** [June 15, 2022, 2:12pm UTC](https://discuss.elastic.co/t/huge-fields-mapping-best-practice/307261/4 "2022-06-15T14:12:57Z")

</div>

Hi @stevesimpson

Maybe mapping [store](https://www.elastic.co/guide/en/elasticsearch/reference/current/mapping-store.html#mapping-store) make sense for you.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [June 21, 2022, 1:57am UTC](https://discuss.elastic.co/t/huge-fields-mapping-best-practice/307261/5 "2022-06-21T01:57:00Z")

</div>

How huge is huge.

---

<div class="post-metadata">

**Author:** ![stevesimpson](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stevesimpson/32/41437_2.png) [@stevesimpson](https://discuss.elastic.co/u/stevesimpson)\
**Post date:** [June 21, 2022, 10:06am UTC](https://discuss.elastic.co/t/huge-fields-mapping-best-practice/307261/6 "2022-06-21T10:06:34Z")

</div>

Hey @warkolm, typically around 50-100kb however they can be around1-2mb in size.

It's a non-typical use case for Elastic I know and not really what it's for, however, having the XML file as part of the event is vital for debugging. It's difficult to fully parse the XML on ingest because they can be unstructured, so XPath, Grok Logstash plugins can be prone to error.

I do think I have a workable solution though which seems to be reasonably performant, and this is actually using the source exclusions and setting the field to be mapped as a non-enabled object type.

We can still retrieve the field from Kibana discover using the "Show single document" button. It's not ideal as I can only look at a single XML doc at a time and we can't use reporting to export the field but it's better than nothing.

If you have any other thoughts or ideas I'm all ears 🙂

---

<div class="post-metadata">

**Author:** ![Tomo\_M](https://avatars.discourse-cdn.com/v4/letter/t/848f3c/32.png) [@Tomo\_M](https://discuss.elastic.co/u/Tomo_M)\
**Post date:** [June 21, 2022, 1:08pm UTC](https://discuss.elastic.co/t/huge-fields-mapping-best-practice/307261/7 "2022-06-21T13:08:05Z")

</div>

> query performance is awful

How awful is it? What is the query and how many it hits? Isn't it just take time to receive some mb messages?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 19, 2022, 1:08pm UTC](https://discuss.elastic.co/t/huge-fields-mapping-best-practice/307261/8 "2022-07-19T13:08:53Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
