# Storing Very Large Text Field in Elasticsearch

**URL:** <https://discuss.elastic.co/t/storing-very-large-text-field-in-elasticsearch/380428>\
**Category:** Elasticsearch\
**Created:** [July 24, 2025, 12:57pm UTC](https://discuss.elastic.co/t/storing-very-large-text-field-in-elasticsearch/380428 "2025-07-24T12:57:26Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![kdwolf](https://avatars.discourse-cdn.com/v4/letter/k/df788c/32.png) [@kdwolf](https://discuss.elastic.co/u/kdwolf)\
**Post date:** [July 24, 2025, 12:57pm UTC](https://discuss.elastic.co/t/storing-very-large-text-field-in-elasticsearch/380428/1 "2025-07-24T12:57:26Z")

</div>

Hello all,

I’d appreciate your input on the following, please:

The business has a requirement to extract a lengthy string (text) from an external object — for example, a MOBI file — and store it in an Elasticsearch index as a searchable field. The challenge is that the string could be quite large, potentially in the region of 40–50 MB (which I think exceeds Lucene limits significantly) , and the business expects to be able to search within this field.

I imagine this isn’t the first time such a scenario has arisen, so I’m keen to understand what best practices are in place for handling this type of requirement.

Many thanks in advance for your insights.

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [July 24, 2025, 4:44pm UTC](https://discuss.elastic.co/t/storing-very-large-text-field-in-elasticsearch/380428/2 "2025-07-24T16:44:55Z")

</div>

Hi @kdwolf

I'm a little confused. Are you referring to a single token that is 40-50MB, or a text field containing multiple strings/tokens (this seems to be what you are implying)? If so, that is fine to store 40-50MB in a single `text` field in Elastic.

Perhaps you are confusing the single token limit in Lucene, which is 32K

Lucene still has a document limit (field limit) of about 2GB.

That said, there are other areas of concern, like actually returning the data due to HTTP limits, etc.

Perhaps take a look at this thread

> [@Maximum characters limit text fields in elasticsearch](https://discuss.elastic.co/t/maximum-characters-limit-text-fields-in-elasticsearch/376616/2):
>
> I think this documentation largely answers your question: [General recommendations | Elasticsearch Guide [8.17] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/general-recommendations.html) Other fields (like keyword) have further limitations, but for text fields I think it's however much data you can send to ES up to the http limit which is about 100MB or if configured higher we eventually disallow blobs or text fields above 2GB. In practice if you consider for instance Moby Dick I think that's about ~2MB of text. So it's probably safe that anything human-con…

Minor Side Note: Most of the text in a .mobi file is binary so not sure if that is just and example ... The Header etc is text but the actual text of the document is binary AFAIK.

---

<div class="post-metadata">

**Author:** ![kdwolf](https://avatars.discourse-cdn.com/v4/letter/k/df788c/32.png) [@kdwolf](https://discuss.elastic.co/u/kdwolf)\
**Post date:** [July 24, 2025, 8:17pm UTC](https://discuss.elastic.co/t/storing-very-large-text-field-in-elasticsearch/380428/3 "2025-07-24T20:17:11Z")

</div>

Thank you, @stephenb  
Yes, I am referring to the extracted text like  
"_Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur. Excepteur sint occaecat cupidatat non proident, sunt in culpa qui officia deserunt mollit anim id est laborum._"

which should be searchable i.e. if the customer is looking for "_Excepteur sint occaecat cupidatat non proident_" phrase it should return all the documents with this phrase, as one would expect.

We face significant performance issue when storing very long text and trying to search amongst millions of documents and I wonder if there is a recommended approach to optimise either the query or the storage (I guess nothing can be done with the latter) when querying this type of a text field?

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [July 24, 2025, 9:50pm UTC](https://discuss.elastic.co/t/storing-very-large-text-field-in-elasticsearch/380428/4 "2025-07-24T21:50:31Z")

</div>

> [@kdwolf](#):
>
> We face significant performance issue when storing very long text and trying to search amongst millions of documents and I wonder if there is a recommended approach to optimise either the query or the storage (I guess nothing can be done with the latter) when querying this type of a text field?

Clearly a lot of detail to dig into ....

A couple of things come to mind...

1st Perhaps it may be that the search is quite fast, finding the document(s) that match the query, but then pulling those Gigantic Documents back out of Lucene, marshaling them up and then sending them back to the client... yeah... that could be quite non-performant.

We are doing on the Semantic search side is chunking up these big docs into digestible parts...

Does your user really want to pull back the entire 50 MBs to find the 3 sentences or paragraph that matches, along with the document name (could be section etc) ?

So perhaps a chunking, perhaps you do not want to do semantic search but you could do similar for normal search

> **[Chunking large documents via ingest pipelines and nested vectors -...](https://www.elastic.co/search-labs/blog/chunking-via-ingest-pipelines)**
>
> Learn how to chunk large documents using ingest pipelines and nested vectors in Elasticsearch for easy passage search in vector search.

`semantic_text`

> **[Semantic text field type | Reference](https://www.elastic.co/docs/reference/elasticsearch/mapping-reference/semantic-text)**
>
> The semantic\_text field type automatically generates embeddings for text content using an inference endpoint. Long passages are automatically chunked...
