# What are the potential drawbacks of changing ( "ignore\_above": 4000) and is there a better way to recommend it?

**URL:** <https://discuss.elastic.co/t/what-are-the-potential-drawbacks-of-changing-ignore-above-4000-and-is-there-a-better-way-to-recommend-it/369951>\
**Category:** Elasticsearch\
**Created:** [November 2, 2024, 9:35pm UTC](https://discuss.elastic.co/t/what-are-the-potential-drawbacks-of-changing-ignore-above-4000-and-is-there-a-better-way-to-recommend-it/369951 "2024-11-02T21:35:12Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![monther](https://avatars.discourse-cdn.com/v4/letter/m/5fc32e/32.png) [@monther](https://discuss.elastic.co/u/monther)\
**Post date:** [November 2, 2024, 9:35pm UTC](https://discuss.elastic.co/t/what-are-the-potential-drawbacks-of-changing-ignore-above-4000-and-is-there-a-better-way-to-recommend-it/369951/1 "2024-11-02T21:35:12Z")

</div>

I have a field in which I store very textual data that may reach 4000 characters, which are symbols, numbers, and (/,\*,-,\_,.). When I store this data in the default state, it is normal and receives the data, but when I search in term, it does not return any value. So I changed "ignore\_above" to 4000 so that I can search accurately using term, but I notice that there is a delay and the use of more space.

Is there another way to store this data with a length of 4000 characters and search for it accurately using term?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 3, 2024, 7:56am UTC](https://discuss.elastic.co/t/what-are-the-potential-drawbacks-of-changing-ignore-above-4000-and-is-there-a-better-way-to-recommend-it/369951/2 "2024-11-03T07:56:56Z")

</div>

> [@monther](#):
>
> but I notice that there is a delay and the use of more space.

I would say that is expected as you are storing a high-cardinality field with very large terms which take up a lot of space. If you are searching based on the full term, might it be an option to store a hash of the field in a separate field and search based on this?

---

<div class="post-metadata">

**Author:** ![monther](https://avatars.discourse-cdn.com/v4/letter/m/5fc32e/32.png) [@monther](https://discuss.elastic.co/u/monther)\
**Post date:** [November 3, 2024, 3:17pm UTC](https://discuss.elastic.co/t/what-are-the-potential-drawbacks-of-changing-ignore-above-4000-and-is-there-a-better-way-to-recommend-it/369951/3 "2024-11-03T15:17:44Z")

</div>

> [@Christian\_Dahlqvist](#):
>
> I would say that is expected as you are storing a high-cardinality field with very large terms which take up a lot of space. If you are searching based on the full term, might it be an option to store a hash of the field in a separate field and search based on this?

No, I cannot divide the data because I rely on it in my search operations to a large extent, and at the same time I want the search to be as fast and accurate as possible.

---

<div class="post-metadata">

**Author:** ![monther](https://avatars.discourse-cdn.com/v4/letter/m/5fc32e/32.png) [@monther](https://discuss.elastic.co/u/monther)\
**Post date:** [November 3, 2024, 3:19pm UTC](https://discuss.elastic.co/t/what-are-the-potential-drawbacks-of-changing-ignore-above-4000-and-is-there-a-better-way-to-recommend-it/369951/4 "2024-11-03T15:19:34Z")

</div>

No, I cannot divide the data because I rely on it in my search operations to a large extent, and at the same time I want the search to be as fast and accurate as possible.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [November 4, 2024, 4:30am UTC](https://discuss.elastic.co/t/what-are-the-potential-drawbacks-of-changing-ignore-above-4000-and-is-there-a-better-way-to-recommend-it/369951/5 "2024-11-04T04:30:16Z")

</div>

I did not suggest dividing it but rather calculate a hash and store that in a separate field and use that for exact term lookup. This requires you to hash the query string as well and may result in hash collisions, although you may reduce the risk of that by selecting an appropriate hash function.

> [@monther](#):
>
> Is there another way to store this data with a length of 4000 characters and search for it accurately using term?

No, not that I am aware of.
