# Elasticsearch Query Error

**URL:** <https://discuss.elastic.co/t/elasticsearch-query-error/332501>\
**Category:** Elasticsearch\
**Created:** [May 4, 2023, 7:10am UTC](https://discuss.elastic.co/t/elasticsearch-query-error/332501 "2023-05-04T07:10:51Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mohsin\_Ashraf](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mohsin_ashraf/32/104333_2.png) [@Mohsin\_Ashraf](https://discuss.elastic.co/u/Mohsin_Ashraf)\
**Post date:** [May 4, 2023, 7:10am UTC](https://discuss.elastic.co/t/elasticsearch-query-error/332501/1 "2023-05-04T07:10:51Z")

</div>

Hi,  
I'm facing this error While querying the data. is there any solution for this?

error:  
query\_shard\_exception

Reason

failed to create query: field expansion for [\*] matches too many fields, limit: 1024, got: 20703

Index uuid

kElrsci\_TC2hgP-SYXoFKw

Index

tsdb-2023.05.04

Caused by type

illegal\_argument\_exception

Caused by reason

field expansion for [\*] matches too many fields, limit: 1024, got: 20703

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 4, 2023, 7:21am UTC](https://discuss.elastic.co/t/elasticsearch-query-error/332501/2 "2023-05-04T07:21:26Z")

</div>

What is the query that is causing the error? What is the mapping of that index?

Which version of Elasticsearch are you using?

---

<div class="post-metadata">

**Author:** ![Mohsin\_Ashraf](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mohsin_ashraf/32/104333_2.png) [@Mohsin\_Ashraf](https://discuss.elastic.co/u/Mohsin_Ashraf)\
**Post date:** [May 4, 2023, 7:37am UTC](https://discuss.elastic.co/t/elasticsearch-query-error/332501/3 "2023-05-04T07:37:37Z")

</div>

Thank you.  
I'm using Elastic 7.17  
here is the index setting: Mapping is dynamic.

> Query  
> host.keyword : "openstack-cluster-non-exposed" AND check\_command : "check\_serversOS" AND check\_result._cpu_.value:\*

> Setting  
> {  
> "tsdb-2023.04.14" : {  
> "settings" : {  
> "index" : {  
> "routing" : {  
> "allocation" : {  
> "include" : {  
> "\_tier\_preference" : "data\_content"  
> }  
> }  
> },  
> "mapping" : {  
> "nested\_fields" : {  
> "limit" : "60000"  
> },  
> "depth" : {  
> "limit" : "60000"  
> },  
> "total\_fields" : {  
> "limit" : "6000000"  
> },  
> "dimension\_fields" : {  
> "limit" : "6000"  
> }  
> },  
> "number\_of\_shards" : "5",  
> "provided\_name" : "tsdb-2023.04.14",  
> "creation\_date" : "1681412458823",  
> "number\_of\_replicas" : "0",  
> "uuid" : "h40EMKtxSDmxAXppDTWHhQ",  
> "version" : {  
> "created" : "7170999"  
> }  
> }  
> }

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 4, 2023, 7:40am UTC](https://discuss.elastic.co/t/elasticsearch-query-error/332501/4 "2023-05-04T07:40:23Z")

</div>

> [@Mohsin\_Ashraf](#):
>
> "mapping" : {  
> "nested\_fields" : {  
> "limit" : "60000"  
> },  
> "depth" : {  
> "limit" : "60000"  
> },  
> "total\_fields" : {  
> "limit" : "6000000"  
> },  
> "dimension\_fields" : {  
> "limit" : "6000"  
> }  
> },

What does a sample document look like? Why have you overridden these parameters?

These type of limits are generally in present for a good reason, so my increasing these dramatically you are asking for trouble. I would recommend you reconsider how you store data in Elasticsearch so you can go back to the default settings.

---

<div class="post-metadata">

**Author:** ![Mohsin\_Ashraf](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mohsin_ashraf/32/104333_2.png) [@Mohsin\_Ashraf](https://discuss.elastic.co/u/Mohsin_Ashraf)\
**Post date:** [May 4, 2023, 7:47am UTC](https://discuss.elastic.co/t/elasticsearch-query-error/332501/5 "2023-05-04T07:47:43Z")

</div>

My use case requires me to increase the field limit. It started off with around 20k fields but has since become 100k, and it's estimated to go to around 600k. The retention policy is for around 2 months.

Since newer data wasn't being written (fields weren't being added/saved), I increased various fields until the issue was resolved.

My current issue, I that I require searching my data using wildcards. I can search just fine without them, but I get errors when I use a wildcard (such as what I mentioned earlier). Even when I test on a search result I know only has a few dozen or so hits......i get the same error.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 4, 2023, 7:50am UTC](https://discuss.elastic.co/t/elasticsearch-query-error/332501/6 "2023-05-04T07:50:39Z")

</div>

Why is the field count growing so much? What does a sample document look like?

If you are going to work with Elasticsearch I believe you need to change how you use it and change the document structure. Your settings are so far beyond what is recommended that I am not surprised that you are running into problems. I would also not be surprised if you start running into cluster stability and performance issues as the cluster state is going to be quite large with mappings that size.

If you can share some sample documents the community might be able to provide some suggestions on how to best restructure the data to align with how Elasticsearch works.

If you can not change the structure of the data Elasticsearch may not be a suitable tool to use. Maybe you should look into using something else?

---

<div class="post-metadata">

**Author:** ![Mohsin\_Ashraf](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mohsin_ashraf/32/104333_2.png) [@Mohsin\_Ashraf](https://discuss.elastic.co/u/Mohsin_Ashraf)\
**Post date:** [May 4, 2023, 8:01am UTC](https://discuss.elastic.co/t/elasticsearch-query-error/332501/7 "2023-05-04T08:01:43Z")

</div>

Fields are being fetched from icinga. Its not growing fast, it more that additional hosts and resources are being added. Moreover, many of the hosts are such that they can add/remove/modify their individual workloads a couple hundred times a day if required, all of which is being monitored, and appropriate perf data being saved on elastic.

I cant reveal confidential sample data, but I can provide a matching pattern. The data below is the data provided by icinga, and written to elastic using elasticwriter.

> Sample data pattern (provided by icinga):  
> rgen\_11\_1=1 rgen\_11\_2=2 rgen\_11\_3=3 rgen\_11\_4=4 rgen\_11\_5=5 rgen\_11\_6=6 rgen\_11\_7=7 rgen\_11\_8=8 rgen\_11\_9=9 rgen\_11\_10=10 rgen\_11\_11=11 rgen\_11\_12=12 rgen\_11\_13=13 rgen\_11\_14=14 rgen\_11\_15=15 rgen\_11\_16=16 rgen\_11\_17=17 rgen\_11\_18=18 rgen\_11\_19=19 rgen\_11\_20=20 rgen\_11\_21=21 rgen\_11\_22=22 rgen\_11\_23=23 rgen\_11\_24=24 rgen\_11\_25=25 rgen\_11\_26=26 rgen\_11\_27=27 rgen\_11\_28=28 rgen\_11\_29=29 rgen\_11\_30=30 rgen\_11\_31=31 rgen\_11\_32=32 rgen\_11\_33=33 rgen\_11\_34=34 rgen\_11\_35=35 rgen\_11\_36=36 rgen\_11\_37=37 rgen\_11\_38=38 rgen\_11\_39=39 rgen\_11\_40=40 rgen\_11\_41=41 rgen\_11\_42=42 rgen\_11\_43=43 rgen\_11\_44=44 rgen\_11\_45=45 rgen\_11\_46=46 rgen\_11\_47=47 rgen\_11\_48=48 rgen\_11\_49=49 rgen\_11\_50=50 rgen\_11\_51=51 rgen\_11\_52=52 rgen\_11\_53=53 rgen\_11\_54=54 rgen\_11\_55=55 rgen\_11\_56=56 rgen\_11\_57=57 rgen\_11\_58=58 rgen\_11\_59=59 rgen\_11\_60=60 rgen\_11\_61=61 rgen\_11\_62=62 rgen\_11\_63=63 rgen\_11\_64=64 rgen\_11\_65=65 rgen\_11\_66=66 rgen\_11\_67=67 rgen\_11\_68=68 rgen\_11\_69=69 rgen\_11\_70=70 rgen\_11\_71=71 rgen\_11\_72=72 rgen\_11\_73=73 rgen\_11\_74=74 rgen\_11\_75=75 rgen\_11\_76=76 rgen\_11\_77=77 rgen\_11\_78=78 rgen\_11\_79=79 rgen\_11\_80=80 rgen\_11\_81=81 rgen\_11\_82=82 rgen\_11\_83=83 rgen\_11\_84=84 rgen\_11\_85=85 rgen\_11\_86=86 rgen\_11\_87=87 rgen\_11\_88=88 rgen\_11\_89=89 rgen\_11\_90=90 rgen\_11\_91=91 rgen\_11\_92=92 rgen\_11\_93=93 rgen\_11\_94=94 rgen\_11\_95=95 rgen\_11\_96=96 rgen\_11\_97=97 rgen\_11\_98=98 rgen\_11\_99=99 rgen\_11\_100=100 rgen\_11\_101=101 rgen\_11\_102=102 rgen\_11\_103=103 rgen\_11\_104=104 rgen\_11\_105=105 rgen\_11\_106=106 rgen\_11\_107=107 rgen\_11\_108=108 rgen\_11\_109=109 rgen\_11\_110=110 rgen\_11\_111=111 rgen\_11\_112=112 rgen\_11\_113=113 rgen\_11\_114=114 rgen\_11\_115=115 rgen\_11\_116=116 rgen\_11\_117=117 rgen\_11\_118=118 rgen\_11\_119=119 rgen\_11\_120=120 rgen\_11\_121=121 rgen\_11\_122=122 rgen\_11\_123=123 rgen\_11\_124=124 rgen\_11\_125=125 rgen\_11\_126=126 rgen\_11\_127=127 rgen\_11\_128=128 rgen\_11\_129=129 rgen\_11\_130=130 rgen\_11\_131=131 rgen\_11\_132=132 rgen\_11\_133=133 rgen\_11\_134=134 rgen\_11\_135=135 rgen\_11\_136=136 rgen\_11\_137=137

I need to run queries along the following lines:

1. Get all data where the field begins with rgen, and ends with 5 (so, rgen\_11\_5, rgen\_10\_5, etc)
2. Get all data that contains _11_ in the middle (so rgen\_11\_1, rgen\_11\_2......so on)
3. Get all data where the value must be 15 (so rgen\_5\_11=15, rgen\_8\_1255=15 ........ so on)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 4, 2023, 8:06am UTC](https://discuss.elastic.co/t/elasticsearch-query-error/332501/8 "2023-05-04T08:06:18Z")

</div>

How many of these data sets are generated per day?

Do you have any other data associated with this data set, e.g. timestamp, source id etc?

Elasticsearch works best with few keys and many values, so one way to solve this would be to break this data set up into many documents looking something like this:

```auto
{
  "set_id": "abcdef1234",
  "key": "rgen_11_99",
  "value": 99,
  "@timestamp": "2023-05-04T08:00:00Z",
  ...
}

```

You can then map the `key` field as `keyword` (for exact match) and have a [wildcard](https://www.elastic.co/guide/en/elasticsearch/reference/8.7/keyword.html#wildcard-field-type) [multi-field](https://www.elastic.co/guide/en/elasticsearch/reference/8.7/mapping-types.html#types-multi-fields) underneath for efficient wildcard matching.

If the values are all related and you do need to get them all in one query you can also consider storing the key-value pair documents above as a [nested](https://www.elastic.co/guide/en/elasticsearch/reference/8.7/nested.html) field. This will make updates more expensive, but if your data is immutable it might be an option although it will complicate query syntax a bit.

---

<div class="post-metadata">

**Author:** ![Mohsin\_Ashraf](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mohsin_ashraf/32/104333_2.png) [@Mohsin\_Ashraf](https://discuss.elastic.co/u/Mohsin_Ashraf)\
**Post date:** [May 4, 2023, 8:21am UTC](https://discuss.elastic.co/t/elasticsearch-query-error/332501/9 "2023-05-04T08:21:35Z")

</div>

I understand what your saying, but my scenario is different. As I'm limited to using elasticwriter to grab data from icinga, my data is being saved in the format below:

> Sample format  
> {  
> "host": "host1",  
> "server": "server1",  
> "service": "service1",  
> "rgen\_11\_99": 1,  
> "rgen\_11\_100": 2,  
> "rgen\_11\_101": 3,  
> "rgen\_12\_45": 4,  
> "@timestamp": "2023-05-04T08:00:00Z",  
> ...  
> }

Essentially the situation is flipped, where I have fewer documents, but more data per document.

On average, i would have around 3k documents per 5 min span, where each document could have between 20 and 1000 values. (only around 100-150 such documents with over 50 values )

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 4, 2023, 8:28am UTC](https://discuss.elastic.co/t/elasticsearch-query-error/332501/10 "2023-05-04T08:28:35Z")

</div>

I do not think that approach will work nor scale. You are likely to face a lot of different issues down the line with that approach. If you are going to use Elasticsearch you will need to find a way to transform the data. Logstash would be able to do this, so if you can find a way to direct your data to Logstash that could solve the problem. You might be able to assign a [default ingest pipeline](https://www.elastic.co/guide/en/elasticsearch/reference/8.7/ingest.html) to the indices through an index template and perform the transformation that way. As far as I know ingest pipelines are not able to split documents, but it may be possible to transform it into a nested structure using e.g. a script processor so it looks something like this:

```auto
{
  "host": "host1",
  "server": "server1",
  "service": "service1",
  "data": [
    {"key": "rgen_11_99", "value": 1},
    {"key": "rgen_11_100", "value": 2},
    {"key": "rgen_11_101", "value": 3},
    {"key": "rgen_12_45", "value": 4}
  ],
  "@timestamp": "2023-05-04T08:00:00Z",
  ...
}

```

This is probably what I would try doing as it still allows your ingest process to write the current data directly to the index in the current format.

I think this should work and also scale and perform quite well. It would also allow you to have a small strict mapping, which avoids a lot of cluster state updates and increases stability.

---

<div class="post-metadata">

**Author:** ![Mohsin\_Ashraf](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mohsin_ashraf/32/104333_2.png) [@Mohsin\_Ashraf](https://discuss.elastic.co/u/Mohsin_Ashraf)\
**Post date:** [May 4, 2023, 10:37am UTC](https://discuss.elastic.co/t/elasticsearch-query-error/332501/11 "2023-05-04T10:37:39Z")

</div>

Noted. Thank you! However, that is an issue for a later time. The current issue is, that I am unable to use wildcards in my search. The large data shouldnt be a problem.

Infact, I have attempted to run the wildcard query on a document I know only has 4 values, and it still failed.

Moreover, I have setup a similar environment (trying multiple versions) on a separate VM, and it accepts wildcard searches. But it doesn't work on the environment I'm currently developing in.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [May 4, 2023, 10:39am UTC](https://discuss.elastic.co/t/elasticsearch-query-error/332501/12 "2023-05-04T10:39:38Z")

</div>

> [@Mohsin\_Ashraf](#):
>
> The large data shouldnt be a problem.

It is.

I would recommend you change the approach immediately. I do not have the time nor energy to help troubleshoot your query issue as it uses an IMHO flawed approach and in my mind is a waste of time. Good luck!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 1, 2023, 10:40am UTC](https://discuss.elastic.co/t/elasticsearch-query-error/332501/13 "2023-06-01T10:40:39Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
