# Error on missing field query

**URL:** <https://discuss.elastic.co/t/error-on-missing-field-query/5688>\
**Category:** Elasticsearch\
**Created:** [October 26, 2011, 12:06pm UTC](https://discuss.elastic.co/t/error-on-missing-field-query/5688 "2011-10-26T12:06:29Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![vpunski](https://avatars.discourse-cdn.com/v4/letter/v/54ee81/32.png) [@vpunski](https://discuss.elastic.co/u/vpunski)\
**Post date:** [October 26, 2011, 12:06pm UTC](https://discuss.elastic.co/t/error-on-missing-field-query/5688/1 "2011-10-26T12:06:29Z")

</div>

Hi,  
6 nodes working cluster, 26M entries, with following configuration:

Index definition:  
curl -XPUT "[http://HOST:9200/my\_index/\_settings](http://HOST:9200/my_index/_settings)" -d '{  
index: {  
number\_of\_shards: 10,  
number\_of\_replicas: 3,  
"analysis": {  
"analyzer": {  
"parent\_hierarchy\_analyzer": {  
"type": "custom",  
"tokenizer": "path\_hierarchy"  
}  
}  
}  
}  
}'

Mapping definition:  
curl -XPUT "[http://HOST:9200/my\_index/my\_object/\_mapping?pretty=true](http://HOST:9200/my_index/my_object/_mapping?pretty=true)" -  
d '  
{  
"infoclone": {  
"properties": {  
"parent\_hierarchy": {  
"type": "string",  
"store": "no",  
"omit\_term\_freq\_and\_positions" : true,  
"analyzer": "parent\_hierarchy\_analyzer",  
"index": "analyzed",  
"omit\_norms" : true,  
"boost" : 1.0,  
"term\_vector" : "no"  
}  
}  
}  
}  
'

I'm trying to add additional field for every object, so I use filter  
below to get all of the objects, without "parent\_hierarchy" field, to  
update it.

curl -XGET [http://HOST:9200/my\_index/my\_object/\_search?pretty=1](http://HOST:9200/my_index/my_object/_search?pretty=1) -d '{  
"from" : 0,  
"size" : 1000,  
"query" : {  
"constant\_score" : {  
"filter" : {  
"bool" : {  
"must" : {  
"missing" : {  
"field" : "parent\_hierarchy"  
}  
}  
}  
}  
}  
}  
}  
'  
I've successfully updated ~16M entries, (from java client, using above  
query, bulk update, refresh=true)

Now, every time I execute the query, I get two errors in response  
body:

1. Every query the IP changes  
{  
"took" : 91,  
"timed\_out" : false,  
"\_shards" : {  
"total" : 10,  
"successful" : 9,  
"failed" : 1,  
"failures" : [ {  
"status" : 500,  
"reason" : "RemoteTransportException[[Schmidt, Johann][inet[/  
10.11.10.74:9300]][search/phase/fetch/id]]; nested:  
FieldReaderException[Invalid numeric type: 38]; "  
} ]  
},  
"hits" : {  
"total" : 686244,  
"max\_score" : 1.0,  
"hits" : []  
}  
}

2. Another version of response  
{  
"took" : 49,  
"timed\_out" : false,  
"\_shards" : {  
"total" : 10,  
"successful" : 9,  
"failed" : 1,  
"failures" : [ {  
"status" : 500,  
"reason" : "FieldReaderException[Invalid numeric type: 38]"  
} ]  
},  
"hits" : {  
"total" : 686244,  
"max\_score" : 1.0,  
"hits" : []  
}  
}

Health status below:  
curl -s -XGET '[http://HOST:9200/\_cluster/health?pretty=1](http://HOST:9200/_cluster/health?pretty=1)'  
{  
"cluster\_name" : "CMWELL\_INDEX\_CLUSTER",  
"status" : "green",  
"timed\_out" : false,  
"number\_of\_nodes" : 12,  
"number\_of\_data\_nodes" : 6,  
"active\_primary\_shards" : 10,  
"active\_shards" : 30,  
"relocating\_shards" : 0,  
"initializing\_shards" : 0,  
"unassigned\_shards" : 0  
}

elsticsearch.yml:

cluster.name : MY\_INDEX

path:  
home: data/elasticsearch  
logs: data/elasticsearch/logs

gateway:  
recover\_after\_nodes: 5  
recover\_after\_time: 1m  
expected\_nodes: 6  
local:  
initial\_shards: 1

#default 10% of memory  
#indices.memory.index\_buffer\_size : 1024m  
#default false  
index.compound\_format : true  
#default 1s  
index.refresh\_interval : 10s  
#defualt 128  
index.term\_index\_interval: 128

#Merging  
#default 10  
index.merge.policy.merge\_factor: 30  
#default 1.6mb  
index.merge.policy.min\_merge\_size: 16mb  
#default unbounded  
#index.merge.policy.max\_merge\_size: 1024mb  
#default unbounded  
#index.merge.policy.maxMergeDocs

#Transaction log settings  
#After how many operations to flush/ Defaults to 20000/  
#index.translog.flush\_threshold\_ops: 20000

#Once the translog hits this size, a flush will happen/ Defaults to  
500mb/  
#index.translog.flush\_threshold\_size

#The period with no flush happening to force a flush/ Defaults to 60m/  
#index.translog.flush\_threshold\_period

#Cache configurations

#defualt 20%  
indices.cache.filter.size: 10%

#defualt -1  
#1 entry ~1MB  
#index.cache.filter.max\_size: 100

#defualt -1  
index.cache.filter.expire: 1m

#defualt -1  
#index.cache.field.max\_size: -1

#default -1  
index.cache.field.expire: 1m

Every idea will be appreciated.  
Thanks

---

<div class="post-metadata">

**Author:** ![vpunski](https://avatars.discourse-cdn.com/v4/letter/v/54ee81/32.png) [@vpunski](https://discuss.elastic.co/u/vpunski)\
**Post date:** [October 26, 2011, 12:07pm UTC](https://discuss.elastic.co/t/error-on-missing-field-query/5688/2 "2011-10-26T12:07:55Z")

</div>

Please note:  
"total" : 10,  
"successful" : 9,

On shard fails.

On Oct 26, 2:06 pm, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:

> Hi,  
> 6 nodes working cluster, 26M entries, with following configuration:
> 
> Index definition:  
> curl -XPUT "[http://HOST:9200/my\_index/\_settings](http://HOST:9200/my_index/_settings)" -d '{  
> index: {  
> number\_of\_shards: 10,  
> number\_of\_replicas: 3,  
> "analysis": {  
> "analyzer": {  
> "parent\_hierarchy\_analyzer": {  
> "type": "custom",  
> "tokenizer": "path\_hierarchy"  
> }  
> }  
> }  
> }
> 
> }'
> 
> Mapping definition:  
> curl -XPUT "[http://HOST:9200/my\_index/my\_object/\_mapping?pretty=true](http://HOST:9200/my_index/my_object/_mapping?pretty=true)" -  
> d '  
> {  
> "infoclone": {  
> "properties": {  
> "parent\_hierarchy": {  
> "type": "string",  
> "store": "no",  
> "omit\_term\_freq\_and\_positions" : true,  
> "analyzer": "parent\_hierarchy\_analyzer",  
> "index": "analyzed",  
> "omit\_norms" : true,  
> "boost" : 1.0,  
> "term\_vector" : "no"  
> }  
> }  
> }}
> 
> '
> 
> I'm trying to add additional field for every object, so I use filter  
> below to get all of the objects, without "parent\_hierarchy" field, to  
> update it.
> 
> curl -XGEThttp://HOST:9200/my\_index/my\_object/\_search?pretty=1-d '{  
> "from" : 0,  
> "size" : 1000,  
> "query" : {  
> "constant\_score" : {  
> "filter" : {  
> "bool" : {  
> "must" : {  
> "missing" : {  
> "field" : "parent\_hierarchy"  
> }  
> }  
> }  
> }  
> }  
> }}
> 
> '  
> I've successfully updated ~16M entries, (from java client, using above  
> query, bulk update, refresh=true)
> 
> Now, every time I execute the query, I get two errors in response  
> body:
> 
> 1. Every query the IP changes  
> {  
> "took" : 91,  
> "timed\_out" : false,  
> "\_shards" : {  
> "total" : 10,  
> "successful" : 9,  
> "failed" : 1,  
> "failures" : [ {  
> "status" : 500,  
> "reason" : "RemoteTransportException[[Schmidt, Johann][inet[/  
> 10.11.10.74:9300]][search/phase/fetch/id]]; nested:  
> FieldReaderException[Invalid numeric type: 38]; "  
> } ]  
> },  
> "hits" : {  
> "total" : 686244,  
> "max\_score" : 1.0,  
> "hits" :   
> }
> 
> }
> 
> 1. Another version of response  
> {  
> "took" : 49,  
> "timed\_out" : false,  
> "\_shards" : {  
> "total" : 10,  
> "successful" : 9,  
> "failed" : 1,  
> "failures" : [ {  
> "status" : 500,  
> "reason" : "FieldReaderException[Invalid numeric type: 38]"  
> } ]  
> },  
> "hits" : {  
> "total" : 686244,  
> "max\_score" : 1.0,  
> "hits" :   
> }
> 
> }
> 
> Health status below:  
> curl -s -XGET '[http://HOST:9200/\_cluster/health?pretty=1](http://HOST:9200/_cluster/health?pretty=1)'  
> {  
> "cluster\_name" : "CMWELL\_INDEX\_CLUSTER",  
> "status" : "green",  
> "timed\_out" : false,  
> "number\_of\_nodes" : 12,  
> "number\_of\_data\_nodes" : 6,  
> "active\_primary\_shards" : 10,  
> "active\_shards" : 30,  
> "relocating\_shards" : 0,  
> "initializing\_shards" : 0,  
> "unassigned\_shards" : 0
> 
> }
> 
> elsticsearch.yml:
> 
> cluster.name : MY\_INDEX
> 
> path:  
> home: data/elasticsearch  
> logs: data/elasticsearch/logs
> 
> gateway:  
> recover\_after\_nodes: 5  
> recover\_after\_time: 1m  
> expected\_nodes: 6  
> local:  
> initial\_shards: 1
> 
> #default 10% of memory  
> #indices.memory.index\_buffer\_size : 1024m  
> #default false  
> index.compound\_format : true  
> #default 1s  
> index.refresh\_interval : 10s  
> #defualt 128  
> index.term\_index\_interval: 128
> 
> #Merging  
> #default 10  
> index.merge.policy.merge\_factor: 30  
> #default 1.6mb  
> index.merge.policy.min\_merge\_size: 16mb  
> #default unbounded  
> #index.merge.policy.max\_merge\_size: 1024mb  
> #default unbounded  
> #index.merge.policy.maxMergeDocs
> 
> #Transaction log settings  
> #After how many operations to flush/ Defaults to 20000/  
> #index.translog.flush\_threshold\_ops: 20000
> 
> #Once the translog hits this size, a flush will happen/ Defaults to  
> 500mb/  
> #index.translog.flush\_threshold\_size
> 
> #The period with no flush happening to force a flush/ Defaults to 60m/  
> #index.translog.flush\_threshold\_period
> 
> #Cache configurations
> 
> #defualt 20%  
> indices.cache.filter.size: 10%
> 
> #defualt -1  
> #1 entry ~1MB  
> #index.cache.filter.max\_size: 100
> 
> #defualt -1  
> index.cache.filter.expire: 1m
> 
> #defualt -1  
> #index.cache.field.max\_size: -1
> 
> #default -1  
> index.cache.field.expire: 1m
> 
> Every idea will be appreciated.  
> Thanks

---

<div class="post-metadata">

**Author:** ![vpunski](https://avatars.discourse-cdn.com/v4/letter/v/54ee81/32.png) [@vpunski](https://discuss.elastic.co/u/vpunski)\
**Post date:** [October 27, 2011, 7:41am UTC](https://discuss.elastic.co/t/error-on-missing-field-query/5688/3 "2011-10-27T07:41:42Z")

</div>

Any ideas?  
Even regarding the error mesage:

"reason" : "RemoteTransportException[[Schmidt, Johann][inet[/  
10.11.10.74:9300]][search/phase/fetch/id]]; nested:  
FieldReaderException[Invalid numeric type: 38]; "

What is FieldReaderException?  
Trying to find out the reason in Lucene code... too low level.  
Can anyone explain the meaning of the error?

Thanks

On Oct 26, 2:07 pm, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:

> Please note:  
> "total" : 10,  
> "successful" : 9,
> 
> On shard fails.
> 
> On Oct 26, 2:06 pm, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:
> 
> > Hi,  
> > 6 nodes working cluster, 26M entries, with following configuration:
> 
> > Index definition:  
> > curl -XPUT "[http://HOST:9200/my\_index/\_settings](http://HOST:9200/my_index/_settings)" -d '{  
> > index: {  
> > number\_of\_shards: 10,  
> > number\_of\_replicas: 3,  
> > "analysis": {  
> > "analyzer": {  
> > "parent\_hierarchy\_analyzer": {  
> > "type": "custom",  
> > "tokenizer": "path\_hierarchy"  
> > }  
> > }  
> > }  
> > }
> 
> > }'
> 
> > Mapping definition:  
> > curl -XPUT "[http://HOST:9200/my\_index/my\_object/\_mapping?pretty=true](http://HOST:9200/my_index/my_object/_mapping?pretty=true)" -  
> > d '  
> > {  
> > "infoclone": {  
> > "properties": {  
> > "parent\_hierarchy": {  
> > "type": "string",  
> > "store": "no",  
> > "omit\_term\_freq\_and\_positions" : true,  
> > "analyzer": "parent\_hierarchy\_analyzer",  
> > "index": "analyzed",  
> > "omit\_norms" : true,  
> > "boost" : 1.0,  
> > "term\_vector" : "no"  
> > }  
> > }  
> > }}
> 
> > '
> 
> > I'm trying to add additional field for every object, so I use filter  
> > below to get all of the objects, without "parent\_hierarchy" field, to  
> > update it.
> 
> > curl -XGEThttp://HOST:9200/my\_index/my\_object/\_search?pretty=1-d'{  
> > "from" : 0,  
> > "size" : 1000,  
> > "query" : {  
> > "constant\_score" : {  
> > "filter" : {  
> > "bool" : {  
> > "must" : {  
> > "missing" : {  
> > "field" : "parent\_hierarchy"  
> > }  
> > }  
> > }  
> > }  
> > }  
> > }}
> 
> > '  
> > I've successfully updated ~16M entries, (from java client, using above  
> > query, bulk update, refresh=true)
> 
> > Now, every time I execute the query, I get two errors in response  
> > body:
> > 
> > 1. Every query the IP changes  
> > {  
> > "took" : 91,  
> > "timed\_out" : false,  
> > "\_shards" : {  
> > "total" : 10,  
> > "successful" : 9,  
> > "failed" : 1,  
> > "failures" : [ {  
> > "status" : 500,  
> > "reason" : "RemoteTransportException[[Schmidt, Johann][inet[/  
> > 10.11.10.74:9300]][search/phase/fetch/id]]; nested:  
> > FieldReaderException[Invalid numeric type: 38]; "  
> > } ]  
> > },  
> > "hits" : {  
> > "total" : 686244,  
> > "max\_score" : 1.0,  
> > "hits" :   
> > }
> 
> > }
> 
> > 1. Another version of response  
> > {  
> > "took" : 49,  
> > "timed\_out" : false,  
> > "\_shards" : {  
> > "total" : 10,  
> > "successful" : 9,  
> > "failed" : 1,  
> > "failures" : [ {  
> > "status" : 500,  
> > "reason" : "FieldReaderException[Invalid numeric type: 38]"  
> > } ]  
> > },  
> > "hits" : {  
> > "total" : 686244,  
> > "max\_score" : 1.0,  
> > "hits" :   
> > }
> 
> > }
> 
> > Health status below:  
> > curl -s -XGET '[http://HOST:9200/\_cluster/health?pretty=1](http://HOST:9200/_cluster/health?pretty=1)'  
> > {  
> > "cluster\_name" : "CMWELL\_INDEX\_CLUSTER",  
> > "status" : "green",  
> > "timed\_out" : false,  
> > "number\_of\_nodes" : 12,  
> > "number\_of\_data\_nodes" : 6,  
> > "active\_primary\_shards" : 10,  
> > "active\_shards" : 30,  
> > "relocating\_shards" : 0,  
> > "initializing\_shards" : 0,  
> > "unassigned\_shards" : 0
> 
> > }
> 
> > elsticsearch.yml:
> 
> > cluster.name : MY\_INDEX
> 
> > path:  
> > home: data/elasticsearch  
> > logs: data/elasticsearch/logs
> 
> > gateway:  
> > recover\_after\_nodes: 5  
> > recover\_after\_time: 1m  
> > expected\_nodes: 6  
> > local:  
> > initial\_shards: 1
> 
> > #default 10% of memory  
> > #indices.memory.index\_buffer\_size : 1024m  
> > #default false  
> > index.compound\_format : true  
> > #default 1s  
> > index.refresh\_interval : 10s  
> > #defualt 128  
> > index.term\_index\_interval: 128
> 
> > #Merging  
> > #default 10  
> > index.merge.policy.merge\_factor: 30  
> > #default 1.6mb  
> > index.merge.policy.min\_merge\_size: 16mb  
> > #default unbounded  
> > #index.merge.policy.max\_merge\_size: 1024mb  
> > #default unbounded  
> > #index.merge.policy.maxMergeDocs
> 
> > #Transaction log settings  
> > #After how many operations to flush/ Defaults to 20000/  
> > #index.translog.flush\_threshold\_ops: 20000
> 
> > #Once the translog hits this size, a flush will happen/ Defaults to  
> > 500mb/  
> > #index.translog.flush\_threshold\_size
> 
> > #The period with no flush happening to force a flush/ Defaults to 60m/  
> > #index.translog.flush\_threshold\_period
> 
> > #Cache configurations
> 
> > #defualt 20%  
> > indices.cache.filter.size: 10%
> 
> > #defualt -1  
> > #1 entry ~1MB  
> > #index.cache.filter.max\_size: 100
> 
> > #defualt -1  
> > index.cache.filter.expire: 1m
> 
> > #defualt -1  
> > #index.cache.field.max\_size: -1
> 
> > #default -1  
> > index.cache.field.expire: 1m
> 
> > Every idea will be appreciated.  
> > Thanks

---

<div class="post-metadata">

**Author:** ![vpunski](https://avatars.discourse-cdn.com/v4/letter/v/54ee81/32.png) [@vpunski](https://discuss.elastic.co/u/vpunski)\
**Post date:** [October 27, 2011, 3:27pm UTC](https://discuss.elastic.co/t/error-on-missing-field-query/5688/4 "2011-10-27T15:27:29Z")

</div>

Still trying to solve the problem, and I have huge advance.  
I noticed that the error returns number of IPs equal to number of  
replicas.  
Using elasticsearch-head it was shard 9.  
The idea that particular shard is broken was checked by opening all  
three replicas in Luke,  
Indeed, the same FieldReaderException exception returned in all  
replicas of 9-th shard, by requesting some documents with  
"problemmatic" IDs.

A number of questions arise:  
How could it happened, that broken shard was replicated on two others?  
Can someone explain the flow of system start up, data consistency  
checking and shard relocation process?  
May initial\_shards=1 be a problem?  
Does the system checks file consistency on startup or during  
relocation to be sure that no corrupted data transferred?

Thanks

On Oct 27, 9:41 am, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:

> Any ideas?  
> Even regarding the error mesage:
> 
> "reason" : "RemoteTransportException[[Schmidt, Johann][inet[/  
> 10.11.10.74:9300]][search/phase/fetch/id]]; nested:  
> FieldReaderException[Invalid numeric type: 38]; "
> 
> What is FieldReaderException?  
> Trying to find out the reason in Lucene code... too low level.  
> Can anyone explain the meaning of the error?
> 
> Thanks
> 
> On Oct 26, 2:07 pm, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:
> 
> > Please note:  
> > "total" : 10,  
> > "successful" : 9,
> 
> > On shard fails.
> 
> > On Oct 26, 2:06 pm, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:
> 
> > > Hi,  
> > > 6 nodes working cluster, 26M entries, with following configuration:
> 
> > > Index definition:  
> > > curl -XPUT "[http://HOST:9200/my\_index/\_settings](http://HOST:9200/my_index/_settings)" -d '{  
> > > index: {  
> > > number\_of\_shards: 10,  
> > > number\_of\_replicas: 3,  
> > > "analysis": {  
> > > "analyzer": {  
> > > "parent\_hierarchy\_analyzer": {  
> > > "type": "custom",  
> > > "tokenizer": "path\_hierarchy"  
> > > }  
> > > }  
> > > }  
> > > }
> 
> > > }'
> 
> > > Mapping definition:  
> > > curl -XPUT "[http://HOST:9200/my\_index/my\_object/\_mapping?pretty=true](http://HOST:9200/my_index/my_object/_mapping?pretty=true)" -  
> > > d '  
> > > {  
> > > "infoclone": {  
> > > "properties": {  
> > > "parent\_hierarchy": {  
> > > "type": "string",  
> > > "store": "no",  
> > > "omit\_term\_freq\_and\_positions" : true,  
> > > "analyzer": "parent\_hierarchy\_analyzer",  
> > > "index": "analyzed",  
> > > "omit\_norms" : true,  
> > > "boost" : 1.0,  
> > > "term\_vector" : "no"  
> > > }  
> > > }  
> > > }}
> 
> > > '
> 
> > > I'm trying to add additional field for every object, so I use filter  
> > > below to get all of the objects, without "parent\_hierarchy" field, to  
> > > update it.
> 
> > > curl -XGEThttp://HOST:9200/my\_index/my\_object/\_search?pretty=1-d'{  
> > > "from" : 0,  
> > > "size" : 1000,  
> > > "query" : {  
> > > "constant\_score" : {  
> > > "filter" : {  
> > > "bool" : {  
> > > "must" : {  
> > > "missing" : {  
> > > "field" : "parent\_hierarchy"  
> > > }  
> > > }  
> > > }  
> > > }  
> > > }  
> > > }}
> 
> > > '  
> > > I've successfully updated ~16M entries, (from java client, using above  
> > > query, bulk update, refresh=true)
> 
> > > Now, every time I execute the query, I get two errors in response  
> > > body:
> > > 
> > > 1. Every query the IP changes  
> > > {  
> > > "took" : 91,  
> > > "timed\_out" : false,  
> > > "\_shards" : {  
> > > "total" : 10,  
> > > "successful" : 9,  
> > > "failed" : 1,  
> > > "failures" : [ {  
> > > "status" : 500,  
> > > "reason" : "RemoteTransportException[[Schmidt, Johann][inet[/  
> > > 10.11.10.74:9300]][search/phase/fetch/id]]; nested:  
> > > FieldReaderException[Invalid numeric type: 38]; "  
> > > } ]  
> > > },  
> > > "hits" : {  
> > > "total" : 686244,  
> > > "max\_score" : 1.0,  
> > > "hits" :   
> > > }
> 
> > > }
> 
> > > 1. Another version of response  
> > > {  
> > > "took" : 49,  
> > > "timed\_out" : false,  
> > > "\_shards" : {  
> > > "total" : 10,  
> > > "successful" : 9,  
> > > "failed" : 1,  
> > > "failures" : [ {  
> > > "status" : 500,  
> > > "reason" : "FieldReaderException[Invalid numeric type: 38]"  
> > > } ]  
> > > },  
> > > "hits" : {  
> > > "total" : 686244,  
> > > "max\_score" : 1.0,  
> > > "hits" :   
> > > }
> 
> > > }
> 
> > > Health status below:  
> > > curl -s -XGET '[http://HOST:9200/\_cluster/health?pretty=1](http://HOST:9200/_cluster/health?pretty=1)'  
> > > {  
> > > "cluster\_name" : "CMWELL\_INDEX\_CLUSTER",  
> > > "status" : "green",  
> > > "timed\_out" : false,  
> > > "number\_of\_nodes" : 12,  
> > > "number\_of\_data\_nodes" : 6,  
> > > "active\_primary\_shards" : 10,  
> > > "active\_shards" : 30,  
> > > "relocating\_shards" : 0,  
> > > "initializing\_shards" : 0,  
> > > "unassigned\_shards" : 0
> 
> > > }
> 
> > > elsticsearch.yml:
> 
> > > cluster.name : MY\_INDEX
> 
> > > path:  
> > > home: data/elasticsearch  
> > > logs: data/elasticsearch/logs
> 
> > > gateway:  
> > > recover\_after\_nodes: 5  
> > > recover\_after\_time: 1m  
> > > expected\_nodes: 6  
> > > local:  
> > > initial\_shards: 1
> 
> > > #default 10% of memory  
> > > #indices.memory.index\_buffer\_size : 1024m  
> > > #default false  
> > > index.compound\_format : true  
> > > #default 1s  
> > > index.refresh\_interval : 10s  
> > > #defualt 128  
> > > index.term\_index\_interval: 128
> 
> > > #Merging  
> > > #default 10  
> > > index.merge.policy.merge\_factor: 30  
> > > #default 1.6mb  
> > > index.merge.policy.min\_merge\_size: 16mb  
> > > #default unbounded  
> > > #index.merge.policy.max\_merge\_size: 1024mb  
> > > #default unbounded  
> > > #index.merge.policy.maxMergeDocs
> 
> > > #Transaction log settings  
> > > #After how many operations to flush/ Defaults to 20000/  
> > > #index.translog.flush\_threshold\_ops: 20000
> 
> > > #Once the translog hits this size, a flush will happen/ Defaults to  
> > > 500mb/  
> > > #index.translog.flush\_threshold\_size
> 
> > > #The period with no flush happening to force a flush/ Defaults to 60m/  
> > > #index.translog.flush\_threshold\_period
> 
> > > #Cache configurations
> 
> > > #defualt 20%  
> > > indices.cache.filter.size: 10%
> 
> > > #defualt -1  
> > > #1 entry ~1MB  
> > > #index.cache.filter.max\_size: 100
> 
> > > #defualt -1  
> > > index.cache.filter.expire: 1m
> 
> > > #defualt -1  
> > > #index.cache.field.max\_size: -1
> 
> > > #default -1  
> > > index.cache.field.expire: 1m
> 
> > > Every idea will be appreciated.  
> > > Thanks

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [October 28, 2011, 5:52am UTC](https://discuss.elastic.co/t/error-on-missing-field-query/5688/5 "2011-10-28T05:52:53Z")

</div>

Is it something that you can recreate? Basically, when an operation occurs  
on a shard, it is also executed on the replica. When shards move around,  
then sync against the primary shard in terms of actual index files.

On Thu, Oct 27, 2011 at 5:27 PM, vadim [vpunski@gmail.com](mailto:vpunski@gmail.com) wrote:

> Still trying to solve the problem, and I have huge advance.  
> I noticed that the error returns number of IPs equal to number of  
> replicas.  
> Using elasticsearch-head it was shard 9.  
> The idea that particular shard is broken was checked by opening all  
> three replicas in Luke,  
> Indeed, the same FieldReaderException exception returned in all  
> replicas of 9-th shard, by requesting some documents with  
> "problemmatic" IDs.
> 
> A number of questions arise:  
> How could it happened, that broken shard was replicated on two others?  
> Can someone explain the flow of system start up, data consistency  
> checking and shard relocation process?  
> May initial\_shards=1 be a problem?  
> Does the system checks file consistency on startup or during  
> relocation to be sure that no corrupted data transferred?
> 
> Thanks
> 
> On Oct 27, 9:41 am, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:
> 
> > Any ideas?  
> > Even regarding the error mesage:
> > 
> > "reason" : "RemoteTransportException[[Schmidt, Johann][inet[/  
> > 10.11.10.74:9300]][search/phase/fetch/id]]; nested:  
> > FieldReaderException[Invalid numeric type: 38]; "
> > 
> > What is FieldReaderException?  
> > Trying to find out the reason in Lucene code... too low level.  
> > Can anyone explain the meaning of the error?
> > 
> > Thanks
> > 
> > On Oct 26, 2:07 pm, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:
> > 
> > > Please note:  
> > > "total" : 10,  
> > > "successful" : 9,
> > 
> > > On shard fails.
> > 
> > > On Oct 26, 2:06 pm, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:
> > 
> > > > Hi,  
> > > > 6 nodes working cluster, 26M entries, with following configuration:
> > 
> > > > Index definition:  
> > > > curl -XPUT "[http://HOST:9200/my\_index/\_settings](http://HOST:9200/my_index/_settings)" -d '{  
> > > > index: {  
> > > > number\_of\_shards: 10,  
> > > > number\_of\_replicas: 3,  
> > > > "analysis": {  
> > > > "analyzer": {  
> > > > "parent\_hierarchy\_analyzer": {  
> > > > "type": "custom",  
> > > > "tokenizer": "path\_hierarchy"  
> > > > }  
> > > > }  
> > > > }  
> > > > }
> > 
> > > > }'
> > 
> > > > Mapping definition:  
> > > > curl -XPUT "[http://HOST:9200/my\_index/my\_object/\_mapping?pretty=true](http://HOST:9200/my_index/my_object/_mapping?pretty=true)"
> 
> - 
> 
> > > > d '  
> > > > {  
> > > > "infoclone": {  
> > > > "properties": {  
> > > > "parent\_hierarchy": {  
> > > > "type": "string",  
> > > > "store": "no",  
> > > > "omit\_term\_freq\_and\_positions" : true,  
> > > > "analyzer": "parent\_hierarchy\_analyzer",  
> > > > "index": "analyzed",  
> > > > "omit\_norms" : true,  
> > > > "boost" : 1.0,  
> > > > "term\_vector" : "no"  
> > > > }  
> > > > }  
> > > > }}
> > 
> > > > '
> > 
> > > > I'm trying to add additional field for every object, so I use filter  
> > > > below to get all of the objects, without "parent\_hierarchy" field, to  
> > > > update it.
> > 
> > > > curl -XGEThttp://HOST:9200/my\_index/my\_object/\_search?pretty=1-d'{  
> > > > "from" : 0,  
> > > > "size" : 1000,  
> > > > "query" : {  
> > > > "constant\_score" : {  
> > > > "filter" : {  
> > > > "bool" : {  
> > > > "must" : {  
> > > > "missing" : {  
> > > > "field" : "parent\_hierarchy"  
> > > > }  
> > > > }  
> > > > }  
> > > > }  
> > > > }  
> > > > }}
> > 
> > > > '  
> > > > I've successfully updated ~16M entries, (from java client, using  
> > > > above  
> > > > query, bulk update, refresh=true)
> > 
> > > > Now, every time I execute the query, I get two errors in response  
> > > > body:
> > > > 
> > > > 1. Every query the IP changes  
> > > > {  
> > > > "took" : 91,  
> > > > "timed\_out" : false,  
> > > > "\_shards" : {  
> > > > "total" : 10,  
> > > > "successful" : 9,  
> > > > "failed" : 1,  
> > > > "failures" : [ {  
> > > > "status" : 500,  
> > > > "reason" : "RemoteTransportException[[Schmidt, Johann][inet[/  
> > > > 10.11.10.74:9300]][search/phase/fetch/id]]; nested:  
> > > > FieldReaderException[Invalid numeric type: 38]; "  
> > > > } ]  
> > > > },  
> > > > "hits" : {  
> > > > "total" : 686244,  
> > > > "max\_score" : 1.0,  
> > > > "hits" :   
> > > > }
> > 
> > > > }
> > 
> > > > 1. Another version of response  
> > > > {  
> > > > "took" : 49,  
> > > > "timed\_out" : false,  
> > > > "\_shards" : {  
> > > > "total" : 10,  
> > > > "successful" : 9,  
> > > > "failed" : 1,  
> > > > "failures" : [ {  
> > > > "status" : 500,  
> > > > "reason" : "FieldReaderException[Invalid numeric type: 38]"  
> > > > } ]  
> > > > },  
> > > > "hits" : {  
> > > > "total" : 686244,  
> > > > "max\_score" : 1.0,  
> > > > "hits" :   
> > > > }
> > 
> > > > }
> > 
> > > > Health status below:  
> > > > curl -s -XGET '[http://HOST:9200/\_cluster/health?pretty=1](http://HOST:9200/_cluster/health?pretty=1)'  
> > > > {  
> > > > "cluster\_name" : "CMWELL\_INDEX\_CLUSTER",  
> > > > "status" : "green",  
> > > > "timed\_out" : false,  
> > > > "number\_of\_nodes" : 12,  
> > > > "number\_of\_data\_nodes" : 6,  
> > > > "active\_primary\_shards" : 10,  
> > > > "active\_shards" : 30,  
> > > > "relocating\_shards" : 0,  
> > > > "initializing\_shards" : 0,  
> > > > "unassigned\_shards" : 0
> > 
> > > > }
> > 
> > > > elsticsearch.yml:
> > 
> > > > cluster.name : MY\_INDEX
> > 
> > > > path:  
> > > > home: data/elasticsearch  
> > > > logs: data/elasticsearch/logs
> > 
> > > > gateway:  
> > > > recover\_after\_nodes: 5  
> > > > recover\_after\_time: 1m  
> > > > expected\_nodes: 6  
> > > > local:  
> > > > initial\_shards: 1
> > 
> > > > #default 10% of memory  
> > > > #indices.memory.index\_buffer\_size : 1024m  
> > > > #default false  
> > > > index.compound\_format : true  
> > > > #default 1s  
> > > > index.refresh\_interval : 10s  
> > > > #defualt 128  
> > > > index.term\_index\_interval: 128
> > 
> > > > #Merging  
> > > > #default 10  
> > > > index.merge.policy.merge\_factor: 30  
> > > > #default 1.6mb  
> > > > index.merge.policy.min\_merge\_size: 16mb  
> > > > #default unbounded  
> > > > #index.merge.policy.max\_merge\_size: 1024mb  
> > > > #default unbounded  
> > > > #index.merge.policy.maxMergeDocs
> > 
> > > > #Transaction log settings  
> > > > #After how many operations to flush/ Defaults to 20000/  
> > > > #index.translog.flush\_threshold\_ops: 20000
> > 
> > > > #Once the translog hits this size, a flush will happen/ Defaults to  
> > > > 500mb/  
> > > > #index.translog.flush\_threshold\_size
> > 
> > > > #The period with no flush happening to force a flush/ Defaults to  
> > > > 60m/  
> > > > #index.translog.flush\_threshold\_period
> > 
> > > > #Cache configurations
> > 
> > > > #defualt 20%  
> > > > indices.cache.filter.size: 10%
> > 
> > > > #defualt -1  
> > > > #1 entry ~1MB  
> > > > #index.cache.filter.max\_size: 100
> > 
> > > > #defualt -1  
> > > > index.cache.filter.expire: 1m
> > 
> > > > #defualt -1  
> > > > #index.cache.field.max\_size: -1
> > 
> > > > #default -1  
> > > > index.cache.field.expire: 1m
> > 
> > > > Every idea will be appreciated.  
> > > > Thanks

---

<div class="post-metadata">

**Author:** ![vpunski](https://avatars.discourse-cdn.com/v4/letter/v/54ee81/32.png) [@vpunski](https://discuss.elastic.co/u/vpunski)\
**Post date:** [October 30, 2011, 7:43am UTC](https://discuss.elastic.co/t/error-on-missing-field-query/5688/6 "2011-10-30T07:43:50Z")

</div>

I don't think it's possible to reproduce... may be by changing the  
file manually, to simulate damage file ...  
I had a lot of full cluster restart during last few weeks ... some of  
them emergency stops with KILL...

Let me understand the flow:  
If some file of primary shard get damaged, it will be replicated over  
the cluster during startup?  
Are there any flow to validate file consistency on start up, for  
example?  
Can you elaborate the life cycle please?  
It's very

Thanks

On Oct 28, 7:52 am, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:

> Is it something that you can recreate? Basically, when an operation occurs  
> on a shard, it is also executed on the replica. When shards move around,  
> then sync against the primary shard in terms of actual index files.

On Oct 28, 7:52 am, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:

> Is it something that you can recreate? Basically, when an operation occurs  
> on a shard, it is also executed on the replica. When shards move around,  
> then sync against the primary shard in terms of actual index files.
> 
> On Thu, Oct 27, 2011 at 5:27 PM, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:
> 
> > Still trying to solve the problem, and I have huge advance.  
> > I noticed that the error returns number of IPs equal to number of  
> > replicas.  
> > Using elasticsearch-head it was shard 9.  
> > The idea that particular shard is broken was checked by opening all  
> > three replicas in Luke,  
> > Indeed, the same FieldReaderException exception returned in all  
> > replicas of 9-th shard, by requesting some documents with  
> > "problemmatic" IDs.
> 
> > A number of questions arise:  
> > How could it happened, that broken shard was replicated on two others?  
> > Can someone explain the flow of system start up, data consistency  
> > checking and shard relocation process?  
> > May initial\_shards=1 be a problem?  
> > Does the system checks file consistency on startup or during  
> > relocation to be sure that no corrupted data transferred?
> 
> > Thanks
> 
> > On Oct 27, 9:41 am, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:
> > 
> > > Any ideas?  
> > > Even regarding the error mesage:
> 
> > > "reason" : "RemoteTransportException[[Schmidt, Johann][inet[/  
> > > 10.11.10.74:9300]][search/phase/fetch/id]]; nested:  
> > > FieldReaderException[Invalid numeric type: 38]; "
> 
> > > What is FieldReaderException?  
> > > Trying to find out the reason in Lucene code... too low level.  
> > > Can anyone explain the meaning of the error?
> 
> > > Thanks
> 
> > > On Oct 26, 2:07 pm, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:
> 
> > > > Please note:  
> > > > "total" : 10,  
> > > > "successful" : 9,
> 
> > > > On shard fails.
> 
> > > > On Oct 26, 2:06 pm, vadim [vpun...@gmail.com](mailto:vpun...@gmail.com) wrote:
> 
> > > > > Hi,  
> > > > > 6 nodes working cluster, 26M entries, with following configuration:
> 
> > > > > Index definition:  
> > > > > curl -XPUT "[http://HOST:9200/my\_index/\_settings](http://HOST:9200/my_index/_settings)" -d '{  
> > > > > index: {  
> > > > > number\_of\_shards: 10,  
> > > > > number\_of\_replicas: 3,  
> > > > > "analysis": {  
> > > > > "analyzer": {  
> > > > > "parent\_hierarchy\_analyzer": {  
> > > > > "type": "custom",  
> > > > > "tokenizer": "path\_hierarchy"  
> > > > > }  
> > > > > }  
> > > > > }  
> > > > > }
> 
> > > > > }'
> 
> > > > > Mapping definition:  
> > > > > curl -XPUT "[http://HOST:9200/my\_index/my\_object/\_mapping?pretty=true](http://HOST:9200/my_index/my_object/_mapping?pretty=true)"
> > 
> > - 
> > 
> > > > > d '  
> > > > > {  
> > > > > "infoclone": {  
> > > > > "properties": {  
> > > > > "parent\_hierarchy": {  
> > > > > "type": "string",  
> > > > > "store": "no",  
> > > > > "omit\_term\_freq\_and\_positions" : true,  
> > > > > "analyzer": "parent\_hierarchy\_analyzer",  
> > > > > "index": "analyzed",  
> > > > > "omit\_norms" : true,  
> > > > > "boost" : 1.0,  
> > > > > "term\_vector" : "no"  
> > > > > }  
> > > > > }  
> > > > > }}
> 
> > > > > '
> 
> > > > > I'm trying to add additional field for every object, so I use filter  
> > > > > below to get all of the objects, without "parent\_hierarchy" field, to  
> > > > > update it.
> 
> > > > > curl -XGEThttp://HOST:9200/my\_index/my\_object/\_search?pretty=1-d'{  
> > > > > "from" : 0,  
> > > > > "size" : 1000,  
> > > > > "query" : {  
> > > > > "constant\_score" : {  
> > > > > "filter" : {  
> > > > > "bool" : {  
> > > > > "must" : {  
> > > > > "missing" : {  
> > > > > "field" : "parent\_hierarchy"  
> > > > > }  
> > > > > }  
> > > > > }  
> > > > > }  
> > > > > }  
> > > > > }}
> 
> > > > > '  
> > > > > I've successfully updated ~16M entries, (from java client, using  
> > > > > above  
> > > > > query, bulk update, refresh=true)
> 
> > > > > Now, every time I execute the query, I get two errors in response  
> > > > > body:
> > > > > 
> > > > > 1. Every query the IP changes  
> > > > > {  
> > > > > "took" : 91,  
> > > > > "timed\_out" : false,  
> > > > > "\_shards" : {  
> > > > > "total" : 10,  
> > > > > "successful" : 9,  
> > > > > "failed" : 1,  
> > > > > "failures" : [ {  
> > > > > "status" : 500,  
> > > > > "reason" : "RemoteTransportException[[Schmidt, Johann][inet[/  
> > > > > 10.11.10.74:9300]][search/phase/fetch/id]]; nested:  
> > > > > FieldReaderException[Invalid numeric type: 38]; "  
> > > > > } ]  
> > > > > },  
> > > > > "hits" : {  
> > > > > "total" : 686244,  
> > > > > "max\_score" : 1.0,  
> > > > > "hits" :   
> > > > > }
> 
> > > > > }
> 
> > > > > 1. Another version of response  
> > > > > {  
> > > > > "took" : 49,  
> > > > > "timed\_out" : false,  
> > > > > "\_shards" : {  
> > > > > "total" : 10,  
> > > > > "successful" : 9,  
> > > > > "failed" : 1,  
> > > > > "failures" : [ {  
> > > > > "status" : 500,  
> > > > > "reason" : "FieldReaderException[Invalid numeric type: 38]"  
> > > > > } ]  
> > > > > },  
> > > > > "hits" : {  
> > > > > "total" : 686244,  
> > > > > "max\_score" : 1.0,  
> > > > > "hits" :   
> > > > > }
> 
> > > > > }
> 
> > > > > Health status below:  
> > > > > curl -s -XGET '[http://HOST:9200/\_cluster/health?pretty=1](http://HOST:9200/_cluster/health?pretty=1)'  
> > > > > {  
> > > > > "cluster\_name" : "CMWELL\_INDEX\_CLUSTER",  
> > > > > "status" : "green",  
> > > > > "timed\_out" : false,  
> > > > > "number\_of\_nodes" : 12,  
> > > > > "number\_of\_data\_nodes" : 6,  
> > > > > "active\_primary\_shards" : 10,  
> > > > > "active\_shards" : 30,  
> > > > > "relocating\_shards" : 0,  
> > > > > "initializing\_shards" : 0,  
> > > > > "unassigned\_shards" : 0
> 
> > > > > }
> 
> > > > > elsticsearch.yml:
> 
> > > > > cluster.name : MY\_INDEX
> 
> > > > > path:  
> > > > > home: data/elasticsearch  
> > > > > logs: data/elasticsearch/logs
> 
> > > > > gateway:  
> > > > > recover\_after\_nodes: 5  
> > > > > recover\_after\_time: 1m  
> > > > > expected\_nodes: 6  
> > > > > local:  
> > > > > initial\_shards: 1
> 
> > > > > #default 10% of memory  
> > > > > #indices.memory.index\_buffer\_size : 1024m  
> > > > > #default false  
> > > > > index.compound\_format : true  
> > > > > #default 1s  
> > > > > index.refresh\_interval : 10s  
> > > > > #defualt 128  
> > > > > index.term\_index\_interval: 128
> 
> > > > > #Merging  
> > > > > #default 10  
> > > > > index.merge.policy.merge\_factor: 30  
> > > > > #default 1.6mb  
> > > > > index.merge.policy.min\_merge\_size: 16mb  
> > > > > #default unbounded  
> > > > > #index.merge.policy.max\_merge\_size: 1024mb  
> > > > > #default unbounded  
> > > > > #index.merge.policy.maxMergeDocs
> 
> > > > > #Transaction log settings  
> > > > > #After how many operations to flush/ Defaults to 20000/  
> > > > > #index.translog.flush\_threshold\_ops: 20000
> 
> > > > > #Once the translog hits this size, a flush will happen/ Defaults to  
> > > > > 500mb/  
> > > > > #index.translog.flush\_threshold\_size
> 
> > > > > #The period with no flush happening to force a flush/ Defaults to  
> > > > > 60m/  
> > > > > #index.translog.flush\_threshold\_period
> 
> > > > > #Cache configurations
> 
> > > > > #defualt 20%  
> > > > > indices.cache.filter.size: 10%
> 
> > > > > #defualt -1  
> > > > > #1 entry ~1MB  
> > > > > #index.cache.filter.max\_size: 100
> 
> > > > > #defualt -1  
> > > > > index.cache.filter.expire: 1m
> 
> > > > > #defualt -1  
> > > > > #index.cache.field.max\_size: -1
> 
> > > > > #default -1  
> > > > > index.cache.field.expire: 1m
> 
> > > > > Every idea will be appreciated.  
> > > > > Thanks

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:50am UTC](https://discuss.elastic.co/t/error-on-missing-field-query/5688/7 "2017-07-06T03:50:28Z")

</div>


