# Unable to read elasticsearch.6.0.1 index data into dataframe using pyspark

**URL:** <https://discuss.elastic.co/t/unable-to-read-elasticsearch-6-0-1-index-data-into-dataframe-using-pyspark/353844>\
**Category:** Elasticsearch\
**Created:** [February 22, 2024, 5:39am UTC](https://discuss.elastic.co/t/unable-to-read-elasticsearch-6-0-1-index-data-into-dataframe-using-pyspark/353844 "2024-02-22T05:39:54Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![vincentnaveen](https://avatars.discourse-cdn.com/v4/letter/v/dc4da7/32.png) [@vincentnaveen](https://discuss.elastic.co/u/vincentnaveen)\
**Post date:** [February 22, 2024, 5:39am UTC](https://discuss.elastic.co/t/unable-to-read-elasticsearch-6-0-1-index-data-into-dataframe-using-pyspark/353844/1 "2024-02-22T05:39:54Z")

</div>

Hello everyone, I am using elastcisearch version 6.0.1 from AWS service. I am trying to read Es index data into spark dataframe using pyspark.

I can read all fields except the fields contain nested arrays.  
The nested array contains another nested array in it.

Mappings in index is below:

```auto
{
	"accounts": {
		"type": "nested",
		"properties": {
			"accountClassificationOne": {
				"type": "text",
				"fields": {
					"keyword": {
						"type": "keyword",
						"ignore_above": 256
					}
				}
			},
			"alternateNames": {
				"type": "nested",
				"properties": {
					"createDate": {
						"properties": {
							"chronology": {
								"type": "object"
							},
							"millis": {
								"type": "long"
							}
						}
					},
					"inactive": {
						"type": "boolean"
					},
					"name": {
						"type": "text",
						"fields": {
							"autocomplete": {
								"type": "text",
								"analyzer": "customer_synonym_autocomplete",
								"search_analyzer": "customer_synonym"
							},
							"de": {
								"type": "text",
								"analyzer": "customer_german_autocomplete",
								"search_analyzer": "german"
							},
							"full": {
								"type": "text",
								"analyzer": "customer_synonym_full"
							},
							"keyword": {
								"type": "keyword",
								"ignore_above": 256
							},
							"normalize": {
								"type": "keyword",
								"normalizer": "lowercase_normalizer"
							}
						},
						"analyzer": "customer_synonym"
					}
				},
				"badDebt": {
					"type": "boolean"
				}
			}
		}
	}
}

```

config in pyspark code trying to read is:

```auto
es_options_read = {
    "es.nodes": es_nodes,
    "es.port": 443,
    "es.resource": "index_name/type",
    "es.query": myquery,
    "es.nodes.wan.only": "true",
    "es.read.field.as.array.include": "accounts",
    "es.read.field.include": "accounts"
}

```

```auto
Error is: 
 org.elasticsearch.hadoop.EsHadoopIllegalStateException: Field 'updateDate.chronology' not found; typically this occurs with arrays which are not mapped as single value

another error sometimes: java.lang.NullPointerException. 

```

i tried multiple combination in arrays.include, struct field, explode and many in read options but no luck. could anyone help me on this.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 21, 2024, 5:40am UTC](https://discuss.elastic.co/t/unable-to-read-elasticsearch-6-0-1-index-data-into-dataframe-using-pyspark/353844/2 "2024-03-21T05:40:51Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
