# Query elasticseatrch with pyspark and nested fields

**URL:** <https://discuss.elastic.co/t/query-elasticseatrch-with-pyspark-and-nested-fields/371745>\
**Category:** Elasticsearch\
**Created:** [December 10, 2024, 9:58am UTC](https://discuss.elastic.co/t/query-elasticseatrch-with-pyspark-and-nested-fields/371745 "2024-12-10T09:58:45Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![xsa\_xsa](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/xsa_xsa/32/131884_2.png) [@xsa\_xsa](https://discuss.elastic.co/u/xsa_xsa)\
**Post date:** [December 10, 2024, 9:58am UTC](https://discuss.elastic.co/t/query-elasticseatrch-with-pyspark-and-nested-fields/371745/1 "2024-12-10T09:58:45Z")

</div>

hello,  
I am trying to query with pyspark an index with documents:

```auto
{
	"field1": "field1_data",
	"field2": [
	{
		"field2_1": "x1",
		"field2_2": "x2"
	},
	{
		"field2_1": "x3",
		"field2_2": "x4"
	}
	]
}

```

if i query from kibana:

```auto
post my_index/_search
{
	"query": {
		"bool": {
			"must": [],
			"filter": [
				"terms" {
					"field2.field2_1": "x1"
				}
			]
		}
	},
	"_source" = ["field1", "field2.field2_1"]
}

```

it retruns everyting fine:

```auto
hits = 
{
	"field1": "field1_data",
	"field2": [
	{
		"field2_1": "x1"
	}
	]
		
}

```

now i try to use pyspark:  
the data frame schema is:

```auto
index1_schema = StructType([
	StructField("field1", StringType(), nullable=True),
	StructField("field2", ArrayType(StructType[
			StructField("field2_1", StringType(), nullable=True),
			StructField("field2_2", StringType(), nullable=True)
		),
		containsNull=False), nullable=True)
])

```

i create spark options for elastic:

```auto
options ={
"es.nodes": ....,
"es.resource": "my_index",
"es.query": '''
{
	"query": {
		"bool": {
			"must": [],
			"filter": [
				"terms" {
					"field2.field2_1": "x1"
				}
			]
		}
	},
	"_source":["field1", "field2.field2_1"]
	}
	'''
]

```

i create a data frame reader

```auto
reader = sparkSession.read.schema(index1_schema).format('org.elasticsearch.spark.sql').options(options)
than i do 

df = reader.load().select(['field1', 'field2']

```

the df looks like:

```auto
field1 | field2
=======================
"field1_data" | [{Null, Null}]

```

where field2 contains an array of pyspark rows with one row object :

```auto
{
	"field2_1":None
}

```

what am i doing wrong?  
anyone? please help
