# Pyspark es.query not working. only default "match\_all" works

**URL:** <https://discuss.elastic.co/t/pyspark-es-query-not-working-only-default-match-all-works/100408>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [September 13, 2017, 7:47pm UTC](https://discuss.elastic.co/t/pyspark-es-query-not-working-only-default-match-all-works/100408 "2017-09-13T19:47:06Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![buster](https://avatars.discourse-cdn.com/v4/letter/b/41988e/32.png) [@buster](https://discuss.elastic.co/u/buster)\
**Post date:** [September 13, 2017, 7:47pm UTC](https://discuss.elastic.co/t/pyspark-es-query-not-working-only-default-match-all-works/100408/1 "2017-09-13T19:47:06Z")

</div>

In pypspark the only way I can get data returned from ES is by leaving es.query default. Why is this?

es\_query = {"match" : {"key" : "value"}}  
es\_conf = {"es.nodes" : "localhost", "es.resource" : "index/type", "es.query" : json.dumps(es\_query)}  
rdd = sc.newAPIHadoopRDD(inputFormatClass="org.elasticsearch.hadoop.mr.EsInputFormat",keyClass="org.apache.hadoop.io.NullWritable",valueClass="org.elasticsearch.hadoop.mr.LinkedMapWritable", conf=es\_conf)

rdd.count()  
...  
0  
rdd.first()  
ValueError: RDD is empty

Yet when,  
es\_query = {"match\_all" : {}}

rdd.first()  
(u'2017-09-01 01:02:03)

\*I have tested the queries by directly querying elastic search and they work so it is something wrong with spark/es-hadoop.

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [October 4, 2017, 3:21am UTC](https://discuss.elastic.co/t/pyspark-es-query-not-working-only-default-match-all-works/100408/2 "2017-10-04T03:21:46Z")

</div>

@buster Are you specifying the query as a string or as a map of maps? It looks like you're omitting the quotes needed to make the query a string in your posted example.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 1, 2017, 3:21am UTC](https://discuss.elastic.co/t/pyspark-es-query-not-working-only-default-match-all-works/100408/3 "2017-11-01T03:21:55Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
