# Spark sql query for empty string

**URL:** <https://discuss.elastic.co/t/spark-sql-query-for-empty-string/141588>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [July 25, 2018, 1:36pm UTC](https://discuss.elastic.co/t/spark-sql-query-for-empty-string/141588 "2018-07-25T13:36:57Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![wangqinghuan](https://avatars.discourse-cdn.com/v4/letter/w/d26b3c/32.png) [@wangqinghuan](https://discuss.elastic.co/u/wangqinghuan)\
**Post date:** [July 25, 2018, 1:36pm UTC](https://discuss.elastic.co/t/spark-sql-query-for-empty-string/141588/1 "2018-07-25T13:36:57Z")

</div>

I am using spark sql to query for an empty string . Documents are returned exactly when I use elasticsearch dsl as following:  
user/\_search  
{  
"query": {  
"term": {  
"type": ""  
}  
}  
}  
But in spark sql, I written a sql like "select \* from user where type = '' ", It does not return any results.  
how to query for empty string in spark sql ?

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [July 25, 2018, 3:07pm UTC](https://discuss.elastic.co/t/spark-sql-query-for-empty-string/141588/2 "2018-07-25T15:07:52Z")

</div>

try setting the `strict` setting on the job settings to `true`. This will instruct the connector to use term queries instead of match queries in the DSL it generates. See [the Spark SQL docs](https://www.elastic.co/guide/en/elasticsearch/hadoop/current/spark.html#spark-data-sources) for more information about those settings.

---

<div class="post-metadata">

**Author:** ![wangqinghuan](https://avatars.discourse-cdn.com/v4/letter/w/d26b3c/32.png) [@wangqinghuan](https://discuss.elastic.co/u/wangqinghuan)\
**Post date:** [July 27, 2018, 10:39am UTC](https://discuss.elastic.co/t/spark-sql-query-for-empty-string/141588/3 "2018-07-27T10:39:38Z")

</div>

thank you. I try to set the param "strict = true", but it does not work exactly. Then I set the param "double.filtering = false", It can work exactly. Why I must disable "double.filtering" when turning strict on?  
The param "double.filtering" is really tricky.

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [July 27, 2018, 3:02pm UTC](https://discuss.elastic.co/t/spark-sql-query-for-empty-string/141588/4 "2018-07-27T15:02:18Z")

</div>

This may be an issue within Spark itself if you have to disable `double.filtering`. Spark likes to double check that pushdown operations have actually filtered everything correctly, so it's possible that Spark is interpreting the filter statements differently than ES-Hadoop does in this case.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 24, 2018, 3:02pm UTC](https://discuss.elastic.co/t/spark-sql-query-for-empty-string/141588/5 "2018-08-24T15:02:19Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
