# Pyspark: from curl to correct settings

**URL:** <https://discuss.elastic.co/t/pyspark-from-curl-to-correct-settings/314999>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [September 23, 2022, 7:58am UTC](https://discuss.elastic.co/t/pyspark-from-curl-to-correct-settings/314999 "2022-09-23T07:58:17Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![obi134](https://avatars.discourse-cdn.com/v4/letter/o/c5a1d2/32.png) [@obi134](https://discuss.elastic.co/u/obi134)\
**Post date:** [September 23, 2022, 7:58am UTC](https://discuss.elastic.co/t/pyspark-from-curl-to-correct-settings/314999/1 "2022-09-23T07:58:17Z")

</div>

Hi there,

I'm trying to push data from databricks/pyspark to elasticsearch following these instructions: [ElasticSearch | Databricks on AWS](https://docs.databricks.com/data/data-sources/elasticsearch.html)

Unfortunately I'm getting this error:

> org.elasticsearch.hadoop.EsHadoopIllegalArgumentException: Cannot detect ES version - typically this happens if the network/Elasticsearch cluster is not accessible or when targeting a WAN/Cloud instance without the proper setting 'es.nodes.wan.only'

Running a curl command from databricks is working and I see the pushed data in kibana. So connection is working in general. But how do I get from curl command to correct settings in python? I already tried to find the right options on [Configuration | Elasticsearch for Apache Hadoop [8.11] | Elastic](https://www.elastic.co/guide/en/elasticsearch/hadoop/current/configuration.html), but currently not successful.

This is the curl command:

```auto
curl --user username:passwd -X PUT -H "Content-Type: application/json" -d '{"name":"John Doe"}' http://my.url.com/elasticsearch/test_databricks/_doc/1

```

Pythoncode I tried:

```auto
df.write
  .format( "org.elasticsearch.spark.sql" )
  .option( "es.nodes", "my.url.com/elasticsearch/")
  .option( "es.net.ssl", "false")
  .option( "es.nodes.wan.only", "true" )
  .option( "es.net.http.auth.user", "username")
  .option( "es.net.http.auth.pass", "passwd")
  .mode( "overwrite" )
  .save( "index/test_databricks" )

```

Thank you in advance

---

<div class="post-metadata">

**Author:** ![Keith\_Massey](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/keith_massey/32/83666_2.png) [@Keith\_Massey](https://discuss.elastic.co/u/Keith_Massey)\
**Post date:** [September 26, 2022, 4:47pm UTC](https://discuss.elastic.co/t/pyspark-from-curl-to-correct-settings/314999/2 "2022-09-26T16:47:38Z")

</div>

You can get that error for a variety of reasons. Look for a `caused by` stack trace that might give you more information. I believe it willl be in your spark driver log.

---

<div class="post-metadata">

**Author:** ![obi134](https://avatars.discourse-cdn.com/v4/letter/o/c5a1d2/32.png) [@obi134](https://discuss.elastic.co/u/obi134)\
**Post date:** [September 29, 2022, 12:58pm UTC](https://discuss.elastic.co/t/pyspark-from-curl-to-correct-settings/314999/3 "2022-09-29T12:58:17Z")

</div>

Using the option `es.nodes.path.prefix` fixed the issue:

```auto
df.write
  .format( "org.elasticsearch.spark.sql" )
  .option( "es.nodes", "my.url.com")
  .option( "es.nodes.path.prefix", "elasticsearch" ) 
  .option( "es.net.ssl", "false")
  .option( "es.nodes.wan.only", "true" )
  .option( "es.net.http.auth.user", "username")
  .option( "es.net.http.auth.pass", "passwd")
  .mode( "overwrite" )
  .save( "/test_databricks" )

```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 27, 2022, 12:58pm UTC](https://discuss.elastic.co/t/pyspark-from-curl-to-correct-settings/314999/4 "2022-10-27T12:58:24Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
