# SparkStreaming to Elasticesrahc ERROR NetworkClient: Connection timed out: connect

**URL:** <https://discuss.elastic.co/t/sparkstreaming-to-elasticesrahc-error-networkclient-connection-timed-out-connect/45834>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [March 30, 2016, 5:46pm UTC](https://discuss.elastic.co/t/sparkstreaming-to-elasticesrahc-error-networkclient-connection-timed-out-connect/45834 "2016-03-30T17:46:46Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![DK\_3](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dk_3/32/48003_2.png) [@DK\_3](https://discuss.elastic.co/u/DK_3)\
**Post date:** [March 30, 2016, 5:46pm UTC](https://discuss.elastic.co/t/sparkstreaming-to-elasticesrahc-error-networkclient-connection-timed-out-connect/45834/1 "2016-03-30T17:46:46Z")

</div>

I'm hitting the following when trying to save a Spark DataFrame to Elasticsearch

`16/03/30 18:28:38 ERROR NetworkClient: Node [172.18.0.2:9200] failed (Connection timed out: connect); selected next node [10.123.45.67:9200]`

I can ping Elastcisearch from the same machine the Spark app is running  
`nc -zv 10.123.45.67 9200_ Connection to 10.123.45.67 9200 port [tcp/*] succeeded!`

My code

```
SparkConf conf = new SparkConf();
conf.setMaster("local[*]");
conf.setAppName("Analyser");
conf.set("es.index.auto.create", "true");
conf.set("es.nodes", "10.123.45.67:9200");

JavaStreamingContext jssc = new JavaStreamingContext(conf, Durations.seconds(1));
JavaDStream<String> lines = jssc.socketTextStream("localhost", 7654);

JavaDStream<String> eventLines = lines.filter((String line) -> line.contains("Event"));
eventLines.foreachRDD((JavaRDD<String> rdd) -> {
    if (!rdd.isEmpty()) {

        SQLContext sqlContext = SQLContext.getOrCreate(rdd.context());
        DataFrame dataFrame = sqlContext.read().json(rdd);
        dataFrame.registerTempTable("Events");

        DataFrame resultDataFrame = sqlContext.sql("Select * from Events");
        resultDataFrame.show(false);
        JavaEsSparkSQL.saveToEs(resultDataFrame, "event/states", ImmutableMap.of("es.mapping.id", "eventId"));
    }
});

jssc.start();
jssc.awaitTermination();
jssc.stop();

```

Full stacktrace from Spark app  
[https://gist.githubusercontent.com/dkirrane/8485d8d6f4c422310ec69d0e89271b35/raw/25f7ee5dc050b8cf88997bf2e9cd1d35d09a7112/45834.log](https://gist.githubusercontent.com/dkirrane/8485d8d6f4c422310ec69d0e89271b35/raw/25f7ee5dc050b8cf88997bf2e9cd1d35d09a7112/45834.log)

---

<div class="post-metadata">

**Author:** ![DK\_3](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dk_3/32/48003_2.png) [@DK\_3](https://discuss.elastic.co/u/DK_3)\
**Post date:** [March 30, 2016, 6:35pm UTC](https://discuss.elastic.co/t/sparkstreaming-to-elasticesrahc-error-networkclient-connection-timed-out-connect/45834/2 "2016-03-30T18:35:03Z")

</div>

Solved the problem by adding

```
    conf.set("es.nodes.discovery", "false");
    conf.set("es.nodes.data.only", "false");
```

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [April 5, 2016, 2:54pm UTC](https://discuss.elastic.co/t/sparkstreaming-to-elasticesrahc-error-networkclient-connection-timed-out-connect/45834/3 "2016-04-05T14:54:35Z")

</div>

"Interesting" fix. It reduces the number of calls to the cluster and it does allow master nodes to be used - considering you are only using one node, it will end up using that one all the time.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:25pm UTC](https://discuss.elastic.co/t/sparkstreaming-to-elasticesrahc-error-networkclient-connection-timed-out-connect/45834/4 "2017-07-06T13:25:28Z")

</div>


