# Indexing 10 million documents using Pyspark

**URL:** <https://discuss.elastic.co/t/indexing-10-million-documents-using-pyspark/77453>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [March 6, 2017, 7:42am UTC](https://discuss.elastic.co/t/indexing-10-million-documents-using-pyspark/77453 "2017-03-06T07:42:41Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![praveen\_S](https://avatars.discourse-cdn.com/v4/letter/p/eb9ed0/32.png) [@praveen\_S](https://discuss.elastic.co/u/praveen_S)\
**Post date:** [March 6, 2017, 7:42am UTC](https://discuss.elastic.co/t/indexing-10-million-documents-using-pyspark/77453/1 "2017-03-06T07:42:41Z")

</div>

I have a ES cluster with 3 dedicated master nodes and 5 dedicated data nodes in my cluster.  
When I try to index using pyspark, everything goes fine till the half way mark. Then, I see warnings reported by spark which says:  
Cannot detect es version- typically this happens when the cluster is not accessible or if es.nodes.wan.only parameter is incorrect.

I am also passing all the nodes in my cluster in es.nodes parameter.  
I have actually set the parameter to true.

What am I missing here?

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [March 12, 2017, 5:46pm UTC](https://discuss.elastic.co/t/indexing-10-million-documents-using-pyspark/77453/2 "2017-03-12T17:46:52Z")

</div>

@praveen_S if you dig into the logs, there should be a corresponding reason why it cannot detect the ES version. This occurs during the initial task start up, so perhaps some of your nodes cannot find a route to the specified Elasticsearch nodes?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 9, 2017, 5:46pm UTC](https://discuss.elastic.co/t/indexing-10-million-documents-using-pyspark/77453/3 "2017-04-09T17:46:53Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
