# Spark Bulk Import Performance Benchmarks

**URL:** <https://discuss.elastic.co/t/spark-bulk-import-performance-benchmarks/75110>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [February 14, 2017, 8:56pm UTC](https://discuss.elastic.co/t/spark-bulk-import-performance-benchmarks/75110 "2017-02-14T20:56:18Z")\
**Posts on this page:** 1\
**Showing post:** 4

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [February 17, 2017, 6:32pm UTC](https://discuss.elastic.co/t/spark-bulk-import-performance-benchmarks/75110/4 "2017-02-17T18:32:02Z")

</div>

> [@jspooner](#):
>
> However in 5.2 "client nodes" are no longer listed as a node type so we swapped out the client nodes for ingest.

Client nodes as an explicit node type have gone away, but they are still around in the sense of nodes that have "master", "data" and "ingest" features turned off. Every node in Elasticsearch is technically a client node, it's just that when we have the option to target client nodes only, we search for nodes that have no roles in 5.x.

That said, if your cluster is transient with no search load, it might make sense to zero in on the default node targeting, which is directly to datanodes. Would you be able to share your job configurations and cluster layout/index settings here? Writing explicitly to datanodes can sometimes be less advantageous when using more complex settings (like skewed shard/node sizes, or multi-index writing).

---

_[View the full topic](https://discuss.elastic.co/t/spark-bulk-import-performance-benchmarks/75110)._
