# Is Spark useful for reindexation?

**URL:** <https://discuss.elastic.co/t/is-spark-useful-for-reindexation/75919>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [February 21, 2017, 4:36pm UTC](https://discuss.elastic.co/t/is-spark-useful-for-reindexation/75919 "2017-02-21T16:36:24Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![geantbrun](https://avatars.discourse-cdn.com/v4/letter/g/ecb155/32.png) [@geantbrun](https://discuss.elastic.co/u/geantbrun)\
**Post date:** [February 21, 2017, 4:36pm UTC](https://discuss.elastic.co/t/is-spark-useful-for-reindexation/75919/1 "2017-02-21T16:36:24Z")

</div>

Hi all,  
I'm at the beginning process of writing an application to reindex all my data (in order to update my mapping, include new fields, etc). I would like to know the best way to go (in terms of speed, I have Tbytes of data). Is Reindex api the good solution? Or would it be useful to use Spark to parallelize the task and make performance gains? Any pointers/links to do such a thing? Any help is greatly appreciated!

---

<div class="post-metadata">

**Author:** ![jspooner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jspooner/32/12984_2.png) [@jspooner](https://discuss.elastic.co/u/jspooner)\
**Post date:** [February 21, 2017, 4:57pm UTC](https://discuss.elastic.co/t/is-spark-useful-for-reindexation/75919/2 "2017-02-21T16:57:48Z")

</div>

Three people on my engineering and ops teams have put a lot of effort into tuning bulk indexing and have had limited success. The best indexing rate we can get is ~70k per second. I have this thread to discuss that issue [Spark Bulk Import Performance Benchmarks](https://discuss.elastic.co/t/spark-bulk-import-performance-benchmarks/75110/4)

What version of Elasticsearch are you using?

This article is a little old but might help

> **[How we reindexed 36 billion documents in 5 days within the same Elasticsearch...](https://thoughts.t37.net/how-we-reindexed-36-billions-documents-in-5-days-within-the-same-elasticsearch-cluster-cd9c054d1db8?gi=1386cfc0c8a4)**
>
> At Synthesio, we use ElasticSearch at various places to run complex queries that fetch up to 50 million rich documents out of tens of…

---

<div class="post-metadata">

**Author:** ![geantbrun](https://avatars.discourse-cdn.com/v4/letter/g/ecb155/32.png) [@geantbrun](https://discuss.elastic.co/u/geantbrun)\
**Post date:** [February 21, 2017, 5:27pm UTC](https://discuss.elastic.co/t/is-spark-useful-for-reindexation/75919/3 "2017-02-21T17:27:03Z")

</div>

Thank you for your answer Jonathan. We're using ES 2.4.2 but eventually will upgrade to 5.0. So, based on your experience, would you say that it's worthwhile to use Spark or do you get the best performance using something more basic like "scroll to retrieve batches of documents from the old index, and the bulk API to push them into the new index" (ref [reindexing your data](https://www.elastic.co/guide/en/elasticsearch/guide/current/reindex.html))?

---

<div class="post-metadata">

**Author:** ![jspooner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jspooner/32/12984_2.png) [@jspooner](https://discuss.elastic.co/u/jspooner)\
**Post date:** [February 21, 2017, 11:01pm UTC](https://discuss.elastic.co/t/is-spark-useful-for-reindexation/75919/4 "2017-02-21T23:01:50Z")

</div>

Yes, I would start off with the technique on [reindexing your data](https://www.elastic.co/guide/en/elasticsearch/guide/current/reindex.html) because it has less setup and overhead than Spark.

---

<div class="post-metadata">

**Author:** ![geantbrun](https://avatars.discourse-cdn.com/v4/letter/g/ecb155/32.png) [@geantbrun](https://discuss.elastic.co/u/geantbrun)\
**Post date:** [February 21, 2017, 11:03pm UTC](https://discuss.elastic.co/t/is-spark-useful-for-reindexation/75919/5 "2017-02-21T23:03:45Z")

</div>

And in terms of performance, how does it compare? Approximately the same?

---

<div class="post-metadata">

**Author:** ![jspooner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jspooner/32/12984_2.png) [@jspooner](https://discuss.elastic.co/u/jspooner)\
**Post date:** [February 21, 2017, 11:26pm UTC](https://discuss.elastic.co/t/is-spark-useful-for-reindexation/75919/6 "2017-02-21T23:26:30Z")

</div>

I don't know the performance of the scroll method but I can say I've spent 6 months optimizing Spark and only get ~70k docs per second. I have 100 TB of data and ended up sampling it before indexing in elasticsearch due to the expensive indexing rate.

Do you have a lot of Spark experience?

---

<div class="post-metadata">

**Author:** ![geantbrun](https://avatars.discourse-cdn.com/v4/letter/g/ecb155/32.png) [@geantbrun](https://discuss.elastic.co/u/geantbrun)\
**Post date:** [February 22, 2017, 12:32am UTC](https://discuss.elastic.co/t/is-spark-useful-for-reindexation/75919/7 "2017-02-22T00:32:11Z")

</div>

Unfortunately no, I'm a beginner only. But your 70K/sec is quite impressive for me who obtain only something like 70K/min! Of course it depends of a lot of things (ex: my average size of doc is 10K, and you?). Sorry for the silly question but what do you mean by "I ended up sampling if before indexing ... due to the expensive indexing rate"? Thanks again for your tips.

---

<div class="post-metadata">

**Author:** ![jspooner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jspooner/32/12984_2.png) [@jspooner](https://discuss.elastic.co/u/jspooner)\
**Post date:** [February 22, 2017, 4:10pm UTC](https://discuss.elastic.co/t/is-spark-useful-for-reindexation/75919/8 "2017-02-22T16:10:28Z")

</div>

Yeah if your a beginner to Spark and don't have the luxury of time I'd use the non-spark method.

70k per sec for a 20 node elasticsearch cluster is not super impressive at all. That results in 6,048,000,000 documents indexed a day and when you have TBs of data your import will run several days. So what I did is sample my original data set by 1/25th. My use case is for data analysis and I have the luxury of producing rough estimates.

---

<div class="post-metadata">

**Author:** ![geantbrun](https://avatars.discourse-cdn.com/v4/letter/g/ecb155/32.png) [@geantbrun](https://discuss.elastic.co/u/geantbrun)\
**Post date:** [February 22, 2017, 4:12pm UTC](https://discuss.elastic.co/t/is-spark-useful-for-reindexation/75919/9 "2017-02-22T16:12:13Z")

</div>

> [@jspooner](#):
>
> dexed

I understand. And, if I may ask again, what is your average document size?

---

<div class="post-metadata">

**Author:** ![jspooner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jspooner/32/12984_2.png) [@jspooner](https://discuss.elastic.co/u/jspooner)\
**Post date:** [February 24, 2017, 12:17am UTC](https://discuss.elastic.co/t/is-spark-useful-for-reindexation/75919/10 "2017-02-24T00:17:40Z")

</div>

Very small. Around 160 bytes.

---

<div class="post-metadata">

**Author:** ![geantbrun](https://avatars.discourse-cdn.com/v4/letter/g/ecb155/32.png) [@geantbrun](https://discuss.elastic.co/u/geantbrun)\
**Post date:** [February 24, 2017, 7:58pm UTC](https://discuss.elastic.co/t/is-spark-useful-for-reindexation/75919/11 "2017-02-24T19:58:38Z")

</div>

Thanks again @jspooner. I took a look at the code you gently provided here ([https://gist.github.com/jspooner/ccba83a5a8f36fe1276350ef838be38d](https://gist.github.com/jspooner/ccba83a5a8f36fe1276350ef838be38d)). Does the spark code only the 19-lines long scala program called push\_to\_es.scala?

---

<div class="post-metadata">

**Author:** ![jspooner](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jspooner/32/12984_2.png) [@jspooner](https://discuss.elastic.co/u/jspooner)\
**Post date:** [February 24, 2017, 9:16pm UTC](https://discuss.elastic.co/t/is-spark-useful-for-reindexation/75919/12 "2017-02-24T21:16:02Z")

</div>

**push\_to\_es.scala** is only a snippet of the code.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 24, 2017, 9:16pm UTC](https://discuss.elastic.co/t/is-spark-useful-for-reindexation/75919/13 "2017-03-24T21:16:22Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
