# Catching exceptions from saveToEs (elasticsearch-spark)

**URL:** <https://discuss.elastic.co/t/catching-exceptions-from-savetoes-elasticsearch-spark/53745>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [June 23, 2016, 7:53am UTC](https://discuss.elastic.co/t/catching-exceptions-from-savetoes-elasticsearch-spark/53745 "2016-06-23T07:53:55Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![nvitucci](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nvitucci/32/22892_2.png) [@nvitucci](https://discuss.elastic.co/u/nvitucci)\
**Post date:** [June 23, 2016, 7:53am UTC](https://discuss.elastic.co/t/catching-exceptions-from-savetoes-elasticsearch-spark/53745/1 "2016-06-23T07:53:55Z")

</div>

Hello,

I am writing an RDD to Elasticsearch using the `saveToEs` method from elasticsearch-spark. The RDD might contain documents that Elasticsearch rejects with a `org.elasticsearch.hadoop.rest.EsHadoopInvalidRequest` exception, and I would like to catch the exception(s) in order to just ignore such malformed documents, so that the job does not get interrupted. How can I do this?

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [June 30, 2016, 7:31pm UTC](https://discuss.elastic.co/t/catching-exceptions-from-savetoes-elasticsearch-spark/53745/2 "2016-06-30T19:31:49Z")

</div>

Right now there's no great way to do this in es-hadoop. I do recommend using the functionality provided by Spark's RDDs to transform or filter out any invalid documents before executing the final `saveToEs`. It's unlikely that we would provide options to filter out data when those options are already present in these frameworks.

---

<div class="post-metadata">

**Author:** ![nvitucci](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nvitucci/32/22892_2.png) [@nvitucci](https://discuss.elastic.co/u/nvitucci)\
**Post date:** [July 6, 2016, 1:04pm UTC](https://discuss.elastic.co/t/catching-exceptions-from-savetoes-elasticsearch-spark/53745/3 "2016-07-06T13:04:13Z")

</div>

Well, if there's a failure on the Elasticsearch side, I'd like to be able to fail gracefully - and not to have my whole job fail. In my specific case there is no simple way to do the checks beforehand, so handling the exception would be easier. Do you see any solutions to this? Maybe a `saveToEs` parameter to "ignore" the exceptions and log them somewhere?

---

<div class="post-metadata">

**Author:** ![coding2012](https://avatars.discourse-cdn.com/v4/letter/c/ce7236/32.png) [@coding2012](https://discuss.elastic.co/u/coding2012)\
**Post date:** [March 8, 2017, 5:17pm UTC](https://discuss.elastic.co/t/catching-exceptions-from-savetoes-elasticsearch-spark/53745/4 "2017-03-08T17:17:55Z")

</div>

I have also faced some issues with this. The JSON was just fine, it was an invalid date in my case. I had to look at the Elasticsearch logs to find out. I do need a way to run the job and just store errors in a different place so they can be re-run later.

---

<div class="post-metadata">

**Author:** ![larghir](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/larghir/32/61526_2.png) [@larghir](https://discuss.elastic.co/u/larghir)\
**Post date:** [March 31, 2017, 8:32am UTC](https://discuss.elastic.co/t/catching-exceptions-from-savetoes-elasticsearch-spark/53745/5 "2017-03-31T08:32:18Z")

</div>

I'd be also interested in this feature. It poses some limitations because it will fail the entire job

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:22pm UTC](https://discuss.elastic.co/t/catching-exceptions-from-savetoes-elasticsearch-spark/53745/6 "2017-07-06T13:22:19Z")

</div>


