# Use Snapshot from Hadoop?

**URL:** <https://discuss.elastic.co/t/use-snapshot-from-hadoop/63773>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [October 24, 2016, 4:23pm UTC](https://discuss.elastic.co/t/use-snapshot-from-hadoop/63773 "2016-10-24T16:23:13Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![ebuildy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ebuildy/32/6070_2.png) [@ebuildy](https://discuss.elastic.co/u/ebuildy)\
**Post date:** [October 24, 2016, 4:23pm UTC](https://discuss.elastic.co/t/use-snapshot-from-hadoop/63773/1 "2016-10-24T16:23:13Z")

</div>

Just curious, is it possible to read a snapshot with Java? (Especially with Hadoop techo. such as Hive or Spark or PIG).

We have a ton of snapshots, would like read data but I don't want to restore them one by one for that.

Many thanks,

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [November 11, 2016, 5:30pm UTC](https://discuss.elastic.co/t/use-snapshot-from-hadoop/63773/2 "2016-11-11T17:30:12Z")

</div>

There was some talk a while ago about potentially supporting this in some fashion, but ultimately it was decided against. When you create a snapshot in Elasticsearch, you're just moving the Lucene indexing files to a block storage location, but those Lucene files can change from release to release and require the same version of reader from ES. It's an idea that has lots of potential for performance improvement, but the current drawbacks we're finding with it means that we're not pursuing it at the moment.

---

<div class="post-metadata">

**Author:** ![ebuildy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ebuildy/32/6070_2.png) [@ebuildy](https://discuss.elastic.co/u/ebuildy)\
**Post date:** [November 11, 2016, 7:58pm UTC](https://discuss.elastic.co/t/use-snapshot-from-hadoop/63773/3 "2016-11-11T19:58:58Z")

</div>

D'ho ;-(

Currently we are using Spark, with fantastic Es4Hadoop plugin, to export ES index into Parquet files, this work fine, but this involves 2 technologies (Spark / Parquet), whereas I would prefer keep ES only.

Thanks you for the explanation.

---

<div class="post-metadata">

**Author:** ![suanmeiguo](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/suanmeiguo/32/11758_2.png) [@suanmeiguo](https://discuss.elastic.co/u/suanmeiguo)\
**Post date:** [June 8, 2017, 6:48pm UTC](https://discuss.elastic.co/t/use-snapshot-from-hadoop/63773/4 "2017-06-08T18:48:45Z")

</div>

I wanna vote on this as well. My spark job read from the cluster directly and it's a big load on the cluster. This feature can definitely help on batch jobs.

---

<div class="post-metadata">

**Author:** ![ebuildy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ebuildy/32/6070_2.png) [@ebuildy](https://discuss.elastic.co/u/ebuildy)\
**Post date:** [June 9, 2017, 10:49am UTC](https://discuss.elastic.co/t/use-snapshot-from-hadoop/63773/5 "2017-06-09T10:49:29Z")

</div>

Yah absolutely, especially in big-data world, it make no sense to use HTTP for data transfert, so slow (also es4hadoop dont use gzip grrr).

In my case, I am going to move from elasticsearch to [druid.io](http://druid.io) , which is more big-data friendly.

---

<div class="post-metadata">

**Author:** ![prabhushrikant](https://avatars.discourse-cdn.com/v4/letter/p/dec6dc/32.png) [@prabhushrikant](https://discuss.elastic.co/u/prabhushrikant)\
**Post date:** [August 4, 2017, 7:46am UTC](https://discuss.elastic.co/t/use-snapshot-from-hadoop/63773/6 "2017-08-04T07:46:48Z")

</div>

Hi , how do you create parquet files from an ES index snapshot?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 1, 2020, 11:18pm UTC](https://discuss.elastic.co/t/use-snapshot-from-hadoop/63773/7 "2020-06-01T23:18:31Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
