# Elasticsearch and Hadoop Questions

**URL:** <https://discuss.elastic.co/t/elasticsearch-and-hadoop-questions/17933>\
**Category:** Elasticsearch\
**Created:** [June 5, 2014, 5:41pm UTC](https://discuss.elastic.co/t/elasticsearch-and-hadoop-questions/17933 "2014-06-05T17:41:34Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![ES\_USER1](https://avatars.discourse-cdn.com/v4/letter/e/74df32/32.png) [@ES\_USER1](https://discuss.elastic.co/u/ES_USER1)\
**Post date:** [June 5, 2014, 5:41pm UTC](https://discuss.elastic.co/t/elasticsearch-and-hadoop-questions/17933/1 "2014-06-05T17:41:34Z")

</div>

Try as I might and I have read all the stuff I can find on ES' website  
about this I understand somewhat how the integration works but not the  
actual nuts and bolts of it.

For example:

Is Hadoop just storing the files that would normally be stored in the local  
filesystem for the ES indexes or is it storing the data that would normally  
be in those indexes and just accessed through es-hadoop?

If it is the latter how do you go about determining whatto set for the  
number of nodes and shards.

If anyone has any information on this or even better yet a place to point  
me to that has better references so that I can research this on my own it  
would be much appreciated.

Thanks.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/b78f2fa6-42c9-4ae7-a4ab-aacbc2c53293%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/b78f2fa6-42c9-4ae7-a4ab-aacbc2c53293%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [June 5, 2014, 10:45pm UTC](https://discuss.elastic.co/t/elasticsearch-and-hadoop-questions/17933/2 "2014-06-05T22:45:02Z")

</div>

Think of es-hadoop as a connector between Hadoop and Elasticsearch. You  
would use it to index data in Hadoop to ES or run queries in ES directly  
from Hadoop.  
Where does ES store the data? That depends on its configuration (completely  
separate from es-hadoop itself). In general (and the default) is to store  
it onto the local file-system. If you want to use it on a shared  
file-system or HDFS you can easily do that by mounting it locally (for  
example, mount HDFS through NFS as a local disk) and point ES to it. ES is  
happy to work with it however the performance will be _significantly_  
degraded and most of the real-time nature of it will go down the window  
since HDFS is a distributed file-system (and thus even basic operations  
like opening a file or closing a file mean at least one call over the  
network) plus you're giving up the amazing OS file-system cache (since the  
fs is not local). If the FS is slow, anything that sits on top of it (like  
ES) will be slow as well.

Hope this helps,

P.S. By the way, if you want/need to snapshot/restore data to/from ES  
from/to HDFS you can use the HDFS repository (more info here:

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

)

On Thu, Jun 5, 2014 at 8:41 PM, ES USER [es.user.2014@gmail.com](mailto:es.user.2014@gmail.com) wrote:

> Try as I might and I have read all the stuff I can find on ES' website  
> about this I understand somewhat how the integration works but not the  
> actual nuts and bolts of it.
> 
> For example:
> 
> Is Hadoop just storing the files that would normally be stored in the  
> local filesystem for the ES indexes or is it storing the data that would  
> normally be in those indexes and just accessed through es-hadoop?
> 
> If it is the latter how do you go about determining whatto set for the  
> number of nodes and shards.
> 
> If anyone has any information on this or even better yet a place to point  
> me to that has better references so that I can research this on my own it  
> would be much appreciated.
> 
> Thanks.
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/b78f2fa6-42c9-4ae7-a4ab-aacbc2c53293%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/b78f2fa6-42c9-4ae7-a4ab-aacbc2c53293%40googlegroups.com)  
> [https://groups.google.com/d/msgid/elasticsearch/b78f2fa6-42c9-4ae7-a4ab-aacbc2c53293%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/b78f2fa6-42c9-4ae7-a4ab-aacbc2c53293%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAJogdmfSrZ49XHgGfnRcfHQTH%3DSy%2B18RQ\_%2BwEqR8MYOuZr%3DjZQ%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAJogdmfSrZ49XHgGfnRcfHQTH%3DSy%2B18RQ_%2BwEqR8MYOuZr%3DjZQ%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![ES\_USER1](https://avatars.discourse-cdn.com/v4/letter/e/74df32/32.png) [@ES\_USER1](https://discuss.elastic.co/u/ES_USER1)\
**Post date:** [June 6, 2014, 11:32am UTC](https://discuss.elastic.co/t/elasticsearch-and-hadoop-questions/17933/3 "2014-06-06T11:32:20Z")

</div>

So if I understand you correctly if the data is stored in Hadoop then  
es-hadoop is really just acting as a job manager? If that is the case what  
is the rule of thumb on how many ES nodes and shard should be set?

On Thursday, June 5, 2014 6:45:09 PM UTC-4, Costin Leau wrote:

> Think of es-hadoop as a connector between Hadoop and Elasticsearch. You  
> would use it to index data in Hadoop to ES or run queries in ES directly  
> from Hadoop.  
> Where does ES store the data? That depends on its configuration  
> (completely separate from es-hadoop itself). In general (and the default)  
> is to store it onto the local file-system. If you want to use it on a  
> shared file-system or HDFS you can easily do that by mounting it locally  
> (for example, mount HDFS through NFS as a local disk) and point ES to it.  
> ES is happy to work with it however the performance will be _significantly_  
> degraded and most of the real-time nature of it will go down the window  
> since HDFS is a distributed file-system (and thus even basic operations  
> like opening a file or closing a file mean at least one call over the  
> network) plus you're giving up the amazing OS file-system cache (since the  
> fs is not local). If the FS is slow, anything that sits on top of it (like  
> ES) will be slow as well.
> 
> Hope this helps,
> 
> P.S. By the way, if you want/need to snapshot/restore data to/from ES  
> from/to HDFS you can use the HDFS repository (more info here:  
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/hadoop/current/setup.html)  
> )
> 
> On Thu, Jun 5, 2014 at 8:41 PM, ES USER \<[es.use...@gmail.com](mailto:es.use...@gmail.com) \<javascript:\>
> 
> > wrote:
> 
> > Try as I might and I have read all the stuff I can find on ES' website  
> > about this I understand somewhat how the integration works but not the  
> > actual nuts and bolts of it.
> > 
> > For example:
> > 
> > Is Hadoop just storing the files that would normally be stored in the  
> > local filesystem for the ES indexes or is it storing the data that would  
> > normally be in those indexes and just accessed through es-hadoop?
> > 
> > If it is the latter how do you go about determining whatto set for the  
> > number of nodes and shards.
> > 
> > If anyone has any information on this or even better yet a place to point  
> > me to that has better references so that I can research this on my own it  
> > would be much appreciated.
> > 
> > Thanks.
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/b78f2fa6-42c9-4ae7-a4ab-aacbc2c53293%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/b78f2fa6-42c9-4ae7-a4ab-aacbc2c53293%40googlegroups.com)  
> > [https://groups.google.com/d/msgid/elasticsearch/b78f2fa6-42c9-4ae7-a4ab-aacbc2c53293%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/b78f2fa6-42c9-4ae7-a4ab-aacbc2c53293%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > .  
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/123c4ed3-077b-4e9f-a838-fa372aea109a%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/123c4ed3-077b-4e9f-a838-fa372aea109a%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [June 6, 2014, 4:39pm UTC](https://discuss.elastic.co/t/elasticsearch-and-hadoop-questions/17933/6 "2014-06-06T16:39:38Z")

</div>

Adding to what Georgi wrote, es-hadoop does not create the shards for you -  
that's up to you or index templates (which I highly recommend). However  
es-hadoop is aware of the target shards and will use them to parallelize  
the reads/writes (such as one task per shard).

On Fri, Jun 6, 2014 at 2:45 PM, Georgi Ivanov [georgi.r.ivanov@gmail.com](mailto:georgi.r.ivanov@gmail.com)  
wrote:

> and i don't think this anyhow related with number of shards and nodes
> 
> On Thursday, June 5, 2014 7:41:34 PM UTC+2, ES USER wrote:
> 
> > Try as I might and I have read all the stuff I can find on ES' website  
> > about this I understand somewhat how the integration works but not the  
> > actual nuts and bolts of it.
> > 
> > For example:
> > 
> > Is Hadoop just storing the files that would normally be stored in the  
> > local filesystem for the ES indexes or is it storing the data that would  
> > normally be in those indexes and just accessed through es-hadoop?
> > 
> > If it is the latter how do you go about determining whatto set for the  
> > number of nodes and shards.
> > 
> > If anyone has any information on this or even better yet a place to point  
> > me to that has better references so that I can research this on my own it  
> > would be much appreciated.
> > 
> > Thanks.
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/90662a91-1557-4f61-86a2-bd2e620aec6f%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/90662a91-1557-4f61-86a2-bd2e620aec6f%40googlegroups.com)  
> [https://groups.google.com/d/msgid/elasticsearch/90662a91-1557-4f61-86a2-bd2e620aec6f%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/90662a91-1557-4f61-86a2-bd2e620aec6f%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .
> 
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAJogdmeDzSDrBLfpTQ3hGxOh1PN4przGkth2-M\_oLdN7VjKYPg%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAJogdmeDzSDrBLfpTQ3hGxOh1PN4przGkth2-M_oLdN7VjKYPg%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![ES\_USER1](https://avatars.discourse-cdn.com/v4/letter/e/74df32/32.png) [@ES\_USER1](https://discuss.elastic.co/u/ES_USER1)\
**Post date:** [June 6, 2014, 6:29pm UTC](https://discuss.elastic.co/t/elasticsearch-and-hadoop-questions/17933/7 "2014-06-06T18:29:32Z")

</div>

I guess the problem I having wrapping my head around is exactly where the  
data is residing and in what format.

If I understand the Georgi's email above is it that you can run map reduce  
jobs against data stored in local ES through by utilizing es-hadoop and you  
can also run ES queries against data in Hadoop utilizing es-hadoop.

Is that correct?

On Friday, June 6, 2014 12:39:44 PM UTC-4, Costin Leau wrote:

> Adding to what Georgi wrote, es-hadoop does not create the shards for you
> 
> - that's up to you or index templates (which I highly recommend). However  
> es-hadoop is aware of the target shards and will use them to parallelize  
> the reads/writes (such as one task per shard).
> 
> On Fri, Jun 6, 2014 at 2:45 PM, Georgi Ivanov \<[georgi....@gmail.com](mailto:georgi....@gmail.com)  
> \<javascript:\>\> wrote:
> 
> > and i don't think this anyhow related with number of shards and nodes
> > 
> > On Thursday, June 5, 2014 7:41:34 PM UTC+2, ES USER wrote:
> > 
> > > Try as I might and I have read all the stuff I can find on ES' website  
> > > about this I understand somewhat how the integration works but not the  
> > > actual nuts and bolts of it.
> > > 
> > > For example:
> > > 
> > > Is Hadoop just storing the files that would normally be stored in the  
> > > local filesystem for the ES indexes or is it storing the data that would  
> > > normally be in those indexes and just accessed through es-hadoop?
> > > 
> > > If it is the latter how do you go about determining whatto set for the  
> > > number of nodes and shards.
> > > 
> > > If anyone has any information on this or even better yet a place to  
> > > point me to that has better references so that I can research this on my  
> > > own it would be much appreciated.
> > > 
> > > Thanks.
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/90662a91-1557-4f61-86a2-bd2e620aec6f%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/90662a91-1557-4f61-86a2-bd2e620aec6f%40googlegroups.com)  
> > [https://groups.google.com/d/msgid/elasticsearch/90662a91-1557-4f61-86a2-bd2e620aec6f%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/90662a91-1557-4f61-86a2-bd2e620aec6f%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > .
> > 
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/ed729795-a7d6-4320-9da2-16b214e653b0%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ed729795-a7d6-4320-9da2-16b214e653b0%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [June 6, 2014, 9:02pm UTC](https://discuss.elastic.co/t/elasticsearch-and-hadoop-questions/17933/8 "2014-06-06T21:02:51Z")

</div>

ES stores data in its own internal format, which typically resides locally.  
What you are stating is partially correct - with the connector you would  
move/copy data between Hadoop and ES since, in order for ES to work with  
data, it needs to actually index it (that is, to see it).  
So you would use es-hadoop to index data from Hadoop in ES or/and query ES  
directly from Hadoop.

On Fri, Jun 6, 2014 at 9:29 PM, ES USER [es.user.2014@gmail.com](mailto:es.user.2014@gmail.com) wrote:

> I guess the problem I having wrapping my head around is exactly where the  
> data is residing and in what format.
> 
> If I understand the Georgi's email above is it that you can run map reduce  
> jobs against data stored in local ES through by utilizing es-hadoop and you  
> can also run ES queries against data in Hadoop utilizing es-hadoop.
> 
> Is that correct?
> 
> On Friday, June 6, 2014 12:39:44 PM UTC-4, Costin Leau wrote:
> 
> > Adding to what Georgi wrote, es-hadoop does not create the shards for you
> > 
> > - that's up to you or index templates (which I highly recommend). However  
> > es-hadoop is aware of the target shards and will use them to parallelize  
> > the reads/writes (such as one task per shard).
> > 
> > On Fri, Jun 6, 2014 at 2:45 PM, Georgi Ivanov [georgi....@gmail.com](mailto:georgi....@gmail.com)  
> > wrote:
> > 
> > > and i don't think this anyhow related with number of shards and nodes
> > > 
> > > On Thursday, June 5, 2014 7:41:34 PM UTC+2, ES USER wrote:
> > > 
> > > > Try as I might and I have read all the stuff I can find on ES' website  
> > > > about this I understand somewhat how the integration works but not the  
> > > > actual nuts and bolts of it.
> > > > 
> > > > For example:
> > > > 
> > > > Is Hadoop just storing the files that would normally be stored in the  
> > > > local filesystem for the ES indexes or is it storing the data that would  
> > > > normally be in those indexes and just accessed through es-hadoop?
> > > > 
> > > > If it is the latter how do you go about determining whatto set for the  
> > > > number of nodes and shards.
> > > > 
> > > > If anyone has any information on this or even better yet a place to  
> > > > point me to that has better references so that I can research this on my  
> > > > own it would be much appreciated.
> > > > 
> > > > Thanks.
> > > 
> > > --  
> > > You received this message because you are subscribed to the Google  
> > > Groups "elasticsearch" group.  
> > > To unsubscribe from this group and stop receiving emails from it, send  
> > > an email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com).  
> > > To view this discussion on the web visit [https://groups.google.com/d/](https://groups.google.com/d/)  
> > > msgid/elasticsearch/90662a91-1557-4f61-86a2-bd2e620aec6f%  
> > > [40googlegroups.com](http://40googlegroups.com)  
> > > [https://groups.google.com/d/msgid/elasticsearch/90662a91-1557-4f61-86a2-bd2e620aec6f%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/90662a91-1557-4f61-86a2-bd2e620aec6f%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > > .
> > > 
> > > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/ed729795-a7d6-4320-9da2-16b214e653b0%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ed729795-a7d6-4320-9da2-16b214e653b0%40googlegroups.com)  
> > [https://groups.google.com/d/msgid/elasticsearch/ed729795-a7d6-4320-9da2-16b214e653b0%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/ed729795-a7d6-4320-9da2-16b214e653b0%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > .
> 
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAJogdmeacuVOXNYcwdYHBg69TAotrqvyuzre\_JeUK-RfAcFBXA%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAJogdmeacuVOXNYcwdYHBg69TAotrqvyuzre_JeUK-RfAcFBXA%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![ES\_USER1](https://avatars.discourse-cdn.com/v4/letter/e/74df32/32.png) [@ES\_USER1](https://discuss.elastic.co/u/ES_USER1)\
**Post date:** [June 9, 2014, 11:47am UTC](https://discuss.elastic.co/t/elasticsearch-and-hadoop-questions/17933/9 "2014-06-09T11:47:14Z")

</div>

Thanks. So just one final question. From what you said above that means  
that you can not run ES queries on data in Hadoop over something like a 6  
month time range without it having to pull in all that data and index it  
first. And I am assuming that the opposite is all correct that Hadoop can  
not run jobs on data in ES without it first pulling in that data to its  
storage first.

On Friday, June 6, 2014 5:03:03 PM UTC-4, Costin Leau wrote:

> ES stores data in its own internal format, which typically resides  
> locally. What you are stating is partially correct - with the connector you  
> would move/copy data between Hadoop and ES since, in order for ES to work  
> with data, it needs to actually index it (that is, to see it).  
> So you would use es-hadoop to index data from Hadoop in ES or/and query ES  
> directly from Hadoop.
> 
> On Fri, Jun 6, 2014 at 9:29 PM, ES USER \<[es.use...@gmail.com](mailto:es.use...@gmail.com) \<javascript:\>
> 
> > wrote:
> 
> > I guess the problem I having wrapping my head around is exactly where the  
> > data is residing and in what format.
> > 
> > If I understand the Georgi's email above is it that you can run map  
> > reduce jobs against data stored in local ES through by utilizing es-hadoop  
> > and you can also run ES queries against data in Hadoop utilizing es-hadoop.
> > 
> > Is that correct?
> > 
> > On Friday, June 6, 2014 12:39:44 PM UTC-4, Costin Leau wrote:
> > 
> > > Adding to what Georgi wrote, es-hadoop does not create the shards for  
> > > you - that's up to you or index templates (which I highly recommend).  
> > > However es-hadoop is aware of the target shards and will use them to  
> > > parallelize the reads/writes (such as one task per shard).
> > > 
> > > On Fri, Jun 6, 2014 at 2:45 PM, Georgi Ivanov [georgi....@gmail.com](mailto:georgi....@gmail.com)  
> > > wrote:
> > > 
> > > > and i don't think this anyhow related with number of shards and nodes
> > > > 
> > > > On Thursday, June 5, 2014 7:41:34 PM UTC+2, ES USER wrote:
> > > > 
> > > > > Try as I might and I have read all the stuff I can find on ES' website  
> > > > > about this I understand somewhat how the integration works but not the  
> > > > > actual nuts and bolts of it.
> > > > > 
> > > > > For example:
> > > > > 
> > > > > Is Hadoop just storing the files that would normally be stored in the  
> > > > > local filesystem for the ES indexes or is it storing the data that would  
> > > > > normally be in those indexes and just accessed through es-hadoop?
> > > > > 
> > > > > If it is the latter how do you go about determining whatto set for the  
> > > > > number of nodes and shards.
> > > > > 
> > > > > If anyone has any information on this or even better yet a place to  
> > > > > point me to that has better references so that I can research this on my  
> > > > > own it would be much appreciated.
> > > > > 
> > > > > Thanks.
> > > > 
> > > > --  
> > > > You received this message because you are subscribed to the Google  
> > > > Groups "elasticsearch" group.  
> > > > To unsubscribe from this group and stop receiving emails from it, send  
> > > > an email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com).  
> > > > To view this discussion on the web visit [https://groups.google.com/d/](https://groups.google.com/d/)  
> > > > msgid/elasticsearch/90662a91-1557-4f61-86a2-bd2e620aec6f%  
> > > > [40googlegroups.com](http://40googlegroups.com)  
> > > > [https://groups.google.com/d/msgid/elasticsearch/90662a91-1557-4f61-86a2-bd2e620aec6f%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/90662a91-1557-4f61-86a2-bd2e620aec6f%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > > > .
> > > > 
> > > > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).
> > > 
> > > --  
> > > You received this message because you are subscribed to the Google Groups  
> > > "elasticsearch" group.  
> > > To unsubscribe from this group and stop receiving emails from it, send an  
> > > email to [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > > To view this discussion on the web visit  
> > > [https://groups.google.com/d/msgid/elasticsearch/ed729795-a7d6-4320-9da2-16b214e653b0%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/ed729795-a7d6-4320-9da2-16b214e653b0%40googlegroups.com)  
> > > [https://groups.google.com/d/msgid/elasticsearch/ed729795-a7d6-4320-9da2-16b214e653b0%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/ed729795-a7d6-4320-9da2-16b214e653b0%40googlegroups.com?utm_medium=email&utm_source=footer)  
> > > .
> > 
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/59691ff7-25b7-4888-bf8b-ca1525637728%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/59691ff7-25b7-4888-bf8b-ca1525637728%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:23am UTC](https://discuss.elastic.co/t/elasticsearch-and-hadoop-questions/17933/11 "2017-07-06T01:23:44Z")

</div>


