# \[hadoop\] Getting elasticsearch-hadoop working with Shark

**URL:** <https://discuss.elastic.co/t/hadoop-getting-elasticsearch-hadoop-working-with-shark/15882>\
**Category:** Elasticsearch\
**Created:** [February 19, 2014, 2:03am UTC](https://discuss.elastic.co/t/hadoop-getting-elasticsearch-hadoop-working-with-shark/15882 "2014-02-19T02:03:02Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![Max\_Lang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/max_lang/32/1782_2.png) [@Max\_Lang](https://discuss.elastic.co/u/Max_Lang)\
**Post date:** [February 19, 2014, 2:03am UTC](https://discuss.elastic.co/t/hadoop-getting-elasticsearch-hadoop-working-with-shark/15882/1 "2014-02-19T02:03:02Z")

</div>

I set everything up using this  
guide: [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2) on an ec2  
cluster. I've copied the elasticsearch-hadoop jars into the hive lib  
directory and I have elasticsearch running on localhost:9200. I'm running  
shark in a screen session with --service screenserver and connecting to it  
at the same time using shark -h localhost.

Unfortunately, when I attempt to write data into elasticsearch, it fails.  
Here's an example:

[localhost:10000] shark\> CREATE EXTERNAL TABLE wiki (id BIGINT, title STRING  
, last\_modified STRING, xml STRING, text STRING) ROW FORMAT DELIMITED  
FIELDS TERMINATED BY '\t' LOCATION 's3n://spark-data/wikipedia-sample/';  
Time taken (including network latency): 0.159 seconds  
14/02/19 01:23:33 INFO CliDriver: Time taken (including network latency):  
0.159 seconds

[localhost:10000] shark\> SELECT title FROM wiki LIMIT 1;  
Alpokalja  
Time taken (including network latency): 2.23 seconds  
14/02/19 01:23:48 INFO CliDriver: Time taken (including network latency):  
2.23 seconds

[localhost:10000] shark\> CREATE EXTERNAL TABLE es\_wiki (id BIGINT, title  
STRING, last\_modified STRING, xml STRING, text STRING) STORED BY  
'org.elasticsearch.hadoop.hive.EsStorageHandler' TBLPROPERTIES('es.resource'  
= 'wikipedia/article');  
Time taken (including network latency): 0.061 seconds  
14/02/19 01:33:51 INFO CliDriver: Time taken (including network latency):  
0.061 seconds

[localhost:10000] shark\> INSERT OVERWRITE TABLE es\_wiki SELECT w.id, w.title  
, w.last\_modified, w.xml, w.text FROM wiki w;  
[Hive Error]: Query returned non-zero code: 9, cause: FAILED: Execution  
Error, return code -101 from shark.execution.SparkTask  
Time taken (including network latency): 3.575 seconds  
14/02/19 01:34:42 INFO CliDriver: Time taken (including network latency):  
3.575 seconds

_The stack trace looks like this:_

org.apache.hadoop.hive.ql.metadata.HiveException  
(org.apache.hadoop.hive.ql.metadata.HiveException: java.io.IOException: Out  
of nodes and retries; caught exception)

org.apache.hadoop.hive.ql.exec.FileSinkOperator.processOp(FileSinkOperator.java:602)  
shark.execution.FileSinkOperator$$anonfun$processPartition$1.apply(FileSinkOperator.scala:84)  
shark.execution.FileSinkOperator$$anonfun$processPartition$1.apply(FileSinkOperator.scala:81)  
scala.collection.Iterator$class.foreach(Iterator.scala:772)  
scala.collection.Iterator$$anon$19.foreach(Iterator.scala:399)  
shark.execution.FileSinkOperator.processPartition(FileSinkOperator.scala:81)  
shark.execution.FileSinkOperator$.writeFiles$1(FileSinkOperator.scala:207)  
shark.execution.FileSinkOperator$$anonfun$executeProcessFileSinkPartition$1.apply(FileSinkOperator.scala:211)  
shark.execution.FileSinkOperator$$anonfun$executeProcessFileSinkPartition$1.apply(FileSinkOperator.scala:211)  
org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:107)  
org.apache.spark.scheduler.Task.run(Task.scala:53)  
org.apache.spark.executor.Executor$TaskRunner$$anonfun$run$1.apply$mcV$sp(Executor.scala:215)  
org.apache.spark.deploy.SparkHadoopUtil.runAsUser(SparkHadoopUtil.scala:50)  
org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:182)  
java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)  
java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)  
java.lang.Thread.run(Thread.java:744  
I should be using Hive 0.9.0, shark 0.8.1, elasticsearch 1.0.0, Hadoop  
1.0.4, and java 1.7.0\_51  
Based on my cursory look at the hadoop and elasticsearch-hadoop sources, it  
looks like hive is just rethrowing an IOException it's getting from Spark,  
and elasticsearch-hadoop is just hitting those exceptions.  
I suppose my questions are: Does this look like an issue with my  
ES/elasticsearch-hadoop config? And has anyone gotten elasticsearch working  
with Spark/Shark?  
Any ideas/insights are appreciated.  
Thanks,Max

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [February 19, 2014, 6:16am UTC](https://discuss.elastic.co/t/hadoop-getting-elasticsearch-hadoop-working-with-shark/15882/2 "2014-02-19T06:16:38Z")

</div>

The error indicates a network error - namely es-hadoop cannot connect to Elasticsearch on the default (localhost:9200)  
HTTP port. Can you double check whether that's indeed the case (using curl or even telnet on that port) - maybe the  
firewall prevents any connections to be made...  
Also you could try using the latest Hive, 0.12 and a more recent Hadoop such as 1.1.2 or 1.2.1.

Additionally, can you enable TRACE logging in your job on es-hadoop packages org.elasticsearch.hadoop.rest and  
org.elasticsearch.hadoop.mr packages and report back ?

Thanks,

On 19/02/2014 4:03 AM, Max Lang wrote:

> I set everything up using this guide: [Running Shark on EC2 · amplab/shark Wiki · GitHub](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2) on an ec2 cluster. I've  
> copied the elasticsearch-hadoop jars into the hive lib directory and I have elasticsearch running on localhost:9200. I'm  
> running shark in a screen session with --service screenserver and connecting to it at the same time using shark -h  
> localhost.
> 
> Unfortunately, when I attempt to write data into elasticsearch, it fails. Here's an example:
> 
> |  
> [localhost:10000]shark\>CREATE EXTERNAL TABLE wiki (id BIGINT,title STRING,last\_modified STRING,xml STRING,text  
> STRING)ROW FORMAT DELIMITED FIELDS TERMINATED BY '\t'LOCATION 's3n://spark-data/wikipedia-sample/';  
> Timetaken (including network latency):0.159seconds  
> 14/02/1901:23:33INFO CliDriver:Timetaken (including network latency):0.159seconds
> 
> [localhost:10000]shark\>SELECT title FROM wiki LIMIT 1;  
> Alpokalja  
> Timetaken (including network latency):2.23seconds  
> 14/02/1901:23:48INFO CliDriver:Timetaken (including network latency):2.23seconds
> 
> [localhost:10000]shark\>CREATE EXTERNAL TABLE es\_wiki (id BIGINT,title STRING,last\_modified STRING,xml STRING,text  
> STRING)STORED BY 'org.elasticsearch.hadoop.hive.EsStorageHandler'TBLPROPERTIES('es.resource'='wikipedia/article');  
> Timetaken (including network latency):0.061seconds  
> 14/02/1901:33:51INFO CliDriver:Timetaken (including network latency):0.061seconds
> 
> [localhost:10000]shark\>INSERT OVERWRITE TABLE es\_wiki SELECT w.id,w.title,w.last\_modified,w.xml,w.text FROM wiki w;  
> [HiveError]:Queryreturned non-zero code:9,cause:FAILED:ExecutionError,returncode -101fromshark.execution.SparkTask  
> Timetaken (including network latency):3.575seconds  
> 14/02/1901:34:42INFO CliDriver:Timetaken (including network latency):3.575seconds  
> |
> 
> _The stack trace looks like this:_
> 
> org.apache.hadoop.hive.ql.metadata.HiveException (org.apache.hadoop.hive.ql.metadata.HiveException: java.io.IOException:  
> Out of nodes and retries; caught exception)
> 
> org.apache.hadoop.hive.ql.exec.FileSinkOperator.processOp(FileSinkOperator.java:602)shark.execution.FileSinkOperator$$anonfun$processPartition$1.apply(FileSinkOperator.scala:84)shark.execution.FileSinkOperator$$anonfun$processPartition$1.apply(FileSinkOperator.scala:81)scala.collection.Iterator$class.foreach(Iterator.scala:772)scala.collection.Iterator$$anon$19.foreach(Iterator.scala:399)shark.execution.FileSinkOperator.processPartition(FileSinkOperator.scala:81)shark.execution.FileSinkOperator$.writeFiles$1(FileSinkOperator.scala:207)shark.execution.FileSinkOperator$$anonfun$executeProcessFileSinkPartition$1.apply(FileSinkOperator.scala:211)shark.execution.FileSinkOperator$$anonfun$executeProcessFileSinkPartition$1.apply(FileSinkOperator.scala:211)org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:107)org.apache.spark.scheduler.Task.run(Task.scala:53)org.apache.spark.executor.Executor$TaskRunner$$anonfun$run$1.apply$mcV$sp(Executor.scala:215)org.apache.spark.deploy.Sp  
> arkHadoopUtil.runAsUser(SparkHadoopUtil.scala:50)org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:182)java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)java.lang.Thread.run(Thread.java:744  
> I should be using Hive 0.9.0, shark 0.8.1, elasticsearch 1.0.0, Hadoop 1.0.4, and java 1.7.0\_51  
> Based on my cursory look at the hadoop and elasticsearch-hadoop sources, it looks like hive is just rethrowing an  
> IOException it's getting from Spark, and elasticsearch-hadoop is just hitting those exceptions.  
> I suppose my questions are: Does this look like an issue with my ES/elasticsearch-hadoop config? And has anyone gotten  
> elasticsearch working with Spark/Shark?  
> Any ideas/insights are appreciated.  
> Thanks,Max
> 
> --  
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an email to  
> [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
Costin

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/53044C46.70807%40gmail.com](https://groups.google.com/d/msgid/elasticsearch/53044C46.70807%40gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Max\_Lang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/max_lang/32/1782_2.png) [@Max\_Lang](https://discuss.elastic.co/u/Max_Lang)\
**Post date:** [February 19, 2014, 10:02pm UTC](https://discuss.elastic.co/t/hadoop-getting-elasticsearch-hadoop-working-with-shark/15882/3 "2014-02-19T22:02:27Z")

</div>

Hey Costin,

Thanks for the swift reply. I abandoned EC2 to take that out of the  
equation and managed to get everything working locally using the latest  
version of everything (though I realized just now I'm still on hive 0.9).  
I'm guessing you're right about some port connection issue because I  
definitely had ES running on that machine.

I changed hive-log4j.properties and added  
#custom logging levels  
#log4j.logger.xxx=DEBUG  
log4j.logger.org.elasticsearch.hadoop.rest=TRACE  
log4j.logger.org.elasticsearch.hadoop.mr=TRACE

But I didn't see any trace logging. Hopefully I can get it working on EC2  
without issue, but, for the future, is this the correct way to set TRACE  
logging?  
Oh and, for reference, I tried running without ES up and I got the  
following, exceptions:

2014-02-19 13:46:08,803 ERROR shark.SharkDriver  
(Logging.scala:logError(64)) - FAILED: Hive Internal Error:  
java.lang.IllegalStateException(Cannot discover Elasticsearch version)  
java.lang.IllegalStateException: Cannot discover Elasticsearch version  
at  
org.elasticsearch.hadoop.hive.EsStorageHandler.init(EsStorageHandler.java:101)

at  
org.elasticsearch.hadoop.hive.EsStorageHandler.configureOutputJobProperties(EsStorageHandler.java:83)

at  
org.apache.hadoop.hive.ql.plan.PlanUtils.configureJobPropertiesForStorageHandler(PlanUtils.java:706)

at  
org.apache.hadoop.hive.ql.plan.PlanUtils.configureOutputJobPropertiesForStorageHandler(PlanUtils.java:675)

at  
org.apache.hadoop.hive.ql.exec.FileSinkOperator.augmentPlan(FileSinkOperator.java:764)

at  
org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.putOpInsertMap(SemanticAnalyzer.java:1518)

at  
org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genFileSinkPlan(SemanticAnalyzer.java:4337)

at  
org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genPostGroupByBodyPlan(SemanticAnalyzer.java:6207)

at  
org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genBodyPlan(SemanticAnalyzer.java:6138)

at  
org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genPlan(SemanticAnalyzer.java:6764)

at  
shark.parse.SharkSemanticAnalyzer.analyzeInternal(SharkSemanticAnalyzer.scala:149)

at  
org.apache.hadoop.hive.ql.parse.BaseSemanticAnalyzer.analyze(BaseSemanticAnalyzer.java:244)

at shark.SharkDriver.compile(SharkDriver.scala:215)  
at org.apache.hadoop.hive.ql.Driver.compile(Driver.java:336)  
at org.apache.hadoop.hive.ql.Driver.run(Driver.java:895)  
at shark.SharkCliDriver.processCmd(SharkCliDriver.scala:324)  
at org.apache.hadoop.hive.cli.CliDriver.processLine(CliDriver.java:406)  
at shark.SharkCliDriver$.main(SharkCliDriver.scala:232)  
at shark.SharkCliDriver.main(SharkCliDriver.scala)  
Caused by: java.io.IOException: Out of nodes and retries; caught exception  
at  
org.elasticsearch.hadoop.rest.NetworkClient.execute(NetworkClient.java:81)  
at org.elasticsearch.hadoop.rest.RestClient.execute(RestClient.java:221)  
at org.elasticsearch.hadoop.rest.RestClient.execute(RestClient.java:205)  
at org.elasticsearch.hadoop.rest.RestClient.execute(RestClient.java:209)  
at org.elasticsearch.hadoop.rest.RestClient.get(RestClient.java:103)  
at org.elasticsearch.hadoop.rest.RestClient.esVersion(RestClient.java:274)  
at  
org.elasticsearch.hadoop.rest.InitializationUtils.discoverEsVersion(InitializationUtils.java:84)

at  
org.elasticsearch.hadoop.hive.EsStorageHandler.init(EsStorageHandler.java:99)

... 18 more  
Caused by: java.net.ConnectException: Connection refused  
at java.net.PlainSocketImpl.socketConnect(Native Method)  
at  
java.net.AbstractPlainSocketImpl.doConnect(AbstractPlainSocketImpl.java:339)

at  
java.net.AbstractPlainSocketImpl.connectToAddress(AbstractPlainSocketImpl.java:200)

at  
java.net.AbstractPlainSocketImpl.connect(AbstractPlainSocketImpl.java:182)  
at java.net.SocksSocketImpl.connect(SocksSocketImpl.java:391)  
at java.net.Socket.connect(Socket.java:579)  
at java.net.Socket.connect(Socket.java:528)  
at java.net.Socket.(Socket.java:425)  
at java.net.Socket.(Socket.java:280)  
at  
org.apache.commons.httpclient.protocol.DefaultProtocolSocketFactory.createSocket(DefaultProtocolSocketFactory.java:80)

at  
org.apache.commons.httpclient.protocol.DefaultProtocolSocketFactory.createSocket(DefaultProtocolSocketFactory.java:122)

at  
org.apache.commons.httpclient.HttpConnection.open(HttpConnection.java:707)  
at  
org.apache.commons.httpclient.HttpMethodDirector.executeWithRetry(HttpMethodDirector.java:387)

at  
org.apache.commons.httpclient.HttpMethodDirector.executeMethod(HttpMethodDirector.java:171)

at  
org.apache.commons.httpclient.HttpClient.executeMethod(HttpClient.java:397)  
at  
org.apache.commons.httpclient.HttpClient.executeMethod(HttpClient.java:323)  
at  
org.elasticsearch.hadoop.rest.commonshttp.CommonsHttpTransport.execute(CommonsHttpTransport.java:160)

at  
org.elasticsearch.hadoop.rest.NetworkClient.execute(NetworkClient.java:74)  
... 25 more

Let me know if there's anything in particular you'd like me to try on EC2.

(For posterity, the versions I used were: hadoop 2.2.0, hive 0.9.0, shark  
8.1, spark 8.1, es-hadoop 1.3.0.M2, java 1.7.0\_15, scala 2.9.3,  
elasticsearch 1.0.0)

Thanks again,  
Max

On Tuesday, February 18, 2014 10:16:38 PM UTC-8, Costin Leau wrote:

> The error indicates a network error - namely es-hadoop cannot connect to  
> Elasticsearch on the default (localhost:9200)  
> HTTP port. Can you double check whether that's indeed the case (using curl  
> or even telnet on that port) - maybe the  
> firewall prevents any connections to be made...  
> Also you could try using the latest Hive, 0.12 and a more recent Hadoop  
> such as 1.1.2 or 1.2.1.
> 
> Additionally, can you enable TRACE logging in your job on es-hadoop  
> packages org.elasticsearch.hadoop.rest and  
> org.elasticsearch.hadoop.mr packages and report back ?
> 
> Thanks,
> 
> On 19/02/2014 4:03 AM, Max Lang wrote:
> 
> > I set everything up using this guide:  
> > [Running Shark on EC2 · amplab/shark Wiki · GitHub](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2) on an ec2  
> > cluster. I've  
> > copied the elasticsearch-hadoop jars into the hive lib directory and I  
> > have elasticsearch running on localhost:9200. I'm  
> > running shark in a screen session with --service screenserver and  
> > connecting to it at the same time using shark -h  
> > localhost.
> > 
> > Unfortunately, when I attempt to write data into elasticsearch, it  
> > fails. Here's an example:
> > 
> > |  
> > [localhost:10000]shark\>CREATE EXTERNAL TABLE wiki (id BIGINT,title  
> > STRING,last\_modified STRING,xml STRING,text  
> > STRING)ROW FORMAT DELIMITED FIELDS TERMINATED BY '\t'LOCATION  
> > 's3n://spark-data/wikipedia-sample/';  
> > Timetaken (including network latency):0.159seconds  
> > 14/02/1901:23:33INFO CliDriver:Timetaken (including network  
> > latency):0.159seconds
> > 
> > [localhost:10000]shark\>SELECT title FROM wiki LIMIT 1;  
> > Alpokalja  
> > Timetaken (including network latency):2.23seconds  
> > 14/02/1901:23:48INFO CliDriver:Timetaken (including network  
> > latency):2.23seconds
> > 
> > [localhost:10000]shark\>CREATE EXTERNAL TABLE es\_wiki (id BIGINT,title  
> > STRING,last\_modified STRING,xml STRING,text  
> > STRING)STORED BY  
> > 'org.elasticsearch.hadoop.hive.EsStorageHandler'TBLPROPERTIES('es.resource'='wikipedia/article');
> 
> > Timetaken (including network latency):0.061seconds  
> > 14/02/1901:33:51INFO CliDriver:Timetaken (including network  
> > latency):0.061seconds
> > 
> > [localhost:10000]shark\>INSERT OVERWRITE TABLE es\_wiki SELECT w.id,w.title,w.last\_modified,w.xml,w.text  
> > FROM wiki w;  
> > [HiveError]:Queryreturned non-zero  
> > code:9,cause:FAILED:ExecutionError,returncode  
> > -101fromshark.execution.SparkTask  
> > Timetaken (including network latency):3.575seconds  
> > 14/02/1901:34:42INFO CliDriver:Timetaken (including network  
> > latency):3.575seconds  
> > |
> > 
> > _The stack trace looks like this:_
> > 
> > org.apache.hadoop.hive.ql.metadata.HiveException  
> > (org.apache.hadoop.hive.ql.metadata.HiveException: java.io.IOException:  
> > Out of nodes and retries; caught exception)
> 
> org.apache.hadoop.hive.ql.exec.FileSinkOperator.processOp(FileSinkOperator.java:602)shark.execution.FileSinkOperator$$anonfun$processPartition$1.apply(FileSinkOperator.scala:84)shark.execution.FileSinkOperator$$anonfun$processPartition$1.apply(FileSinkOperator.scala:81)scala.collection.Iterator$class.foreach(Iterator.scala:772)scala.collection.Iterator$$anon$19.foreach(Iterator.scala:399)shark.execution.FileSinkOperator.processPartition(FileSinkOperator.scala:81)shark.execution.FileSinkOperator$.writeFiles$1(FileSinkOperator.scala:207)shark.execution.FileSinkOperator$$anonfun$executeProcessFileSinkPartition$1.apply(FileSinkOperator.scala:211)shark.execution.FileSinkOperator$$anonfun$executeProcessFileSinkPartition$1.apply(FileSinkOperator.scala:211)org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:107)org.apache.spark.scheduler.Task.run(Task.scala:53)org.apache.spark.executor.Executor$TaskRunner$$anonfun$run$1.apply$mcV$sp(Executor.scala:215)org.apache.spark.deploy.Sp
> 
> arkHadoopUtil.runAsUser(SparkHadoopUtil.scala:50)org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:182)java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)java.lang.Thread.run(Thread.java:744
> 
> > I should be using Hive 0.9.0, shark 0.8.1, elasticsearch 1.0.0, Hadoop  
> > 1.0.4, and java 1.7.0\_51  
> > Based on my cursory look at the hadoop and elasticsearch-hadoop sources,  
> > it looks like hive is just rethrowing an  
> > IOException it's getting from Spark, and elasticsearch-hadoop is just  
> > hitting those exceptions.  
> > I suppose my questions are: Does this look like an issue with my  
> > ES/elasticsearch-hadoop config? And has anyone gotten  
> > elasticsearch working with Spark/Shark?  
> > Any ideas/insights are appreciated.  
> > Thanks,Max
> > 
> > --  
> > You received this message because you are subscribed to the Google  
> > Groups "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send  
> > an email to  
> > [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > To view this discussion on the web visit
> 
> [https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com).
> 
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).
> 
> --  
> Costin

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [February 19, 2014, 10:35pm UTC](https://discuss.elastic.co/t/hadoop-getting-elasticsearch-hadoop-working-with-shark/15882/4 "2014-02-19T22:35:40Z")

</div>

Hi,

Setting logging in Hive/Hadoop can be tricky since the log4j needs to be picked up by the running JVM otherwise you  
won't see anything.  
Take a look at this link on how to tell Hive to use your logging settings [1].

For the next release, we might introduce dedicated exceptions for the simple fact that some libraries, like Hive,  
swallow the stack trace and it's unclear what the issue is which makes the exception (IllegalStateException) ambiguous.

Let me know how it goes and whether you will encounter any issues with Shark. Or if you don't 🙂

Thanks!

[1] [GettingStarted - Apache Hive - Apache Software Foundation](https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs)

On 20/02/2014 12:02 AM, Max Lang wrote:

> Hey Costin,
> 
> Thanks for the swift reply. I abandoned EC2 to take that out of the equation and managed to get everything working  
> locally using the latest version of everything (though I realized just now I'm still on hive 0.9). I'm guessing you're  
> right about some port connection issue because I definitely had ES running on that machine.
> 
> I changed hive-log4j.properties and added  
> |  
> #custom logging levels  
> #log4j.logger.xxx=DEBUG  
> log4j.logger.org.elasticsearch.hadoop.rest=TRACE  
> log4j.logger.org.elasticsearch.hadoop.mr=TRACE  
> |
> 
> But I didn't see any trace logging. Hopefully I can get it working on EC2 without issue, but, for the future, is this  
> the correct way to set TRACE logging?
> 
> Oh and, for reference, I tried running without ES up and I got the following, exceptions:
> 
> 2014-02-19 13:46:08,803 ERROR shark.SharkDriver (Logging.scala:logError(64)) - FAILED: Hive Internal Error:  
> java.lang.IllegalStateException(Cannot discover Elasticsearch version)  
> java.lang.IllegalStateException: Cannot discover Elasticsearch version  
> at org.elasticsearch.hadoop.hive.EsStorageHandler.init(EsStorageHandler.java:101)  
> at org.elasticsearch.hadoop.hive.EsStorageHandler.configureOutputJobProperties(EsStorageHandler.java:83)  
> at org.apache.hadoop.hive.ql.plan.PlanUtils.configureJobPropertiesForStorageHandler(PlanUtils.java:706)  
> at org.apache.hadoop.hive.ql.plan.PlanUtils.configureOutputJobPropertiesForStorageHandler(PlanUtils.java:675)  
> at org.apache.hadoop.hive.ql.exec.FileSinkOperator.augmentPlan(FileSinkOperator.java:764)  
> at org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.putOpInsertMap(SemanticAnalyzer.java:1518)  
> at org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genFileSinkPlan(SemanticAnalyzer.java:4337)  
> at org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genPostGroupByBodyPlan(SemanticAnalyzer.java:6207)  
> at org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genBodyPlan(SemanticAnalyzer.java:6138)  
> at org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genPlan(SemanticAnalyzer.java:6764)  
> at shark.parse.SharkSemanticAnalyzer.analyzeInternal(SharkSemanticAnalyzer.scala:149)  
> at org.apache.hadoop.hive.ql.parse.BaseSemanticAnalyzer.analyze(BaseSemanticAnalyzer.java:244)  
> at shark.SharkDriver.compile(SharkDriver.scala:215)  
> at org.apache.hadoop.hive.ql.Driver.compile(Driver.java:336)  
> at org.apache.hadoop.hive.ql.Driver.run(Driver.java:895)  
> at shark.SharkCliDriver.processCmd(SharkCliDriver.scala:324)  
> at org.apache.hadoop.hive.cli.CliDriver.processLine(CliDriver.java:406)  
> at shark.SharkCliDriver$.main(SharkCliDriver.scala:232)  
> at shark.SharkCliDriver.main(SharkCliDriver.scala)  
> Caused by: java.io.IOException: Out of nodes and retries; caught exception  
> at org.elasticsearch.hadoop.rest.NetworkClient.execute(NetworkClient.java:81)  
> at org.elasticsearch.hadoop.rest.RestClient.execute(RestClient.java:221)  
> at org.elasticsearch.hadoop.rest.RestClient.execute(RestClient.java:205)  
> at org.elasticsearch.hadoop.rest.RestClient.execute(RestClient.java:209)  
> at org.elasticsearch.hadoop.rest.RestClient.get(RestClient.java:103)  
> at org.elasticsearch.hadoop.rest.RestClient.esVersion(RestClient.java:274)  
> at org.elasticsearch.hadoop.rest.InitializationUtils.discoverEsVersion(InitializationUtils.java:84)  
> at org.elasticsearch.hadoop.hive.EsStorageHandler.init(EsStorageHandler.java:99)  
> ... 18 more  
> Caused by: java.net.ConnectException: Connection refused  
> at java.net.PlainSocketImpl.socketConnect(Native Method)  
> at java.net.AbstractPlainSocketImpl.doConnect(AbstractPlainSocketImpl.java:339)  
> at java.net.AbstractPlainSocketImpl.connectToAddress(AbstractPlainSocketImpl.java:200)  
> at java.net.AbstractPlainSocketImpl.connect(AbstractPlainSocketImpl.java:182)  
> at java.net.SocksSocketImpl.connect(SocksSocketImpl.java:391)  
> at java.net.Socket.connect(Socket.java:579)  
> at java.net.Socket.connect(Socket.java:528)  
> at java.net.Socket.(Socket.java:425)  
> at java.net.Socket.(Socket.java:280)  
> at org.apache.commons.httpclient.protocol.DefaultProtocolSocketFactory.createSocket(DefaultProtocolSocketFactory.java:80)  
> at org.apache.commons.httpclient.protocol.DefaultProtocolSocketFactory.createSocket(DefaultProtocolSocketFactory.java:122)  
> at org.apache.commons.httpclient.HttpConnection.open(HttpConnection.java:707)  
> at org.apache.commons.httpclient.HttpMethodDirector.executeWithRetry(HttpMethodDirector.java:387)  
> at org.apache.commons.httpclient.HttpMethodDirector.executeMethod(HttpMethodDirector.java:171)  
> at org.apache.commons.httpclient.HttpClient.executeMethod(HttpClient.java:397)  
> at org.apache.commons.httpclient.HttpClient.executeMethod(HttpClient.java:323)  
> at org.elasticsearch.hadoop.rest.commonshttp.CommonsHttpTransport.execute(CommonsHttpTransport.java:160)  
> at org.elasticsearch.hadoop.rest.NetworkClient.execute(NetworkClient.java:74)  
> ... 25 more
> 
> Let me know if there's anything in particular you'd like me to try on EC2.
> 
> (For posterity, the versions I used were: hadoop 2.2.0, hive 0.9.0, shark 8.1, spark 8.1, es-hadoop 1.3.0.M2, java  
> 1.7.0\_15, scala 2.9.3, elasticsearch 1.0.0)
> 
> Thanks again,  
> Max
> 
> On Tuesday, February 18, 2014 10:16:38 PM UTC-8, Costin Leau wrote:
> 
> ```
> The error indicates a network error - namely es-hadoop cannot connect to Elasticsearch on the default (localhost:9200)
> HTTP port. Can you double check whether that's indeed the case (using curl or even telnet on that port) - maybe the
> firewall prevents any connections to be made...
> Also you could try using the latest Hive, 0.12 and a more recent Hadoop such as 1.1.2 or 1.2.1.
> 
> Additionally, can you enable TRACE logging in your job on es-hadoop packages org.elasticsearch.hadoop.rest and
> org.elasticsearch.hadoop.mr <http://org.elasticsearch.hadoop.mr> packages and report back ?
> 
> Thanks,
> 
> On 19/02/2014 4:03 AM, Max Lang wrote:
> > I set everything up using this guide:https://github.com/amplab/shark/wiki/Running-Shark-on-EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2> on an ec2 cluster. I've
> > copied the elasticsearch-hadoop jars into the hive lib directory and I have elasticsearch running on localhost:9200. I'm
> > running shark in a screen session with --service screenserver and connecting to it at the same time using shark -h
> > localhost.
> >
> > Unfortunately, when I attempt to write data into elasticsearch, it fails. Here's an example:
> >
> > |
> > [localhost:10000]shark>CREATE EXTERNAL TABLE wiki (id BIGINT,title STRING,last_modified STRING,xml STRING,text
> > STRING)ROW FORMAT DELIMITED FIELDS TERMINATED BY '\t'LOCATION 's3n://spark-data/wikipedia-sample/';
> > Timetaken (including network latency):0.159seconds
> > 14/02/1901:23:33INFO CliDriver:Timetaken (including network latency):0.159seconds
> >
> > [localhost:10000]shark>SELECT title FROM wiki LIMIT 1;
> > Alpokalja
> > Timetaken (including network latency):2.23seconds
> > 14/02/1901:23:48INFO CliDriver:Timetaken (including network latency):2.23seconds
> >
> > [localhost:10000]shark>CREATE EXTERNAL TABLE es_wiki (id BIGINT,title STRING,last_modified STRING,xml STRING,text
> > STRING)STORED BY 'org.elasticsearch.hadoop.hive.EsStorageHandler'TBLPROPERTIES('es.resource'='wikipedia/article');
> > Timetaken (including network latency):0.061seconds
> > 14/02/1901:33:51INFO CliDriver:Timetaken (including network latency):0.061seconds
> >
> > [localhost:10000]shark>INSERT OVERWRITE TABLE es_wiki SELECTw.id <http://w.id>,w.title,w.last_modified,w.xml,w.text FROM wiki w;
> > [HiveError]:Queryreturned non-zero code:9,cause:FAILED:ExecutionError,returncode -101fromshark.execution.SparkTask
> > Timetaken (including network latency):3.575seconds
> > 14/02/1901:34:42INFO CliDriver:Timetaken (including network latency):3.575seconds
> > |
> >
> > *The stack trace looks like this:*
> >
> > org.apache.hadoop.hive.ql.metadata.HiveException (org.apache.hadoop.hive.ql.metadata.HiveException: java.io.IOException:
> > Out of nodes and retries; caught exception)
> >
> > org.apache.hadoop.hive.ql.exec.FileSinkOperator.processOp(FileSinkOperator.java:602)shark.execution.FileSinkOperator$$anonfun$processPartition$1.apply(FileSinkOperator.scala:84)shark.execution.FileSinkOperator$$anonfun$processPartition$1.apply(FileSinkOperator.scala:81)scala.collection.Iterator$class.foreach(Iterator.scala:772)scala.collection.Iterator$$anon$19.foreach(Iterator.scala:399)shark.execution.FileSinkOperator.processPartition(FileSinkOperator.scala:81)shark.execution.FileSinkOperator$.writeFiles$1(FileSinkOperator.scala:207)shark.execution.FileSinkOperator$$anonfun$executeProcessFileSinkPartition$1.apply(FileSinkOperator.scala:211)shark.execution.FileSinkOperator$$anonfun$executeProcessFileSinkPartition$1.apply(FileSinkOperator.scala:211)org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:107)org.apache.spark.scheduler.Task.run(Task.scala:53)org.apache.spark.executor.Executor$TaskRunner$$anonfun$run$1.apply$mcV$sp(Executor.scala:215)org.apache.spark.dep
> 
> ```

loy.Sp

> ```
> arkHadoopUtil.runAsUser(SparkHadoopUtil.scala:50)org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:182)java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)java.lang.Thread.run(Thread.java:744
> 
> > I should be using Hive 0.9.0, shark 0.8.1, elasticsearch 1.0.0, Hadoop 1.0.4, and java 1.7.0_51
> > Based on my cursory look at the hadoop and elasticsearch-hadoop sources, it looks like hive is just rethrowing an
> > IOException it's getting from Spark, and elasticsearch-hadoop is just hitting those exceptions.
> > I suppose my questions are: Does this look like an issue with my ES/elasticsearch-hadoop config? And has anyone gotten
> > elasticsearch working with Spark/Shark?
> > Any ideas/insights are appreciated.
> > Thanks,Max
> >
> > --
> > You received this message because you are subscribed to the Google Groups "elasticsearch" group.
> > To unsubscribe from this group and stop receiving emails from it, send an email to
> >elasticsearc...@googlegroups.com <javascript:>.
> > To view this discussion on the web visit
> >https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>.
> > For more options, visithttps://groups.google.com/groups/opt_out <https://groups.google.com/groups/opt_out>.
> 
> --
> Costin
> 
> ```
> 
> --  
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an email to  
> [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
Costin

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/530531BC.80807%40gmail.com](https://groups.google.com/d/msgid/elasticsearch/530531BC.80807%40gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Max\_Lang](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/max_lang/32/1782_2.png) [@Max\_Lang](https://discuss.elastic.co/u/Max_Lang)\
**Post date:** [February 21, 2014, 11:06pm UTC](https://discuss.elastic.co/t/hadoop-getting-elasticsearch-hadoop-working-with-shark/15882/5 "2014-02-21T23:06:19Z")

</div>

I managed to get it working on ec2 without issue this time. I'd say the  
biggest difference was that this time I set up a dedicated ES machine. Is  
it possible that, because I was using a cluster with slaves, when I used  
"localhost" the slaves couldn't find the ES instance running on the master?  
Or do all the requests go through the master?

On Wednesday, February 19, 2014 2:35:40 PM UTC-8, Costin Leau wrote:

> Hi,
> 
> Setting logging in Hive/Hadoop can be tricky since the log4j needs to be  
> picked up by the running JVM otherwise you  
> won't see anything.  
> Take a look at this link on how to tell Hive to use your logging settings  
> [1].
> 
> For the next release, we might introduce dedicated exceptions for the  
> simple fact that some libraries, like Hive,  
> swallow the stack trace and it's unclear what the issue is which makes the  
> exception (IllegalStateException) ambiguous.
> 
> Let me know how it goes and whether you will encounter any issues with  
> Shark. Or if you don't 🙂
> 
> Thanks!
> 
> [1]  
> [GettingStarted - Apache Hive - Apache Software Foundation](https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs)
> 
> On 20/02/2014 12:02 AM, Max Lang wrote:
> 
> > Hey Costin,
> > 
> > Thanks for the swift reply. I abandoned EC2 to take that out of the  
> > equation and managed to get everything working  
> > locally using the latest version of everything (though I realized just  
> > now I'm still on hive 0.9). I'm guessing you're  
> > right about some port connection issue because I definitely had ES  
> > running on that machine.
> > 
> > I changed hive-log4j.properties and added  
> > |  
> > #custom logging levels  
> > #log4j.logger.xxx=DEBUG  
> > log4j.logger.org.elasticsearch.hadoop.rest=TRACE  
> > log4j.logger.org.elasticsearch.hadoop.mr=TRACE  
> > |
> > 
> > But I didn't see any trace logging. Hopefully I can get it working on  
> > EC2 without issue, but, for the future, is this  
> > the correct way to set TRACE logging?
> > 
> > Oh and, for reference, I tried running without ES up and I got the  
> > following, exceptions:
> > 
> > 2014-02-19 13:46:08,803 ERROR shark.SharkDriver  
> > (Logging.scala:logError(64)) - FAILED: Hive Internal Error:  
> > java.lang.IllegalStateException(Cannot discover Elasticsearch version)  
> > java.lang.IllegalStateException: Cannot discover Elasticsearch version  
> > at  
> > org.elasticsearch.hadoop.hive.EsStorageHandler.init(EsStorageHandler.java:101)
> 
> > at  
> > org.elasticsearch.hadoop.hive.EsStorageHandler.configureOutputJobProperties(EsStorageHandler.java:83)
> 
> > at  
> > org.apache.hadoop.hive.ql.plan.PlanUtils.configureJobPropertiesForStorageHandler(PlanUtils.java:706)
> 
> > at  
> > org.apache.hadoop.hive.ql.plan.PlanUtils.configureOutputJobPropertiesForStorageHandler(PlanUtils.java:675)
> 
> > at  
> > org.apache.hadoop.hive.ql.exec.FileSinkOperator.augmentPlan(FileSinkOperator.java:764)
> 
> > at  
> > org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.putOpInsertMap(SemanticAnalyzer.java:1518)
> 
> > at  
> > org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genFileSinkPlan(SemanticAnalyzer.java:4337)
> 
> > at  
> > org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genPostGroupByBodyPlan(SemanticAnalyzer.java:6207)
> 
> > at  
> > org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genBodyPlan(SemanticAnalyzer.java:6138)
> 
> > at  
> > org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genPlan(SemanticAnalyzer.java:6764)
> 
> > at  
> > shark.parse.SharkSemanticAnalyzer.analyzeInternal(SharkSemanticAnalyzer.scala:149)
> 
> > at  
> > org.apache.hadoop.hive.ql.parse.BaseSemanticAnalyzer.analyze(BaseSemanticAnalyzer.java:244)
> 
> > at shark.SharkDriver.compile(SharkDriver.scala:215)  
> > at org.apache.hadoop.hive.ql.Driver.compile(Driver.java:336)  
> > at org.apache.hadoop.hive.ql.Driver.run(Driver.java:895)  
> > at shark.SharkCliDriver.processCmd(SharkCliDriver.scala:324)  
> > at org.apache.hadoop.hive.cli.CliDriver.processLine(CliDriver.java:406)  
> > at shark.SharkCliDriver$.main(SharkCliDriver.scala:232)  
> > at shark.SharkCliDriver.main(SharkCliDriver.scala)  
> > Caused by: java.io.IOException: Out of nodes and retries; caught  
> > exception  
> > at  
> > org.elasticsearch.hadoop.rest.NetworkClient.execute(NetworkClient.java:81)  
> > at org.elasticsearch.hadoop.rest.RestClient.execute(RestClient.java:221)  
> > at org.elasticsearch.hadoop.rest.RestClient.execute(RestClient.java:205)  
> > at org.elasticsearch.hadoop.rest.RestClient.execute(RestClient.java:209)  
> > at org.elasticsearch.hadoop.rest.RestClient.get(RestClient.java:103)  
> > at  
> > org.elasticsearch.hadoop.rest.RestClient.esVersion(RestClient.java:274)  
> > at  
> > org.elasticsearch.hadoop.rest.InitializationUtils.discoverEsVersion(InitializationUtils.java:84)
> 
> > at  
> > org.elasticsearch.hadoop.hive.EsStorageHandler.init(EsStorageHandler.java:99)
> 
> > ... 18 more  
> > Caused by: java.net.ConnectException: Connection refused  
> > at java.net.PlainSocketImpl.socketConnect(Native Method)  
> > at  
> > java.net.AbstractPlainSocketImpl.doConnect(AbstractPlainSocketImpl.java:339)
> 
> > at  
> > java.net.AbstractPlainSocketImpl.connectToAddress(AbstractPlainSocketImpl.java:200)
> 
> > at  
> > java.net.AbstractPlainSocketImpl.connect(AbstractPlainSocketImpl.java:182)  
> > at java.net.SocksSocketImpl.connect(SocksSocketImpl.java:391)  
> > at java.net.Socket.connect(Socket.java:579)  
> > at java.net.Socket.connect(Socket.java:528)  
> > at java.net.Socket.(Socket.java:425)  
> > at java.net.Socket.(Socket.java:280)  
> > at  
> > org.apache.commons.httpclient.protocol.DefaultProtocolSocketFactory.createSocket(DefaultProtocolSocketFactory.java:80)
> 
> > at  
> > org.apache.commons.httpclient.protocol.DefaultProtocolSocketFactory.createSocket(DefaultProtocolSocketFactory.java:122)
> 
> > at  
> > org.apache.commons.httpclient.HttpConnection.open(HttpConnection.java:707)  
> > at  
> > org.apache.commons.httpclient.HttpMethodDirector.executeWithRetry(HttpMethodDirector.java:387)
> 
> > at  
> > org.apache.commons.httpclient.HttpMethodDirector.executeMethod(HttpMethodDirector.java:171)
> 
> > at  
> > org.apache.commons.httpclient.HttpClient.executeMethod(HttpClient.java:397)  
> > at  
> > org.apache.commons.httpclient.HttpClient.executeMethod(HttpClient.java:323)  
> > at  
> > org.elasticsearch.hadoop.rest.commonshttp.CommonsHttpTransport.execute(CommonsHttpTransport.java:160)
> 
> > at  
> > org.elasticsearch.hadoop.rest.NetworkClient.execute(NetworkClient.java:74)  
> > ... 25 more
> > 
> > Let me know if there's anything in particular you'd like me to try on  
> > EC2.
> > 
> > (For posterity, the versions I used were: hadoop 2.2.0, hive 0.9.0,  
> > shark 8.1, spark 8.1, es-hadoop 1.3.0.M2, java  
> > 1.7.0\_15, scala 2.9.3, elasticsearch 1.0.0)
> > 
> > Thanks again,  
> > Max
> > 
> > On Tuesday, February 18, 2014 10:16:38 PM UTC-8, Costin Leau wrote:
> > 
> > ```
> > The error indicates a network error - namely es-hadoop cannot 
> > 
> > ```
> 
> connect to Elasticsearch on the default (localhost:9200)
> 
> > ```
> > HTTP port. Can you double check whether that's indeed the case 
> > 
> > ```
> 
> (using curl or even telnet on that port) - maybe the
> 
> > ```
> > firewall prevents any connections to be made... 
> > Also you could try using the latest Hive, 0.12 and a more recent 
> > 
> > ```
> 
> Hadoop such as 1.1.2 or 1.2.1.
> 
> > ```
> > Additionally, can you enable TRACE logging in your job on es-hadoop 
> > 
> > ```
> 
> packages org.elasticsearch.hadoop.rest and
> 
> > ```
> > org.elasticsearch.hadoop.mr <http://org.elasticsearch.hadoop.mr> 
> > 
> > ```
> 
> packages and report back ?
> 
> > ```
> > Thanks, 
> > 
> > On 19/02/2014 4:03 AM, Max Lang wrote: 
> > > I set everything up using this guide:
> > 
> > ```
> 
> [Running Shark on EC2 · amplab/shark Wiki · GitHub](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)
> 
> > ```
> > <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2> on an 
> > 
> > ```
> 
> ec2 cluster. I've
> 
> > ```
> > > copied the elasticsearch-hadoop jars into the hive lib directory 
> > 
> > ```
> 
> and I have elasticsearch running on localhost:9200. I'm
> 
> > ```
> > > running shark in a screen session with --service screenserver and 
> > 
> > ```
> 
> connecting to it at the same time using shark -h
> 
> > ```
> > > localhost. 
> > > 
> > > Unfortunately, when I attempt to write data into elasticsearch, it 
> > 
> > ```
> 
> fails. Here's an example:
> 
> > ```
> > > 
> > > | 
> > > [localhost:10000]shark>CREATE EXTERNAL TABLE wiki (id BIGINT,title 
> > 
> > ```
> 
> STRING,last\_modified STRING,xml STRING,text
> 
> > ```
> > > STRING)ROW FORMAT DELIMITED FIELDS TERMINATED BY '\t'LOCATION 
> > 
> > ```
> 
> 's3n://spark-data/wikipedia-sample/';
> 
> > ```
> > > Timetaken (including network latency):0.159seconds 
> > > 14/02/1901:23:33INFO CliDriver:Timetaken (including network 
> > 
> > ```
> 
> latency):0.159seconds
> 
> > ```
> > > 
> > > [localhost:10000]shark>SELECT title FROM wiki LIMIT 1; 
> > > Alpokalja 
> > > Timetaken (including network latency):2.23seconds 
> > > 14/02/1901:23:48INFO CliDriver:Timetaken (including network 
> > 
> > ```
> 
> latency):2.23seconds
> 
> > ```
> > > 
> > > [localhost:10000]shark>CREATE EXTERNAL TABLE es_wiki (id 
> > 
> > ```
> 
> BIGINT,title STRING,last\_modified STRING,xml STRING,text
> 
> > ```
> > > STRING)STORED BY 
> > 
> > ```
> 
> 'org.elasticsearch.hadoop.hive.EsStorageHandler'TBLPROPERTIES('es.resource'='wikipedia/article');
> 
> > ```
> > > Timetaken (including network latency):0.061seconds 
> > > 14/02/1901:33:51INFO CliDriver:Timetaken (including network 
> > 
> > ```
> 
> latency):0.061seconds
> 
> > ```
> > > 
> > > [localhost:10000]shark>INSERT OVERWRITE TABLE es_wiki SELECTw.id <
> > 
> > ```
> 
> [http://w.id](http://w.id)\>,w.title,w.last\_modified,w.xml,w.text FROM wiki w;
> 
> > ```
> > > [HiveError]:Queryreturned non-zero 
> > 
> > ```
> 
> code:9,cause:FAILED:ExecutionError,returncode  
> -101fromshark.execution.SparkTask
> 
> > ```
> > > Timetaken (including network latency):3.575seconds 
> > > 14/02/1901:34:42INFO CliDriver:Timetaken (including network 
> > 
> > ```
> 
> latency):3.575seconds
> 
> > ```
> > > | 
> > > 
> > > *The stack trace looks like this:* 
> > > 
> > > org.apache.hadoop.hive.ql.metadata.HiveException 
> > 
> > ```
> 
> (org.apache.hadoop.hive.ql.metadata.HiveException: java.io.IOException:
> 
> > ```
> > > Out of nodes and retries; caught exception) 
> > > 
> > > 
> > 
> > ```
> 
> org.apache.hadoop.hive.ql.exec.FileSinkOperator.processOp(FileSinkOperator.java:602)shark.execution.FileSinkOperator$$anonfun$processPartition$1.apply(FileSinkOperator.scala:84)shark.execution.FileSinkOperator$$anonfun$processPartition$1.apply(FileSinkOperator.scala:81)scala.collection.Iterator$class.foreach(Iterator.scala:772)scala.collection.Iterator$$anon$19.foreach(Iterator.scala:399)shark.execution.FileSinkOperator.processPartition(FileSinkOperator.scala:81)shark.execution.FileSinkOperator$.writeFiles$1(FileSinkOperator.scala:207)shark.execution.FileSinkOperator$$anonfun$executeProcessFileSinkPartition$1.apply(FileSinkOperator.scala:211)shark.execution.FileSinkOperator$$anonfun$executeProcessFileSinkPartition$1.apply(FileSinkOperator.scala:211)org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:107)org.apache.spark.scheduler.Task.run(Task.scala:53)org.apache.spark.executor.Executor$TaskRunner$$anonfun$run$1.apply$mcV$sp(Executor.scala:215)org.apache.spark.dep
> 
> loy.Sp
> 
> > 
> 
> arkHadoopUtil.runAsUser(SparkHadoopUtil.scala:50)org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:182)java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)java.lang.Thread.run(Thread.java:744
> 
> > ```
> > > I should be using Hive 0.9.0, shark 0.8.1, elasticsearch 1.0.0, 
> > 
> > ```
> 
> Hadoop 1.0.4, and java 1.7.0\_51
> 
> > ```
> > > Based on my cursory look at the hadoop and elasticsearch-hadoop 
> > 
> > ```
> 
> sources, it looks like hive is just rethrowing an
> 
> > ```
> > > IOException it's getting from Spark, and elasticsearch-hadoop is 
> > 
> > ```
> 
> just hitting those exceptions.
> 
> > ```
> > > I suppose my questions are: Does this look like an issue with my 
> > 
> > ```
> 
> ES/elasticsearch-hadoop config? And has anyone gotten
> 
> > ```
> > > elasticsearch working with Spark/Shark? 
> > > Any ideas/insights are appreciated. 
> > > Thanks,Max 
> > > 
> > > -- 
> > > You received this message because you are subscribed to the Google 
> > 
> > ```
> 
> Groups "elasticsearch" group.
> 
> > ```
> > > To unsubscribe from this group and stop receiving emails from it, 
> > 
> > ```
> 
> send an email to
> 
> > ```
> > >elasticsearc...@googlegroups.com <javascript:>. 
> > > To view this discussion on the web visit 
> > >
> > 
> > ```
> 
> [https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com)
> 
> > ```
> > <
> > 
> > ```
> 
> [https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com)\>.
> 
> > ```
> > > For more options, visithttps://groups.google.com/groups/opt_out <
> > 
> > ```
> 
> [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)\>.
> 
> > ```
> > -- 
> > Costin 
> > 
> > ```
> > 
> > --  
> > You received this message because you are subscribed to the Google  
> > Groups "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send  
> > an email to  
> > [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > To view this discussion on the web visit
> 
> [https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com).
> 
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).
> 
> --  
> Costin

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/e29e342d-de74-4ed6-93d4-875fc728c5a5%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/e29e342d-de74-4ed6-93d4-875fc728c5a5%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [February 22, 2014, 10:31am UTC](https://discuss.elastic.co/t/hadoop-getting-elasticsearch-hadoop-working-with-shark/15882/6 "2014-02-22T10:31:21Z")

</div>

Yeah, it might have been some sort of network configuration issue where services where running on different machines and  
localhost pointed to a different location.

Either way, I'm glad to hear things have are moving forward.

Cheers,

On 22/02/2014 1:06 AM, Max Lang wrote:

> I managed to get it working on ec2 without issue this time. I'd say the biggest difference was that this time I set up a  
> dedicated ES machine. Is it possible that, because I was using a cluster with slaves, when I used "localhost" the slaves  
> couldn't find the ES instance running on the master? Or do all the requests go through the master?
> 
> On Wednesday, February 19, 2014 2:35:40 PM UTC-8, Costin Leau wrote:
> 
> ```
> Hi,
> 
> Setting logging in Hive/Hadoop can be tricky since the log4j needs to be picked up by the running JVM otherwise you
> won't see anything.
> Take a look at this link on how to tell Hive to use your logging settings [1].
> 
> For the next release, we might introduce dedicated exceptions for the simple fact that some libraries, like Hive,
> swallow the stack trace and it's unclear what the issue is which makes the exception (IllegalStateException) ambiguous.
> 
> Let me know how it goes and whether you will encounter any issues with Shark. Or if you don't :)
> 
> Thanks!
> 
> [1] https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs
> <https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs>
> 
> On 20/02/2014 12:02 AM, Max Lang wrote:
> > Hey Costin,
> >
> > Thanks for the swift reply. I abandoned EC2 to take that out of the equation and managed to get everything working
> > locally using the latest version of everything (though I realized just now I'm still on hive 0.9). I'm guessing you're
> > right about some port connection issue because I definitely had ES running on that machine.
> >
> > I changed hive-log4j.properties and added
> > |
> > #custom logging levels
> > #log4j.logger.xxx=DEBUG
> > log4j.logger.org.elasticsearch.hadoop.rest=TRACE
> >log4j.logger.org.elasticsearch.hadoop.mr <http://log4j.logger.org.elasticsearch.hadoop.mr>=TRACE
> > |
> >
> > But I didn't see any trace logging. Hopefully I can get it working on EC2 without issue, but, for the future, is this
> > the correct way to set TRACE logging?
> >
> > Oh and, for reference, I tried running without ES up and I got the following, exceptions:
> >
> > 2014-02-19 13:46:08,803 ERROR shark.SharkDriver (Logging.scala:logError(64)) - FAILED: Hive Internal Error:
> > java.lang.IllegalStateException(Cannot discover Elasticsearch version)
> > java.lang.IllegalStateException: Cannot discover Elasticsearch version
> > at org.elasticsearch.hadoop.hive.EsStorageHandler.init(EsStorageHandler.java:101)
> > at org.elasticsearch.hadoop.hive.EsStorageHandler.configureOutputJobProperties(EsStorageHandler.java:83)
> > at org.apache.hadoop.hive.ql.plan.PlanUtils.configureJobPropertiesForStorageHandler(PlanUtils.java:706)
> > at org.apache.hadoop.hive.ql.plan.PlanUtils.configureOutputJobPropertiesForStorageHandler(PlanUtils.java:675)
> > at org.apache.hadoop.hive.ql.exec.FileSinkOperator.augmentPlan(FileSinkOperator.java:764)
> > at org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.putOpInsertMap(SemanticAnalyzer.java:1518)
> > at org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genFileSinkPlan(SemanticAnalyzer.java:4337)
> > at org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genPostGroupByBodyPlan(SemanticAnalyzer.java:6207)
> > at org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genBodyPlan(SemanticAnalyzer.java:6138)
> > at org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genPlan(SemanticAnalyzer.java:6764)
> > at shark.parse.SharkSemanticAnalyzer.analyzeInternal(SharkSemanticAnalyzer.scala:149)
> > at org.apache.hadoop.hive.ql.parse.BaseSemanticAnalyzer.analyze(BaseSemanticAnalyzer.java:244)
> > at shark.SharkDriver.compile(SharkDriver.scala:215)
> > at org.apache.hadoop.hive.ql.Driver.compile(Driver.java:336)
> > at org.apache.hadoop.hive.ql.Driver.run(Driver.java:895)
> > at shark.SharkCliDriver.processCmd(SharkCliDriver.scala:324)
> > at org.apache.hadoop.hive.cli.CliDriver.processLine(CliDriver.java:406)
> > at shark.SharkCliDriver$.main(SharkCliDriver.scala:232)
> > at shark.SharkCliDriver.main(SharkCliDriver.scala)
> > Caused by: java.io.IOException: Out of nodes and retries; caught exception
> > at org.elasticsearch.hadoop.rest.NetworkClient.execute(NetworkClient.java:81)
> > at org.elasticsearch.hadoop.rest.RestClient.execute(RestClient.java:221)
> > at org.elasticsearch.hadoop.rest.RestClient.execute(RestClient.java:205)
> > at org.elasticsearch.hadoop.rest.RestClient.execute(RestClient.java:209)
> > at org.elasticsearch.hadoop.rest.RestClient.get(RestClient.java:103)
> > at org.elasticsearch.hadoop.rest.RestClient.esVersion(RestClient.java:274)
> > at org.elasticsearch.hadoop.rest.InitializationUtils.discoverEsVersion(InitializationUtils.java:84)
> > at org.elasticsearch.hadoop.hive.EsStorageHandler.init(EsStorageHandler.java:99)
> > ... 18 more
> > Caused by: java.net.ConnectException: Connection refused
> > at java.net.PlainSocketImpl.socketConnect(Native Method)
> > at java.net.AbstractPlainSocketImpl.doConnect(AbstractPlainSocketImpl.java:339)
> > at java.net.AbstractPlainSocketImpl.connectToAddress(AbstractPlainSocketImpl.java:200)
> > at java.net.AbstractPlainSocketImpl.connect(AbstractPlainSocketImpl.java:182)
> > at java.net.SocksSocketImpl.connect(SocksSocketImpl.java:391)
> > at java.net.Socket.connect(Socket.java:579)
> > at java.net.Socket.connect(Socket.java:528)
> > at java.net.Socket.<init>(Socket.java:425)
> > at java.net.Socket.<init>(Socket.java:280)
> > at org.apache.commons.httpclient.protocol.DefaultProtocolSocketFactory.createSocket(DefaultProtocolSocketFactory.java:80)
> > at org.apache.commons.httpclient.protocol.DefaultProtocolSocketFactory.createSocket(DefaultProtocolSocketFactory.java:122)
> > at org.apache.commons.httpclient.HttpConnection.open(HttpConnection.java:707)
> > at org.apache.commons.httpclient.HttpMethodDirector.executeWithRetry(HttpMethodDirector.java:387)
> > at org.apache.commons.httpclient.HttpMethodDirector.executeMethod(HttpMethodDirector.java:171)
> > at org.apache.commons.httpclient.HttpClient.executeMethod(HttpClient.java:397)
> > at org.apache.commons.httpclient.HttpClient.executeMethod(HttpClient.java:323)
> > at org.elasticsearch.hadoop.rest.commonshttp.CommonsHttpTransport.execute(CommonsHttpTransport.java:160)
> > at org.elasticsearch.hadoop.rest.NetworkClient.execute(NetworkClient.java:74)
> > ... 25 more
> >
> > Let me know if there's anything in particular you'd like me to try on EC2.
> >
> > (For posterity, the versions I used were: hadoop 2.2.0, hive 0.9.0, shark 8.1, spark 8.1, es-hadoop 1.3.0.M2, java
> > 1.7.0_15, scala 2.9.3, elasticsearch 1.0.0)
> >
> > Thanks again,
> > Max
> >
> > On Tuesday, February 18, 2014 10:16:38 PM UTC-8, Costin Leau wrote:
> >
> > The error indicates a network error - namely es-hadoop cannot connect to Elasticsearch on the default (localhost:9200)
> > HTTP port. Can you double check whether that's indeed the case (using curl or even telnet on that port) - maybe the
> > firewall prevents any connections to be made...
> > Also you could try using the latest Hive, 0.12 and a more recent Hadoop such as 1.1.2 or 1.2.1.
> >
> > Additionally, can you enable TRACE logging in your job on es-hadoop packages org.elasticsearch.hadoop.rest and
> >org.elasticsearch.hadoop.mr <http://org.elasticsearch.hadoop.mr> <http://org.elasticsearch.hadoop.mr
> <http://org.elasticsearch.hadoop.mr>> packages and report back ?
> >
> > Thanks,
> >
> > On 19/02/2014 4:03 AM, Max Lang wrote:
> > > I set everything up using this guide:https://github.com/amplab/shark/wiki/Running-Shark-on-EC2 <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>
> > <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>> on an ec2 cluster. I've
> > > copied the elasticsearch-hadoop jars into the hive lib directory and I have elasticsearch running on localhost:9200. I'm
> > > running shark in a screen session with --service screenserver and connecting to it at the same time using shark -h
> > > localhost.
> > >
> > > Unfortunately, when I attempt to write data into elasticsearch, it fails. Here's an example:
> > >
> > > |
> > > [localhost:10000]shark>CREATE EXTERNAL TABLE wiki (id BIGINT,title STRING,last_modified STRING,xml STRING,text
> > > STRING)ROW FORMAT DELIMITED FIELDS TERMINATED BY '\t'LOCATION 's3n://spark-data/wikipedia-sample/';
> > > Timetaken (including network latency):0.159seconds
> > > 14/02/1901:23:33INFO CliDriver:Timetaken (including network latency):0.159seconds
> > >
> > > [localhost:10000]shark>SELECT title FROM wiki LIMIT 1;
> > > Alpokalja
> > > Timetaken (including network latency):2.23seconds
> > > 14/02/1901:23:48INFO CliDriver:Timetaken (including network latency):2.23seconds
> > >
> > > [localhost:10000]shark>CREATE EXTERNAL TABLE es_wiki (id BIGINT,title STRING,last_modified STRING,xml STRING,text
> > > STRING)STORED BY 'org.elasticsearch.hadoop.hive.EsStorageHandler'TBLPROPERTIES('es.resource'='wikipedia/article');
> > > Timetaken (including network latency):0.061seconds
> > > 14/02/1901:33:51INFO CliDriver:Timetaken (including network latency):0.061seconds
> > >
> > > [localhost:10000]shark>INSERT OVERWRITE TABLE es_wiki SELECTw.id <http://w.id>,w.title,w.last_modified,w.xml,w.text FROM wiki w;
> > > [HiveError]:Queryreturned non-zero code:9,cause:FAILED:ExecutionError,returncode -101fromshark.execution.SparkTask
> > > Timetaken (including network latency):3.575seconds
> > > 14/02/1901:34:42INFO CliDriver:Timetaken (including network latency):3.575seconds
> > > |
> > >
> > > *The stack trace looks like this:*
> > >
> > > org.apache.hadoop.hive.ql.metadata.HiveException (org.apache.hadoop.hive.ql.metadata.HiveException: java.io.IOException:
> > > Out of nodes and retries; caught exception)
> > >
> > > org.apache.hadoop.hive.ql.exec.FileSinkOperator.processOp(FileSinkOperator.java:602)shark.execution.FileSinkOperator$$anonfun$processPartition$1.apply(FileSinkOperator.scala:84)shark.execution.FileSinkOperator$$anonfun$processPartition$1.apply(FileSinkOperator.scala:81)scala.collection.Iterator$class.foreach(Iterator.scala:772)scala.collection.Iterator$$anon$19.foreach(Iterator.scala:399)shark.execution.FileSinkOperator.processPartition(FileSinkOperator.scala:81)shark.execution.FileSinkOperator$.writeFiles$1(FileSinkOperator.scala:207)shark.execution.FileSinkOperator$$anonfun$executeProcessFileSinkPartition$1.apply(FileSinkOperator.scala:211)shark.execution.FileSinkOperator$$anonfun$executeProcessFileSinkPartition$1.apply(FileSinkOperator.scala:211)org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:107)org.apache.spark.scheduler.Task.run(Task.scala:53)org.apache.spark.executor.Executor$TaskRunner$$anonfun$run$1.apply$mcV$sp(Executor.scala:215)org.apache.spa
> 
> ```

rk.dep

> ```
> loy.Sp
> >
> > arkHadoopUtil.runAsUser(SparkHadoopUtil.scala:50)org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:182)java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)java.lang.Thread.run(Thread.java:744
> 
> >
> > > I should be using Hive 0.9.0, shark 0.8.1, elasticsearch 1.0.0, Hadoop 1.0.4, and java 1.7.0_51
> > > Based on my cursory look at the hadoop and elasticsearch-hadoop sources, it looks like hive is just rethrowing an
> > > IOException it's getting from Spark, and elasticsearch-hadoop is just hitting those exceptions.
> > > I suppose my questions are: Does this look like an issue with my ES/elasticsearch-hadoop config? And has anyone gotten
> > > elasticsearch working with Spark/Shark?
> > > Any ideas/insights are appreciated.
> > > Thanks,Max
> > >
> > > --
> > > You received this message because you are subscribed to the Google Groups "elasticsearch" group.
> > > To unsubscribe from this group and stop receiving emails from it, send an email to
> > >elasticsearc...@googlegroups.com <javascript:>.
> > > To view this discussion on the web visit
> > >https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>
> > <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>.
> > > For more options, visithttps://groups.google.com/groups/opt_out <http://groups.google.com/groups/opt_out> <https://groups.google.com/groups/opt_out
> <https://groups.google.com/groups/opt_out>>.
> >
> > --
> > Costin
> >
> > --
> > You received this message because you are subscribed to the Google Groups "elasticsearch" group.
> > To unsubscribe from this group and stop receiving emails from it, send an email to
> >elasticsearc...@googlegroups.com <javascript:>.
> > To view this discussion on the web visit
> >https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com>.
> > For more options, visithttps://groups.google.com/groups/opt_out <https://groups.google.com/groups/opt_out>.
> 
> --
> Costin
> 
> ```
> 
> --  
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an email to  
> [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/e29e342d-de74-4ed6-93d4-875fc728c5a5%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/e29e342d-de74-4ed6-93d4-875fc728c5a5%40googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

--  
Costin

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/53087C79.4030109%40gmail.com](https://groups.google.com/d/msgid/elasticsearch/53087C79.4030109%40gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Nick\_Pentreath](https://avatars.discourse-cdn.com/v4/letter/n/7ea924/32.png) [@Nick\_Pentreath](https://discuss.elastic.co/u/Nick_Pentreath)\
**Post date:** [March 20, 2014, 10:00am UTC](https://discuss.elastic.co/t/hadoop-getting-elasticsearch-hadoop-working-with-shark/15882/7 "2014-03-20T10:00:54Z")

</div>

Hi

I am struggling to get this working too. I'm just trying locally for now,  
running Shark 0.8.1, Hive 0.9.0 and ES 1.0.1 with ES-hadoop 1.3.0.M2.

I managed to get a basic example working with WRITING into an index. But  
I'm really after READING and index.

I believe I have set everything up correctly, I've added the jar to Shark:  
ADD JAR /path/to/es-hadoop.jar;

created a table:  
CREATE EXTERNAL TABLE test\_read (name string, price double)

STORED BY 'org.elasticsearch.hadoop.hive.EsStorageHandler'

TBLPROPERTIES('es.resource' = 'test\_index/test\_type/\_search?q=\*');

And then trying to 'SELECT \* FROM test \_read' gives me :

org.apache.spark.SparkException: Job aborted: Task 3.0:0 failed more than 0  
times; aborting job java.lang.ClassCastException:  
org.elasticsearch.hadoop.hive.EsHiveInputFormat$ESHiveSplit cannot be cast  
to org.elasticsearch.hadoop.hive.EsHiveInputFormat$ESHiveSplit

at  
org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:827)

at  
org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:825)

at  
scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:60)

at scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:47)

at  
org.apache.spark.scheduler.DAGScheduler.abortStage(DAGScheduler.scala:825)

at  
org.apache.spark.scheduler.DAGScheduler.processEvent(DAGScheduler.scala:440)

at  
org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$run(DAGScheduler.scala:502)

at  
org.apache.spark.scheduler.DAGScheduler$$anon$1.run(DAGScheduler.scala:157)

FAILED: Execution Error, return code -101 from shark.execution.SparkTask

In fact I get the same error thrown when trying to READ from the table that  
I successfully WROTE to...  
On Saturday, 22 February 2014 12:31:21 UTC+2, Costin Leau wrote:

> Yeah, it might have been some sort of network configuration issue where  
> services where running on different machines and  
> localhost pointed to a different location.
> 
> Either way, I'm glad to hear things have are moving forward.
> 
> Cheers,
> 
> On 22/02/2014 1:06 AM, Max Lang wrote:
> 
> > I managed to get it working on ec2 without issue this time. I'd say the  
> > biggest difference was that this time I set up a  
> > dedicated ES machine. Is it possible that, because I was using a cluster  
> > with slaves, when I used "localhost" the slaves  
> > couldn't find the ES instance running on the master? Or do all the  
> > requests go through the master?
> > 
> > On Wednesday, February 19, 2014 2:35:40 PM UTC-8, Costin Leau wrote:
> > 
> > ```
> > Hi, 
> > 
> > Setting logging in Hive/Hadoop can be tricky since the log4j needs 
> > 
> > ```
> 
> to be picked up by the running JVM otherwise you
> 
> > ```
> > won't see anything. 
> > Take a look at this link on how to tell Hive to use your logging 
> > 
> > ```
> 
> settings [1].
> 
> > ```
> > For the next release, we might introduce dedicated exceptions for 
> > 
> > ```
> 
> the simple fact that some libraries, like Hive,
> 
> > ```
> > swallow the stack trace and it's unclear what the issue is which 
> > 
> > ```
> 
> makes the exception (IllegalStateException) ambiguous.
> 
> > ```
> > Let me know how it goes and whether you will encounter any issues 
> > 
> > ```
> 
> with Shark. Or if you don't 🙂
> 
> > ```
> > Thanks! 
> > 
> > [1] 
> > 
> > ```
> 
> [GettingStarted - Apache Hive - Apache Software Foundation](https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs)
> 
> > ```
> > <
> > 
> > ```
> 
> [GettingStarted - Apache Hive - Apache Software Foundation](https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs)\>
> 
> > ```
> > On 20/02/2014 12:02 AM, Max Lang wrote: 
> > > Hey Costin, 
> > > 
> > > Thanks for the swift reply. I abandoned EC2 to take that out of 
> > 
> > ```
> 
> the equation and managed to get everything working
> 
> > ```
> > > locally using the latest version of everything (though I realized 
> > 
> > ```
> 
> just now I'm still on hive 0.9). I'm guessing you're
> 
> > ```
> > > right about some port connection issue because I definitely had ES 
> > 
> > ```
> 
> running on that machine.
> 
> > ```
> > > 
> > > I changed hive-log4j.properties and added 
> > > | 
> > > #custom logging levels 
> > > #log4j.logger.xxx=DEBUG 
> > > log4j.logger.org.elasticsearch.hadoop.rest=TRACE 
> > >log4j.logger.org.elasticsearch.hadoop.mr <
> > 
> > ```
> 
> [http://log4j.logger.org.elasticsearch.hadoop.mr](http://log4j.logger.org.elasticsearch.hadoop.mr)\>=TRACE
> 
> > ```
> > > | 
> > > 
> > > But I didn't see any trace logging. Hopefully I can get it working 
> > 
> > ```
> 
> on EC2 without issue, but, for the future, is this
> 
> > ```
> > > the correct way to set TRACE logging? 
> > > 
> > > Oh and, for reference, I tried running without ES up and I got the 
> > 
> > ```
> 
> following, exceptions:
> 
> > ```
> > > 
> > > 2014-02-19 13:46:08,803 ERROR shark.SharkDriver 
> > 
> > ```
> 
> (Logging.scala:logError(64)) - FAILED: Hive Internal Error:
> 
> > ```
> > > java.lang.IllegalStateException(Cannot discover Elasticsearch 
> > 
> > ```
> 
> version)
> 
> > ```
> > > java.lang.IllegalStateException: Cannot discover Elasticsearch 
> > 
> > ```
> 
> version
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.elasticsearch.hadoop.hive.EsStorageHandler.init(EsStorageHandler.java:101)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.elasticsearch.hadoop.hive.EsStorageHandler.configureOutputJobProperties(EsStorageHandler.java:83)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.apache.hadoop.hive.ql.plan.PlanUtils.configureJobPropertiesForStorageHandler(PlanUtils.java:706)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.apache.hadoop.hive.ql.plan.PlanUtils.configureOutputJobPropertiesForStorageHandler(PlanUtils.java:675)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.apache.hadoop.hive.ql.exec.FileSinkOperator.augmentPlan(FileSinkOperator.java:764)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.putOpInsertMap(SemanticAnalyzer.java:1518)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genFileSinkPlan(SemanticAnalyzer.java:4337)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genPostGroupByBodyPlan(SemanticAnalyzer.java:6207)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genBodyPlan(SemanticAnalyzer.java:6138)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genPlan(SemanticAnalyzer.java:6764)
> 
> > ```
> > > at 
> > 
> > ```
> 
> shark.parse.SharkSemanticAnalyzer.analyzeInternal(SharkSemanticAnalyzer.scala:149)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.apache.hadoop.hive.ql.parse.BaseSemanticAnalyzer.analyze(BaseSemanticAnalyzer.java:244)
> 
> > ```
> > > at shark.SharkDriver.compile(SharkDriver.scala:215) 
> > > at org.apache.hadoop.hive.ql.Driver.compile(Driver.java:336) 
> > > at org.apache.hadoop.hive.ql.Driver.run(Driver.java:895) 
> > > at shark.SharkCliDriver.processCmd(SharkCliDriver.scala:324) 
> > > at 
> > 
> > ```
> 
> org.apache.hadoop.hive.cli.CliDriver.processLine(CliDriver.java:406)
> 
> > ```
> > > at shark.SharkCliDriver$.main(SharkCliDriver.scala:232) 
> > > at shark.SharkCliDriver.main(SharkCliDriver.scala) 
> > > Caused by: java.io.IOException: Out of nodes and retries; caught 
> > 
> > ```
> 
> exception
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.elasticsearch.hadoop.rest.NetworkClient.execute(NetworkClient.java:81)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.elasticsearch.hadoop.rest.RestClient.execute(RestClient.java:221)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.elasticsearch.hadoop.rest.RestClient.execute(RestClient.java:205)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.elasticsearch.hadoop.rest.RestClient.execute(RestClient.java:209)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.elasticsearch.hadoop.rest.RestClient.get(RestClient.java:103)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.elasticsearch.hadoop.rest.RestClient.esVersion(RestClient.java:274)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.elasticsearch.hadoop.rest.InitializationUtils.discoverEsVersion(InitializationUtils.java:84)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.elasticsearch.hadoop.hive.EsStorageHandler.init(EsStorageHandler.java:99)
> 
> > ```
> > > ... 18 more 
> > > Caused by: java.net.ConnectException: Connection refused 
> > > at java.net.PlainSocketImpl.socketConnect(Native Method) 
> > > at 
> > 
> > ```
> 
> java.net.AbstractPlainSocketImpl.doConnect(AbstractPlainSocketImpl.java:339)
> 
> > ```
> > > at 
> > 
> > ```
> 
> java.net.AbstractPlainSocketImpl.connectToAddress(AbstractPlainSocketImpl.java:200)
> 
> > ```
> > > at 
> > 
> > ```
> 
> java.net.AbstractPlainSocketImpl.connect(AbstractPlainSocketImpl.java:182)
> 
> > ```
> > > at java.net.SocksSocketImpl.connect(SocksSocketImpl.java:391) 
> > > at java.net.Socket.connect(Socket.java:579) 
> > > at java.net.Socket.connect(Socket.java:528) 
> > > at java.net.Socket.<init>(Socket.java:425) 
> > > at java.net.Socket.<init>(Socket.java:280) 
> > > at 
> > 
> > ```
> 
> org.apache.commons.httpclient.protocol.DefaultProtocolSocketFactory.createSocket(DefaultProtocolSocketFactory.java:80)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.apache.commons.httpclient.protocol.DefaultProtocolSocketFactory.createSocket(DefaultProtocolSocketFactory.java:122)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.apache.commons.httpclient.HttpConnection.open(HttpConnection.java:707)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.apache.commons.httpclient.HttpMethodDirector.executeWithRetry(HttpMethodDirector.java:387)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.apache.commons.httpclient.HttpMethodDirector.executeMethod(HttpMethodDirector.java:171)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.apache.commons.httpclient.HttpClient.executeMethod(HttpClient.java:397)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.apache.commons.httpclient.HttpClient.executeMethod(HttpClient.java:323)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.elasticsearch.hadoop.rest.commonshttp.CommonsHttpTransport.execute(CommonsHttpTransport.java:160)
> 
> > ```
> > > at 
> > 
> > ```
> 
> org.elasticsearch.hadoop.rest.NetworkClient.execute(NetworkClient.java:74)
> 
> > ```
> > > ... 25 more 
> > > 
> > > Let me know if there's anything in particular you'd like me to try 
> > 
> > ```
> 
> on EC2.
> 
> > ```
> > > 
> > > (For posterity, the versions I used were: hadoop 2.2.0, hive 
> > 
> > ```
> 
> 0.9.0, shark 8.1, spark 8.1, es-hadoop 1.3.0.M2, java
> 
> > ```
> > > 1.7.0_15, scala 2.9.3, elasticsearch 1.0.0) 
> > > 
> > > Thanks again, 
> > > Max 
> > > 
> > > On Tuesday, February 18, 2014 10:16:38 PM UTC-8, Costin Leau 
> > 
> > ```
> 
> wrote:
> 
> > ```
> > > 
> > > The error indicates a network error - namely es-hadoop cannot 
> > 
> > ```
> 
> connect to Elasticsearch on the default (localhost:9200)
> 
> > ```
> > > HTTP port. Can you double check whether that's indeed the case 
> > 
> > ```
> 
> (using curl or even telnet on that port) - maybe the
> 
> > ```
> > > firewall prevents any connections to be made... 
> > > Also you could try using the latest Hive, 0.12 and a more 
> > 
> > ```
> 
> recent Hadoop such as 1.1.2 or 1.2.1.
> 
> > ```
> > > 
> > > Additionally, can you enable TRACE logging in your job on 
> > 
> > ```
> 
> es-hadoop packages org.elasticsearch.hadoop.rest and
> 
> > ```
> > >org.elasticsearch.hadoop.mr <http://org.elasticsearch.hadoop.mr> <
> > 
> > ```
> 
> [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)
> 
> > ```
> > <http://org.elasticsearch.hadoop.mr>> packages and report back ? 
> > > 
> > > Thanks, 
> > > 
> > > On 19/02/2014 4:03 AM, Max Lang wrote: 
> > > > I set everything up using this guide:
> > 
> > ```
> 
> [Running Shark on EC2 · amplab/shark Wiki · GitHub](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2) \<  
> [Running Shark on EC2 · amplab/shark Wiki · GitHub](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>
> 
> > ```
> > > <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2 
> > <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>> on an 
> > 
> > ```
> 
> ec2 cluster. I've
> 
> > ```
> > > > copied the elasticsearch-hadoop jars into the hive lib 
> > 
> > ```
> 
> directory and I have elasticsearch running on localhost:9200. I'm
> 
> > ```
> > > > running shark in a screen session with --service 
> > 
> > ```
> 
> screenserver and connecting to it at the same time using shark -h
> 
> > ```
> > > > localhost. 
> > > > 
> > > > Unfortunately, when I attempt to write data into 
> > 
> > ```
> 
> elasticsearch, it fails. Here's an example:
> 
> > ```
> > > > 
> > > > | 
> > > > [localhost:10000]shark>CREATE EXTERNAL TABLE wiki (id 
> > 
> > ```
> 
> BIGINT,title STRING,last\_modified STRING,xml STRING,text
> 
> > ```
> > > > STRING)ROW FORMAT DELIMITED FIELDS TERMINATED BY 
> > 
> > ```
> 
> '\t'LOCATION 's3n://spark-data/wikipedia-sample/';
> 
> > ```
> > > > Timetaken (including network latency):0.159seconds 
> > > > 14/02/1901:23:33INFO CliDriver:Timetaken (including network 
> > 
> > ```
> 
> latency):0.159seconds
> 
> > ```
> > > > 
> > > > [localhost:10000]shark>SELECT title FROM wiki LIMIT 1; 
> > > > Alpokalja 
> > > > Timetaken (including network latency):2.23seconds 
> > > > 14/02/1901:23:48INFO CliDriver:Timetaken (including network 
> > 
> > ```
> 
> latency):2.23seconds
> 
> > ```
> > > > 
> > > > [localhost:10000]shark>CREATE EXTERNAL TABLE es_wiki (id 
> > 
> > ```
> 
> BIGINT,title STRING,last\_modified STRING,xml STRING,text
> 
> > ```
> > > > STRING)STORED BY 
> > 
> > ```
> 
> 'org.elasticsearch.hadoop.hive.EsStorageHandler'TBLPROPERTIES('es.resource'='wikipedia/article');
> 
> > ```
> > > > Timetaken (including network latency):0.061seconds 
> > > > 14/02/1901:33:51INFO CliDriver:Timetaken (including network 
> > 
> > ```
> 
> latency):0.061seconds
> 
> > ```
> > > > 
> > > > [localhost:10000]shark>INSERT OVERWRITE TABLE es_wiki 
> > 
> > ```
> 
> SELECTw.id [http://w.id](http://w.id),w.title,w.last\_modified,w.xml,w.text FROM wiki  
> w;
> 
> > ```
> > > > [HiveError]:Queryreturned non-zero 
> > 
> > ```
> 
> code:9,cause:FAILED:ExecutionError,returncode  
> -101fromshark.execution.SparkTask
> 
> > ```
> > > > Timetaken (including network latency):3.575seconds 
> > > > 14/02/1901:34:42INFO CliDriver:Timetaken (including network 
> > 
> > ```
> 
> latency):3.575seconds
> 
> > ```
> > > > | 
> > > > 
> > > > *The stack trace looks like this:* 
> > > > 
> > > > org.apache.hadoop.hive.ql.metadata.HiveException 
> > 
> > ```
> 
> (org.apache.hadoop.hive.ql.metadata.HiveException: java.io.IOException:
> 
> > ```
> > > > Out of nodes and retries; caught exception) 
> > > > 
> > > > 
> > 
> > ```
> 
> org.apache.hadoop.hive.ql.exec.FileSinkOperator.processOp(FileSinkOperator.java:602)shark.execution.FileSinkOperator$$anonfun$processPartition$1.apply(FileSinkOperator.scala:84)shark.execution.FileSinkOperator$$anonfun$processPartition$1.apply(FileSinkOperator.scala:81)scala.collection.Iterator$class.foreach(Iterator.scala:772)scala.collection.Iterator$$anon$19.foreach(Iterator.scala:399)shark.execution.FileSinkOperator.processPartition(FileSinkOperator.scala:81)shark.execution.FileSinkOperator$.writeFiles$1(FileSinkOperator.scala:207)shark.execution.FileSinkOperator$$anonfun$executeProcessFileSinkPartition$1.apply(FileSinkOperator.scala:211)shark.execution.FileSinkOperator$$anonfun$executeProcessFileSinkPartition$1.apply(FileSinkOperator.scala:211)org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:107)org.apache.spark.scheduler.Task.run(Task.scala:53)org.apache.spark.executor.Executor$TaskRunner$$anonfun$run$1.apply$mcV$sp(Executor.scala:215)org.apache.spa
> 
> rk.dep
> 
> > ```
> > loy.Sp 
> > > 
> > >     
> > 
> > ```
> 
> arkHadoopUtil.runAsUser(SparkHadoopUtil.scala:50)org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:182)java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)java.lang.Thread.run(Thread.java:744
> 
> > ```
> > > 
> > > > I should be using Hive 0.9.0, shark 0.8.1, elasticsearch 
> > 
> > ```
> 
> 1.0.0, Hadoop 1.0.4, and java 1.7.0\_51
> 
> > ```
> > > > Based on my cursory look at the hadoop and 
> > 
> > ```
> 
> elasticsearch-hadoop sources, it looks like hive is just rethrowing an
> 
> > ```
> > > > IOException it's getting from Spark, and 
> > 
> > ```
> 
> elasticsearch-hadoop is just hitting those exceptions.
> 
> > ```
> > > > I suppose my questions are: Does this look like an issue 
> > 
> > ```
> 
> with my ES/elasticsearch-hadoop config? And has anyone gotten
> 
> > ```
> > > > elasticsearch working with Spark/Shark? 
> > > > Any ideas/insights are appreciated. 
> > > > Thanks,Max 
> > > > 
> > > > -- 
> > > > You received this message because you are subscribed to the 
> > 
> > ```
> 
> Google Groups "elasticsearch" group.
> 
> > ```
> > > > To unsubscribe from this group and stop receiving emails 
> > 
> > ```
> 
> from it, send an email to
> 
> > ```
> > > >elasticsearc...@googlegroups.com <javascript:>. 
> > > > To view this discussion on the web visit 
> > > >
> > 
> > ```
> 
> [https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com)
> 
> > ```
> > <
> > 
> > ```
> 
> [https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com)\>
> 
> > ```
> > > <
> > 
> > ```
> 
> [https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com)
> 
> > ```
> > <
> > 
> > ```
> 
> [https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com)\>\>.
> 
> > ```
> > > > For more options, visithttps://
> > 
> > ```
> 
> [groups.google.com/groups/opt\_out](http://groups.google.com/groups/opt_out) [http://groups.google.com/groups/opt\_out](http://groups.google.com/groups/opt_out)  
> \<[https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)
> 
> > ```
> > <https://groups.google.com/groups/opt_out>>. 
> > > 
> > > -- 
> > > Costin 
> > > 
> > > -- 
> > > You received this message because you are subscribed to the Google 
> > 
> > ```
> 
> Groups "elasticsearch" group.
> 
> > ```
> > > To unsubscribe from this group and stop receiving emails from it, 
> > 
> > ```
> 
> send an email to
> 
> > ```
> > >elasticsearc...@googlegroups.com <javascript:>. 
> > > To view this discussion on the web visit 
> > >
> > 
> > ```
> 
> [https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com)
> 
> > ```
> > <
> > 
> > ```
> 
> [https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com)\>.
> 
> > ```
> > > For more options, visithttps://groups.google.com/groups/opt_out <
> > 
> > ```
> 
> [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)\>.
> 
> > ```
> > -- 
> > Costin 
> > 
> > ```
> > 
> > --  
> > You received this message because you are subscribed to the Google  
> > Groups "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send  
> > an email to  
> > [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > To view this discussion on the web visit
> 
> [https://groups.google.com/d/msgid/elasticsearch/e29e342d-de74-4ed6-93d4-875fc728c5a5%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/e29e342d-de74-4ed6-93d4-875fc728c5a5%40googlegroups.com).
> 
> > For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).
> 
> --  
> Costin

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/c1081bf2-117a-4af2-ba90-2c38a4572782%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/c1081bf2-117a-4af2-ba90-2c38a4572782%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [March 20, 2014, 12:44pm UTC](https://discuss.elastic.co/t/hadoop-getting-elasticsearch-hadoop-working-with-shark/15882/8 "2014-03-20T12:44:37Z")

</div>

I recommend using master - there are several improvements done in this area. Also using the latest Shark (0.9.0) and  
Hive (0.12) will help.

On 3/20/14 12:00 PM, Nick Pentreath wrote:

> Hi
> 
> I am struggling to get this working too. I'm just trying locally for now, running Shark 0.8.1, Hive 0.9.0 and ES 1.0.1  
> with ES-hadoop 1.3.0.M2.
> 
> I managed to get a basic example working with WRITING into an index. But I'm really after READING and index.
> 
> I believe I have set everything up correctly, I've added the jar to Shark:  
> ADD JAR /path/to/es-hadoop.jar;
> 
> created a table:  
> CREATE EXTERNAL TABLE test\_read (name string, price double)
> 
> STORED BY 'org.elasticsearch.hadoop.hive.EsStorageHandler'
> 
> TBLPROPERTIES('es.resource' = 'test\_index/test\_type/\_search?q=\*');
> 
> And then trying to 'SELECT \* FROM test \_read' gives me :
> 
> org.apache.spark.SparkException: Job aborted: Task 3.0:0 failed more than 0 times; aborting job  
> java.lang.ClassCastException: org.elasticsearch.hadoop.hive.EsHiveInputFormat$ESHiveSplit cannot be cast to  
> org.elasticsearch.hadoop.hive.EsHiveInputFormat$ESHiveSplit
> 
> at org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:827)
> 
> at org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(DAGScheduler.scala:825)
> 
> at scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:60)
> 
> at scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:47)
> 
> at org.apache.spark.scheduler.DAGScheduler.abortStage(DAGScheduler.scala:825)
> 
> at org.apache.spark.scheduler.DAGScheduler.processEvent(DAGScheduler.scala:440)
> 
> at org.apache.spark.scheduler.DAGScheduler.org$apache$spark$scheduler$DAGScheduler$$run(DAGScheduler.scala:502)
> 
> at org.apache.spark.scheduler.DAGScheduler$$anon$1.run(DAGScheduler.scala:157)
> 
> FAILED: Execution Error, return code -101 from shark.execution.SparkTask
> 
> In fact I get the same error thrown when trying to READ from the table that I successfully WROTE to...
> 
> On Saturday, 22 February 2014 12:31:21 UTC+2, Costin Leau wrote:
> 
> ```
> Yeah, it might have been some sort of network configuration issue where services where running on different machines
> and
> localhost pointed to a different location.
> 
> Either way, I'm glad to hear things have are moving forward.
> 
> Cheers,
> 
> On 22/02/2014 1:06 AM, Max Lang wrote:
> > I managed to get it working on ec2 without issue this time. I'd say the biggest difference was that this time I set up a
> > dedicated ES machine. Is it possible that, because I was using a cluster with slaves, when I used "localhost" the slaves
> > couldn't find the ES instance running on the master? Or do all the requests go through the master?
> >
> >
> > On Wednesday, February 19, 2014 2:35:40 PM UTC-8, Costin Leau wrote:
> >
> > Hi,
> >
> > Setting logging in Hive/Hadoop can be tricky since the log4j needs to be picked up by the running JVM otherwise you
> > won't see anything.
> > Take a look at this link on how to tell Hive to use your logging settings [1].
> >
> > For the next release, we might introduce dedicated exceptions for the simple fact that some libraries, like Hive,
> > swallow the stack trace and it's unclear what the issue is which makes the exception (IllegalStateException) ambiguous.
> >
> > Let me know how it goes and whether you will encounter any issues with Shark. Or if you don't :)
> >
> > Thanks!
> >
> > [1]https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs
> <https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs>
> > <https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs
> <https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs>>
> >
> > On 20/02/2014 12:02 AM, Max Lang wrote:
> > > Hey Costin,
> > >
> > > Thanks for the swift reply. I abandoned EC2 to take that out of the equation and managed to get everything working
> > > locally using the latest version of everything (though I realized just now I'm still on hive 0.9). I'm guessing you're
> > > right about some port connection issue because I definitely had ES running on that machine.
> > >
> > > I changed hive-log4j.properties and added
> > > |
> > > #custom logging levels
> > > #log4j.logger.xxx=DEBUG
> > > log4j.logger.org.elasticsearch.hadoop.rest=TRACE
> > >log4j.logger.org.elasticsearch.hadoop.mr <http://log4j.logger.org.elasticsearch.hadoop.mr>
> <http://log4j.logger.org.elasticsearch.hadoop.mr <http://log4j.logger.org.elasticsearch.hadoop.mr>>=TRACE
> > > |
> > >
> > > But I didn't see any trace logging. Hopefully I can get it working on EC2 without issue, but, for the future, is this
> > > the correct way to set TRACE logging?
> > >
> > > Oh and, for reference, I tried running without ES up and I got the following, exceptions:
> > >
> > > 2014-02-19 13:46:08,803 ERROR shark.SharkDriver (Logging.scala:logError(64)) - FAILED: Hive Internal Error:
> > > java.lang.IllegalStateException(Cannot discover Elasticsearch version)
> > > java.lang.IllegalStateException: Cannot discover Elasticsearch version
> > > at org.elasticsearch.hadoop.hive.EsStorageHandler.init(EsStorageHandler.java:101)
> > > at org.elasticsearch.hadoop.hive.EsStorageHandler.configureOutputJobProperties(EsStorageHandler.java:83)
> > > at org.apache.hadoop.hive.ql.plan.PlanUtils.configureJobPropertiesForStorageHandler(PlanUtils.java:706)
> > > at org.apache.hadoop.hive.ql.plan.PlanUtils.configureOutputJobPropertiesForStorageHandler(PlanUtils.java:675)
> > > at org.apache.hadoop.hive.ql.exec.FileSinkOperator.augmentPlan(FileSinkOperator.java:764)
> > > at org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.putOpInsertMap(SemanticAnalyzer.java:1518)
> > > at org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genFileSinkPlan(SemanticAnalyzer.java:4337)
> > > at org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genPostGroupByBodyPlan(SemanticAnalyzer.java:6207)
> > > at org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genBodyPlan(SemanticAnalyzer.java:6138)
> > > at org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.genPlan(SemanticAnalyzer.java:6764)
> > > at shark.parse.SharkSemanticAnalyzer.analyzeInternal(SharkSemanticAnalyzer.scala:149)
> > > at org.apache.hadoop.hive.ql.parse.BaseSemanticAnalyzer.analyze(BaseSemanticAnalyzer.java:244)
> > > at shark.SharkDriver.compile(SharkDriver.scala:215)
> > > at org.apache.hadoop.hive.ql.Driver.compile(Driver.java:336)
> > > at org.apache.hadoop.hive.ql.Driver.run(Driver.java:895)
> > > at shark.SharkCliDriver.processCmd(SharkCliDriver.scala:324)
> > > at org.apache.hadoop.hive.cli.CliDriver.processLine(CliDriver.java:406)
> > > at shark.SharkCliDriver$.main(SharkCliDriver.scala:232)
> > > at shark.SharkCliDriver.main(SharkCliDriver.scala)
> > > Caused by: java.io.IOException: Out of nodes and retries; caught exception
> > > at org.elasticsearch.hadoop.rest.NetworkClient.execute(NetworkClient.java:81)
> > > at org.elasticsearch.hadoop.rest.RestClient.execute(RestClient.java:221)
> > > at org.elasticsearch.hadoop.rest.RestClient.execute(RestClient.java:205)
> > > at org.elasticsearch.hadoop.rest.RestClient.execute(RestClient.java:209)
> > > at org.elasticsearch.hadoop.rest.RestClient.get(RestClient.java:103)
> > > at org.elasticsearch.hadoop.rest.RestClient.esVersion(RestClient.java:274)
> > > at org.elasticsearch.hadoop.rest.InitializationUtils.discoverEsVersion(InitializationUtils.java:84)
> > > at org.elasticsearch.hadoop.hive.EsStorageHandler.init(EsStorageHandler.java:99)
> > > ... 18 more
> > > Caused by: java.net.ConnectException: Connection refused
> > > at java.net.PlainSocketImpl.socketConnect(Native Method)
> > > at java.net.AbstractPlainSocketImpl.doConnect(AbstractPlainSocketImpl.java:339)
> > > at java.net.AbstractPlainSocketImpl.connectToAddress(AbstractPlainSocketImpl.java:200)
> > > at java.net.AbstractPlainSocketImpl.connect(AbstractPlainSocketImpl.java:182)
> > > at java.net.SocksSocketImpl.connect(SocksSocketImpl.java:391)
> > > at java.net.Socket.connect(Socket.java:579)
> > > at java.net.Socket.connect(Socket.java:528)
> > > at java.net.Socket.<init>(Socket.java:425)
> > > at java.net.Socket.<init>(Socket.java:280)
> > > at org.apache.commons.httpclient.protocol.DefaultProtocolSocketFactory.createSocket(DefaultProtocolSocketFactory.java:80)
> > > at org.apache.commons.httpclient.protocol.DefaultProtocolSocketFactory.createSocket(DefaultProtocolSocketFactory.java:122)
> > > at org.apache.commons.httpclient.HttpConnection.open(HttpConnection.java:707)
> > > at org.apache.commons.httpclient.HttpMethodDirector.executeWithRetry(HttpMethodDirector.java:387)
> > > at org.apache.commons.httpclient.HttpMethodDirector.executeMethod(HttpMethodDirector.java:171)
> > > at org.apache.commons.httpclient.HttpClient.executeMethod(HttpClient.java:397)
> > > at org.apache.commons.httpclient.HttpClient.executeMethod(HttpClient.java:323)
> > > at org.elasticsearch.hadoop.rest.commonshttp.CommonsHttpTransport.execute(CommonsHttpTransport.java:160)
> > > at org.elasticsearch.hadoop.rest.NetworkClient.execute(NetworkClient.java:74)
> > > ... 25 more
> > >
> > > Let me know if there's anything in particular you'd like me to try on EC2.
> > >
> > > (For posterity, the versions I used were: hadoop 2.2.0, hive 0.9.0, shark 8.1, spark 8.1, es-hadoop 1.3.0.M2, java
> > > 1.7.0_15, scala 2.9.3, elasticsearch 1.0.0)
> > >
> > > Thanks again,
> > > Max
> > >
> > > On Tuesday, February 18, 2014 10:16:38 PM UTC-8, Costin Leau wrote:
> > >
> > > The error indicates a network error - namely es-hadoop cannot connect to Elasticsearch on the default (localhost:9200)
> > > HTTP port. Can you double check whether that's indeed the case (using curl or even telnet on that port) - maybe the
> > > firewall prevents any connections to be made...
> > > Also you could try using the latest Hive, 0.12 and a more recent Hadoop such as 1.1.2 or 1.2.1.
> > >
> > > Additionally, can you enable TRACE logging in your job on es-hadoop packages org.elasticsearch.hadoop.rest and
> > >org.elasticsearch.hadoop.mr <http://org.elasticsearch.hadoop.mr> <http://org.elasticsearch.hadoop.mr
> <http://org.elasticsearch.hadoop.mr>> <http://org.elasticsearch.hadoop.mr <http://org.elasticsearch.hadoop.mr>
> > <http://org.elasticsearch.hadoop.mr <http://org.elasticsearch.hadoop.mr>>> packages and report back ?
> > >
> > > Thanks,
> > >
> > > On 19/02/2014 4:03 AM, Max Lang wrote:
> > > > I set everything up using this guide:https://github.com/amplab/shark/wiki/Running-Shark-on-EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2 <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>
> > > <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2 <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>
> > <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>> on an ec2 cluster. I've
> > > > copied the elasticsearch-hadoop jars into the hive lib directory and I have elasticsearch running on localhost:9200. I'm
> > > > running shark in a screen session with --service screenserver and connecting to it at the same time using shark -h
> > > > localhost.
> > > >
> > > > Unfortunately, when I attempt to write data into elasticsearch, it fails. Here's an example:
> > > >
> > > > |
> > > > [localhost:10000]shark>CREATE EXTERNAL TABLE wiki (id BIGINT,title STRING,last_modified STRING,xml STRING,text
> > > > STRING)ROW FORMAT DELIMITED FIELDS TERMINATED BY '\t'LOCATION 's3n://spark-data/wikipedia-sample/';
> > > > Timetaken (including network latency):0.159seconds
> > > > 14/02/1901:23:33INFO CliDriver:Timetaken (including network latency):0.159seconds
> > > >
> > > > [localhost:10000]shark>SELECT title FROM wiki LIMIT 1;
> > > > Alpokalja
> > > > Timetaken (including network latency):2.23seconds
> > > > 14/02/1901:23:48INFO CliDriver:Timetaken (including network latency):2.23seconds
> > > >
> > > > [localhost:10000]shark>CREATE EXTERNAL TABLE es_wiki (id BIGINT,title STRING,last_modified STRING,xml STRING,text
> > > > STRING)STORED BY 'org.elasticsearch.hadoop.hive.EsStorageHandler'TBLPROPERTIES('es.resource'='wikipedia/article');
> > > > Timetaken (including network latency):0.061seconds
> > > > 14/02/1901:33:51INFO CliDriver:Timetaken (including network latency):0.061seconds
> > > >
> > > > [localhost:10000]shark>INSERT OVERWRITE TABLE es_wiki SELECTw.id <http://w.id>,w.title,w.last_modified,w.xml,w.text FROM wiki w;
> > > > [HiveError]:Queryreturned non-zero code:9,cause:FAILED:ExecutionError,returncode -101fromshark.execution.SparkTask
> > > > Timetaken (including network latency):3.575seconds
> > > > 14/02/1901:34:42INFO CliDriver:Timetaken (including network latency):3.575seconds
> > > > |
> > > >
> > > > *The stack trace looks like this:*
> > > >
> > > > org.apache.hadoop.hive.ql.metadata.HiveException (org.apache.hadoop.hive.ql.metadata.HiveException: java.io.IOException:
> > > > Out of nodes and retries; caught exception)
> > > >
> > > > org.apache.hadoop.hive.ql.exec.FileSinkOperator.processOp(FileSinkOperator.java:602)shark.execution.FileSinkOperator$$anonfun$processPartition$1.apply(FileSinkOperator.scala:84)shark.execution.FileSinkOperator$$anonfun$processPartition$1.apply(FileSinkOperator.scala:81)scala.collection.Iterator$class.foreach(Iterator.scala:772)scala.collection.Iterator$$anon$19.foreach(Iterator.scala:399)shark.execution.FileSinkOperator.processPartition(FileSinkOperator.scala:81)shark.execution.FileSinkOperator$.writeFiles$1(FileSinkOperator.scala:207)shark.execution.FileSinkOperator$$anonfun$executeProcessFileSinkPartition$1.apply(FileSinkOperator.scala:211)shark.execution.FileSinkOperator$$anonfun$executeProcessFileSinkPartition$1.apply(FileSinkOperator.scala:211)org.apache.spark.scheduler.ResultTask.runTask(ResultTask.scala:107)org.apache.spark.scheduler.Task.run(Task.scala:53)org.apache.spark.executor.Executor$TaskRunner$$anonfun$run$1.apply$mcV$sp(Executor.scala:215)org.apac
> 
> ```

he.spa

> ```
> rk.dep
> >
> > loy.Sp
> > >
> > > arkHadoopUtil.runAsUser(SparkHadoopUtil.scala:50)org.apache.spark.executor.Executor$TaskRunner.run(Executor.scala:182)java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1145)java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:615)java.lang.Thread.run(Thread.java:744
> 
> >
> > >
> > > > I should be using Hive 0.9.0, shark 0.8.1, elasticsearch 1.0.0, Hadoop 1.0.4, and java 1.7.0_51
> > > > Based on my cursory look at the hadoop and elasticsearch-hadoop sources, it looks like hive is just rethrowing an
> > > > IOException it's getting from Spark, and elasticsearch-hadoop is just hitting those exceptions.
> > > > I suppose my questions are: Does this look like an issue with my ES/elasticsearch-hadoop config? And has anyone gotten
> > > > elasticsearch working with Spark/Shark?
> > > > Any ideas/insights are appreciated.
> > > > Thanks,Max
> > > >
> > > > --
> > > > You received this message because you are subscribed to the Google Groups "elasticsearch" group.
> > > > To unsubscribe from this group and stop receiving emails from it, send an email to
> > > >elasticsearc...@googlegroups.com <javascript:>.
> > > > To view this discussion on the web visit
> > > >https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>
> > <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>
> > > <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>
> > <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>>.
> > > > For more options, visithttps://groups.google.com/groups/opt_out <http://groups.google.com/groups/opt_out> <http://groups.google.com/groups/opt_out
> <http://groups.google.com/groups/opt_out>> <https://groups.google.com/groups/opt_out
> <https://groups.google.com/groups/opt_out>
> > <https://groups.google.com/groups/opt_out <https://groups.google.com/groups/opt_out>>>.
> > >
> > > --
> > > Costin
> > >
> > > --
> > > You received this message because you are subscribed to the Google Groups "elasticsearch" group.
> > > To unsubscribe from this group and stop receiving emails from it, send an email to
> > >elasticsearc...@googlegroups.com <javascript:>.
> > > To view this discussion on the web visit
> > >https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com>
> > <https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com>>.
> > > For more options, visithttps://groups.google.com/groups/opt_out <http://groups.google.com/groups/opt_out> <https://groups.google.com/groups/opt_out
> <https://groups.google.com/groups/opt_out>>.
> >
> > --
> > Costin
> >
> > --
> > You received this message because you are subscribed to the Google Groups "elasticsearch" group.
> > To unsubscribe from this group and stop receiving emails from it, send an email to
> >elasticsearc...@googlegroups.com <javascript:>.
> > To view this discussion on the web visit
> >https://groups.google.com/d/msgid/elasticsearch/e29e342d-de74-4ed6-93d4-875fc728c5a5%40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/e29e342d-de74-4ed6-93d4-875fc728c5a5%40googlegroups.com>.
> > For more options, visithttps://groups.google.com/groups/opt_out <https://groups.google.com/groups/opt_out>.
> 
> --
> Costin
> 
> ```
> 
> --  
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an email to  
> [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com) [mailto:elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/c1081bf2-117a-4af2-ba90-2c38a4572782%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/c1081bf2-117a-4af2-ba90-2c38a4572782%40googlegroups.com)  
> [https://groups.google.com/d/msgid/elasticsearch/c1081bf2-117a-4af2-ba90-2c38a4572782%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/c1081bf2-117a-4af2-ba90-2c38a4572782%40googlegroups.com?utm_medium=email&utm_source=footer).  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
Costin

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/532AE2B5.8080004%40gmail.com](https://groups.google.com/d/msgid/elasticsearch/532AE2B5.8080004%40gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Nick\_Pentreath](https://avatars.discourse-cdn.com/v4/letter/n/7ea924/32.png) [@Nick\_Pentreath](https://discuss.elastic.co/u/Nick_Pentreath)\
**Post date:** [March 27, 2014, 2:29pm UTC](https://discuss.elastic.co/t/hadoop-getting-elasticsearch-hadoop-working-with-shark/15882/9 "2014-03-27T14:29:15Z")

</div>

Thanks for the response.

I tried latest Shark (cdh4 version of 0.9.1 here  
[http://cloudera.rst.im/shark/](http://cloudera.rst.im/shark/) ) - this uses hadoop 1.0.4 and hive 0.11 I  
believe, and build elasticsearch-hadoop from github master.

Still getting same error:  
org.elasticsearch.hadoop.hive.EsHiveInputFormat$EsHiveSplit cannot be cast  
to org.elasticsearch.hadoop.hive.EsHiveInputFormat$EsHiveSplit

Will using hive 0.11 / hadoop 1.0.4 vs hive 0.12 / hadoop 1.2.1 in  
es-hadoop master make a difference?

Anyone else actually got this working?

On Thu, Mar 20, 2014 at 2:44 PM, Costin Leau [costin.leau@gmail.com](mailto:costin.leau@gmail.com) wrote:

> I recommend using master - there are several improvements done in this  
> area. Also using the latest Shark (0.9.0) and Hive (0.12) will help.
> 
> On 3/20/14 12:00 PM, Nick Pentreath wrote:
> 
> > Hi
> > 
> > I am struggling to get this working too. I'm just trying locally for now,  
> > running Shark 0.8.1, Hive 0.9.0 and ES 1.0.1  
> > with ES-hadoop 1.3.0.M2.
> > 
> > I managed to get a basic example working with WRITING into an index. But  
> > I'm really after READING and index.
> > 
> > I believe I have set everything up correctly, I've added the jar to Shark:  
> > ADD JAR /path/to/es-hadoop.jar;
> > 
> > created a table:  
> > CREATE EXTERNAL TABLE test\_read (name string, price double)
> > 
> > STORED BY 'org.elasticsearch.hadoop.hive.EsStorageHandler'
> > 
> > TBLPROPERTIES('es.resource' = 'test\_index/test\_type/\_search?q=\*');
> > 
> > And then trying to 'SELECT \* FROM test \_read' gives me :
> > 
> > org.apache.spark.SparkException: Job aborted: Task 3.0:0 failed more  
> > than 0 times; aborting job  
> > java.lang.ClassCastException: org.elasticsearch.hadoop.hive.EsHiveInputFormat$ESHiveSplit  
> > cannot be cast to  
> > org.elasticsearch.hadoop.hive.EsHiveInputFormat$ESHiveSplit
> > 
> > at org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(  
> > DAGScheduler.scala:827)
> > 
> > at org.apache.spark.scheduler.DAGScheduler$$anonfun$abortStage$1.apply(  
> > DAGScheduler.scala:825)
> > 
> > at scala.collection.mutable.ResizableArray$class.foreach(  
> > ResizableArray.scala:60)
> > 
> > at scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:47)
> > 
> > at org.apache.spark.scheduler.DAGScheduler.abortStage(  
> > DAGScheduler.scala:825)
> > 
> > at org.apache.spark.scheduler.DAGScheduler.processEvent(  
> > DAGScheduler.scala:440)
> > 
> > at org.apache.spark.scheduler.DAGScheduler.org$apache$spark$  
> > scheduler$DAGScheduler$$run(DAGScheduler.scala:502)
> > 
> > at org.apache.spark.scheduler.DAGScheduler$$anon$1.run(  
> > DAGScheduler.scala:157)
> > 
> > FAILED: Execution Error, return code -101 from shark.execution.SparkTask
> > 
> > In fact I get the same error thrown when trying to READ from the table  
> > that I successfully WROTE to...
> > 
> > On Saturday, 22 February 2014 12:31:21 UTC+2, Costin Leau wrote:
> > 
> > ```
> > Yeah, it might have been some sort of network configuration issue
> > 
> > ```
> > 
> > where services where running on different machines  
> > and  
> > localhost pointed to a different location.
> > 
> > ```
> > Either way, I'm glad to hear things have are moving forward.
> > 
> > Cheers,
> > 
> > On 22/02/2014 1:06 AM, Max Lang wrote:
> > > I managed to get it working on ec2 without issue this time. I'd say
> > 
> > ```
> > 
> > the biggest difference was that this time I set up a  
> > \> dedicated ES machine. Is it possible that, because I was using a  
> > cluster with slaves, when I used "localhost" the slaves  
> > \> couldn't find the ES instance running on the master? Or do all the  
> > requests go through the master?  
> > \>  
> > \>  
> > \> On Wednesday, February 19, 2014 2:35:40 PM UTC-8, Costin Leau wrote:  
> > \>  
> > \> Hi,  
> > \>  
> > \> Setting logging in Hive/Hadoop can be tricky since the log4j  
> > needs to be picked up by the running JVM otherwise you  
> > \> won't see anything.  
> > \> Take a look at this link on how to tell Hive to use your  
> > logging settings [1].  
> > \>  
> > \> For the next release, we might introduce dedicated exceptions  
> > for the simple fact that some libraries, like Hive,  
> > \> swallow the stack trace and it's unclear what the issue is  
> > which makes the exception (IllegalStateException) ambiguous.  
> > \>  
> > \> Let me know how it goes and whether you will encounter any  
> > issues with Shark. Or if you don't 🙂  
> > \>  
> > \> Thanks!  
> > \>  
> > \> [1][Home - Apache Hive - Apache Software Foundation](https://cwiki.apache.org/confluence/display/Hive/)  
> > GettingStarted#GettingStarted-ErrorLogs  
> > \<[Home - Apache Hive - Apache Software Foundation](https://cwiki.apache.org/confluence/display/Hive/)  
> > GettingStarted#GettingStarted-ErrorLogs\>  
> > \> \<[Home - Apache Hive - Apache Software Foundation](https://cwiki.apache.org/confluence/display/Hive/)  
> > GettingStarted#GettingStarted-ErrorLogs  
> > \<[Home - Apache Hive - Apache Software Foundation](https://cwiki.apache.org/confluence/display/Hive/)  
> > GettingStarted#GettingStarted-ErrorLogs\>\>  
> > \>  
> > \> On 20/02/2014 12:02 AM, Max Lang wrote:  
> > \> \> Hey Costin,  
> > \> \>  
> > \> \> Thanks for the swift reply. I abandoned EC2 to take that out  
> > of the equation and managed to get everything working  
> > \> \> locally using the latest version of everything (though I  
> > realized just now I'm still on hive 0.9). I'm guessing you're  
> > \> \> right about some port connection issue because I definitely  
> > had ES running on that machine.  
> > \> \>  
> > \> \> I changed hive-log4j.properties and added  
> > \> \> |  
> > \> \> #custom logging levels  
> > \> \> #log4j.logger.xxx=DEBUG  
> > \> \> log4j.logger.org.elasticsearch.hadoop.rest=TRACE  
> > \> \>log4j.logger.org.elasticsearch.hadoop.mr \<  
> > [http://log4j.logger.org.elasticsearch.hadoop.mr](http://log4j.logger.org.elasticsearch.hadoop.mr)\>  
> > \<[http://log4j.logger.org.elasticsearch.hadoop.mr](http://log4j.logger.org.elasticsearch.hadoop.mr) \<  
> > [http://log4j.logger.org.elasticsearch.hadoop.mr](http://log4j.logger.org.elasticsearch.hadoop.mr)\>\>=TRACE
> > 
> > ```
> > > > |
> > > >
> > > > But I didn't see any trace logging. Hopefully I can get it
> > 
> > ```
> > 
> > working on EC2 without issue, but, for the future, is this  
> > \> \> the correct way to set TRACE logging?  
> > \> \>  
> > \> \> Oh and, for reference, I tried running without ES up and I  
> > got the following, exceptions:  
> > \> \>  
> > \> \> 2014-02-19 13:46:08,803 ERROR shark.SharkDriver  
> > (Logging.scala:logError(64)) - FAILED: Hive Internal Error:  
> > \> \> java.lang.IllegalStateException(Cannot discover  
> > Elasticsearch version)  
> > \> \> java.lang.IllegalStateException: Cannot discover  
> > Elasticsearch version  
> > \> \> at org.elasticsearch.hadoop.hive.EsStorageHandler.init(  
> > EsStorageHandler.java:101)  
> > \> \> at org.elasticsearch.hadoop.hive.EsStorageHandler.  
> > configureOutputJobProperties(EsStorageHandler.java:83)  
> > \> \> at org.apache.hadoop.hive.ql.plan.PlanUtils.  
> > configureJobPropertiesForStorageHandler(PlanUtils.java:706)  
> > \> \> at org.apache.hadoop.hive.ql.plan.PlanUtils.  
> > configureOutputJobPropertiesForStorageHandler(PlanUtils.java:675)  
> > \> \> at org.apache.hadoop.hive.ql.exec.FileSinkOperator.  
> > augmentPlan(FileSinkOperator.java:764)  
> > \> \> at org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.  
> > putOpInsertMap(SemanticAnalyzer.java:1518)  
> > \> \> at org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.  
> > genFileSinkPlan(SemanticAnalyzer.java:4337)  
> > \> \> at org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.  
> > genPostGroupByBodyPlan(SemanticAnalyzer.java:6207)  
> > \> \> at org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.  
> > genBodyPlan(SemanticAnalyzer.java:6138)  
> > \> \> at org.apache.hadoop.hive.ql.parse.SemanticAnalyzer.  
> > genPlan(SemanticAnalyzer.java:6764)  
> > \> \> at shark.parse.SharkSemanticAnalyzer.analyzeInternal(  
> > SharkSemanticAnalyzer.scala:149)  
> > \> \> at org.apache.hadoop.hive.ql.parse.BaseSemanticAnalyzer.  
> > analyze(BaseSemanticAnalyzer.java:244)  
> > \> \> at shark.SharkDriver.compile(SharkDriver.scala:215)  
> > \> \> at org.apache.hadoop.hive.ql.Driver.compile(Driver.java:336)  
> > \> \> at org.apache.hadoop.hive.ql.Driver.run(Driver.java:895)  
> > \> \> at shark.SharkCliDriver.processCmd(SharkCliDriver.scala:324)  
> > \> \> at org.apache.hadoop.hive.cli.CliDriver.processLine(  
> > CliDriver.java:406)  
> > \> \> at shark.SharkCliDriver$.main(SharkCliDriver.scala:232)  
> > \> \> at shark.SharkCliDriver.main(SharkCliDriver.scala)  
> > \> \> Caused by: java.io.IOException: Out of nodes and retries;  
> > caught exception  
> > \> \> at org.elasticsearch.hadoop.rest.NetworkClient.execute(  
> > NetworkClient.java:81)  
> > \> \> at org.elasticsearch.hadoop.rest.  
> > RestClient.execute(RestClient.java:221)  
> > \> \> at org.elasticsearch.hadoop.rest.  
> > RestClient.execute(RestClient.java:205)  
> > \> \> at org.elasticsearch.hadoop.rest.  
> > RestClient.execute(RestClient.java:209)  
> > \> \> at org.elasticsearch.hadoop.rest.RestClient.get(RestClient.  
> > java:103)  
> > \> \> at org.elasticsearch.hadoop.rest.RestClient.esVersion(  
> > RestClient.java:274)  
> > \> \> at org.elasticsearch.hadoop.rest.InitializationUtils.  
> > discoverEsVersion(InitializationUtils.java:84)  
> > \> \> at org.elasticsearch.hadoop.hive.EsStorageHandler.init(  
> > EsStorageHandler.java:99)  
> > \> \> ... 18 more  
> > \> \> Caused by: java.net.ConnectException: Connection refused  
> > \> \> at java.net.PlainSocketImpl.socketConnect(Native Method)  
> > \> \> at java.net.AbstractPlainSocketImpl.doConnect(  
> > AbstractPlainSocketImpl.java:339)  
> > \> \> at java.net.AbstractPlainSocketImpl.connectToAddress(  
> > AbstractPlainSocketImpl.java:200)  
> > \> \> at java.net.AbstractPlainSocketImpl.connect(  
> > AbstractPlainSocketImpl.java:182)  
> > \> \> at java.net.SocksSocketImpl.connect(SocksSocketImpl.java:391)  
> > \> \> at java.net.Socket.connect(Socket.java:579)  
> > \> \> at java.net.Socket.connect(Socket.java:528)  
> > \> \> at java.net.Socket.(Socket.java:425)  
> > \> \> at java.net.Socket.(Socket.java:280)  
> > \> \> at org.apache.commons.httpclient.protocol.  
> > DefaultProtocolSocketFactory.createSocket(DefaultProtocolSocketFactory.  
> > java:80)  
> > \> \> at org.apache.commons.httpclient.protocol.  
> > DefaultProtocolSocketFactory.createSocket(DefaultProtocolSocketFactory.  
> > java:122)  
> > \> \> at org.apache.commons.httpclient.HttpConnection.open(  
> > HttpConnection.java:707)  
> > \> \> at org.apache.commons.httpclient.HttpMethodDirector.  
> > executeWithRetry(HttpMethodDirector.java:387)  
> > \> \> at org.apache.commons.httpclient.HttpMethodDirector.  
> > executeMethod(HttpMethodDirector.java:171)  
> > \> \> at org.apache.commons.httpclient.HttpClient.executeMethod(  
> > HttpClient.java:397)  
> > \> \> at org.apache.commons.httpclient.HttpClient.executeMethod(  
> > HttpClient.java:323)  
> > \> \> at org.elasticsearch.hadoop.rest.commonshttp.  
> > CommonsHttpTransport.execute(CommonsHttpTransport.java:160)  
> > \> \> at org.elasticsearch.hadoop.rest.NetworkClient.execute(  
> > NetworkClient.java:74)  
> > \> \> ... 25 more  
> > \> \>  
> > \> \> Let me know if there's anything in particular you'd like me  
> > to try on EC2.  
> > \> \>  
> > \> \> (For posterity, the versions I used were: hadoop 2.2.0, hive  
> > 0.9.0, shark 8.1, spark 8.1, es-hadoop 1.3.0.M2, java  
> > \> \> 1.7.0\_15, scala 2.9.3, elasticsearch 1.0.0)  
> > \> \>  
> > \> \> Thanks again,  
> > \> \> Max  
> > \> \>  
> > \> \> On Tuesday, February 18, 2014 10:16:38 PM UTC-8, Costin Leau  
> > wrote:  
> > \> \>  
> > \> \> The error indicates a network error - namely es-hadoop  
> > cannot connect to Elasticsearch on the default (localhost:9200)  
> > \> \> HTTP port. Can you double check whether that's indeed the  
> > case (using curl or even telnet on that port) - maybe the  
> > \> \> firewall prevents any connections to be made...  
> > \> \> Also you could try using the latest Hive, 0.12 and a more  
> > recent Hadoop such as 1.1.2 or 1.2.1.  
> > \> \>  
> > \> \> Additionally, can you enable TRACE logging in your job on  
> > es-hadoop packages org.elasticsearch.hadoop.rest and  
> > \> \>org.elasticsearch.hadoop.mr \<[http://org.elasticsearch](http://org.elasticsearch).  
> > hadoop.mr\> \<[http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)  
> > [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)\> \<[http://org.elasticsearch](http://org.elasticsearch).  
> > hadoop.mr [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)
> > 
> > ```
> > > <http://org.elasticsearch.hadoop.mr <http://org.elasticsearch.
> > 
> > ```
> > 
> > hadoop.mr\>\>\> packages and report back ?  
> > \> \>  
> > \> \> Thanks,  
> > \> \>  
> > \> \> On 19/02/2014 4:03 AM, Max Lang wrote:  
> > \> \> \> I set everything up using this guide:  
> > [Running Shark on EC2 · amplab/shark Wiki · GitHub](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)  
> > \<[Running Shark on EC2 · amplab/shark Wiki · GitHub](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2) \<  
> > [Running Shark on EC2 · amplab/shark Wiki · GitHub](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>\>  
> > \> \> \<[Create new page · amplab/shark Wiki · GitHub](https://github.com/amplab/shark/wiki/Running-Shark-on-)  
> > EC2 [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)  
> > \> \<[Running Shark on EC2 · amplab/shark Wiki · GitHub](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>\> on an  
> > ec2 cluster. I've  
> > \> \> \> copied the elasticsearch-hadoop jars into the hive lib  
> > directory and I have elasticsearch running on localhost:9200. I'm  
> > \> \> \> running shark in a screen session with --service  
> > screenserver and connecting to it at the same time using shark -h  
> > \> \> \> localhost.  
> > \> \> \>  
> > \> \> \> Unfortunately, when I attempt to write data into  
> > elasticsearch, it fails. Here's an example:  
> > \> \> \>  
> > \> \> \> |  
> > \> \> \> [localhost:10000]shark\>CREATE EXTERNAL TABLE wiki (id  
> > BIGINT,title STRING,last\_modified STRING,xml STRING,text  
> > \> \> \> STRING)ROW FORMAT DELIMITED FIELDS TERMINATED BY  
> > '\t'LOCATION 's3n://spark-data/wikipedia-sample/';  
> > \> \> \> Timetaken (including network latency):0.159seconds  
> > \> \> \> 14/02/1901:23:33INFO CliDriver:Timetaken (including  
> > network latency):0.159seconds  
> > \> \> \>  
> > \> \> \> [localhost:10000]shark\>SELECT title FROM wiki LIMIT 1;  
> > \> \> \> Alpokalja  
> > \> \> \> Timetaken (including network latency):2.23seconds  
> > \> \> \> 14/02/1901:23:48INFO CliDriver:Timetaken (including  
> > network latency):2.23seconds  
> > \> \> \>  
> > \> \> \> [localhost:10000]shark\>CREATE EXTERNAL TABLE es\_wiki  
> > (id BIGINT,title STRING,last\_modified STRING,xml STRING,text  
> > \> \> \> STRING)STORED BY 'org.elasticsearch.hadoop.  
> > hive.EsStorageHandler'TBLPROPERTIES('es.resource'='wikipedia/article');  
> > \> \> \> Timetaken (including network latency):0.061seconds  
> > \> \> \> 14/02/1901:33:51INFO CliDriver:Timetaken (including  
> > network latency):0.061seconds  
> > \> \> \>  
> > \> \> \> [localhost:10000]shark\>INSERT OVERWRITE TABLE es\_wiki  
> > SELECTw.id [http://w.id](http://w.id),w.title,w.last\_modified,w.xml,w.text FROM wiki  
> > w;  
> > \> \> \> [HiveError]:Queryreturned non-zero code:9,cause:FAILED:ExecutionError,returncode  
> > -101fromshark.execution.SparkTask  
> > \> \> \> Timetaken (including network latency):3.575seconds  
> > \> \> \> 14/02/1901:34:42INFO CliDriver:Timetaken (including  
> > network latency):3.575seconds  
> > \> \> \> |  
> > \> \> \>  
> > \> \> \> _The stack trace looks like this:_  
> > \> \> \>  
> > \> \> \> org.apache.hadoop.hive.ql.metadata.HiveException  
> > (org.apache.hadoop.hive.ql.metadata.HiveException: java.io.IOException:  
> > \> \> \> Out of nodes and retries; caught exception)  
> > \> \> \>  
> > \> \> \> org.apache.hadoop.hive.ql.exec.FileSinkOperator.  
> > processOp(FileSinkOperator.java:602)shark.execution.  
> > FileSinkOperator$$anonfun$processPartition$1.apply(  
> > FileSinkOperator.scala:84)shark.execution.FileSinkOperator$$anonfun$  
> > processPartition$1.apply(FileSinkOperator.scala:81)  
> > scala.collection.Iterator$class.foreach(Iterator.scala:  
> > 772)scala.collection.Iterator$$anon$19.foreach(Iterator.  
> > scala:399)shark.execution.FileSinkOperator.processPartition(  
> > FileSinkOperator.scala:81)shark.execution.FileSinkOperator$.writeFiles$  
> > 1(FileSinkOperator.scala:207)shark.execution.FileSinkOperator$$anonfun$  
> > executeProcessFileSinkPartition$1.apply(FileSinkOperator.  
> > scala:211)shark.execution.FileSinkOperator$$anonfun$  
> > executeProcessFileSinkPartition$1.apply(FileSinkOperator.  
> > scala:211)org.apache.spark.scheduler.ResultTask.runTask(  
> > ResultTask.scala:107)org.apache.spark.scheduler.Task.  
> > run(Task.scala:53)org.apache.spark.executor.Executor$  
> > TaskRunner$$anonfun$run$1.apply$mcV$sp(Executor.scala:215)org.apac
> 
> he.spa
> 
> > ```
> > rk.dep
> > >
> > > loy.Sp
> > > >
> > > > arkHadoopUtil.runAsUser(SparkHadoopUtil.scala:50)org.
> > 
> > ```
> > 
> > apache.spark.executor.Executor$TaskRunner.run(  
> > Executor.scala:182)java.util.concurrent.ThreadPoolExecutor.  
> > runWorker(ThreadPoolExecutor.java:1145)java.util.  
> > concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.  
> > java:615)java.lang.Thread.run(Thread.java:744
> > 
> > ```
> > >
> > > >
> > > > > I should be using Hive 0.9.0, shark 0.8.1,
> > 
> > ```
> > 
> > elasticsearch 1.0.0, Hadoop 1.0.4, and java 1.7.0\_51  
> > \> \> \> Based on my cursory look at the hadoop and  
> > elasticsearch-hadoop sources, it looks like hive is just rethrowing an  
> > \> \> \> IOException it's getting from Spark, and  
> > elasticsearch-hadoop is just hitting those exceptions.  
> > \> \> \> I suppose my questions are: Does this look like an  
> > issue with my ES/elasticsearch-hadoop config? And has anyone gotten  
> > \> \> \> elasticsearch working with Spark/Shark?  
> > \> \> \> Any ideas/insights are appreciated.  
> > \> \> \> Thanks,Max  
> > \> \> \>  
> > \> \> \> --  
> > \> \> \> You received this message because you are subscribed to  
> > the Google Groups "elasticsearch" group.  
> > \> \> \> To unsubscribe from this group and stop receiving  
> > emails from it, send an email to  
> > \> \> \>[elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > \> \> \> To view this discussion on the web visit  
> > \> \> \>[https://groups.google.com/d/](https://groups.google.com/d/)  
> > msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%  
> > [40googlegroups.com](http://40googlegroups.com)  
> > \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)\>  
> > \> \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)  
> > \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)\>\>  
> > \> \> \<[https://groups.google.com/d/](https://groups.google.com/d/)  
> > msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%  
> > [40googlegroups.com](http://40googlegroups.com)  
> > \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)\>  
> > \> \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)  
> > \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)\>\>\>.  
> > \> \> \> For more options, visithttps://groups.google.  
> > com/groups/opt\_out [http://groups.google.com/groups/opt\_out](http://groups.google.com/groups/opt_out) \<  
> > [http://groups.google.com/groups/opt\_out](http://groups.google.com/groups/opt_out)
> > 
> > ```
> > <http://groups.google.com/groups/opt_out>> <
> > 
> > ```
> > 
> > [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)  
> > [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)  
> > \> \<[https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out) \<  
> > [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)\>\>\>.  
> > \> \>  
> > \> \> --  
> > \> \> Costin  
> > \> \>  
> > \> \> --  
> > \> \> You received this message because you are subscribed to the  
> > Google Groups "elasticsearch" group.  
> > \> \> To unsubscribe from this group and stop receiving emails from  
> > it, send an email to  
> > \> \>[elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > \> \> To view this discussion on the web visit  
> > \> \>[https://groups.google.com/d/msgid/elasticsearch/86187c3a-](https://groups.google.com/d/msgid/elasticsearch/86187c3a-)  
> > 0974-4d10-9689-e83da788c04a%[40googlegroups.com](http://40googlegroups.com)  
> > \<[https://groups.google.com/d/msgid/elasticsearch/86187c3a-](https://groups.google.com/d/msgid/elasticsearch/86187c3a-)  
> > 0974-4d10-9689-e83da788c04a%[40googlegroups.com](http://40googlegroups.com)\>  
> > \> \<[https://groups.google.com/d/msgid/elasticsearch/86187c3a-](https://groups.google.com/d/msgid/elasticsearch/86187c3a-)  
> > 0974-4d10-9689-e83da788c04a%[40googlegroups.com](http://40googlegroups.com)  
> > \<[https://groups.google.com/d/msgid/elasticsearch/86187c3a-](https://groups.google.com/d/msgid/elasticsearch/86187c3a-)  
> > 0974-4d10-9689-e83da788c04a%[40googlegroups.com](http://40googlegroups.com)\>\>.  
> > \> \> For more options, visithttps://groups.google.  
> > com/groups/opt\_out [http://groups.google.com/groups/opt\_out](http://groups.google.com/groups/opt_out) \<  
> > [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)  
> > [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)\>.  
> > \>  
> > \> --  
> > \> Costin  
> > \>  
> > \> --  
> > \> You received this message because you are subscribed to the Google  
> > Groups "elasticsearch" group.  
> > \> To unsubscribe from this group and stop receiving emails from it,  
> > send an email to  
> > \>[elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com) \<javascript:\>.  
> > \> To view this discussion on the web visit  
> > \>[https://groups.google.com/d/msgid/elasticsearch/e29e342d-](https://groups.google.com/d/msgid/elasticsearch/e29e342d-)  
> > de74-4ed6-93d4-875fc728c5a5%[40googlegroups.com](http://40googlegroups.com)  
> > \<[https://groups.google.com/d/msgid/elasticsearch/e29e342d-](https://groups.google.com/d/msgid/elasticsearch/e29e342d-)  
> > de74-4ed6-93d4-875fc728c5a5%[40googlegroups.com](http://40googlegroups.com)\>.
> > 
> > ```
> > > For more options, visithttps://groups.google.com/groups/opt_out <
> > 
> > ```
> > 
> > [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)\>.
> > 
> > ```
> > --
> > Costin
> > 
> > ```
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to  
> > [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com) \<mailto:elasticsearch+  
> > [unsubscribe@googlegroups.com](mailto:unsubscribe@googlegroups.com)\>.
> > 
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/c1081bf2-](https://groups.google.com/d/msgid/elasticsearch/c1081bf2-)  
> > 117a-4af2-ba90-2c38a4572782%[40googlegroups.com](http://40googlegroups.com)  
> > \<[https://groups.google.com/d/msgid/elasticsearch/c1081bf2-](https://groups.google.com/d/msgid/elasticsearch/c1081bf2-)  
> > 117a-4af2-ba90-2c38a4572782%[40GGGROUPS CASINO – Real Slot Casino for 10,000+ Senior Players](http://40googlegroups.com?utm_medium=)  
> > email&utm\_source=footer\>.
> > 
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).
> 
> --  
> Costin
> 
> --  
> You received this message because you are subscribed to a topic in the  
> Google Groups "elasticsearch" group.  
> To unsubscribe from this topic, visit [https://groups.google.com/d/](https://groups.google.com/d/)  
> topic/elasticsearch/S-BrzwUHJbM/unsubscribe.  
> To unsubscribe from this group and all its topics, send an email to  
> [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit [https://groups.google.com/d/](https://groups.google.com/d/)  
> msgid/elasticsearch/532AE2B5.8080004%[40gmail.com](http://40gmail.com).
> 
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GNJD0wMJPzXwQqvfL4%2B0nZmw4XzFrPdEc%2BOPLZVeNuZpw%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GNJD0wMJPzXwQqvfL4%2B0nZmw4XzFrPdEc%2BOPLZVeNuZpw%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [March 27, 2014, 2:50pm UTC](https://discuss.elastic.co/t/hadoop-getting-elasticsearch-hadoop-working-with-shark/15882/10 "2014-03-27T14:50:14Z")

</div>

Using the latest hive and hadoop is preferred as they contain various bug fixes.  
The error suggests a classpath issue - namely the same class is loaded twice for some reason and hence the casting fails.

Let's connect on IRC - give me a ping when you're available (user is costin).

Cheers,

On 3/27/14 4:29 PM, Nick Pentreath wrote:

> Thanks for the response.
> 
> I tried latest Shark (cdh4 version of 0.9.1 here [http://cloudera.rst.im/shark/](http://cloudera.rst.im/shark/) ) - this uses hadoop 1.0.4 and hive 0.11  
> I believe, and build elasticsearch-hadoop from github master.
> 
> Still getting same error:  
> org.elasticsearch.hadoop.hive.EsHiveInputFormat$EsHiveSplit cannot be cast to  
> org.elasticsearch.hadoop.hive.EsHiveInputFormat$EsHiveSplit
> 
> Will using hive 0.11 / hadoop 1.0.4 vs hive 0.12 / hadoop 1.2.1 in es-hadoop master make a difference?
> 
> Anyone else actually got this working?
> 
> On Thu, Mar 20, 2014 at 2:44 PM, Costin Leau \<[costin.leau@gmail.com](mailto:costin.leau@gmail.com) [mailto:costin.leau@gmail.com](mailto:costin.leau@gmail.com)\> wrote:
> 
> ```
> I recommend using master - there are several improvements done in this area. Also using the latest Shark (0.9.0) and
> Hive (0.12) will help.
> 
> On 3/20/14 12:00 PM, Nick Pentreath wrote:
> 
> Hi
> 
> I am struggling to get this working too. I'm just trying locally for now, running Shark 0.8.1, Hive 0.9.0 and ES
> 1.0.1
> with ES-hadoop 1.3.0.M2.
> 
> I managed to get a basic example working with WRITING into an index. But I'm really after READING and index.
> 
> I believe I have set everything up correctly, I've added the jar to Shark:
> ADD JAR /path/to/es-hadoop.jar;
> 
> created a table:
> CREATE EXTERNAL TABLE test_read (name string, price double)
> 
> STORED BY 'org.elasticsearch.hadoop.__hive.EsStorageHandler'
> 
> TBLPROPERTIES('es.resource' = 'test_index/test_type/_search?__q=*');
> 
> And then trying to 'SELECT * FROM test _read' gives me :
> 
> org.apache.spark.__SparkException: Job aborted: Task 3.0:0 failed more than 0 times; aborting job
> java.lang.ClassCastException: org.elasticsearch.hadoop.hive.__EsHiveInputFormat$ESHiveSplit cannot be cast to
> org.elasticsearch.hadoop.hive.__EsHiveInputFormat$ESHiveSplit
> 
> at org.apache.spark.scheduler. __DAGScheduler$$anonfun$__ abortStage$1.apply(__DAGScheduler.scala:827)
> 
> at org.apache.spark.scheduler. __DAGScheduler$$anonfun$__ abortStage$1.apply(__DAGScheduler.scala:825)
> 
> at scala.collection.mutable.__ResizableArray$class.foreach(__ResizableArray.scala:60)
> 
> at scala.collection.mutable.__ArrayBuffer.foreach(__ArrayBuffer.scala:47)
> 
> at org.apache.spark.scheduler.__DAGScheduler.abortStage(__DAGScheduler.scala:825)
> 
> at org.apache.spark.scheduler.__DAGScheduler.processEvent(__DAGScheduler.scala:440)
> 
> at org.apache.spark.scheduler.__DAGScheduler.org
> <http://org.apache.spark.scheduler.DAGScheduler.org>$apache$spark$__scheduler$DAGScheduler$$run(__DAGScheduler.scala:502)
> 
> at org.apache.spark.scheduler.__DAGScheduler$$anon$1.run(__DAGScheduler.scala:157)
> 
> FAILED: Execution Error, return code -101 from shark.execution.SparkTask
> 
> In fact I get the same error thrown when trying to READ from the table that I successfully WROTE to...
> 
> On Saturday, 22 February 2014 12:31:21 UTC+2, Costin Leau wrote:
> 
> Yeah, it might have been some sort of network configuration issue where services where running on different
> machines
> and
> localhost pointed to a different location.
> 
> Either way, I'm glad to hear things have are moving forward.
> 
> Cheers,
> 
> On 22/02/2014 1:06 AM, Max Lang wrote:
> > I managed to get it working on ec2 without issue this time. I'd say the biggest difference was that this
> time I set up a
> > dedicated ES machine. Is it possible that, because I was using a cluster with slaves, when I used
> "localhost" the slaves
> > couldn't find the ES instance running on the master? Or do all the requests go through the master?
> >
> >
> > On Wednesday, February 19, 2014 2:35:40 PM UTC-8, Costin Leau wrote:
> >
> > Hi,
> >
> > Setting logging in Hive/Hadoop can be tricky since the log4j needs to be picked up by the running JVM
> otherwise you
> > won't see anything.
> > Take a look at this link on how to tell Hive to use your logging settings [1].
> >
> > For the next release, we might introduce dedicated exceptions for the simple fact that some
> libraries, like Hive,
> > swallow the stack trace and it's unclear what the issue is which makes the exception
> (IllegalStateException) ambiguous.
> >
> > Let me know how it goes and whether you will encounter any issues with Shark. Or if you don't :)
> >
> > Thanks!
> >
> > [1]https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs
> <https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs>
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs
> <https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs>>
> > <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs
> <https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs>
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs
> <https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs>>>
> >
> > On 20/02/2014 12:02 AM, Max Lang wrote:
> > > Hey Costin,
> > >
> > > Thanks for the swift reply. I abandoned EC2 to take that out of the equation and managed to get
> everything working
> > > locally using the latest version of everything (though I realized just now I'm still on hive 0.9).
> I'm guessing you're
> > > right about some port connection issue because I definitely had ES running on that machine.
> > >
> > > I changed hive-log4j.properties and added
> > > |
> > > #custom logging levels
> > > #log4j.logger.xxx=DEBUG
> > > log4j.logger.org <http://log4j.logger.org>. __elasticsearch.hadoop.rest=__ TRACE
> > >log4j.logger.org.__elasticsearch.hadoop.mr <http://log4j.logger.org.elasticsearch.hadoop.mr>
> <http://log4j.logger.org.__elasticsearch.hadoop.mr <http://log4j.logger.org.elasticsearch.hadoop.mr>>
> <http://log4j.logger.org.__elasticsearch.hadoop.mr <http://log4j.logger.org.elasticsearch.hadoop.mr>
> <http://log4j.logger.org. __elasticsearch.hadoop.mr <http://log4j.logger.org.elasticsearch.hadoop.mr>>>=__ TRACE
> 
> > > |
> > >
> > > But I didn't see any trace logging. Hopefully I can get it working on EC2 without issue, but, for
> the future, is this
> > > the correct way to set TRACE logging?
> > >
> > > Oh and, for reference, I tried running without ES up and I got the following, exceptions:
> > >
> > > 2014-02-19 13:46:08,803 ERROR shark.SharkDriver (Logging.scala:logError(64)) - FAILED: Hive
> Internal Error:
> > > java.lang.__IllegalStateException(Cannot discover Elasticsearch version)
> > > java.lang.__IllegalStateException: Cannot discover Elasticsearch version
> > > at org.elasticsearch.hadoop.hive.__EsStorageHandler.init(__EsStorageHandler.java:101)
> > > at
> org.elasticsearch.hadoop.hive. __EsStorageHandler.__ configureOutputJobProperties(__EsStorageHandler.java:83)
> > > at
> org.apache.hadoop.hive.ql. __plan.PlanUtils.__ configureJobPropertiesForStora__geHandler(PlanUtils.java:706)
> > > at
> org.apache.hadoop.hive.ql. __plan.PlanUtils.__ configureOutputJobPropertiesFo__rStorageHandler(PlanUtils.__java:675)
> > > at org.apache.hadoop.hive.ql. __exec.FileSinkOperator.__ augmentPlan(FileSinkOperator.__java:764)
> > > at org.apache.hadoop.hive.ql. __parse.SemanticAnalyzer.__ putOpInsertMap(__SemanticAnalyzer.java:1518)
> > > at org.apache.hadoop.hive.ql. __parse.SemanticAnalyzer.__ genFileSinkPlan(__SemanticAnalyzer.java:4337)
> > > at
> org.apache.hadoop.hive.ql. __parse.SemanticAnalyzer.__ genPostGroupByBodyPlan(__SemanticAnalyzer.java:6207)
> > > at org.apache.hadoop.hive.ql. __parse.SemanticAnalyzer.__ genBodyPlan(SemanticAnalyzer.__java:6138)
> > > at org.apache.hadoop.hive.ql. __parse.SemanticAnalyzer.__ genPlan(SemanticAnalyzer.java:__6764)
> > > at shark.parse. __SharkSemanticAnalyzer.__ analyzeInternal( __SharkSemanticAnalyzer.scala:__ 149)
> > > at org.apache.hadoop.hive.ql. __parse.BaseSemanticAnalyzer.__ analyze(BaseSemanticAnalyzer.__java:244)
> > > at shark.SharkDriver.compile(__SharkDriver.scala:215)
> > > at org.apache.hadoop.hive.ql.__Driver.compile(Driver.java:__336)
> > > at org.apache.hadoop.hive.ql.__Driver.run(Driver.java:895)
> > > at shark.SharkCliDriver.__processCmd(SharkCliDriver.__scala:324)
> > > at org.apache.hadoop.hive.cli.__CliDriver.processLine(__CliDriver.java:406)
> > > at shark.SharkCliDriver$.main(__SharkCliDriver.scala:232)
> > > at shark.SharkCliDriver.main(__SharkCliDriver.scala)
> > > Caused by: java.io.IOException: Out of nodes and retries; caught exception
> > > at org.elasticsearch.hadoop.rest.__NetworkClient.execute(__NetworkClient.java:81)
> > > at org.elasticsearch.hadoop.rest.__RestClient.execute(RestClient.__java:221)
> > > at org.elasticsearch.hadoop.rest.__RestClient.execute(RestClient.__java:205)
> > > at org.elasticsearch.hadoop.rest.__RestClient.execute(RestClient.__java:209)
> > > at org.elasticsearch.hadoop.rest.__RestClient.get(RestClient.__java:103)
> > > at org.elasticsearch.hadoop.rest.__RestClient.esVersion(__RestClient.java:274)
> > > at
> org.elasticsearch.hadoop.rest. __InitializationUtils.__ discoverEsVersion(__InitializationUtils.java:84)
> > > at org.elasticsearch.hadoop.hive.__EsStorageHandler.init(__EsStorageHandler.java:99)
> > > ... 18 more
> > > Caused by: java.net.ConnectException: Connection refused
> > > at java.net.PlainSocketImpl.__socketConnect(Native Method)
> > > at java.net
> <http://java.net>. __AbstractPlainSocketImpl.__ doConnect( __AbstractPlainSocketImpl.java:__ 339)
> > > at java.net
> <http://java.net>. __AbstractPlainSocketImpl.__ connectToAddress( __AbstractPlainSocketImpl.java:__ 200)
> > > at java.net <http://java.net>. __AbstractPlainSocketImpl.__ connect( __AbstractPlainSocketImpl.java:__ 182)
> > > at java.net.SocksSocketImpl.__connect(SocksSocketImpl.java:__391)
> > > at java.net.Socket.connect(__Socket.java:579)
> > > at java.net.Socket.connect(__Socket.java:528)
> > > at java.net.Socket.<init>(Socket.__java:425)
> > > at java.net.Socket.<init>(Socket.__java:280)
> > > at
> org.apache.commons.httpclient. __protocol.__ DefaultProtocolSocketFactory.__createSocket(__DefaultProtocolSocketFactory.__java:80)
> > > at
> org.apache.commons.httpclient. __protocol.__ DefaultProtocolSocketFactory.__createSocket(__DefaultProtocolSocketFactory.__java:122)
> > > at org.apache.commons.httpclient.__HttpConnection.open(__HttpConnection.java:707)
> > > at org.apache.commons.httpclient. __HttpMethodDirector.__ executeWithRetry(__HttpMethodDirector.java:387)
> > > at org.apache.commons.httpclient. __HttpMethodDirector.__ executeMethod(__HttpMethodDirector.java:171)
> > > at org.apache.commons.httpclient.__HttpClient.executeMethod(__HttpClient.java:397)
> > > at org.apache.commons.httpclient.__HttpClient.executeMethod(__HttpClient.java:323)
> > > at
> org.elasticsearch.hadoop.rest. __commonshttp.__ CommonsHttpTransport.execute(__CommonsHttpTransport.java:160)
> > > at org.elasticsearch.hadoop.rest.__NetworkClient.execute(__NetworkClient.java:74)
> > > ... 25 more
> > >
> > > Let me know if there's anything in particular you'd like me to try on EC2.
> > >
> > > (For posterity, the versions I used were: hadoop 2.2.0, hive 0.9.0, shark 8.1, spark 8.1, es-hadoop
> 1.3.0.M2, java
> > > 1.7.0_15, scala 2.9.3, elasticsearch 1.0.0)
> > >
> > > Thanks again,
> > > Max
> > >
> > > On Tuesday, February 18, 2014 10:16:38 PM UTC-8, Costin Leau wrote:
> > >
> > > The error indicates a network error - namely es-hadoop cannot connect to Elasticsearch on the
> default (localhost:9200)
> > > HTTP port. Can you double check whether that's indeed the case (using curl or even telnet on
> that port) - maybe the
> > > firewall prevents any connections to be made...
> > > Also you could try using the latest Hive, 0.12 and a more recent Hadoop such as 1.1.2 or 1.2.1.
> > >
> > > Additionally, can you enable TRACE logging in your job on es-hadoop packages
> org.elasticsearch.hadoop.rest and
> > >org.elasticsearch.hadoop.mr <http://org.elasticsearch.hadoop.mr>
> <http://org.elasticsearch.__hadoop.mr <http://org.elasticsearch.hadoop.mr>>
> <http://org.elasticsearch.__hadoop.mr <http://org.elasticsearch.hadoop.mr>
> <http://org.elasticsearch.__hadoop.mr <http://org.elasticsearch.hadoop.mr>>>
> <http://org.elasticsearch. __hadoop.mr <http://org.elasticsearch.hadoop.mr> <http://org.elasticsearch.__ hadoop.mr
> <http://org.elasticsearch.hadoop.mr>>
> 
> > <http://org.elasticsearch.__hadoop.mr <http://org.elasticsearch.hadoop.mr>
> <http://org.elasticsearch.__hadoop.mr <http://org.elasticsearch.hadoop.mr>>>> packages and report back ?
> > >
> > > Thanks,
> > >
> > > On 19/02/2014 4:03 AM, Max Lang wrote:
> > > > I set everything up using this
> guide:https://github.com/ __amplab/shark/wiki/Running-__ Shark-on-EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>>
> > > <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>
> > <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>>> on an ec2 cluster. I've
> > > > copied the elasticsearch-hadoop jars into the hive lib directory and I have elasticsearch
> running on localhost:9200. I'm
> > > > running shark in a screen session with --service screenserver and connecting to it at the
> same time using shark -h
> > > > localhost.
> > > >
> > > > Unfortunately, when I attempt to write data into elasticsearch, it fails. Here's an example:
> > > >
> > > > |
> > > > [localhost:10000]shark>CREATE EXTERNAL TABLE wiki (id BIGINT,title STRING,last_modified
> STRING,xml STRING,text
> > > > STRING)ROW FORMAT DELIMITED FIELDS TERMINATED BY '\t'LOCATION
> 's3n://spark-data/wikipedia-__sample/';
> > > > Timetaken (including network latency):0.159seconds
> > > > 14/02/1901:23:33INFO CliDriver:Timetaken (including network latency):0.159seconds
> > > >
> > > > [localhost:10000]shark>SELECT title FROM wiki LIMIT 1;
> > > > Alpokalja
> > > > Timetaken (including network latency):2.23seconds
> > > > 14/02/1901:23:48INFO CliDriver:Timetaken (including network latency):2.23seconds
> > > >
> > > > [localhost:10000]shark>CREATE EXTERNAL TABLE es_wiki (id BIGINT,title STRING,last_modified
> STRING,xml STRING,text
> > > > STRING)STORED BY
> 'org.elasticsearch.hadoop. __hive.EsStorageHandler'__ TBLPROPERTIES('es.resource'='__wikipedia/article');
> > > > Timetaken (including network latency):0.061seconds
> > > > 14/02/1901:33:51INFO CliDriver:Timetaken (including network latency):0.061seconds
> > > >
> > > > [localhost:10000]shark>INSERT OVERWRITE TABLE es_wiki SELECTw.id
> <http://w.id>,w.title,w.last___modified,w.xml,w.text FROM wiki w;
> > > > [HiveError]:Queryreturned non-zero code:9,cause:FAILED:__ExecutionError,returncode
> -101fromshark.execution.__SparkTask
> > > > Timetaken (including network latency):3.575seconds
> > > > 14/02/1901:34:42INFO CliDriver:Timetaken (including network latency):3.575seconds
> > > > |
> > > >
> > > > *The stack trace looks like this:*
> > > >
> > > > org.apache.hadoop.hive.ql.__metadata.HiveException
> (org.apache.hadoop.hive.ql.__metadata.HiveException: java.io.IOException:
> > > > Out of nodes and retries; caught exception)
> > > >
> > > >
> org.apache.hadoop.hive.ql. __exec.FileSinkOperator.__ processOp(FileSinkOperator.__java:602)shark.execution.__FileSinkOperator$$anonfun$__processPartition$1.apply(__FileSinkOperator.scala:84) __shark.execution.__ FileSinkOperator$$anonfun$__processPartition$1.apply(__FileSinkOperator.scala:81) __scala.collection.Iterator$__ class.foreach(Iterator.scala:__772)scala.collection.Iterator$__$anon$19.foreach(Iterator.__scala:399)shark.execution.__FileSinkOperator.__processPartition(__FileSinkOperator.scala:81) __shark.execution.__ FileSinkOperator$.writeFiles$__1(FileSinkOperator.scala:207)__shark.execution. __FileSinkOperator$$anonfun$__ executeProcessFileSinkPartitio__n$1.apply(FileSinkOperator.__scala:211)shark.execution. __FileSinkOperator$$anonfun$__ executeProcessFileSinkPartitio__n$1.apply(FileSinkOperator.__scala:211)org.apache.spark.__scheduler.ResultTask.runTask(__ResultTask.scala:107)org. __apache.spark.scheduler.Task.__ run(Task.scala:53)org.apache. __spark.executor.Executor$__ Task
> 
> ```

Runner$$anonfun$run$1.\_\_apply$mcV$sp(Executor.scala:\_\_215)org.apac

> ```
> he.spa
> 
> rk.dep
> >
> > loy.Sp
> > >
> > >
> arkHadoopUtil.runAsUser(__SparkHadoopUtil.scala:50)org.__apache.spark.executor.__Executor$TaskRunner.run(__Executor.scala:182)java.util. __concurrent.ThreadPoolExecutor.__ runWorker(ThreadPoolExecutor.__java:1145)java.util.__concurrent.ThreadPoolExecutor$__Worker.run(ThreadPoolExecutor.__java:615)java.lang.Thread.run(__Thread.java:744
> 
> >
> > >
> > > > I should be using Hive 0.9.0, shark 0.8.1, elasticsearch 1.0.0, Hadoop 1.0.4, and java 1.7.0_51
> > > > Based on my cursory look at the hadoop and elasticsearch-hadoop sources, it looks like hive
> is just rethrowing an
> > > > IOException it's getting from Spark, and elasticsearch-hadoop is just hitting those exceptions.
> > > > I suppose my questions are: Does this look like an issue with my ES/elasticsearch-hadoop
> config? And has anyone gotten
> > > > elasticsearch working with Spark/Shark?
> > > > Any ideas/insights are appreciated.
> > > > Thanks,Max
> > > >
> > > > --
> > > > You received this message because you are subscribed to the Google Groups "elasticsearch" group.
> > > > To unsubscribe from this group and stop receiving emails from it, send an email to
> > > >elasticsearc...@googlegroups.__com <mailto:elasticsearc...@googlegroups.com> <javascript:>.
> > > > To view this discussion on the web visit
> > >
> >https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>
> >
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>>
> > >
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>
> >
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>>>.
> > > > For more options, visithttps://groups.google.__com/groups/opt_out
> <http://groups.google.com/groups/opt_out> <http://groups.google.com/__groups/opt_out
> <http://groups.google.com/groups/opt_out>> <http://groups.google.com/__groups/opt_out
> <http://groups.google.com/groups/opt_out>
> 
> <http://groups.google.com/__groups/opt_out <http://groups.google.com/groups/opt_out>>>
> <https://groups.google.com/__groups/opt_out <https://groups.google.com/groups/opt_out>
> <https://groups.google.com/__groups/opt_out <https://groups.google.com/groups/opt_out>>
> > <https://groups.google.com/__groups/opt_out <https://groups.google.com/groups/opt_out>
> <https://groups.google.com/__groups/opt_out <https://groups.google.com/groups/opt_out>>>>.
> > >
> > > --
> > > Costin
> > >
> > > --
> > > You received this message because you are subscribed to the Google Groups "elasticsearch" group.
> > > To unsubscribe from this group and stop receiving emails from it, send an email to
> > >elasticsearc...@googlegroups.__com <mailto:elasticsearc...@googlegroups.com> <javascript:>.
> > > To view this discussion on the web visit
> >
> >https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com>>
> >
> <https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com>>>.
> > > For more options, visithttps://groups.google.__com/groups/opt_out
> <http://groups.google.com/groups/opt_out> <http://groups.google.com/__groups/opt_out
> <http://groups.google.com/groups/opt_out>> <https://groups.google.com/__groups/opt_out
> <https://groups.google.com/groups/opt_out>
> <https://groups.google.com/__groups/opt_out <https://groups.google.com/groups/opt_out>>>.
> >
> > --
> > Costin
> >
> > --
> > You received this message because you are subscribed to the Google Groups "elasticsearch" group.
> > To unsubscribe from this group and stop receiving emails from it, send an email to
> >elasticsearc...@googlegroups.__com <mailto:elasticsearc...@googlegroups.com> <javascript:>.
> > To view this discussion on the web visit
> 
> >https://groups.google.com/d/ __msgid/elasticsearch/e29e342d-__ de74-4ed6-93d4-875fc728c5a5%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/e29e342d-de74-4ed6-93d4-875fc728c5a5%40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/e29e342d-__ de74-4ed6-93d4-875fc728c5a5%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/e29e342d-de74-4ed6-93d4-875fc728c5a5%40googlegroups.com>>.
> 
> > For more options, visithttps://groups.google.__com/groups/opt_out
> <http://groups.google.com/groups/opt_out> <https://groups.google.com/__groups/opt_out
> <https://groups.google.com/groups/opt_out>>.
> 
> --
> Costin
> 
> --
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.
> To unsubscribe from this group and stop receiving emails from it, send an email to
> elasticsearch+unsubscribe@__googlegroups.com <mailto:elasticsearch%2Bunsubscribe@googlegroups.com>
> <mailto:elasticsearch+__unsubscribe@googlegroups.com <mailto:elasticsearch%2Bunsubscribe@googlegroups.com>>.
> 
> To view this discussion on the web visit
> https://groups.google.com/d/ __msgid/elasticsearch/c1081bf2-__ 117a-4af2-ba90-2c38a4572782%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/c1081bf2-117a-4af2-ba90-2c38a4572782%40googlegroups.com>
> <https://groups.google.com/d/ __msgid/elasticsearch/c1081bf2-__ 117a-4af2-ba90-2c38a4572782% __40googlegroups.com?utm_medium=__ email&utm_source=footer
> <https://groups.google.com/d/msgid/elasticsearch/c1081bf2-117a-4af2-ba90-2c38a4572782%40googlegroups.com?utm_medium=email&utm_source=footer>>.
> 
> For more options, visit https://groups.google.com/d/__optout <https://groups.google.com/d/optout>.
> 
> --
> Costin
> 
> --
> You received this message because you are subscribed to a topic in the Google Groups "elasticsearch" group.
> To unsubscribe from this topic, visit https://groups.google.com/d/ __topic/elasticsearch/S-__ BrzwUHJbM/unsubscribe
> <https://groups.google.com/d/topic/elasticsearch/S-BrzwUHJbM/unsubscribe>.
> To unsubscribe from this group and all its topics, send an email to elasticsearch+unsubscribe@__googlegroups.com
> <mailto:elasticsearch%2Bunsubscribe@googlegroups.com>.
> To view this discussion on the web visit
> https://groups.google.com/d/ __msgid/elasticsearch/532AE2B5.__ 8080004%40gmail.com
> <https://groups.google.com/d/msgid/elasticsearch/532AE2B5.8080004%40gmail.com>.
> 
> For more options, visit https://groups.google.com/d/__optout <https://groups.google.com/d/optout>.
> 
> ```
> 
> --  
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an email to  
> [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com) [mailto:elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GNJD0wMJPzXwQqvfL4%2B0nZmw4XzFrPdEc%2BOPLZVeNuZpw%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GNJD0wMJPzXwQqvfL4%2B0nZmw4XzFrPdEc%2BOPLZVeNuZpw%40mail.gmail.com)  
> [https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GNJD0wMJPzXwQqvfL4%2B0nZmw4XzFrPdEc%2BOPLZVeNuZpw%40mail.gmail.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GNJD0wMJPzXwQqvfL4%2B0nZmw4XzFrPdEc%2BOPLZVeNuZpw%40mail.gmail.com?utm_medium=email&utm_source=footer).  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
Costin

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/53343AA6.1000405%40gmail.com](https://groups.google.com/d/msgid/elasticsearch/53343AA6.1000405%40gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Nick\_Pentreath](https://avatars.discourse-cdn.com/v4/letter/n/7ea924/32.png) [@Nick\_Pentreath](https://discuss.elastic.co/u/Nick_Pentreath)\
**Post date:** [May 13, 2014, 5:25pm UTC](https://discuss.elastic.co/t/hadoop-getting-elasticsearch-hadoop-working-with-shark/15882/11 "2014-05-13T17:25:17Z")

</div>

Hi Costin

Sorry for the silence on this issue. This went a bit quiet.

But the good news is I've come back to it and managed to get it all working  
with the new shark 0.9.1 release and 2.0.0RC1. Actually if I used ADD JAR I  
got the same exception but when I just put the JAR into the shark lib/  
folder it worked fine (which seems to point to the classpath issue you  
mention).

However, I seem to have an issue with date \<-\> timestamp conversion.

I have a field in ES called "\_ts" that has type "date" and the default  
format "dateOptionalTime". When I do a query that includes the timestamp it  
comes back NULL:

select ts from table ...  
(note I use a correct es.mapping.names to map the \_ts field in ES to ts  
field in Hive/Shark that has timestamp type).

below is some of the debug-level output:

14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the TIMESTAMP data  
type range so converted to null. Given data is :96997506-06-30  
19:08:168:16.768  
14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the TIMESTAMP data  
type range so converted to null. Given data is :96997605-06-28  
19:08:168:16.768  
14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the TIMESTAMP data  
type range so converted to null. Given data is :96997624-06-28  
19:08:168:16.768  
14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the TIMESTAMP data  
type range so converted to null. Given data is :96997629-06-28  
19:08:168:16.768  
14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the TIMESTAMP data  
type range so converted to null. Given data is :96997634-06-29  
19:08:168:16.768  
NULL  
NULL  
NULL  
NULL  
NULL

The data that I index in the \_ts field is timestamp in ms (long). It  
doesn't seem to be converted correctly but the data is correct (in ms at  
least) and I can query against it using date formats and date math in ES.

Example snippet from debug log from above:  
,"\_ts":1397130475607}}]}}"

Any ideas or am I doing something silly?

I do see that the Hive timestamp expects either seconds since epoch of a  
string-based format that has nanosecond granularity. Is this the issue with  
just ms long timestamp data?

Thanks  
Nick

On Thu, Mar 27, 2014 at 4:50 PM, Costin Leau [costin.leau@gmail.com](mailto:costin.leau@gmail.com) wrote:

> Using the latest hive and hadoop is preferred as they contain various bug  
> fixes.  
> The error suggests a classpath issue - namely the same class is loaded  
> twice for some reason and hence the casting fails.
> 
> Let's connect on IRC - give me a ping when you're available (user is  
> costin).
> 
> Cheers,
> 
> On 3/27/14 4:29 PM, Nick Pentreath wrote:
> 
> > Thanks for the response.
> > 
> > I tried latest Shark (cdh4 version of 0.9.1 here  
> > [http://cloudera.rst.im/shark/](http://cloudera.rst.im/shark/) ) - this uses hadoop 1.0.4 and hive 0.11  
> > I believe, and build elasticsearch-hadoop from github master.
> > 
> > Still getting same error:  
> > org.elasticsearch.hadoop.hive.EsHiveInputFormat$EsHiveSplit cannot be  
> > cast to  
> > org.elasticsearch.hadoop.hive.EsHiveInputFormat$EsHiveSplit
> > 
> > Will using hive 0.11 / hadoop 1.0.4 vs hive 0.12 / hadoop 1.2.1 in  
> > es-hadoop master make a difference?
> > 
> > Anyone else actually got this working?
> > 
> > On Thu, Mar 20, 2014 at 2:44 PM, Costin Leau \<[costin.leau@gmail.com](mailto:costin.leau@gmail.com)\<mailto:  
> > [costin.leau@gmail.com](mailto:costin.leau@gmail.com)\>\> wrote:
> > 
> > ```
> > I recommend using master - there are several improvements done in
> > 
> > ```
> > 
> > this area. Also using the latest Shark (0.9.0) and  
> > Hive (0.12) will help.
> > 
> > ```
> > On 3/20/14 12:00 PM, Nick Pentreath wrote:
> > 
> > Hi
> > 
> > I am struggling to get this working too. I'm just trying locally
> > 
> > ```
> > 
> > for now, running Shark 0.8.1, Hive 0.9.0 and ES  
> > 1.0.1  
> > with ES-hadoop 1.3.0.M2.
> > 
> > ```
> > I managed to get a basic example working with WRITING into an
> > 
> > ```
> > 
> > index. But I'm really after READING and index.
> > 
> > ```
> > I believe I have set everything up correctly, I've added the jar
> > 
> > ```
> > 
> > to Shark:  
> > ADD JAR /path/to/es-hadoop.jar;
> > 
> > ```
> > created a table:
> > CREATE EXTERNAL TABLE test_read (name string, price double)
> > 
> > STORED BY 'org.elasticsearch.hadoop.__hive.EsStorageHandler'
> > 
> > TBLPROPERTIES('es.resource' = 'test_index/test_type/_search?
> > 
> > ```
> > 
> > \_\_q=\*');
> > 
> > ```
> > And then trying to 'SELECT * FROM test _read' gives me :
> > 
> > org.apache.spark.__SparkException: Job aborted: Task 3.0:0
> > 
> > ```
> > 
> > failed more than 0 times; aborting job  
> > java.lang.ClassCastException: org.elasticsearch.hadoop.hive.  
> > \_\_EsHiveInputFormat$ESHiveSplit cannot be cast to  
> > org.elasticsearch.hadoop.hive.\_\_EsHiveInputFormat$ESHiveSplit
> > 
> > ```
> > at org.apache.spark.scheduler. __DAGScheduler$$anonfun$__
> > 
> > ```
> > 
> > abortStage$1.apply(\_\_DAGScheduler.scala:827)
> > 
> > ```
> > at org.apache.spark.scheduler. __DAGScheduler$$anonfun$__
> > 
> > ```
> > 
> > abortStage$1.apply(\_\_DAGScheduler.scala:825)
> > 
> > ```
> > at scala.collection.mutable.__ResizableArray$class.foreach(_
> > 
> > ```
> > 
> > \_ResizableArray.scala:60)
> > 
> > ```
> > at scala.collection.mutable.__ArrayBuffer.foreach(__
> > 
> > ```
> > 
> > ArrayBuffer.scala:47)
> > 
> > ```
> > at org.apache.spark.scheduler.__DAGScheduler.abortStage(__
> > 
> > ```
> > 
> > DAGScheduler.scala:825)
> > 
> > ```
> > at org.apache.spark.scheduler.__DAGScheduler.processEvent(__
> > 
> > ```
> > 
> > DAGScheduler.scala:440)
> > 
> > ```
> > at org.apache.spark.scheduler.__DAGScheduler.org
> > <http://org.apache.spark.scheduler.DAGScheduler.org>$
> > 
> > ```
> > 
> > apache$spark$\_\_scheduler$DAGScheduler$$run(\_\_DAGScheduler.scala:502)
> > 
> > ```
> > at org.apache.spark.scheduler.__DAGScheduler$$anon$1.run(__
> > 
> > ```
> > 
> > DAGScheduler.scala:157)
> > 
> > ```
> > FAILED: Execution Error, return code -101 from
> > 
> > ```
> > 
> > shark.execution.SparkTask
> > 
> > ```
> > In fact I get the same error thrown when trying to READ from the
> > 
> > ```
> > 
> > table that I successfully WROTE to...
> > 
> > ```
> > On Saturday, 22 February 2014 12:31:21 UTC+2, Costin Leau wrote:
> > 
> > Yeah, it might have been some sort of network configuration
> > 
> > ```
> > 
> > issue where services where running on different  
> > machines  
> > and  
> > localhost pointed to a different location.
> > 
> > ```
> > Either way, I'm glad to hear things have are moving forward.
> > 
> > Cheers,
> > 
> > On 22/02/2014 1:06 AM, Max Lang wrote:
> > > I managed to get it working on ec2 without issue this
> > 
> > ```
> > 
> > time. I'd say the biggest difference was that this  
> > time I set up a  
> > \> dedicated ES machine. Is it possible that, because I was  
> > using a cluster with slaves, when I used  
> > "localhost" the slaves  
> > \> couldn't find the ES instance running on the master? Or do  
> > all the requests go through the master?  
> > \>  
> > \>  
> > \> On Wednesday, February 19, 2014 2:35:40 PM UTC-8, Costin  
> > Leau wrote:  
> > \>  
> > \> Hi,  
> > \>  
> > \> Setting logging in Hive/Hadoop can be tricky since the  
> > log4j needs to be picked up by the running JVM  
> > otherwise you  
> > \> won't see anything.  
> > \> Take a look at this link on how to tell Hive to use  
> > your logging settings [1].  
> > \>  
> > \> For the next release, we might introduce dedicated  
> > exceptions for the simple fact that some  
> > libraries, like Hive,  
> > \> swallow the stack trace and it's unclear what the  
> > issue is which makes the exception  
> > (IllegalStateException) ambiguous.  
> > \>  
> > \> Let me know how it goes and whether you will encounter  
> > any issues with Shark. Or if you don't 🙂  
> > \>  
> > \> Thanks!  
> > \>  
> > \> [1][https://cwiki.apache.org/\_\_](https://cwiki.apache.org/__)  
> > confluence/display/Hive/\_\_GettingStarted#GettingStarted-\_\_ErrorLogs  
> > \<[Home - Apache Hive - Apache Software Foundation](https://cwiki.apache.org/confluence/display/Hive/)  
> > GettingStarted#GettingStarted-ErrorLogs\>  
> > \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-\_\_ErrorLogs  
> > \<[Home - Apache Hive - Apache Software Foundation](https://cwiki.apache.org/confluence/display/Hive/)  
> > GettingStarted#GettingStarted-ErrorLogs\>\>  
> > \> \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-\_\_ErrorLogs  
> > \<[Home - Apache Hive - Apache Software Foundation](https://cwiki.apache.org/confluence/display/Hive/)  
> > GettingStarted#GettingStarted-ErrorLogs\>  
> > \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-\_\_ErrorLogs
> > 
> > ```
> > <https://cwiki.apache.org/confluence/display/Hive/
> > 
> > ```
> > 
> > GettingStarted#GettingStarted-ErrorLogs\>\>\>  
> > \>  
> > \> On 20/02/2014 12:02 AM, Max Lang wrote:  
> > \> \> Hey Costin,  
> > \> \>  
> > \> \> Thanks for the swift reply. I abandoned EC2 to take  
> > that out of the equation and managed to get  
> > everything working  
> > \> \> locally using the latest version of everything  
> > (though I realized just now I'm still on hive 0.9).  
> > I'm guessing you're  
> > \> \> right about some port connection issue because I  
> > definitely had ES running on that machine.  
> > \> \>  
> > \> \> I changed hive-log4j.properties and added  
> > \> \> |  
> > \> \> #custom logging levels  
> > \> \> #log4j.logger.xxx=DEBUG  
> > \> \> [log4j.logger.org](http://log4j.logger.org) [http://log4j.logger.org](http://log4j.logger.org).\_\_  
> > elasticsearch.hadoop.rest=\_\_TRACE  
> > \> \>[log4j.logger.org](http://log4j.logger.org).\_\_elasticsearch.hadoop.mr \<  
> > [http://log4j.logger.org.elasticsearch.hadoop.mr](http://log4j.logger.org.elasticsearch.hadoop.mr)\>  
> > \<[http://log4j.logger.org](http://log4j.logger.org).\_\_elasticsearch.hadoop.mr \<  
> > [http://log4j.logger.org.elasticsearch.hadoop.mr](http://log4j.logger.org.elasticsearch.hadoop.mr)\>\>  
> > \<[http://log4j.logger.org](http://log4j.logger.org).\_\_elasticsearch.hadoop.mr \<  
> > [http://log4j.logger.org.elasticsearch.hadoop.mr](http://log4j.logger.org.elasticsearch.hadoop.mr)\>  
> > \<[http://log4j.logger.org](http://log4j.logger.org).\_\_elasticsearch.hadoop.mr \<  
> > [http://log4j.logger.org.elasticsearch.hadoop.mr](http://log4j.logger.org.elasticsearch.hadoop.mr)\>\>\>=\_\_TRACE
> > 
> > ```
> > > > |
> > > >
> > > > But I didn't see any trace logging. Hopefully I can
> > 
> > ```
> > 
> > get it working on EC2 without issue, but, for  
> > the future, is this  
> > \> \> the correct way to set TRACE logging?  
> > \> \>  
> > \> \> Oh and, for reference, I tried running without ES up  
> > and I got the following, exceptions:  
> > \> \>  
> > \> \> 2014-02-19 13:46:08,803 ERROR shark.SharkDriver  
> > (Logging.scala:logError(64)) - FAILED: Hive  
> > Internal Error:  
> > \> \> java.lang.\_\_IllegalStateException(Cannot discover  
> > Elasticsearch version)  
> > \> \> java.lang.\_\_IllegalStateException: Cannot discover  
> > Elasticsearch version  
> > \> \> at org.elasticsearch.hadoop.hive.  
> > \_\_EsStorageHandler.init(\_\_EsStorageHandler.java:101)  
> > \> \> at  
> > org.elasticsearch.hadoop.hive. **EsStorageHandler.**  
> > configureOutputJobProperties(\_\_EsStorageHandler.java:83)  
> > \> \> at  
> > org.apache.hadoop.hive.ql. **plan.PlanUtils.**  
> > configureJobPropertiesForStora\_\_geHandler(PlanUtils.java:706)  
> > \> \> at  
> > org.apache.hadoop.hive.ql. **plan.PlanUtils.**  
> > configureOutputJobPropertiesFo\_\_rStorageHandler(PlanUtils.**java:675)  
> > \> \> at org.apache.hadoop.hive.ql.**  
> > exec.FileSinkOperator.\_\_augmentPlan(FileSinkOperator.**java:764)  
> > \> \> at org.apache.hadoop.hive.ql.**  
> > parse.SemanticAnalyzer.\_\_putOpInsertMap(**SemanticAnalyzer.java:1518)  
> > \> \> at org.apache.hadoop.hive.ql.**  
> > parse.SemanticAnalyzer.\_\_genFileSinkPlan(\_\_SemanticAnalyzer.java:4337)  
> > \> \> at  
> > org.apache.hadoop.hive.ql. **parse.SemanticAnalyzer.**  
> > genPostGroupByBodyPlan(**SemanticAnalyzer.java:6207)  
> > \> \> at org.apache.hadoop.hive.ql.**  
> > parse.SemanticAnalyzer.\_\_genBodyPlan(SemanticAnalyzer.**java:6138)  
> > \> \> at org.apache.hadoop.hive.ql.**  
> > parse.SemanticAnalyzer.\_\_genPlan(SemanticAnalyzer.java:\_\_6764)  
> > \> \> at shark.parse. **SharkSemanticAnalyzer.**  
> > analyzeInternal(**SharkSemanticAnalyzer.scala:149)  
> > \> \> at org.apache.hadoop.hive.ql.  
> > parse.BaseSemanticAnalyzer.analyze(BaseSemanticAnalyzer.java:244)  
> > \> \> at shark.SharkDriver.compile(  
> > SharkDriver.scala:215)  
> > \> \> at org.apache.hadoop.hive.ql.  
> > Driver.compile(Driver.java:336)  
> > \> \> at org.apache.hadoop.hive.ql.  
> > Driver.run(Driver.java:895)  
> > \> \> at shark.SharkCliDriver.**  
> > processCmd(SharkCliDriver.**scala:324)  
> > \> \> at org.apache.hadoop.hive.cli.**  
> > CliDriver.processLine(**CliDriver.java:406)  
> > \> \> at shark.SharkCliDriver$.main(**  
> > SharkCliDriver.scala:232)  
> > \> \> at shark.SharkCliDriver.main(\_\_SharkCliDriver.scala)
> > 
> > ```
> > > > Caused by: java.io.IOException: Out of nodes and
> > 
> > ```
> > 
> > retries; caught exception  
> > \> \> at org.elasticsearch.hadoop.rest.  
> > \_\_NetworkClient.execute(\_\_NetworkClient.java:81)  
> > \> \> at org.elasticsearch.hadoop.rest.  
> > \_\_RestClient.execute(RestClient.\_\_java:221)  
> > \> \> at org.elasticsearch.hadoop.rest.  
> > \_\_RestClient.execute(RestClient.\_\_java:205)  
> > \> \> at org.elasticsearch.hadoop.rest.  
> > \_\_RestClient.execute(RestClient.\_\_java:209)  
> > \> \> at org.elasticsearch.hadoop.rest.  
> > \_\_RestClient.get(RestClient.\_\_java:103)  
> > \> \> at org.elasticsearch.hadoop.rest.  
> > \_\_RestClient.esVersion(\_\_RestClient.java:274)  
> > \> \> at  
> > org.elasticsearch.hadoop.rest. **InitializationUtils.**  
> > discoverEsVersion(\_\_InitializationUtils.java:84)  
> > \> \> at org.elasticsearch.hadoop.hive.  
> > \_\_EsStorageHandler.init(\_\_EsStorageHandler.java:99)
> > 
> > ```
> > > > ... 18 more
> > > > Caused by: java.net.ConnectException: Connection
> > 
> > ```
> > 
> > refused  
> > \> \> at java.net.PlainSocketImpl.\_\_socketConnect(Native  
> > Method)  
> > \> \> at [java.net](http://java.net)  
> > [http://java.net](http://java.net).\_\_AbstractPlainSocketImpl.**doConnect(**  
> > AbstractPlainSocketImpl.java:\_\_339)  
> > \> \> at [java.net](http://java.net)  
> > [http://java.net](http://java.net).\_\_AbstractPlainSocketImpl.**connectToAddress(**  
> > AbstractPlainSocketImpl.java:**200)  
> > \> \> at [java.net](http://java.net) [http://java.net](http://java.net).**  
> > AbstractPlainSocketImpl.\_\_connect(\_\_AbstractPlainSocketImpl.java:**182)  
> > \> \> at java.net.SocksSocketImpl.**  
> > connect(SocksSocketImpl.java:\_\_391)  
> > \> \> at java.net.Socket.connect(\_\_Socket.java:579)  
> > \> \> at java.net.Socket.connect(\_\_Socket.java:528)  
> > \> \> at java.net.Socket.(Socket.\_\_java:425)  
> > \> \> at java.net.Socket.(Socket.\_\_java:280)  
> > \> \> at  
> > org.apache.commons.httpclient. **protocol.**  
> > DefaultProtocolSocketFactory.**createSocket(**  
> > DefaultProtocolSocketFactory.\_\_java:80)  
> > \> \> at  
> > org.apache.commons.httpclient. **protocol.**  
> > DefaultProtocolSocketFactory.**createSocket(**  
> > DefaultProtocolSocketFactory.\_\_java:122)  
> > \> \> at org.apache.commons.httpclient.  
> > \_\_HttpConnection.open(\_\_HttpConnection.java:707)  
> > \> \> at org.apache.commons.httpclient.  
> > \_\_HttpMethodDirector.\_\_executeWithRetry(\_\_HttpMethodDirector.java:387)  
> > \> \> at org.apache.commons.httpclient.  
> > \_\_HttpMethodDirector.\_\_executeMethod(\_\_HttpMethodDirector.java:171)  
> > \> \> at org.apache.commons.httpclient.  
> > \_\_HttpClient.executeMethod(\_\_HttpClient.java:397)  
> > \> \> at org.apache.commons.httpclient.  
> > \_\_HttpClient.executeMethod(\_\_HttpClient.java:323)  
> > \> \> at  
> > org.elasticsearch.hadoop.rest. **commonshttp.**  
> > CommonsHttpTransport.execute(\_\_CommonsHttpTransport.java:160)  
> > \> \> at org.elasticsearch.hadoop.rest.  
> > \_\_NetworkClient.execute(\_\_NetworkClient.java:74)
> > 
> > ```
> > > > ... 25 more
> > > >
> > > > Let me know if there's anything in particular you'd
> > 
> > ```
> > 
> > like me to try on EC2.  
> > \> \>  
> > \> \> (For posterity, the versions I used were: hadoop  
> > 2.2.0, hive 0.9.0, shark 8.1, spark 8.1, es-hadoop  
> > 1.3.0.M2, java  
> > \> \> 1.7.0\_15, scala 2.9.3, elasticsearch 1.0.0)  
> > \> \>  
> > \> \> Thanks again,  
> > \> \> Max  
> > \> \>  
> > \> \> On Tuesday, February 18, 2014 10:16:38 PM UTC-8,  
> > Costin Leau wrote:  
> > \> \>  
> > \> \> The error indicates a network error - namely  
> > es-hadoop cannot connect to Elasticsearch on the  
> > default (localhost:9200)  
> > \> \> HTTP port. Can you double check whether that's  
> > indeed the case (using curl or even telnet on  
> > that port) - maybe the  
> > \> \> firewall prevents any connections to be made...  
> > \> \> Also you could try using the latest Hive, 0.12  
> > and a more recent Hadoop such as 1.1.2 or 1.2.1.  
> > \> \>  
> > \> \> Additionally, can you enable TRACE logging in  
> > your job on es-hadoop packages  
> > org.elasticsearch.hadoop.rest and  
> > \> \>org.elasticsearch.hadoop.mr \<  
> > [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)\>  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_hadoop.mr \<[http://org.elasticsearch](http://org.elasticsearch).  
> > hadoop.mr\>\>  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_hadoop.mr \<[http://org.elasticsearch](http://org.elasticsearch).  
> > hadoop.mr\>  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_hadoop.mr \<  
> > [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)\>\>\>  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_hadoop.mr \<[http://org.elasticsearch](http://org.elasticsearch).  
> > hadoop.mr\> \<[http://org.elasticsearch](http://org.elasticsearch).\_\_hadoop.mr  
> > [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)\>
> > 
> > ```
> > > <http://org.elasticsearch.__hadoop.mr <
> > 
> > ```
> > 
> > [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)\>  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_hadoop.mr \<[http://org.elasticsearch](http://org.elasticsearch).  
> > hadoop.mr\>\>\>\> packages and report back ?
> > 
> > ```
> > > >
> > > > Thanks,
> > > >
> > > > On 19/02/2014 4:03 AM, Max Lang wrote:
> > > > > I set everything up using this
> > guide:https://github.com/ __amplab/shark/wiki/Running-__
> > 
> > ```
> > 
> > Shark-on-EC2  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)  
> > \<[https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ )  
> > EC2  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>  
> > \<[https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ )  
> > EC2  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)  
> > \<[https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>\>  
> > \> \> \<[https://github.com/amplab/\_\_](https://github.com/amplab/__)  
> > shark/wiki/Running-Shark-on-\_\_EC2  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)  
> > \<[https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>  
> > \> \<[https://github.com/amplab/\_\_](https://github.com/amplab/__)  
> > shark/wiki/Running-Shark-on-\_\_EC2  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)  
> > \<[https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ )  
> > EC2
> > 
> > ```
> > <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>>>
> > 
> > ```
> > 
> > on an ec2 cluster. I've  
> > \> \> \> copied the elasticsearch-hadoop jars into the  
> > hive lib directory and I have elasticsearch  
> > running on localhost:9200. I'm  
> > \> \> \> running shark in a screen session with  
> > --service screenserver and connecting to it at the  
> > same time using shark -h  
> > \> \> \> localhost.  
> > \> \> \>  
> > \> \> \> Unfortunately, when I attempt to write data  
> > into elasticsearch, it fails. Here's an example:  
> > \> \> \>  
> > \> \> \> |  
> > \> \> \> [localhost:10000]shark\>CREATE EXTERNAL TABLE  
> > wiki (id BIGINT,title STRING,last\_modified  
> > STRING,xml STRING,text  
> > \> \> \> STRING)ROW FORMAT DELIMITED FIELDS TERMINATED  
> > BY '\t'LOCATION  
> > 's3n://spark-data/wikipedia-\_\_sample/';
> > 
> > ```
> > > > > Timetaken (including network
> > 
> > ```
> > 
> > latency):0.159seconds  
> > \> \> \> 14/02/1901:23:33INFO CliDriver:Timetaken  
> > (including network latency):0.159seconds  
> > \> \> \>  
> > \> \> \> [localhost:10000]shark\>SELECT title FROM wiki  
> > LIMIT 1;  
> > \> \> \> Alpokalja  
> > \> \> \> Timetaken (including network  
> > latency):2.23seconds  
> > \> \> \> 14/02/1901:23:48INFO CliDriver:Timetaken  
> > (including network latency):2.23seconds  
> > \> \> \>  
> > \> \> \> [localhost:10000]shark\>CREATE EXTERNAL TABLE  
> > es\_wiki (id BIGINT,title STRING,last\_modified  
> > STRING,xml STRING,text  
> > \> \> \> STRING)STORED BY  
> > 'org.elasticsearch.hadoop. **hive.EsStorageHandler'**  
> > TBLPROPERTIES('es.resource'='\_\_wikipedia/article');
> > 
> > ```
> > > > > Timetaken (including network
> > 
> > ```
> > 
> > latency):0.061seconds  
> > \> \> \> 14/02/1901:33:51INFO CliDriver:Timetaken  
> > (including network latency):0.061seconds  
> > \> \> \>  
> > \> \> \> [localhost:10000]shark\>INSERT OVERWRITE TABLE  
> > es\_wiki SELECTw.id  
> > [http://w.id](http://w.id),w.title,w.last\_\_\_modified,w.xml,w.text FROM wiki w;  
> > \> \> \> [HiveError]:Queryreturned non-zero  
> > code:9,cause:FAILED:\_\_ExecutionError,returncode  
> > -101fromshark.execution.\_\_SparkTask
> > 
> > ```
> > > > > Timetaken (including network
> > 
> > ```
> > 
> > latency):3.575seconds  
> > \> \> \> 14/02/1901:34:42INFO CliDriver:Timetaken  
> > (including network latency):3.575seconds  
> > \> \> \> |  
> > \> \> \>  
> > \> \> \> _The stack trace looks like this:_  
> > \> \> \>  
> > \> \> \> org.apache.hadoop.hive.ql.\_\_  
> > metadata.HiveException  
> > (org.apache.hadoop.hive.ql.\_\_metadata.HiveException:  
> > java.io.IOException:
> > 
> > ```
> > > > > Out of nodes and retries; caught exception)
> > > > >
> > > > >
> > org.apache.hadoop.hive.ql. __exec.FileSinkOperator.__
> > 
> > ```
> > 
> > processOp(FileSinkOperator.**java:602)shark.execution.**  
> > FileSinkOperator$$anonfun$**processPartition$1.apply(**  
> > FileSinkOperator.scala:84) **shark.execution.**  
> > FileSinkOperator$$anonfun$**processPartition$1.apply(**  
> > FileSinkOperator.scala:81) **scala.collection.Iterator$**  
> > class.foreach(Iterator.scala:**772)scala.collection.  
> > Iterator$**$anon$19.foreach(Iterator.\_\_scala:399)shark.  
> > execution.\_\_FileSinkOperator.**processPartition(**  
> > FileSinkOperator.scala:81) **shark.execution.**  
> > FileSinkOperator$.writeFiles$\_\_1(FileSinkOperator.scala:207)  
> > \__shark.execution. **FileSinkOperator$$anonfun$**  
> > executeProcessFileSinkPartitio\_\_n$1.apply(FileSinkOperator._  
> > _scala:211)shark.execution. **FileSinkOperator$$anonfun$**  
> > executeProcessFileSinkPartitio\_\_n$1.apply(FileSinkOperator._  
> > \_scala:211)org.apache.spark.\__scheduler.ResultTask.runTask(_  
> > \_ResultTask.scala:107)org. **apache.spark.scheduler.Task.**  
> > run(Task.scala:53)org.apache.\_\_spark.executor.Executor$\_\_Task
> 
> Runner$$anonfun$run$1.\_\_apply$mcV$sp(Executor.scala:\_\_215)org.apac
> 
> > ```
> > he.spa
> > 
> > rk.dep
> > >
> > > loy.Sp
> > > >
> > > >
> > arkHadoopUtil.runAsUser(__SparkHadoopUtil.scala:50)org._
> > 
> > ```
> > 
> > \_apache.spark.executor.**Executor$TaskRunner.run(**  
> > Executor.scala:182)java.util. **concurrent.ThreadPoolExecutor.**  
> > runWorker(ThreadPoolExecutor.**java:1145)java.util.**  
> > concurrent.ThreadPoolExecutor$\_\_Worker.run(ThreadPoolExecutor.\_\_java:615)  
> > java.lang.Thread.run(\_\_Thread.java:744
> > 
> > ```
> > >
> > > >
> > > > > I should be using Hive 0.9.0, shark 0.8.1,
> > 
> > ```
> > 
> > elasticsearch 1.0.0, Hadoop 1.0.4, and java 1.7.0\_51  
> > \> \> \> Based on my cursory look at the hadoop and  
> > elasticsearch-hadoop sources, it looks like hive  
> > is just rethrowing an  
> > \> \> \> IOException it's getting from Spark, and  
> > elasticsearch-hadoop is just hitting those exceptions.  
> > \> \> \> I suppose my questions are: Does this look  
> > like an issue with my ES/elasticsearch-hadoop  
> > config? And has anyone gotten  
> > \> \> \> elasticsearch working with Spark/Shark?  
> > \> \> \> Any ideas/insights are appreciated.  
> > \> \> \> Thanks,Max  
> > \> \> \>  
> > \> \> \> --  
> > \> \> \> You received this message because you are  
> > subscribed to the Google Groups "elasticsearch" group.  
> > \> \> \> To unsubscribe from this group and stop  
> > receiving emails from it, send an email to  
> > \> \> \>elasticsearc...@googlegroups.\_\_com \<mailto:  
> > [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com)\> \<javascript:\>.
> > 
> > ```
> > > > > To view this discussion on the web visit
> > > >
> > >https://groups.google.com/d/__msgid/elasticsearch/9486faff-
> > 
> > ```
> > 
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/9486faff-
> > 
> > ```
> > 
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)\>\>  
> > \>  
> > \<[https://groups.google.com/d/\_\_msgid/elasticsearch/9486faff-](https://groups.google.com/d/__msgid/elasticsearch/9486faff-)  
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/9486faff-
> > 
> > ```
> > 
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)\>\>\>  
> > \> \>  
> > \<[https://groups.google.com/d/\_\_msgid/elasticsearch/9486faff-](https://groups.google.com/d/__msgid/elasticsearch/9486faff-)  
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/9486faff-
> > 
> > ```
> > 
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)\>\>  
> > \>  
> > \<[https://groups.google.com/d/\_\_msgid/elasticsearch/9486faff-](https://groups.google.com/d/__msgid/elasticsearch/9486faff-)  
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/9486faff-
> > 
> > ```
> > 
> > \_\_3eaf-4344-8931-3121bbc5d9c7%**[40googlegroups.com](http://40googlegroups.com)  
> > \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)\>\>\>\>.  
> > \> \> \> For more options, visithttps://groups.google.**  
> > com/groups/opt\_out  
> > [http://groups.google.com/groups/opt\_out](http://groups.google.com/groups/opt_out) \<  
> > [http://groups.google.com/\_\_groups/opt\_out](http://groups.google.com/__groups/opt_out)  
> > [http://groups.google.com/groups/opt\_out](http://groups.google.com/groups/opt_out)\> \<  
> > [http://groups.google.com/\_\_groups/opt\_out](http://groups.google.com/__groups/opt_out)  
> > [http://groups.google.com/groups/opt\_out](http://groups.google.com/groups/opt_out)
> > 
> > ```
> > <http://groups.google.com/__groups/opt_out <
> > 
> > ```
> > 
> > [http://groups.google.com/groups/opt\_out](http://groups.google.com/groups/opt_out)\>\>\>  
> > \<[https://groups.google.com/\_\_groups/opt\_out](https://groups.google.com/__groups/opt_out) \<  
> > [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)\>  
> > \<[https://groups.google.com/\_\_groups/opt\_out](https://groups.google.com/__groups/opt_out) \<  
> > [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)\>\>  
> > \> \<[https://groups.google.com/\_\_groups/opt\_out](https://groups.google.com/__groups/opt_out) \<  
> > [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)\>  
> > \<[https://groups.google.com/\_\_groups/opt\_out](https://groups.google.com/__groups/opt_out) \<  
> > [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)\>\>\>\>.
> > 
> > ```
> > > >
> > > > --
> > > > Costin
> > > >
> > > > --
> > > > You received this message because you are subscribed
> > 
> > ```
> > 
> > to the Google Groups "elasticsearch" group.  
> > \> \> To unsubscribe from this group and stop receiving  
> > emails from it, send an email to  
> > \> \>elasticsearc...@googlegroups.\_\_com \<mailto:  
> > [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com)\> \<javascript:\>.
> > 
> > ```
> > > > To view this discussion on the web visit
> > >
> > >https://groups.google.com/d/__msgid/elasticsearch/86187c3a-
> > 
> > ```
> > 
> > \_\_0974-4d10-9689-e83da788c04a%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/86187c3a-](https://groups.google.com/d/msgid/elasticsearch/86187c3a-)  
> > 0974-4d10-9689-e83da788c04a%[40googlegroups.com](http://40googlegroups.com)\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/86187c3a-
> > 
> > ```
> > 
> > \_\_0974-4d10-9689-e83da788c04a%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/86187c3a-](https://groups.google.com/d/msgid/elasticsearch/86187c3a-)  
> > 0974-4d10-9689-e83da788c04a%[40googlegroups.com](http://40googlegroups.com)\>\>  
> > \>  
> > \<[https://groups.google.com/d/\_\_msgid/elasticsearch/86187c3a-](https://groups.google.com/d/__msgid/elasticsearch/86187c3a-)  
> > \_\_0974-4d10-9689-e83da788c04a%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/86187c3a-](https://groups.google.com/d/msgid/elasticsearch/86187c3a-)  
> > 0974-4d10-9689-e83da788c04a%[40googlegroups.com](http://40googlegroups.com)\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/86187c3a-
> > 
> > ```
> > 
> > \_\_0974-4d10-9689-e83da788c04a%**[40googlegroups.com](http://40googlegroups.com)  
> > \<[https://groups.google.com/d/msgid/elasticsearch/86187c3a-](https://groups.google.com/d/msgid/elasticsearch/86187c3a-)  
> > 0974-4d10-9689-e83da788c04a%[40googlegroups.com](http://40googlegroups.com)\>\>\>.  
> > \> \> For more options, visithttps://groups.google.**  
> > com/groups/opt\_out  
> > [http://groups.google.com/groups/opt\_out](http://groups.google.com/groups/opt_out) \<  
> > [http://groups.google.com/\_\_groups/opt\_out](http://groups.google.com/__groups/opt_out)  
> > [http://groups.google.com/groups/opt\_out](http://groups.google.com/groups/opt_out)\> \<  
> > [https://groups.google.com/\_\_groups/opt\_out](https://groups.google.com/__groups/opt_out)  
> > [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)  
> > \<[https://groups.google.com/\_\_groups/opt\_out](https://groups.google.com/__groups/opt_out) \<  
> > [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)\>\>\>.
> > 
> > ```
> > >
> > > --
> > > Costin
> > >
> > > --
> > > You received this message because you are subscribed to
> > 
> > ```
> > 
> > the Google Groups "elasticsearch" group.  
> > \> To unsubscribe from this group and stop receiving emails  
> > from it, send an email to  
> > \>elasticsearc...@googlegroups.\_\_com \<mailto:elasticsearc...@  
> > [googlegroups.com](http://googlegroups.com)\> \<javascript:\>.
> > 
> > ```
> > > To view this discussion on the web visit
> > 
> > >https://groups.google.com/d/__msgid/elasticsearch/e29e342d-
> > 
> > ```
> > 
> > \_\_de74-4ed6-93d4-875fc728c5a5%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/e29e342d-](https://groups.google.com/d/msgid/elasticsearch/e29e342d-)  
> > de74-4ed6-93d4-875fc728c5a5%[40googlegroups.com](http://40googlegroups.com)\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/e29e342d-
> > 
> > ```
> > 
> > \_\_de74-4ed6-93d4-875fc728c5a5%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/e29e342d-](https://groups.google.com/d/msgid/elasticsearch/e29e342d-)  
> > de74-4ed6-93d4-875fc728c5a5%[40googlegroups.com](http://40googlegroups.com)\>\>.
> > 
> > ```
> > > For more options, visithttps://groups.google.__
> > 
> > ```
> > 
> > com/groups/opt\_out  
> > [http://groups.google.com/groups/opt\_out](http://groups.google.com/groups/opt_out) \<  
> > [https://groups.google.com/\_\_groups/opt\_out](https://groups.google.com/__groups/opt_out)
> > 
> > ```
> > <https://groups.google.com/groups/opt_out>>.
> > 
> > --
> > Costin
> > 
> > --
> > You received this message because you are subscribed to the
> > 
> > ```
> > 
> > Google Groups "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it,  
> > send an email to  
> > elasticsearch+unsubscribe@\_\_googlegroups.com \<mailto:  
> > elasticsearch%2Bunsubscribe@googlegroups.com\>  
> > \<[mailto:elasticsearch+\_\_unsubscribe@googlegroups.com](mailto:elasticsearch+__unsubscribe@googlegroups.com) \<mailto:  
> > elasticsearch%2Bunsubscribe@googlegroups.com\>\>.
> > 
> > ```
> > To view this discussion on the web visit
> > https://groups.google.com/d/__msgid/elasticsearch/c1081bf2-_
> > 
> > ```
> > 
> > \_117a-4af2-ba90-2c38a4572782%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/c1081bf2-](https://groups.google.com/d/msgid/elasticsearch/c1081bf2-)  
> > 117a-4af2-ba90-2c38a4572782%[40googlegroups.com](http://40googlegroups.com)\>  
> > \<[https://groups.google.com/d/\_\_msgid/elasticsearch/c1081bf2-](https://groups.google.com/d/__msgid/elasticsearch/c1081bf2-)  
> > \_\_117a-4af2-ba90-2c38a4572782%\__[40googlegroups.com?utm](http://40googlegroups.com?utm)_  
> > medium=\_\_email&utm\_source=footer  
> > \<[https://groups.google.com/d/msgid/elasticsearch/c1081bf2-](https://groups.google.com/d/msgid/elasticsearch/c1081bf2-)  
> > 117a-4af2-ba90-2c38a4572782%[40googlegroups.com?utm\_medium=](http://40googlegroups.com?utm_medium=)  
> > email&utm\_source=footer\>\>.
> > 
> > ```
> > For more options, visit https://groups.google.com/d/__optout <
> > 
> > ```
> > 
> > [https://groups.google.com/d/optout](https://groups.google.com/d/optout)\>.
> > 
> > ```
> > --
> > Costin
> > 
> > --
> > 
> > You received this message because you are subscribed to a topic in
> > 
> > ```
> > 
> > the Google Groups "elasticsearch" group.  
> > To unsubscribe from this topic, visit [https://groups.google.com/d/\_\_](https://groups.google.com/d/__)  
> > topic/elasticsearch/S-\_\_BrzwUHJbM/unsubscribe  
> > \<[https://groups.google.com/d/topic/elasticsearch/S-](https://groups.google.com/d/topic/elasticsearch/S-)  
> > BrzwUHJbM/unsubscribe\>.  
> > To unsubscribe from this group and all its topics, send an email to  
> > elasticsearch+unsubscribe@\_\_googlegroups.com  
> > [mailto:elasticsearch%2Bunsubscribe@googlegroups.com](mailto:elasticsearch%2Bunsubscribe@googlegroups.com).
> > 
> > ```
> > To view this discussion on the web visit
> > https://groups.google.com/d/__msgid/elasticsearch/532AE2B5._
> > 
> > ```
> > 
> > \_8080004%[40gmail.com](http://40gmail.com)  
> > \<[https://groups.google.com/d/msgid/elasticsearch/532AE2B5](https://groups.google.com/d/msgid/elasticsearch/532AE2B5).  
> > 8080004%[40gmail.com](http://40gmail.com)\>.
> > 
> > ```
> > For more options, visit https://groups.google.com/d/__optout <
> > 
> > ```
> > 
> > [https://groups.google.com/d/optout](https://groups.google.com/d/optout)\>.
> > 
> > --  
> > You received this message because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails from it, send an  
> > email to  
> > [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com) \<mailto:elasticsearch+  
> > [unsubscribe@googlegroups.com](mailto:unsubscribe@googlegroups.com)\>.  
> > To view this discussion on the web visit  
> > [https://groups.google.com/d/msgid/elasticsearch/CALD%](https://groups.google.com/d/msgid/elasticsearch/CALD%25)  
> > 2B6GNJD0wMJPzXwQqvfL4%2B0nZmw4XzFrPdEc%2BOPLZVeNuZpw%[40mail.gmail.com](http://40mail.gmail.com)  
> > \<[https://groups.google.com/d/msgid/elasticsearch/CALD%](https://groups.google.com/d/msgid/elasticsearch/CALD%25)  
> > 2B6GNJD0wMJPzXwQqvfL4%2B0nZmw4XzFrPdEc%2BOPLZVeNuZpw%40mail.gmail.  
> > com?utm\_medium=email&utm\_source=footer\>.
> > 
> > For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).
> 
> --  
> Costin
> 
> --  
> You received this message because you are subscribed to a topic in the  
> Google Groups "elasticsearch" group.  
> To unsubscribe from this topic, visit [https://groups.google.com/d/](https://groups.google.com/d/)  
> topic/elasticsearch/S-BrzwUHJbM/unsubscribe.  
> To unsubscribe from this group and all its topics, send an email to  
> [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit [https://groups.google.com/d/](https://groups.google.com/d/)  
> msgid/elasticsearch/53343AA6.1000405%[40gmail.com](http://40gmail.com).
> 
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GMsgCB2Yqs2LLsbGinXSBOhB4ULVX1eaMm0vTvGpgLY7A%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GMsgCB2Yqs2LLsbGinXSBOhB4ULVX1eaMm0vTvGpgLY7A%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [May 13, 2014, 6:18pm UTC](https://discuss.elastic.co/t/hadoop-getting-elasticsearch-hadoop-working-with-shark/15882/12 "2014-05-13T18:18:59Z")

</div>

Hi Nick,

I'm glad to see you are making progress. This week I'm mainly on the road but maybe we can meet on the IRC next week, my  
invitation still stands 🙂  
Timestamp is relatively new type and doesn't handle timezones properly - it is backed by java.sq.Timestamp so it  
inherits a lot of its issues.  
For some reason the year in your date is rather off so it's worth checking the data read by es-hadoop before passing it  
to Hive (see [1]).  
I've had issues myself with it and it the moment the cluster is in a different timezone than the dataset itself things  
get buggy.  
Try using a UDF to do the conversion from the long to a timestamp - I've tried doing something similar in our conversion  
but since we don't know the timezones  
used, it's easy for things to get mixed.

Cheers,

[1] [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/hadoop/current/troubleshooting.html)

On 5/13/14 8:25 PM, Nick Pentreath wrote:

> Hi Costin
> 
> Sorry for the silence on this issue. This went a bit quiet.
> 
> But the good news is I've come back to it and managed to get it all working with the new shark 0.9.1 release and  
> 2.0.0RC1. Actually if I used ADD JAR I got the same exception but when I just put the JAR into the shark lib/ folder it  
> worked fine (which seems to point to the classpath issue you mention).
> 
> However, I seem to have an issue with date \<-\> timestamp conversion.
> 
> I have a field in ES called "\_ts" that has type "date" and the default format "dateOptionalTime". When I do a query that  
> includes the timestamp it comes back NULL:
> 
> select ts from table ...  
> (note I use a correct es.mapping.names to map the \_ts field in ES to ts field in Hive/Shark that has timestamp type).
> 
> below is some of the debug-level output:
> 
> 14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the TIMESTAMP data type range so converted to null. Given data  
> is :96997506-06-30 19:08:168:16.768  
> 14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the TIMESTAMP data type range so converted to null. Given data  
> is :96997605-06-28 19:08:168:16.768  
> 14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the TIMESTAMP data type range so converted to null. Given data  
> is :96997624-06-28 19:08:168:16.768  
> 14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the TIMESTAMP data type range so converted to null. Given data  
> is :96997629-06-28 19:08:168:16.768  
> 14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the TIMESTAMP data type range so converted to null. Given data  
> is :96997634-06-29 19:08:168:16.768  
> NULL  
> NULL  
> NULL  
> NULL  
> NULL
> 
> The data that I index in the \_ts field is timestamp in ms (long). It doesn't seem to be converted correctly but the data  
> is correct (in ms at least) and I can query against it using date formats and date math in ES.
> 
> Example snippet from debug log from above:  
> ,"\_ts":1397130475607}}]}}"
> 
> Any ideas or am I doing something silly?
> 
> I do see that the Hive timestamp expects either seconds since epoch of a string-based format that has nanosecond  
> granularity. Is this the issue with just ms long timestamp data?
> 
> Thanks  
> Nick
> 
> On Thu, Mar 27, 2014 at 4:50 PM, Costin Leau \<[costin.leau@gmail.com](mailto:costin.leau@gmail.com) [mailto:costin.leau@gmail.com](mailto:costin.leau@gmail.com)\> wrote:
> 
> ```
> Using the latest hive and hadoop is preferred as they contain various bug fixes.
> The error suggests a classpath issue - namely the same class is loaded twice for some reason and hence the casting
> fails.
> 
> Let's connect on IRC - give me a ping when you're available (user is costin).
> 
> Cheers,
> 
> On 3/27/14 4:29 PM, Nick Pentreath wrote:
> 
> Thanks for the response.
> 
> I tried latest Shark (cdh4 version of 0.9.1 here http://cloudera.rst.im/shark/ ) - this uses hadoop 1.0.4 and
> hive 0.11
> I believe, and build elasticsearch-hadoop from github master.
> 
> Still getting same error:
> org.elasticsearch.hadoop.hive.__EsHiveInputFormat$EsHiveSplit cannot be cast to
> org.elasticsearch.hadoop.hive.__EsHiveInputFormat$EsHiveSplit
> 
> Will using hive 0.11 / hadoop 1.0.4 vs hive 0.12 / hadoop 1.2.1 in es-hadoop master make a difference?
> 
> Anyone else actually got this working?
> 
> On Thu, Mar 20, 2014 at 2:44 PM, Costin Leau <costin.leau@gmail.com <mailto:costin.leau@gmail.com>
> <mailto:costin.leau@gmail.com <mailto:costin.leau@gmail.com>>__> wrote:
> 
> I recommend using master - there are several improvements done in this area. Also using the latest Shark
> (0.9.0) and
> Hive (0.12) will help.
> 
> On 3/20/14 12:00 PM, Nick Pentreath wrote:
> 
> Hi
> 
> I am struggling to get this working too. I'm just trying locally for now, running Shark 0.8.1, Hive
> 0.9.0 and ES
> 1.0.1
> with ES-hadoop 1.3.0.M2.
> 
> I managed to get a basic example working with WRITING into an index. But I'm really after READING and
> index.
> 
> I believe I have set everything up correctly, I've added the jar to Shark:
> ADD JAR /path/to/es-hadoop.jar;
> 
> created a table:
> CREATE EXTERNAL TABLE test_read (name string, price double)
> 
> STORED BY 'org.elasticsearch.hadoop. ____ hive.EsStorageHandler'
> 
> TBLPROPERTIES('es.resource' = 'test_index/test_type/_search? ____ q=*');
> 
> And then trying to 'SELECT * FROM test _read' gives me :
> 
> org.apache.spark. ____ SparkException: Job aborted: Task 3.0:0 failed more than 0 times; aborting job
> java.lang.ClassCastException: org.elasticsearch.hadoop.hive. ____EsHiveInputFormat$__ ESHiveSplit cannot
> be cast to
> org.elasticsearch.hadoop.hive. ____EsHiveInputFormat$__ ESHiveSplit
> 
> at org.apache.spark.scheduler. ____DAGScheduler$$anonfun$____ abortStage$1.apply( ____ DAGScheduler.scala:827)
> 
> at org.apache.spark.scheduler. ____DAGScheduler$$anonfun$____ abortStage$1.apply( ____ DAGScheduler.scala:825)
> 
> at scala.collection.mutable.____ResizableArray$class.foreach(____ResizableArray.scala:60)
> 
> at scala.collection.mutable.____ArrayBuffer.foreach(____ArrayBuffer.scala:47)
> 
> at org.apache.spark.scheduler.____DAGScheduler.abortStage(____DAGScheduler.scala:825)
> 
> at org.apache.spark.scheduler.____DAGScheduler.processEvent(____DAGScheduler.scala:440)
> 
> at org.apache.spark.scheduler. ____ DAGScheduler.org
> <http://org.apache.spark.__scheduler.DAGScheduler.org
> <http://org.apache.spark.scheduler.DAGScheduler.org>>$ __apache$spark$__ scheduler$__DAGScheduler$$run(____DAGScheduler.scala:502)
> 
> at org.apache.spark.scheduler.____DAGScheduler$$anon$1.run(____DAGScheduler.scala:157)
> 
> FAILED: Execution Error, return code -101 from shark.execution.SparkTask
> 
> In fact I get the same error thrown when trying to READ from the table that I successfully WROTE to...
> 
> On Saturday, 22 February 2014 12:31:21 UTC+2, Costin Leau wrote:
> 
> Yeah, it might have been some sort of network configuration issue where services where running on
> different
> machines
> and
> localhost pointed to a different location.
> 
> Either way, I'm glad to hear things have are moving forward.
> 
> Cheers,
> 
> On 22/02/2014 1:06 AM, Max Lang wrote:
> > I managed to get it working on ec2 without issue this time. I'd say the biggest difference was
> that this
> time I set up a
> > dedicated ES machine. Is it possible that, because I was using a cluster with slaves, when I used
> "localhost" the slaves
> > couldn't find the ES instance running on the master? Or do all the requests go through the master?
> >
> >
> > On Wednesday, February 19, 2014 2:35:40 PM UTC-8, Costin Leau wrote:
> >
> > Hi,
> >
> > Setting logging in Hive/Hadoop can be tricky since the log4j needs to be picked up by the
> running JVM
> otherwise you
> > won't see anything.
> > Take a look at this link on how to tell Hive to use your logging settings [1].
> >
> > For the next release, we might introduce dedicated exceptions for the simple fact that some
> libraries, like Hive,
> > swallow the stack trace and it's unclear what the issue is which makes the exception
> (IllegalStateException) ambiguous.
> >
> > Let me know how it goes and whether you will encounter any issues with Shark. Or if you don't :)
> >
> > Thanks!
> >
> >
> [1]https://cwiki.apache.org/ ____confluence/display/Hive/____ GettingStarted#GettingStarted- ____ ErrorLogs
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs>
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs
> <https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs>>
> 
> <https://cwiki.apache.org/ ____confluence/display/Hive/____ GettingStarted#GettingStarted- ____ ErrorLogs
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs>
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs
> <https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs>>>
> >
> <https://cwiki.apache.org/ ____confluence/display/Hive/____ GettingStarted#GettingStarted- ____ ErrorLogs
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs>
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs
> <https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs>>
> 
> <https://cwiki.apache.org/ ____confluence/display/Hive/____ GettingStarted#GettingStarted- ____ ErrorLogs
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs>
> 
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs
> <https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs>>>>
> >
> > On 20/02/2014 12:02 AM, Max Lang wrote:
> > > Hey Costin,
> > >
> > > Thanks for the swift reply. I abandoned EC2 to take that out of the equation and managed
> to get
> everything working
> > > locally using the latest version of everything (though I realized just now I'm still on
> hive 0.9).
> I'm guessing you're
> > > right about some port connection issue because I definitely had ES running on that machine.
> > >
> > > I changed hive-log4j.properties and added
> > > |
> > > #custom logging levels
> > > #log4j.logger.xxx=DEBUG
> > > log4j.logger.org <http://log4j.logger.org>
> <http://log4j.logger.org>. ____elasticsearch.hadoop.rest=____ TRACE
> > >log4j.logger.org. __elasticsea__ rch.hadoop.mr <http://elasticsearch.hadoop.mr>
> <http://log4j.logger.org.__elasticsearch.hadoop.mr <http://log4j.logger.org.elasticsearch.hadoop.mr>>
> <http://log4j.logger.org. __ela__ sticsearch.hadoop.mr <http://elasticsearch.hadoop.mr>
> <http://log4j.logger.org.__elasticsearch.hadoop.mr <http://log4j.logger.org.elasticsearch.hadoop.mr>>>
> <http://log4j.logger.org. __ela__ sticsearch.hadoop.mr <http://elasticsearch.hadoop.mr>
> <http://log4j.logger.org.__elasticsearch.hadoop.mr <http://log4j.logger.org.elasticsearch.hadoop.mr>>
> <http://log4j.logger.org. __ela__ sticsearch.hadoop.mr <http://elasticsearch.hadoop.mr>
> <http://log4j.logger.org. __elasticsearch.hadoop.mr <http://log4j.logger.org.elasticsearch.hadoop.mr>>>>=____ TRACE
> 
> > > |
> > >
> > > But I didn't see any trace logging. Hopefully I can get it working on EC2 without issue,
> but, for
> the future, is this
> > > the correct way to set TRACE logging?
> > >
> > > Oh and, for reference, I tried running without ES up and I got the following, exceptions:
> > >
> > > 2014-02-19 13:46:08,803 ERROR shark.SharkDriver (Logging.scala:logError(64)) - FAILED: Hive
> Internal Error:
> > > java.lang. ____ IllegalStateException(Cannot discover Elasticsearch version)
> > > java.lang. ____ IllegalStateException: Cannot discover Elasticsearch version
> > > at org.elasticsearch.hadoop.hive.____EsStorageHandler.init(____EsStorageHandler.java:101)
> > > at
> 
> org.elasticsearch.hadoop.hive. ____EsStorageHandler.____ configureOutputJobProperties( ____ EsStorageHandler.java:83)
> > > at
> 
> org.apache.hadoop.hive.ql. ____plan.PlanUtils.____ configureJobPropertiesForStora____geHandler(PlanUtils.java:__706)
> > > at
> 
> org.apache.hadoop.hive.ql. ____plan.PlanUtils.____ configureOutputJobPropertiesFo____rStorageHandler(PlanUtils.____java:675)
> > > at
> org.apache.hadoop.hive.ql. ____exec.FileSinkOperator.____ augmentPlan(FileSinkOperator. ____ java:764)
> > > at
> org.apache.hadoop.hive.ql. ____parse.SemanticAnalyzer.____ putOpInsertMap( ____ SemanticAnalyzer.java:1518)
> > > at
> org.apache.hadoop.hive.ql. ____parse.SemanticAnalyzer.____ genFileSinkPlan( ____ SemanticAnalyzer.java:4337)
> > > at
> 
> org.apache.hadoop.hive.ql. ____parse.SemanticAnalyzer.____ genPostGroupByBodyPlan( ____ SemanticAnalyzer.java:6207)
> > > at
> org.apache.hadoop.hive.ql. ____parse.SemanticAnalyzer.____ genBodyPlan(SemanticAnalyzer. ____ java:6138)
> > > at
> org.apache.hadoop.hive.ql. ____parse.SemanticAnalyzer.____ genPlan(SemanticAnalyzer.java: ____ 6764)
> > > at
> shark.parse. ____SharkSemanticAnalyzer.____ analyzeInternal( ____SharkSemanticAnalyzer.scala:____ 149)
> > > at
> org.apache.hadoop.hive.ql. ____parse.BaseSemanticAnalyzer.____ analyze(BaseSemanticAnalyzer. ____ java:244)
> > > at shark.SharkDriver.compile( ____ SharkDriver.scala:215)
> > > at org.apache.hadoop.hive.ql.____Driver.compile(Driver.java:____336)
> > > at org.apache.hadoop.hive.ql. ____ Driver.run(Driver.java:895)
> > > at shark.SharkCliDriver.____processCmd(SharkCliDriver.____scala:324)
> > > at org.apache.hadoop.hive.cli.____CliDriver.processLine(____CliDriver.java:406)
> > > at shark.SharkCliDriver$.main( ____ SharkCliDriver.scala:232)
> > > at shark.SharkCliDriver.main( ____ SharkCliDriver.scala)
> 
> > > Caused by: java.io.IOException: Out of nodes and retries; caught exception
> > > at org.elasticsearch.hadoop.rest.____NetworkClient.execute(____NetworkClient.java:81)
> > > at org.elasticsearch.hadoop.rest.____RestClient.execute(__RestClient.__java:221)
> > > at org.elasticsearch.hadoop.rest.____RestClient.execute(__RestClient.__java:205)
> > > at org.elasticsearch.hadoop.rest.____RestClient.execute(__RestClient.__java:209)
> > > at org.elasticsearch.hadoop.rest.____RestClient.get(RestClient.____java:103)
> > > at org.elasticsearch.hadoop.rest.____RestClient.esVersion(____RestClient.java:274)
> > > at
> 
> org.elasticsearch.hadoop.rest. ____InitializationUtils.____ discoverEsVersion( ____ InitializationUtils.java:84)
> > > at org.elasticsearch.hadoop.hive.____EsStorageHandler.init(____EsStorageHandler.java:99)
> 
> > > ... 18 more
> > > Caused by: java.net.ConnectException: Connection refused
> > > at java.net.PlainSocketImpl. ____ socketConnect(Native Method)
> > > at java.net <http://java.net>
> <http://java.net>. ____AbstractPlainSocketImpl.____ doConnect( ____AbstractPlainSocketImpl.java:____ 339)
> > > at java.net <http://java.net>
> 
> <http://java.net>. ____AbstractPlainSocketImpl.____ connectToAddress( ____AbstractPlainSocketImpl.java:____ 200)
> > > at java.net <http://java.net>
> <http://java.net>. ____AbstractPlainSocketImpl.____ connect( ____AbstractPlainSocketImpl.java:____ 182)
> > > at java.net.SocksSocketImpl.____connect(SocksSocketImpl.java:____391)
> > > at java.net.Socket.connect( ____ Socket.java:579)
> > > at java.net.Socket.connect( ____ Socket.java:528)
> > > at java.net.Socket.<init>(Socket. ____ java:425)
> > > at java.net.Socket.<init>(Socket. ____ java:280)
> > > at
> 
> org.apache.commons.httpclient. ____protocol.____ DefaultProtocolSocketFactory.____createSocket(____DefaultProtocolSocketFactory. ____ java:80)
> > > at
> 
> org.apache.commons.httpclient. ____protocol.____ DefaultProtocolSocketFactory.____createSocket(____DefaultProtocolSocketFactory. ____ java:122)
> > > at org.apache.commons.httpclient.____HttpConnection.open(____HttpConnection.java:707)
> > > at
> org.apache.commons.httpclient. ____HttpMethodDirector.____ executeWithRetry( ____ HttpMethodDirector.java:387)
> > > at
> org.apache.commons.httpclient. ____HttpMethodDirector.____ executeMethod( ____ HttpMethodDirector.java:171)
> > > at org.apache.commons.httpclient.____HttpClient.executeMethod(____HttpClient.java:397)
> > > at org.apache.commons.httpclient.____HttpClient.executeMethod(____HttpClient.java:323)
> > > at
> 
> org.elasticsearch.hadoop.rest. ____commonshttp.____ CommonsHttpTransport.execute( ____CommonsHttpTransport.java:__ 160)
> > > at org.elasticsearch.hadoop.rest.____NetworkClient.execute(____NetworkClient.java:74)
> 
> > > ... 25 more
> > >
> > > Let me know if there's anything in particular you'd like me to try on EC2.
> > >
> > > (For posterity, the versions I used were: hadoop 2.2.0, hive 0.9.0, shark 8.1, spark 8.1,
> es-hadoop
> 1.3.0.M2, java
> > > 1.7.0_15, scala 2.9.3, elasticsearch 1.0.0)
> > >
> > > Thanks again,
> > > Max
> > >
> > > On Tuesday, February 18, 2014 10:16:38 PM UTC-8, Costin Leau wrote:
> > >
> > > The error indicates a network error - namely es-hadoop cannot connect to Elasticsearch
> on the
> default (localhost:9200)
> > > HTTP port. Can you double check whether that's indeed the case (using curl or even
> telnet on
> that port) - maybe the
> > > firewall prevents any connections to be made...
> > > Also you could try using the latest Hive, 0.12 and a more recent Hadoop such as 1.1.2
> or 1.2.1.
> > >
> > > Additionally, can you enable TRACE logging in your job on es-hadoop packages
> org.elasticsearch.hadoop.rest and
> > >org.elasticsearch.hadoop.mr <http://org.elasticsearch.hadoop.mr>
> <http://org.elasticsearch.__hadoop.mr <http://org.elasticsearch.hadoop.mr>>
> <http://org.elasticsearch. __ha__ doop.mr <http://hadoop.mr> <http://org.elasticsearch.__hadoop.mr
> <http://org.elasticsearch.hadoop.mr>>>
> <http://org.elasticsearch. __ha__ doop.mr <http://hadoop.mr> <http://org.elasticsearch.__hadoop.mr
> <http://org.elasticsearch.hadoop.mr>>
> <http://org.elasticsearch. __ha__ doop.mr <http://hadoop.mr> <http://org.elasticsearch.__hadoop.mr
> <http://org.elasticsearch.hadoop.mr>>>>
> <http://org.elasticsearch. __ha__ doop.mr <http://hadoop.mr> <http://org.elasticsearch.__hadoop.mr
> <http://org.elasticsearch.hadoop.mr>> <http://org.elasticsearch. __ha__ doop.mr <http://hadoop.mr>
> <http://org.elasticsearch.__hadoop.mr <http://org.elasticsearch.hadoop.mr>>>
> 
> > <http://org.elasticsearch. __ha__ doop.mr <http://hadoop.mr>
> <http://org.elasticsearch.__hadoop.mr <http://org.elasticsearch.hadoop.mr>>
> <http://org.elasticsearch. __ha__ doop.mr <http://hadoop.mr> <http://org.elasticsearch.__hadoop.mr
> <http://org.elasticsearch.hadoop.mr>>>>> packages and report back ?
> 
> > >
> > > Thanks,
> > >
> > > On 19/02/2014 4:03 AM, Max Lang wrote:
> > > > I set everything up using this
> guide:https://github.com/ ____amplab/shark/wiki/Running-____ Shark-on-EC2
> <https://github.com/ __amplab/shark/wiki/Running-__ Shark-on-EC2>
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>>
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>>>
> > > <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>>
> > <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>
> 
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>>>> on an ec2 cluster. I've
> > > > copied the elasticsearch-hadoop jars into the hive lib directory and I have
> elasticsearch
> running on localhost:9200. I'm
> > > > running shark in a screen session with --service screenserver and connecting to it
> at the
> same time using shark -h
> > > > localhost.
> > > >
> > > > Unfortunately, when I attempt to write data into elasticsearch, it fails. Here's an
> example:
> > > >
> > > > |
> > > > [localhost:10000]shark>CREATE EXTERNAL TABLE wiki (id BIGINT,title STRING,last_modified
> STRING,xml STRING,text
> > > > STRING)ROW FORMAT DELIMITED FIELDS TERMINATED BY '\t'LOCATION
> 's3n://spark-data/wikipedia- ____ sample/';
> 
> > > > Timetaken (including network latency):0.159seconds
> > > > 14/02/1901:23:33INFO CliDriver:Timetaken (including network latency):0.159seconds
> > > >
> > > > [localhost:10000]shark>SELECT title FROM wiki LIMIT 1;
> > > > Alpokalja
> > > > Timetaken (including network latency):2.23seconds
> > > > 14/02/1901:23:48INFO CliDriver:Timetaken (including network latency):2.23seconds
> > > >
> > > > [localhost:10000]shark>CREATE EXTERNAL TABLE es_wiki (id BIGINT,title
> STRING,last_modified
> STRING,xml STRING,text
> > > > STRING)STORED BY
> 
> 'org.elasticsearch.hadoop. ____hive.EsStorageHandler'____ TBLPROPERTIES('es.resource'=' ____ wikipedia/article');
> 
> > > > Timetaken (including network latency):0.061seconds
> > > > 14/02/1901:33:51INFO CliDriver:Timetaken (including network latency):0.061seconds
> > > >
> > > > [localhost:10000]shark>INSERT OVERWRITE TABLE es_wiki SELECTw.id
> <http://w.id>,w.title,w.last _____ modified,w.xml,w.text FROM wiki w;
> > > > [HiveError]:Queryreturned non-zero code:9,cause:FAILED: ____ ExecutionError,returncode
> -101fromshark.execution. ____ SparkTask
> 
> > > > Timetaken (including network latency):3.575seconds
> > > > 14/02/1901:34:42INFO CliDriver:Timetaken (including network latency):3.575seconds
> > > > |
> > > >
> > > > *The stack trace looks like this:*
> > > >
> > > > org.apache.hadoop.hive.ql. ____ metadata.HiveException
> (org.apache.hadoop.hive.ql. ____ metadata.HiveException: java.io.IOException:
> 
> > > > Out of nodes and retries; caught exception)
> > > >
> > > >
> 
> org.apache.hadoop.hive.ql. ____exec.FileSinkOperator.____ processOp(FileSinkOperator.____java:602)shark.execution.____FileSinkOperator$$anonfun$____processPartition$1.apply(____FileSinkOperator.scala:84) ____shark.execution.____ FileSinkOperator$$anonfun$____processPartition$1.apply(____FileSinkOperator.scala:81) ____scala.collection.Iterator$____ class.foreach(Iterator.scala:____772)scala.collection.__Iterator$__$anon$19.foreach(__Iterator.__scala:399)shark.__execution. __FileSinkOperator.____ processPartition(____FileSinkOperator.scala:81)____shark.execution. ____FileSinkOperator$.writeFiles$____ 1(FileSinkOperator.scala:207) ____shark.execution.____ FileSinkOperator$$anonfun$ ____executeProcessFileSinkPartitio____ n$1.apply(FileSinkOperator.____scala:211)shark.execution.____FileSinkOperator$$anonfun$ ____executeProcessFileSinkPartitio____ n$1.apply(FileSinkOperator.____scala:211)org.apache.spark.____scheduler.ResultTask.runTask(____ResultTask.scala:107)org.____apache.spark.scheduler.Ta
> 
> ```

sk.\_\_\_\_run(Task.scala:53)org.apache.\_\_\_\_spark.executor.Executor$\_\_\_\_Task

> ```
> Runner$$anonfun$run$1. __apply$__ mcV$sp(Executor.scala:__215)__org.apac
> 
> he.spa
> 
> rk.dep
> >
> > loy.Sp
> > >
> > >
> 
> arkHadoopUtil.runAsUser(____SparkHadoopUtil.scala:50)org.____apache.spark.executor.____Executor$TaskRunner.run(____Executor.scala:182)java.util. ____concurrent.__ ThreadPoolExecutor.____runWorker(ThreadPoolExecutor.____java:1145)java.util. ____concurrent.ThreadPoolExecutor$____ Worker.run( __ThreadPoolExecutor.__ java:615)__java.lang.Thread.run(__Thread.__java:744
> 
> >
> > >
> > > > I should be using Hive 0.9.0, shark 0.8.1, elasticsearch 1.0.0, Hadoop 1.0.4, and
> java 1.7.0_51
> > > > Based on my cursory look at the hadoop and elasticsearch-hadoop sources, it looks
> like hive
> is just rethrowing an
> > > > IOException it's getting from Spark, and elasticsearch-hadoop is just hitting those
> exceptions.
> > > > I suppose my questions are: Does this look like an issue with my ES/elasticsearch-hadoop
> config? And has anyone gotten
> > > > elasticsearch working with Spark/Shark?
> > > > Any ideas/insights are appreciated.
> > > > Thanks,Max
> > > >
> > > > --
> > > > You received this message because you are subscribed to the Google Groups
> "elasticsearch" group.
> > > > To unsubscribe from this group and stop receiving emails from it, send an email to
> > > >elasticsearc...@googlegroups. ____com <mailto:elasticsearc...@__ googlegroups.com
> <mailto:elasticsearc...@googlegroups.com>> <javascript:>.
> 
> > > > To view this discussion on the web visit
> > >
> 
> >https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____40googlegroups.com <https://groups.google.com/d/__ msgid/elasticsearch/9486faff- __3eaf-4344-8931-3121bbc5d9c7%__ 40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>>
> >
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>>>
> > >
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>>
> >
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>>>>.
> > > > For more options, visithttps://groups.google. ____ com/groups/opt_out
> <http://groups.google.com/__groups/opt_out <http://groups.google.com/groups/opt_out>>
> <http://groups.google.com/ ____groups/opt_out <http://groups.google.com/__ groups/opt_out>
> <http://groups.google.com/__groups/opt_out <http://groups.google.com/groups/opt_out>>>
> <http://groups.google.com/ ____groups/opt_out <http://groups.google.com/__ groups/opt_out>
> <http://groups.google.com/__groups/opt_out <http://groups.google.com/groups/opt_out>>
> 
> <http://groups.google.com/ ____groups/opt_out <http://groups.google.com/__ groups/opt_out>
> <http://groups.google.com/__groups/opt_out <http://groups.google.com/groups/opt_out>>>>
> <https://groups.google.com/ ____groups/opt_out <https://groups.google.com/__ groups/opt_out>
> <https://groups.google.com/__groups/opt_out <https://groups.google.com/groups/opt_out>>
> <https://groups.google.com/ ____groups/opt_out <https://groups.google.com/__ groups/opt_out>
> <https://groups.google.com/__groups/opt_out <https://groups.google.com/groups/opt_out>>>
> > <https://groups.google.com/ ____groups/opt_out <https://groups.google.com/__ groups/opt_out>
> <https://groups.google.com/__groups/opt_out <https://groups.google.com/groups/opt_out>>
> <https://groups.google.com/ ____groups/opt_out <https://groups.google.com/__ groups/opt_out>
> <https://groups.google.com/__groups/opt_out <https://groups.google.com/groups/opt_out>>>>>.
> 
> > >
> > > --
> > > Costin
> > >
> > > --
> > > You received this message because you are subscribed to the Google Groups "elasticsearch"
> group.
> > > To unsubscribe from this group and stop receiving emails from it, send an email to
> > >elasticsearc...@googlegroups. ____com <mailto:elasticsearc...@__ googlegroups.com
> <mailto:elasticsearc...@googlegroups.com>> <javascript:>.
> 
> > > To view this discussion on the web visit
> >
> 
> >https://groups.google.com/d/ ____msgid/elasticsearch/86187c3a-____ 0974-4d10-9689-e83da788c04a% ____40googlegroups.com <https://groups.google.com/d/__ msgid/elasticsearch/86187c3a- __0974-4d10-9689-e83da788c04a%__ 40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/86187c3a-____ 0974-4d10-9689-e83da788c04a% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com>>>
> >
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/86187c3a-____ 0974-4d10-9689-e83da788c04a% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/86187c3a-____ 0974-4d10-9689-e83da788c04a% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com>>>>.
> > > For more options, visithttps://groups.google. ____ com/groups/opt_out
> <http://groups.google.com/__groups/opt_out <http://groups.google.com/groups/opt_out>>
> <http://groups.google.com/ ____groups/opt_out <http://groups.google.com/__ groups/opt_out>
> <http://groups.google.com/__groups/opt_out <http://groups.google.com/groups/opt_out>>>
> <https://groups.google.com/ ____groups/opt_out <https://groups.google.com/__ groups/opt_out>
> <https://groups.google.com/__groups/opt_out <https://groups.google.com/groups/opt_out>>
> <https://groups.google.com/ ____groups/opt_out <https://groups.google.com/__ groups/opt_out>
> <https://groups.google.com/__groups/opt_out <https://groups.google.com/groups/opt_out>>>>.
> 
> >
> > --
> > Costin
> >
> > --
> > You received this message because you are subscribed to the Google Groups "elasticsearch" group.
> > To unsubscribe from this group and stop receiving emails from it, send an email to
> >elasticsearc...@googlegroups. ____com <mailto:elasticsearc...@__ googlegroups.com
> <mailto:elasticsearc...@googlegroups.com>> <javascript:>.
> 
> > To view this discussion on the web visit
> 
> >https://groups.google.com/d/ ____msgid/elasticsearch/e29e342d-____ de74-4ed6-93d4-875fc728c5a5% ____40googlegroups.com <https://groups.google.com/d/__ msgid/elasticsearch/e29e342d- __de74-4ed6-93d4-875fc728c5a5%__ 40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/e29e342d-__ de74-4ed6-93d4-875fc728c5a5%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/e29e342d-de74-4ed6-93d4-875fc728c5a5%40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/e29e342d-____ de74-4ed6-93d4-875fc728c5a5% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/e29e342d-__ de74-4ed6-93d4-875fc728c5a5%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/e29e342d-__ de74-4ed6-93d4-875fc728c5a5%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/e29e342d-de74-4ed6-93d4-875fc728c5a5%40googlegroups.com>>>.
> 
> > For more options, visithttps://groups.google. ____ com/groups/opt_out
> <http://groups.google.com/__groups/opt_out <http://groups.google.com/groups/opt_out>>
> <https://groups.google.com/ ____groups/opt_out <https://groups.google.com/__ groups/opt_out>
> 
> <https://groups.google.com/__groups/opt_out <https://groups.google.com/groups/opt_out>>>.
> 
> --
> Costin
> 
> --
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.
> To unsubscribe from this group and stop receiving emails from it, send an email to
> elasticsearch+unsubscribe@ __go__ oglegroups.com <http://googlegroups.com>
> <mailto:elasticsearch% __2Bunsubscribe@googlegroups.com <mailto:elasticsearch%252Bunsubscribe@googlegroups.com>__ >
> <mailto:elasticsearch+ ____ unsubscribe@googlegroups.com
> <mailto:elasticsearch%2B __unsubscribe@googlegroups.com> <mailto:elasticsearch%__ 2Bunsubscribe@googlegroups.com
> <mailto:elasticsearch%252Bunsubscribe@googlegroups.com>__>>.
> 
> To view this discussion on the web visit
> https://groups.google.com/d/ ____msgid/elasticsearch/c1081bf2-____ 117a-4af2-ba90-2c38a4572782% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/c1081bf2-__ 117a-4af2-ba90-2c38a4572782%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/c1081bf2-__ 117a-4af2-ba90-2c38a4572782%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/c1081bf2-117a-4af2-ba90-2c38a4572782%40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/c1081bf2-____ 117a-4af2-ba90-2c38a4572782% ____40googlegroups.com?utm___ medium= __email&utm_source=__ footer
> <https://groups.google.com/d/ __msgid/elasticsearch/c1081bf2-__ 117a-4af2-ba90-2c38a4572782% __40googlegroups.com?utm_medium=__ email&utm_source=footer>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/c1081bf2-__ 117a-4af2-ba90-2c38a4572782% __40googlegroups.com?utm_medium=__ email&utm_source=footer
> <https://groups.google.com/d/msgid/elasticsearch/c1081bf2-117a-4af2-ba90-2c38a4572782%40googlegroups.com?utm_medium=email&utm_source=footer>>>.
> 
> For more options, visit https://groups.google.com/d/ ____optout <https://groups.google.com/d/__ optout>
> <https://groups.google.com/d/__optout <https://groups.google.com/d/optout>>.
> 
> --
> Costin
> 
> --
> 
> You received this message because you are subscribed to a topic in the Google Groups "elasticsearch" group.
> To unsubscribe from this topic, visit
> https://groups.google.com/d/ ____topic/elasticsearch/S-____ BrzwUHJbM/unsubscribe
> <https://groups.google.com/d/ __topic/elasticsearch/S-__ BrzwUHJbM/unsubscribe>
> <https://groups.google.com/d/ __topic/elasticsearch/S-__ BrzwUHJbM/unsubscribe
> <https://groups.google.com/d/topic/elasticsearch/S-BrzwUHJbM/unsubscribe>>.
> To unsubscribe from this group and all its topics, send an email to
> elasticsearch+unsubscribe@ __go__ oglegroups.com <http://googlegroups.com>
> <mailto:elasticsearch%__2Bunsubscribe@googlegroups.com
> <mailto:elasticsearch%252Bunsubscribe@googlegroups.com>__>.
> 
> To view this discussion on the web visit
> https://groups.google.com/d/ ____msgid/elasticsearch/532AE2B5.____ 8080004%40gmail.com
> <https://groups.google.com/d/ __msgid/elasticsearch/532AE2B5.__ 8080004%40gmail.com>
> <https://groups.google.com/d/ __msgid/elasticsearch/532AE2B5.__ 8080004%40gmail.com
> <https://groups.google.com/d/msgid/elasticsearch/532AE2B5.8080004%40gmail.com>>.
> 
> For more options, visit https://groups.google.com/d/ ____optout <https://groups.google.com/d/__ optout>
> <https://groups.google.com/d/__optout <https://groups.google.com/d/optout>>.
> 
> --
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.
> To unsubscribe from this group and stop receiving emails from it, send an email to
> elasticsearch+unsubscribe@__googlegroups.com <mailto:elasticsearch%2Bunsubscribe@googlegroups.com>
> <mailto:elasticsearch+__unsubscribe@googlegroups.com <mailto:elasticsearch%2Bunsubscribe@googlegroups.com>>.
> To view this discussion on the web visit
> https://groups.google.com/d/ __msgid/elasticsearch/CALD%__ 2B6GNJD0wMJPzXwQqvfL4% __2B0nZmw4XzFrPdEc%__ 2BOPLZVeNuZpw%40mail.gmail.com
> <https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GNJD0wMJPzXwQqvfL4%2B0nZmw4XzFrPdEc%2BOPLZVeNuZpw%40mail.gmail.com>
> <https://groups.google.com/d/ __msgid/elasticsearch/CALD%__ 2B6GNJD0wMJPzXwQqvfL4% __2B0nZmw4XzFrPdEc%__ 2BOPLZVeNuZpw%40mail.gmail. __com?utm_medium=email&utm___ source=footer
> <https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GNJD0wMJPzXwQqvfL4%2B0nZmw4XzFrPdEc%2BOPLZVeNuZpw%40mail.gmail.com?utm_medium=email&utm_source=footer>>.
> 
> For more options, visit https://groups.google.com/d/__optout <https://groups.google.com/d/optout>.
> 
> --
> Costin
> 
> --
> You received this message because you are subscribed to a topic in the Google Groups "elasticsearch" group.
> To unsubscribe from this topic, visit https://groups.google.com/d/ __topic/elasticsearch/S-__ BrzwUHJbM/unsubscribe
> <https://groups.google.com/d/topic/elasticsearch/S-BrzwUHJbM/unsubscribe>.
> To unsubscribe from this group and all its topics, send an email to elasticsearch+unsubscribe@__googlegroups.com
> <mailto:elasticsearch%2Bunsubscribe@googlegroups.com>.
> To view this discussion on the web visit
> https://groups.google.com/d/ __msgid/elasticsearch/53343AA6.__ 1000405%40gmail.com
> <https://groups.google.com/d/msgid/elasticsearch/53343AA6.1000405%40gmail.com>.
> 
> For more options, visit https://groups.google.com/d/__optout <https://groups.google.com/d/optout>.
> 
> ```
> 
> --  
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an email to  
> [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com) [mailto:elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GMsgCB2Yqs2LLsbGinXSBOhB4ULVX1eaMm0vTvGpgLY7A%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GMsgCB2Yqs2LLsbGinXSBOhB4ULVX1eaMm0vTvGpgLY7A%40mail.gmail.com)  
> [https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GMsgCB2Yqs2LLsbGinXSBOhB4ULVX1eaMm0vTvGpgLY7A%40mail.gmail.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GMsgCB2Yqs2LLsbGinXSBOhB4ULVX1eaMm0vTvGpgLY7A%40mail.gmail.com?utm_medium=email&utm_source=footer).  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
Costin

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/53726213.8040606%40gmail.com](https://groups.google.com/d/msgid/elasticsearch/53726213.8040606%40gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Nick\_Pentreath](https://avatars.discourse-cdn.com/v4/letter/n/7ea924/32.png) [@Nick\_Pentreath](https://discuss.elastic.co/u/Nick_Pentreath)\
**Post date:** [May 13, 2014, 7:14pm UTC](https://discuss.elastic.co/t/hadoop-getting-elasticsearch-hadoop-working-with-shark/15882/13 "2014-05-13T19:14:48Z")

</div>

Ok - well let me know when you're around.

The mapreduce inputformat works fine. I'm using it with Spark to access the  
ES data via ESInputFormat and run analytics and machine learning jobs on  
that data, and the same \_ts field works and is the correct data (though it  
comes through as org.apache.hadoop.io.Text, which I convert to Long or a  
DateTime as required).

Perhaps I'm missing it somewhere but is it possible to force a field to be  
a type? i.e. similar the es.field.mapping could I tell it that it must  
parse the field as a string (since then I can take it and do whatever  
parsing / casting I want).

I could just use the new Spark SQL module (which I'm seriously considering  
right now having explored it a bit in the last few days), but some of the  
stuff we do requires a SQL Console and JDBC, so having Shark able to just  
pull in ES data is definitely very useful...

On Tue, May 13, 2014 at 8:18 PM, Costin Leau [costin.leau@gmail.com](mailto:costin.leau@gmail.com) wrote:

> Hi Nick,
> 
> I'm glad to see you are making progress. This week I'm mainly on the road  
> but maybe we can meet on the IRC next week, my invitation still stands 🙂  
> Timestamp is relatively new type and doesn't handle timezones properly -  
> it is backed by java.sq.Timestamp so it inherits a lot of its issues.  
> For some reason the year in your date is rather off so it's worth checking  
> the data read by es-hadoop before passing it to Hive (see [1]).  
> I've had issues myself with it and it the moment the cluster is in a  
> different timezone than the dataset itself things get buggy.  
> Try using a UDF to do the conversion from the long to a timestamp - I've  
> tried doing something similar in our conversion but since we don't know the  
> timezones  
> used, it's easy for things to get mixed.
> 
> Cheers,
> 
> [1] [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/guide/en/elasticsearch/hadoop/)  
> current/troubleshooting.html
> 
> On 5/13/14 8:25 PM, Nick Pentreath wrote:
> 
> > Hi Costin
> > 
> > Sorry for the silence on this issue. This went a bit quiet.
> > 
> > But the good news is I've come back to it and managed to get it all  
> > working with the new shark 0.9.1 release and  
> > 2.0.0RC1. Actually if I used ADD JAR I got the same exception but when I  
> > just put the JAR into the shark lib/ folder it  
> > worked fine (which seems to point to the classpath issue you mention).
> > 
> > However, I seem to have an issue with date \<-\> timestamp conversion.
> > 
> > I have a field in ES called "\_ts" that has type "date" and the default  
> > format "dateOptionalTime". When I do a query that  
> > includes the timestamp it comes back NULL:
> > 
> > select ts from table ...  
> > (note I use a correct es.mapping.names to map the \_ts field in ES to ts  
> > field in Hive/Shark that has timestamp type).
> > 
> > below is some of the debug-level output:
> > 
> > 14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the TIMESTAMP  
> > data type range so converted to null. Given data  
> > is :96997506-06-30 19:08:168:16.768  
> > 14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the TIMESTAMP  
> > data type range so converted to null. Given data  
> > is :96997605-06-28 19:08:168:16.768  
> > 14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the TIMESTAMP  
> > data type range so converted to null. Given data  
> > is :96997624-06-28 19:08:168:16.768  
> > 14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the TIMESTAMP  
> > data type range so converted to null. Given data  
> > is :96997629-06-28 19:08:168:16.768  
> > 14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the TIMESTAMP  
> > data type range so converted to null. Given data  
> > is :96997634-06-29 19:08:168:16.768  
> > NULL  
> > NULL  
> > NULL  
> > NULL  
> > NULL
> > 
> > The data that I index in the \_ts field is timestamp in ms (long). It  
> > doesn't seem to be converted correctly but the data  
> > is correct (in ms at least) and I can query against it using date formats  
> > and date math in ES.
> > 
> > Example snippet from debug log from above:  
> > ,"\_ts":1397130475607}}]}}"
> > 
> > Any ideas or am I doing something silly?
> > 
> > I do see that the Hive timestamp expects either seconds since epoch of a  
> > string-based format that has nanosecond  
> > granularity. Is this the issue with just ms long timestamp data?
> > 
> > Thanks  
> > Nick
> > 
> > On Thu, Mar 27, 2014 at 4:50 PM, Costin Leau \<[costin.leau@gmail.com](mailto:costin.leau@gmail.com)\<mailto:  
> > [costin.leau@gmail.com](mailto:costin.leau@gmail.com)\>\> wrote:
> > 
> > ```
> > Using the latest hive and hadoop is preferred as they contain various
> > 
> > ```
> > 
> > bug fixes.  
> > The error suggests a classpath issue - namely the same class is  
> > loaded twice for some reason and hence the casting  
> > fails.
> > 
> > ```
> > Let's connect on IRC - give me a ping when you're available (user is
> > 
> > ```
> > 
> > costin).
> > 
> > ```
> > Cheers,
> > 
> > On 3/27/14 4:29 PM, Nick Pentreath wrote:
> > 
> > Thanks for the response.
> > 
> > I tried latest Shark (cdh4 version of 0.9.1 here
> > 
> > ```
> > 
> > [http://cloudera.rst.im/shark/](http://cloudera.rst.im/shark/) ) - this uses hadoop 1.0.4 and  
> > hive 0.11  
> > I believe, and build elasticsearch-hadoop from github master.
> > 
> > ```
> > Still getting same error:
> > org.elasticsearch.hadoop.hive.__EsHiveInputFormat$EsHiveSplit
> > 
> > ```
> > 
> > cannot be cast to  
> > org.elasticsearch.hadoop.hive.\_\_EsHiveInputFormat$EsHiveSplit
> > 
> > ```
> > Will using hive 0.11 / hadoop 1.0.4 vs hive 0.12 / hadoop 1.2.1
> > 
> > ```
> > 
> > in es-hadoop master make a difference?
> > 
> > ```
> > Anyone else actually got this working?
> > 
> > On Thu, Mar 20, 2014 at 2:44 PM, Costin Leau <
> > 
> > ```
> > 
> > [costin.leau@gmail.com](mailto:costin.leau@gmail.com) [mailto:costin.leau@gmail.com](mailto:costin.leau@gmail.com)  
> > \<[mailto:costin.leau@gmail.com](mailto:costin.leau@gmail.com) [mailto:costin.leau@gmail.com](mailto:costin.leau@gmail.com)\>\_\_\>  
> > wrote:
> > 
> > ```
> > I recommend using master - there are several improvements
> > 
> > ```
> > 
> > done in this area. Also using the latest Shark  
> > (0.9.0) and  
> > Hive (0.12) will help.
> > 
> > ```
> > On 3/20/14 12:00 PM, Nick Pentreath wrote:
> > 
> > Hi
> > 
> > I am struggling to get this working too. I'm just trying
> > 
> > ```
> > 
> > locally for now, running Shark 0.8.1, Hive  
> > 0.9.0 and ES  
> > 1.0.1  
> > with ES-hadoop 1.3.0.M2.
> > 
> > ```
> > I managed to get a basic example working with WRITING
> > 
> > ```
> > 
> > into an index. But I'm really after READING and  
> > index.
> > 
> > ```
> > I believe I have set everything up correctly, I've added
> > 
> > ```
> > 
> > the jar to Shark:  
> > ADD JAR /path/to/es-hadoop.jar;
> > 
> > ```
> > created a table:
> > CREATE EXTERNAL TABLE test_read (name string, price
> > 
> > ```
> > 
> > double)
> > 
> > ```
> > STORED BY 'org.elasticsearch.hadoop. ____
> > 
> > ```
> > 
> > hive.EsStorageHandler'
> > 
> > ```
> > TBLPROPERTIES('es.resource' =
> > 
> > ```
> > 
> > 'test\_index/test\_type/\_search?\_\_\_\_q=\*');
> > 
> > ```
> > And then trying to 'SELECT * FROM test _read' gives me :
> > 
> > org.apache.spark. ____ SparkException: Job aborted: Task
> > 
> > ```
> > 
> > 3.0:0 failed more than 0 times; aborting job  
> > java.lang.ClassCastException:  
> > org.elasticsearch.hadoop.hive.\_\_\_\_EsHiveInputFormat$\_\_ESHiveSplit cannot  
> > be cast to  
> > org.elasticsearch.hadoop.hive.\_\_ **EsHiveInputFormat$**  
> > ESHiveSplit
> > 
> > ```
> > at org.apache.spark.scheduler.___
> > 
> > ```
> > 
> > \_DAGScheduler$$anonfun$\_\_\_\_abortStage$1.apply(\_\_\_\_DAGScheduler.scala:827)
> > 
> > ```
> > at org.apache.spark.scheduler.___
> > 
> > ```
> > 
> > \_DAGScheduler$$anonfun$\_\_\_\_abortStage$1.apply(\_\_\_\_DAGScheduler.scala:825)
> > 
> > ```
> > at scala.collection.mutable. ____
> > 
> > ```
> > 
> > ResizableArray$class.foreach(\_\_\_\_ResizableArray.scala:60)
> > 
> > ```
> > at scala.collection.mutable.____ArrayBuffer.foreach(____
> > 
> > ```
> > 
> > ArrayBuffer.scala:47)
> > 
> > ```
> > at org.apache.spark.scheduler.___
> > 
> > ```
> > 
> > \_DAGScheduler.abortStage(\_\_\_\_DAGScheduler.scala:825)
> > 
> > ```
> > at org.apache.spark.scheduler.___
> > 
> > ```
> > 
> > \_DAGScheduler.processEvent(\_\_\_\_DAGScheduler.scala:440)
> > 
> > ```
> > at org.apache.spark.scheduler. ____ DAGScheduler.org
> > <http://org.apache.spark.__scheduler.DAGScheduler.org
> > <http://org.apache.spark.scheduler.DAGScheduler.org>>$_
> > 
> > ```
> > 
> > \_apache$spark$\_\_scheduler$\_\_DAGScheduler$$run(\_\_\_\_DAGScheduler.scala:502)
> > 
> > ```
> > at org.apache.spark.scheduler.___
> > 
> > ```
> > 
> > \_DAGScheduler$$anon$1.run(\_\_\_\_DAGScheduler.scala:157)
> > 
> > ```
> > FAILED: Execution Error, return code -101 from
> > 
> > ```
> > 
> > shark.execution.SparkTask
> > 
> > ```
> > In fact I get the same error thrown when trying to READ
> > 
> > ```
> > 
> > from the table that I successfully WROTE to...
> > 
> > ```
> > On Saturday, 22 February 2014 12:31:21 UTC+2, Costin
> > 
> > ```
> > 
> > Leau wrote:
> > 
> > ```
> > Yeah, it might have been some sort of network
> > 
> > ```
> > 
> > configuration issue where services where running on  
> > different  
> > machines  
> > and  
> > localhost pointed to a different location.
> > 
> > ```
> > Either way, I'm glad to hear things have are moving
> > 
> > ```
> > 
> > forward.
> > 
> > ```
> > Cheers,
> > 
> > On 22/02/2014 1:06 AM, Max Lang wrote:
> > > I managed to get it working on ec2 without issue
> > 
> > ```
> > 
> > this time. I'd say the biggest difference was  
> > that this  
> > time I set up a  
> > \> dedicated ES machine. Is it possible that,  
> > because I was using a cluster with slaves, when I used  
> > "localhost" the slaves  
> > \> couldn't find the ES instance running on the  
> > master? Or do all the requests go through the master?  
> > \>  
> > \>  
> > \> On Wednesday, February 19, 2014 2:35:40 PM UTC-8,  
> > Costin Leau wrote:  
> > \>  
> > \> Hi,  
> > \>  
> > \> Setting logging in Hive/Hadoop can be tricky  
> > since the log4j needs to be picked up by the  
> > running JVM  
> > otherwise you  
> > \> won't see anything.  
> > \> Take a look at this link on how to tell Hive  
> > to use your logging settings [1].  
> > \>  
> > \> For the next release, we might introduce  
> > dedicated exceptions for the simple fact that some  
> > libraries, like Hive,  
> > \> swallow the stack trace and it's unclear what  
> > the issue is which makes the exception  
> > (IllegalStateException) ambiguous.  
> > \>  
> > \> Let me know how it goes and whether you will  
> > encounter any issues with Shark. Or if you don't 🙂  
> > \>  
> > \> Thanks!  
> > \>  
> > \>  
> > [1][https://cwiki.apache.org/\_\_\_\_confluence/display/Hive/\_\_\_\_](https://cwiki.apache.org/ ____confluence/display/Hive/____ )  
> > GettingStarted#GettingStarted-\_\_\_\_ErrorLogs  
> > \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-\_\_ErrorLogs\>  
> > \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-\_\_ErrorLogs  
> > \<[Home - Apache Hive - Apache Software Foundation](https://cwiki.apache.org/confluence/display/Hive/)  
> > GettingStarted#GettingStarted-ErrorLogs\>\>
> > 
> > ```
> > <https://cwiki.apache.org/ ____confluence/display/Hive/____
> > 
> > ```
> > 
> > GettingStarted#GettingStarted-\_\_\_\_ErrorLogs  
> > \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-\_\_ErrorLogs\>  
> > \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-\_\_ErrorLogs  
> > \<[Home - Apache Hive - Apache Software Foundation](https://cwiki.apache.org/confluence/display/Hive/)  
> > GettingStarted#GettingStarted-ErrorLogs\>\>\>  
> > \>  
> > \<[https://cwiki.apache.org/\_\_\_\_confluence/display/Hive/\_\_\_\_](https://cwiki.apache.org/ ____confluence/display/Hive/____ )  
> > GettingStarted#GettingStarted-\_\_\_\_ErrorLogs  
> > \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-\_\_ErrorLogs\>  
> > \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-\_\_ErrorLogs  
> > \<[Home - Apache Hive - Apache Software Foundation](https://cwiki.apache.org/confluence/display/Hive/)  
> > GettingStarted#GettingStarted-ErrorLogs\>\>
> > 
> > ```
> > <https://cwiki.apache.org/ ____confluence/display/Hive/____
> > 
> > ```
> > 
> > GettingStarted#GettingStarted-\_\_\_\_ErrorLogs  
> > \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-\_\_ErrorLogs\>
> > 
> > ```
> > <https://cwiki.apache.org/ __confluence/display/Hive/__
> > 
> > ```
> > 
> > GettingStarted#GettingStarted-\_\_ErrorLogs  
> > \<[Home - Apache Hive - Apache Software Foundation](https://cwiki.apache.org/confluence/display/Hive/)  
> > GettingStarted#GettingStarted-ErrorLogs\>\>\>\>  
> > \>  
> > \> On 20/02/2014 12:02 AM, Max Lang wrote:  
> > \> \> Hey Costin,  
> > \> \>  
> > \> \> Thanks for the swift reply. I abandoned EC2  
> > to take that out of the equation and managed  
> > to get  
> > everything working  
> > \> \> locally using the latest version of  
> > everything (though I realized just now I'm still on  
> > hive 0.9).  
> > I'm guessing you're  
> > \> \> right about some port connection issue  
> > because I definitely had ES running on that machine.  
> > \> \>  
> > \> \> I changed hive-log4j.properties and added  
> > \> \> |  
> > \> \> #custom logging levels  
> > \> \> #log4j.logger.xxx=DEBUG  
> > \> \> [log4j.logger.org](http://log4j.logger.org) [http://log4j.logger.org](http://log4j.logger.org)  
> > [http://log4j.logger.org](http://log4j.logger.org).\_\_\_\_elasticsearch.hadoop.rest=\_\_\_\_TRACE  
> > \> \>[log4j.logger.org](http://log4j.logger.org).\_\_elasticsea\_\_rch.hadoop.mr\<  
> > [http://elasticsearch.hadoop.mr](http://elasticsearch.hadoop.mr)\>  
> > \<[http://log4j.logger.org](http://log4j.logger.org).\_\_elasticsearch.hadoop.mr \<  
> > [http://log4j.logger.org.elasticsearch.hadoop.mr](http://log4j.logger.org.elasticsearch.hadoop.mr)\>\>  
> > \<[http://log4j.logger.org](http://log4j.logger.org).\_\_ela\_\_sticsearch.hadoop.mr \<  
> > [http://elasticsearch.hadoop.mr](http://elasticsearch.hadoop.mr)\>  
> > \<[http://log4j.logger.org](http://log4j.logger.org).\_\_elasticsearch.hadoop.mr \<  
> > [http://log4j.logger.org.elasticsearch.hadoop.mr](http://log4j.logger.org.elasticsearch.hadoop.mr)\>\>\>  
> > \<[http://log4j.logger.org](http://log4j.logger.org).\_\_ela  
> > \_\_sticsearch.hadoop.mr [http://elasticsearch.hadoop.mr](http://elasticsearch.hadoop.mr)  
> > \<[http://log4j.logger.org](http://log4j.logger.org).\_\_elasticsearch.hadoop.mr \<  
> > [http://log4j.logger.org.elasticsearch.hadoop.mr](http://log4j.logger.org.elasticsearch.hadoop.mr)\>\>  
> > \<[http://log4j.logger.org](http://log4j.logger.org).\_\_ela\_\_sticsearch.hadoop.mr \<  
> > [http://elasticsearch.hadoop.mr](http://elasticsearch.hadoop.mr)\>  
> > \<[http://log4j.logger.org](http://log4j.logger.org).\_\_elasticsearch.hadoop.mr \<  
> > [http://log4j.logger.org.elasticsearch.hadoop.mr](http://log4j.logger.org.elasticsearch.hadoop.mr)\>\>\>\>=\_\_\_\_TRACE
> > 
> > ```
> > > > |
> > > >
> > > > But I didn't see any trace logging.
> > 
> > ```
> > 
> > Hopefully I can get it working on EC2 without issue,  
> > but, for  
> > the future, is this  
> > \> \> the correct way to set TRACE logging?  
> > \> \>  
> > \> \> Oh and, for reference, I tried running  
> > without ES up and I got the following, exceptions:  
> > \> \>  
> > \> \> 2014-02-19 13:46:08,803 ERROR  
> > shark.SharkDriver (Logging.scala:logError(64)) - FAILED: Hive  
> > Internal Error:  
> > \> \> java.lang.\_\_\_\_IllegalStateException(Cannot  
> > discover Elasticsearch version)  
> > \> \> java.lang.\_\_\_\_IllegalStateException:  
> > Cannot discover Elasticsearch version  
> > \> \> at org.elasticsearch.hadoop.hive.  
> > \_\_\_\_EsStorageHandler.init(\_\_\_\_EsStorageHandler.java:101)  
> > \> \> at
> > 
> > ```
> > org.elasticsearch.hadoop.hive. ____EsStorageHandler.____
> > 
> > ```
> > 
> > configureOutputJobProperties(\_\_\_\_EsStorageHandler.java:83)  
> > \> \> at
> > 
> > ```
> > org.apache.hadoop.hive.ql. ____plan.PlanUtils.____
> > 
> > ```
> > 
> > configureJobPropertiesForStora\_\_\_\_geHandler(PlanUtils.java:\_\_706)  
> > \> \> at
> > 
> > ```
> > org.apache.hadoop.hive.ql. ____plan.PlanUtils.____
> > 
> > ```
> > 
> > configureOutputJobPropertiesFo\_\_\_\_rStorageHandler(PlanUtils.\_\_\_\_java:675)  
> > \> \> at  
> > org.apache.hadoop.hive.ql. **exec.FileSinkOperator.**  
> > augmentPlan(FileSinkOperator.\_\_\_\_java:764)  
> > \> \> at  
> > org.apache.hadoop.hive.ql. **parse.SemanticAnalyzer.**  
> > putOpInsertMap(\_\_\_\_SemanticAnalyzer.java:1518)  
> > \> \> at  
> > org.apache.hadoop.hive.ql. **parse.SemanticAnalyzer.**  
> > genFileSinkPlan(\_\_\_\_SemanticAnalyzer.java:4337)  
> > \> \> at
> > 
> > ```
> > org.apache.hadoop.hive.ql. ____parse.SemanticAnalyzer.____
> > 
> > ```
> > 
> > genPostGroupByBodyPlan(\_\_\_\_SemanticAnalyzer.java:6207)  
> > \> \> at  
> > org.apache.hadoop.hive.ql. **parse.SemanticAnalyzer.**  
> > genBodyPlan(SemanticAnalyzer.\_\_\_\_java:6138)  
> > \> \> at  
> > org.apache.hadoop.hive.ql. **parse.SemanticAnalyzer.**  
> > genPlan(SemanticAnalyzer.java:\_\_\_\_6764)  
> > \> \> at  
> > shark.parse.\_\_**SharkSemanticAnalyzer.analyzeInternal(  
> > SharkSemanticAnalyzer.scala:149)  
> > \> \> at  
> > org.apache.hadoop.hive.ql._parse.BaseSemanticAnalyzer.  
> > analyze(BaseSemanticAnalyzer.java:244)  
> > \> \> at shark.SharkDriver.compile(  
> > SharkDriver.scala:215)  
> > \> \> at org.apache.hadoop.hive.ql._  
> > Driver.compile(Driver.java:336)  
> > \> \> at org.apache.hadoop.hive.ql.  
> > Driver.run(Driver.java:895)  
> > \> \> at shark.SharkCliDriver.**  
> > processCmd(SharkCliDriver._**scala:324)  
> > \> \> at org.apache.hadoop.hive.cli.**  
> > CliDriver.processLine(**CliDriver.java:406)  
> > \> \> at shark.SharkCliDriver$.main(**  
> > SharkCliDriver.scala:232)  
> > \> \> at shark.SharkCliDriver.main(_  
> > SharkCliDriver.scala)
> > 
> > ```
> > > > Caused by: java.io.IOException: Out of
> > 
> > ```
> > 
> > nodes and retries; caught exception  
> > \> \> at org.elasticsearch.hadoop.rest.  
> > \_\_\_\_NetworkClient.execute(\_\_\_\_NetworkClient.java:81)  
> > \> \> at org.elasticsearch.hadoop.rest.  
> > \_\_\_\_RestClient.execute(\_\_RestClient.\_\_java:221)  
> > \> \> at org.elasticsearch.hadoop.rest.  
> > \_\_\_\_RestClient.execute(\_\_RestClient.\_\_java:205)  
> > \> \> at org.elasticsearch.hadoop.rest.  
> > \_\_\_\_RestClient.execute(\_\_RestClient.\_\_java:209)  
> > \> \> at org.elasticsearch.hadoop.rest.  
> > \_\_\_\_RestClient.get(RestClient.\_\_\_\_java:103)  
> > \> \> at org.elasticsearch.hadoop.rest.  
> > \_\_\_\_RestClient.esVersion(\_\_\_\_RestClient.java:274)  
> > \> \> at
> > 
> > ```
> > org.elasticsearch.hadoop.rest. ____InitializationUtils.____
> > 
> > ```
> > 
> > discoverEsVersion(\_\_\_\_InitializationUtils.java:84)  
> > \> \> at org.elasticsearch.hadoop.hive.  
> > \_\_\_\_EsStorageHandler.init(\_\_\_\_EsStorageHandler.java:99)
> > 
> > ```
> > > > ... 18 more
> > > > Caused by: java.net.ConnectException:
> > 
> > ```
> > 
> > Connection refused  
> > \> \> at java.net.PlainSocketImpl.\_\_\_\_socketConnect(Native  
> > Method)  
> > \> \> at [java.net](http://java.net) [http://java.net](http://java.net)  
> > [http://java.net](http://java.net). **AbstractPlainSocketImpl.**  
> > doConnect(\_\_\_\_AbstractPlainSocketImpl.java:\_\_\_\_339)  
> > \> \> at [java.net](http://java.net) [http://java.net](http://java.net)
> > 
> > ```
> > <http://java.net>. ____AbstractPlainSocketImpl.____
> > 
> > ```
> > 
> > connectToAddress(\_\_\_\_AbstractPlainSocketImpl.java:\_\_\_\_200)  
> > \> \> at [java.net](http://java.net) [http://java.net](http://java.net)  
> > [http://java.net](http://java.net).**AbstractPlainSocketImpl.connect(  
> > AbstractPlainSocketImpl.java:182)  
> > \> \> at java.net.SocksSocketImpl.  
> > connect(SocksSocketImpl.java:391)  
> > \> \> at java.net.Socket.connect(  
> > Socket.java:579)  
> > \> \> at java.net.Socket.connect(**  
> > Socket.java:528)  
> > \> \> at java.net.Socket.(Socket.  
> > \_\_\_\_java:425)  
> > \> \> at java.net.Socket.(Socket.  
> > \_\_\_\_java:280)  
> > \> \> at
> > 
> > ```
> > org.apache.commons.httpclient. ____protocol.____
> > 
> > ```
> > 
> > DefaultProtocolSocketFactory.**createSocket(**  
> > DefaultProtocolSocketFactory.\_\_\_\_java:80)  
> > \> \> at
> > 
> > ```
> > org.apache.commons.httpclient. ____protocol.____
> > 
> > ```
> > 
> > DefaultProtocolSocketFactory.**createSocket(**  
> > DefaultProtocolSocketFactory.\_\_\_\_java:122)  
> > \> \> at org.apache.commons.httpclient.  
> > \_\_\_\_HttpConnection.open(\_\_\_\_HttpConnection.java:707)  
> > \> \> at  
> > org.apache.commons.httpclient. **HttpMethodDirector.**  
> > executeWithRetry(\_\_\_\_HttpMethodDirector.java:387)  
> > \> \> at  
> > org.apache.commons.httpclient. **HttpMethodDirector.**  
> > executeMethod(\_\_\_\_HttpMethodDirector.java:171)  
> > \> \> at org.apache.commons.httpclient.  
> > \_\_\_\_HttpClient.executeMethod(\_\_\_\_HttpClient.java:397)  
> > \> \> at org.apache.commons.httpclient.  
> > \_\_\_\_HttpClient.executeMethod(\_\_\_\_HttpClient.java:323)  
> > \> \> at
> > 
> > ```
> > org.elasticsearch.hadoop.rest. ____commonshttp.____
> > 
> > ```
> > 
> > CommonsHttpTransport.execute(\_\_\_\_CommonsHttpTransport.java:\_\_160)  
> > \> \> at org.elasticsearch.hadoop.rest.  
> > \_\_\_\_NetworkClient.execute(\_\_\_\_NetworkClient.java:74)
> > 
> > ```
> > > > ... 25 more
> > > >
> > > > Let me know if there's anything in
> > 
> > ```
> > 
> > particular you'd like me to try on EC2.  
> > \> \>  
> > \> \> (For posterity, the versions I used were:  
> > hadoop 2.2.0, hive 0.9.0, shark 8.1, spark 8.1,  
> > es-hadoop  
> > 1.3.0.M2, java  
> > \> \> 1.7.0\_15, scala 2.9.3, elasticsearch 1.0.0)  
> > \> \>  
> > \> \> Thanks again,  
> > \> \> Max  
> > \> \>  
> > \> \> On Tuesday, February 18, 2014 10:16:38 PM  
> > UTC-8, Costin Leau wrote:  
> > \> \>  
> > \> \> The error indicates a network error -  
> > namely es-hadoop cannot connect to Elasticsearch  
> > on the  
> > default (localhost:9200)  
> > \> \> HTTP port. Can you double check whether  
> > that's indeed the case (using curl or even  
> > telnet on  
> > that port) - maybe the  
> > \> \> firewall prevents any connections to be  
> > made...  
> > \> \> Also you could try using the latest  
> > Hive, 0.12 and a more recent Hadoop such as 1.1.2  
> > or 1.2.1.  
> > \> \>  
> > \> \> Additionally, can you enable TRACE  
> > logging in your job on es-hadoop packages  
> > org.elasticsearch.hadoop.rest and  
> > \> \>org.elasticsearch.hadoop.mr \<  
> > [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)\>  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_hadoop.mr \<[http://org.elasticsearch](http://org.elasticsearch).  
> > hadoop.mr\>\>  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_ha\_\_doop.mr \<  
> > [http://hadoop.mr](http://hadoop.mr)\> \<[http://org.elasticsearch](http://org.elasticsearch).\_\_hadoop.mr  
> > [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)\>\>  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_ha\_\_doop.mr \<  
> > [http://hadoop.mr](http://hadoop.mr)\> \<[http://org.elasticsearch](http://org.elasticsearch).\_\_hadoop.mr  
> > [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)\>  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_ha\_\_doop.mr \<  
> > [http://hadoop.mr](http://hadoop.mr)\> \<[http://org.elasticsearch](http://org.elasticsearch).\_\_hadoop.mr  
> > [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)\>\>\>  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_ha\_\_doop.mr \<  
> > [http://hadoop.mr](http://hadoop.mr)\> \<[http://org.elasticsearch](http://org.elasticsearch).\_\_hadoop.mr  
> > [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)\> \<[http://org.elasticsearch](http://org.elasticsearch).  
> > \_\_ha\_\_doop.mr [http://hadoop.mr](http://hadoop.mr)  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_hadoop.mr \<  
> > [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)\>\>\>
> > 
> > ```
> > > <http://org.elasticsearch. __ha__ doop.mr <
> > 
> > ```
> > 
> > [http://hadoop.mr](http://hadoop.mr)\>  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_hadoop.mr \<[http://org.elasticsearch](http://org.elasticsearch).  
> > hadoop.mr\>\>  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_ha\_\_doop.mr \<  
> > [http://hadoop.mr](http://hadoop.mr)\> \<[http://org.elasticsearch](http://org.elasticsearch).\_\_hadoop.mr  
> > [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)\>\>\>\> packages and report  
> > back ?
> > 
> > ```
> > > >
> > > > Thanks,
> > > >
> > > > On 19/02/2014 4:03 AM, Max Lang wrote:
> > > > > I set everything up using this
> > guide:https://github.com/ ____
> > 
> > ```
> > 
> > amplab/shark/wiki/Running-\_\_\_\_Shark-on-EC2  
> > [https://github.com/\_\_amplab/shark/wiki/Running-\_\_Shark-on-EC2](https://github.com/ __amplab/shark/wiki/Running-__ Shark-on-EC2)  
> > \<[https://github.com/amplab/\_\_](https://github.com/amplab/__)  
> > shark/wiki/Running-Shark-on-\_\_EC2  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>  
> > \<[https://github.com/amplab/\_\_\_](https://github.com/amplab/___)  
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)  
> > \<[https://github.com/amplab/\_\_](https://github.com/amplab/__)  
> > shark/wiki/Running-Shark-on-\_\_EC2  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>\>  
> > \<[https://github.com/amplab/\_\_\_](https://github.com/amplab/___)  
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)  
> > \<[https://github.com/amplab/\_\_](https://github.com/amplab/__)  
> > shark/wiki/Running-Shark-on-\_\_EC2  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>  
> > \<[https://github.com/amplab/\_\_\_](https://github.com/amplab/___)  
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)  
> > \<[https://github.com/amplab/\_\_](https://github.com/amplab/__)  
> > shark/wiki/Running-Shark-on-\_\_EC2  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>\>\>  
> > \> \> \<[https://github.com/amplab/\_\_\_](https://github.com/amplab/___)  
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)  
> > \<[https://github.com/amplab/\_\_](https://github.com/amplab/__)  
> > shark/wiki/Running-Shark-on-\_\_EC2  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>  
> > \<[https://github.com/amplab/\_\_\_](https://github.com/amplab/___)  
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)  
> > \<[https://github.com/amplab/\_\_](https://github.com/amplab/__)  
> > shark/wiki/Running-Shark-on-\_\_EC2  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>\>  
> > \> \<[https://github.com/amplab/\_\_\_](https://github.com/amplab/___)  
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)  
> > \<[https://github.com/amplab/\_\_](https://github.com/amplab/__)  
> > shark/wiki/Running-Shark-on-\_\_EC2  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>  
> > \<[https://github.com/amplab/\_\_\_](https://github.com/amplab/___)  
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)
> > 
> > ```
> > <https://github.com/amplab/__
> > 
> > ```
> > 
> > shark/wiki/Running-Shark-on-\_\_EC2  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>\>\>\>  
> > on an ec2 cluster. I've  
> > \> \> \> copied the elasticsearch-hadoop jars  
> > into the hive lib directory and I have  
> > elasticsearch  
> > running on localhost:9200. I'm  
> > \> \> \> running shark in a screen session  
> > with --service screenserver and connecting to it  
> > at the  
> > same time using shark -h  
> > \> \> \> localhost.  
> > \> \> \>  
> > \> \> \> Unfortunately, when I attempt to  
> > write data into elasticsearch, it fails. Here's an  
> > example:  
> > \> \> \>  
> > \> \> \> |  
> > \> \> \> [localhost:10000]shark\>CREATE  
> > EXTERNAL TABLE wiki (id BIGINT,title STRING,last\_modified  
> > STRING,xml STRING,text  
> > \> \> \> STRING)ROW FORMAT DELIMITED FIELDS  
> > TERMINATED BY '\t'LOCATION  
> > 's3n://spark-data/wikipedia-\_\_\_\_sample/';
> > 
> > ```
> > > > > Timetaken (including network
> > 
> > ```
> > 
> > latency):0.159seconds  
> > \> \> \> 14/02/1901:23:33INFO  
> > CliDriver:Timetaken (including network latency):0.159seconds  
> > \> \> \>  
> > \> \> \> [localhost:10000]shark\>SELECT title  
> > FROM wiki LIMIT 1;  
> > \> \> \> Alpokalja  
> > \> \> \> Timetaken (including network  
> > latency):2.23seconds  
> > \> \> \> 14/02/1901:23:48INFO  
> > CliDriver:Timetaken (including network latency):2.23seconds  
> > \> \> \>  
> > \> \> \> [localhost:10000]shark\>CREATE  
> > EXTERNAL TABLE es\_wiki (id BIGINT,title  
> > STRING,last\_modified  
> > STRING,xml STRING,text  
> > \> \> \> STRING)STORED BY
> > 
> > ```
> > 'org.elasticsearch.hadoop. ____hive.EsStorageHandler'____
> > 
> > ```
> > 
> > TBLPROPERTIES('es.resource'='\_\_\_\_wikipedia/article');
> > 
> > ```
> > > > > Timetaken (including network
> > 
> > ```
> > 
> > latency):0.061seconds  
> > \> \> \> 14/02/1901:33:51INFO  
> > CliDriver:Timetaken (including network latency):0.061seconds  
> > \> \> \>  
> > \> \> \> [localhost:10000]shark\>INSERT  
> > OVERWRITE TABLE es\_wiki SELECTw.id  
> > [http://w.id](http://w.id),w.title,w.last\_\_\_\_\_modified,w.xml,w.text  
> > FROM wiki w;  
> > \> \> \> [HiveError]:Queryreturned non-zero  
> > code:9,cause:FAILED:\_\_\_\_ExecutionError,returncode  
> > -101fromshark.execution.\_\_\_\_SparkTask
> > 
> > ```
> > > > > Timetaken (including network
> > 
> > ```
> > 
> > latency):3.575seconds  
> > \> \> \> 14/02/1901:34:42INFO  
> > CliDriver:Timetaken (including network latency):3.575seconds  
> > \> \> \> |  
> > \> \> \>  
> > \> \> \> _The stack trace looks like this:_  
> > \> \> \>  
> > \> \> \> org.apache.hadoop.hive.ql.\_\_\_\_  
> > metadata.HiveException  
> > (org.apache.hadoop.hive.ql.\_\_\_\_metadata.HiveException:  
> > java.io.IOException:
> > 
> > ```
> > > > > Out of nodes and retries; caught
> > 
> > ```
> > 
> > exception)  
> > \> \> \>  
> > \> \> \>
> > 
> > ```
> > org.apache.hadoop.hive.ql. ____exec.FileSinkOperator.____
> > 
> > ```
> > 
> > processOp(FileSinkOperator.**java:602)shark.execution.**  
> > FileSinkOperator$$anonfun$**processPartition$1.apply(**  
> > FileSinkOperator.scala:84) **shark.execution.**  
> > FileSinkOperator$$anonfun$**processPartition$1.apply(**  
> > FileSinkOperator.scala:81) **scala.collection.Iterator$**  
> > class.foreach(Iterator.scala:\_\_**772)scala.collection.**  
> > Iterator$\_\_$anon$19.foreach(\_\_Iterator.**scala:399)shark.**  
> > execution.\_\_FileSinkOperator.**processPartition(**  
> > FileSinkOperator.scala:81) **shark.execution.**  
> > FileSinkOperator$.writeFiles$\_\_\_\_1(FileSinkOperator.scala:  
> > 207)\_\_\_\_shark.execution. **FileSinkOperator$$anonfun$**  
> > executeProcessFileSinkPartitio\_\_\_\_n$1.apply(FileSinkOperator.\_\_\_\_scala:  
> > 211)shark.execution. **FileSinkOperator$$anonfun$**  
> > executeProcessFileSinkPartitio\_\_\_\_n$1.apply(FileSinkOperator.\_\_\_\_scala:  
> > 211)org.apache.spark.\_\_\__scheduler.ResultTask.runTask(_  
> > \_\_\_ResultTask.scala:107)org.\_\_\_\_apache.spark.scheduler.Ta
> 
> sk.\_\_\_\_run(Task.scala:53)org.apache.\_\_\_\_spark.executor.Executor$\_\_\_\_Task
> 
> > ```
> > Runner$$anonfun$run$1. __apply$__ mcV$sp(Executor.scala:__215)
> > 
> > ```
> > 
> > \_\_org.apac
> > 
> > ```
> > he.spa
> > 
> > rk.dep
> > >
> > > loy.Sp
> > > >
> > > >
> > 
> > arkHadoopUtil.runAsUser( ____ SparkHadoopUtil.scala:50)org._
> > 
> > ```
> > 
> > \_\_\_apache.spark.executor.**Executor$TaskRunner.run(**  
> > Executor.scala:182)java.util.\_\_ **concurrent.ThreadPoolExecutor.**  
> > runWorker(ThreadPoolExecutor.**java:1145)java.util.**  
> > concurrent.ThreadPoolExecutor$\_\_**Worker.run(**  
> > ThreadPoolExecutor.\_\_java:615)\_\_java.lang.Thread.run(\_\_Thread.\_\_java:744
> > 
> > ```
> > >
> > > >
> > > > > I should be using Hive 0.9.0, shark
> > 
> > ```
> > 
> > 0.8.1, elasticsearch 1.0.0, Hadoop 1.0.4, and  
> > java 1.7.0\_51  
> > \> \> \> Based on my cursory look at the  
> > hadoop and elasticsearch-hadoop sources, it looks  
> > like hive  
> > is just rethrowing an  
> > \> \> \> IOException it's getting from Spark,  
> > and elasticsearch-hadoop is just hitting those  
> > exceptions.  
> > \> \> \> I suppose my questions are: Does this  
> > look like an issue with my ES/elasticsearch-hadoop  
> > config? And has anyone gotten  
> > \> \> \> elasticsearch working with  
> > Spark/Shark?  
> > \> \> \> Any ideas/insights are appreciated.  
> > \> \> \> Thanks,Max  
> > \> \> \>  
> > \> \> \> --  
> > \> \> \> You received this message because you  
> > are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > \> \> \> To unsubscribe from this group and  
> > stop receiving emails from it, send an email to  
> > \> \> \>elasticsearc...@googlegroups.\_\_\_\_com  
> > \<mailto:elasticsearc...@\_\_googlegroups.com  
> > [mailto:elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com)\> \<javascript:\>.
> > 
> > ```
> > > > > To view this discussion on the web
> > 
> > ```
> > 
> > visit  
> > \> \>
> > 
> > ```
> > >https://groups.google.com/d/ ____ msgid/elasticsearch/
> > 
> > ```
> > 
> > 9486faff-\_\_\_\_3eaf-4344-8931-3121bbc5d9c7%\_\_\_\_40googlegroups.com \<  
> > [https://groups.google.com/d/\_\_msgid/elasticsearch/9486faff-](https://groups.google.com/d/__msgid/elasticsearch/9486faff-)  
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/9486faff-
> > 
> > ```
> > 
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)\>\>
> > 
> > ```
> > <https://groups.google.com/d/ ____ msgid/elasticsearch/
> > 
> > ```
> > 
> > 9486faff-\_\_\_\_3eaf-4344-8931-3121bbc5d9c7%\_\_\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/\_\_msgid/elasticsearch/9486faff-](https://groups.google.com/d/__msgid/elasticsearch/9486faff-)  
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/9486faff-
> > 
> > ```
> > 
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)\>\>\>  
> > \>
> > 
> > ```
> > <https://groups.google.com/d/ ____ msgid/elasticsearch/
> > 
> > ```
> > 
> > 9486faff-\_\_\_\_3eaf-4344-8931-3121bbc5d9c7%\_\_\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/\_\_msgid/elasticsearch/9486faff-](https://groups.google.com/d/__msgid/elasticsearch/9486faff-)  
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/9486faff-
> > 
> > ```
> > 
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)\>\>
> > 
> > ```
> > <https://groups.google.com/d/ ____ msgid/elasticsearch/
> > 
> > ```
> > 
> > 9486faff-\_\_\_\_3eaf-4344-8931-3121bbc5d9c7%\_\_\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/\_\_msgid/elasticsearch/9486faff-](https://groups.google.com/d/__msgid/elasticsearch/9486faff-)  
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/9486faff-
> > 
> > ```
> > 
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)\>\>\>\>  
> > \> \>
> > 
> > ```
> > <https://groups.google.com/d/ ____ msgid/elasticsearch/
> > 
> > ```
> > 
> > 9486faff-\_\_\_\_3eaf-4344-8931-3121bbc5d9c7%\_\_\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/\_\_msgid/elasticsearch/9486faff-](https://groups.google.com/d/__msgid/elasticsearch/9486faff-)  
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/9486faff-
> > 
> > ```
> > 
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)\>\>
> > 
> > ```
> > <https://groups.google.com/d/ ____ msgid/elasticsearch/
> > 
> > ```
> > 
> > 9486faff-\_\_\_\_3eaf-4344-8931-3121bbc5d9c7%\_\_\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/\_\_msgid/elasticsearch/9486faff-](https://groups.google.com/d/__msgid/elasticsearch/9486faff-)  
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/9486faff-
> > 
> > ```
> > 
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)\>\>\>  
> > \>
> > 
> > ```
> > <https://groups.google.com/d/ ____ msgid/elasticsearch/
> > 
> > ```
> > 
> > 9486faff-\_\_\_\_3eaf-4344-8931-3121bbc5d9c7%\_\_\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/\_\_msgid/elasticsearch/9486faff-](https://groups.google.com/d/__msgid/elasticsearch/9486faff-)  
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/9486faff-
> > 
> > ```
> > 
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)\>\>
> > 
> > ```
> > <https://groups.google.com/d/ ____ msgid/elasticsearch/
> > 
> > ```
> > 
> > 9486faff-\_\_\_\_3eaf-4344-8931-3121bbc5d9c7%\_\_\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/\_\_msgid/elasticsearch/9486faff-](https://groups.google.com/d/__msgid/elasticsearch/9486faff-)  
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/9486faff-
> > 
> > ```
> > 
> > \_\_3eaf-4344-8931-3121bbc5d9c7%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/9486faff-](https://groups.google.com/d/msgid/elasticsearch/9486faff-)  
> > 3eaf-4344-8931-3121bbc5d9c7%[40googlegroups.com](http://40googlegroups.com)\>\>\>\>\>.  
> > \> \> \> For more options,  
> > visithttps://groups.google.\_\_\_\_com/groups/opt\_out  
> > \<[http://groups.google.com/\_\_groups/opt\_out](http://groups.google.com/__groups/opt_out) \<  
> > [http://groups.google.com/groups/opt\_out](http://groups.google.com/groups/opt_out)\>\>  
> > \<[http://groups.google.com/\_\_\_\_groups/opt\_out](http://groups.google.com/ ____ groups/opt_out) \<  
> > [http://groups.google.com/\_\_groups/opt\_out](http://groups.google.com/__groups/opt_out)\>  
> > \<[http://groups.google.com/\_\_groups/opt\_out](http://groups.google.com/__groups/opt_out) \<  
> > [http://groups.google.com/groups/opt\_out](http://groups.google.com/groups/opt_out)\>\>\>  
> > \<[http://groups.google.com/\_\_\_\_groups/opt\_out](http://groups.google.com/ ____ groups/opt_out) \<  
> > [http://groups.google.com/\_\_groups/opt\_out](http://groups.google.com/__groups/opt_out)\>  
> > \<[http://groups.google.com/\_\_groups/opt\_out](http://groups.google.com/__groups/opt_out) \<  
> > [http://groups.google.com/groups/opt\_out](http://groups.google.com/groups/opt_out)\>\>
> > 
> > ```
> > <http://groups.google.com/ ____ groups/opt_out <
> > 
> > ```
> > 
> > [http://groups.google.com/\_\_groups/opt\_out](http://groups.google.com/__groups/opt_out)\>  
> > \<[http://groups.google.com/\_\_groups/opt\_out](http://groups.google.com/__groups/opt_out) \<  
> > [http://groups.google.com/groups/opt\_out](http://groups.google.com/groups/opt_out)\>\>\>\>  
> > \<[https://groups.google.com/\_\_\_\_groups/opt\_out](https://groups.google.com/ ____ groups/opt_out) \<  
> > [https://groups.google.com/\_\_groups/opt\_out](https://groups.google.com/__groups/opt_out)\>  
> > \<[https://groups.google.com/\_\_groups/opt\_out](https://groups.google.com/__groups/opt_out) \<  
> > [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)\>\>  
> > \<[https://groups.google.com/\_\_\_\_groups/opt\_out](https://groups.google.com/ ____ groups/opt_out) \<  
> > [https://groups.google.com/\_\_groups/opt\_out](https://groups.google.com/__groups/opt_out)\>  
> > \<[https://groups.google.com/\_\_groups/opt\_out](https://groups.google.com/__groups/opt_out) \<  
> > [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)\>\>\>  
> > \> \<[https://groups.google.com/\_\_\_\_groups/opt\_out](https://groups.google.com/ ____ groups/opt_out)\<  
> > [https://groups.google.com/\_\_groups/opt\_out](https://groups.google.com/__groups/opt_out)\>  
> > \<[https://groups.google.com/\_\_groups/opt\_out](https://groups.google.com/__groups/opt_out) \<  
> > [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)\>\>  
> > \<[https://groups.google.com/\_\_\_\_groups/opt\_out](https://groups.google.com/ ____ groups/opt_out) \<  
> > [https://groups.google.com/\_\_groups/opt\_out](https://groups.google.com/__groups/opt_out)\>  
> > \<[https://groups.google.com/\_\_groups/opt\_out](https://groups.google.com/__groups/opt_out) \<  
> > [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)\>\>\>\>\>.
> > 
> > ```
> > > >
> > > > --
> > > > Costin
> > > >
> > > > --
> > > > You received this message because you are
> > 
> > ```
> > 
> > subscribed to the Google Groups "elasticsearch"  
> > group.  
> > \> \> To unsubscribe from this group and stop  
> > receiving emails from it, send an email to  
> > \> \>elasticsearc...@googlegroups.\_\_\_\_com  
> > \<mailto:elasticsearc...@\_\_googlegroups.com  
> > [mailto:elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com)\> \<javascript:\>.
> > 
> > ```
> > > > To view this discussion on the web visit
> > >
> > 
> > >https://groups.google.com/d/ ____ msgid/elasticsearch/
> > 
> > ```
> > 
> > 86187c3a-\_\_\_\_0974-4d10-9689-e83da788c04a%\_\_\_\_40googlegroups.com \<  
> > [https://groups.google.com/d/\_\_msgid/elasticsearch/86187c3a-](https://groups.google.com/d/__msgid/elasticsearch/86187c3a-)  
> > \_\_0974-4d10-9689-e83da788c04a%\_\_40googlegroups.com\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/86187c3a-
> > 
> > ```
> > 
> > \_\_0974-4d10-9689-e83da788c04a%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/86187c3a-](https://groups.google.com/d/msgid/elasticsearch/86187c3a-)  
> > 0974-4d10-9689-e83da788c04a%[40googlegroups.com](http://40googlegroups.com)\>\>
> > 
> > ```
> > <https://groups.google.com/d/ ____ msgid/elasticsearch/
> > 
> > ```
> > 
> > 86187c3a-\_\_\_\_0974-4d10-9689-e83da788c04a%\_\_\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/\_\_msgid/elasticsearch/86187c3a-](https://groups.google.com/d/__msgid/elasticsearch/86187c3a-)  
> > \_\_0974-4d10-9689-e83da788c04a%\_\_40googlegroups.com\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/86187c3a-
> > 
> > ```
> > 
> > \_\_0974-4d10-9689-e83da788c04a%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/86187c3a-](https://groups.google.com/d/msgid/elasticsearch/86187c3a-)  
> > 0974-4d10-9689-e83da788c04a%[40googlegroups.com](http://40googlegroups.com)\>\>\>  
> > \>
> > 
> > ```
> > <https://groups.google.com/d/ ____ msgid/elasticsearch/
> > 
> > ```
> > 
> > 86187c3a-\_\_\_\_0974-4d10-9689-e83da788c04a%\_\_\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/\_\_msgid/elasticsearch/86187c3a-](https://groups.google.com/d/__msgid/elasticsearch/86187c3a-)  
> > \_\_0974-4d10-9689-e83da788c04a%\_\_40googlegroups.com\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/86187c3a-
> > 
> > ```
> > 
> > \_\_0974-4d10-9689-e83da788c04a%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/86187c3a-](https://groups.google.com/d/msgid/elasticsearch/86187c3a-)  
> > 0974-4d10-9689-e83da788c04a%[40googlegroups.com](http://40googlegroups.com)\>\>
> > 
> > ```
> > <https://groups.google.com/d/ ____ msgid/elasticsearch/
> > 
> > ```
> > 
> > 86187c3a-\_\_\_\_0974-4d10-9689-e83da788c04a%\_\_\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/\_\_msgid/elasticsearch/86187c3a-](https://groups.google.com/d/__msgid/elasticsearch/86187c3a-)  
> > \_\_0974-4d10-9689-e83da788c04a%\_\_40googlegroups.com\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/86187c3a-
> > 
> > ```
> > 
> > \_\_0974-4d10-9689-e83da788c04a%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/86187c3a-](https://groups.google.com/d/msgid/elasticsearch/86187c3a-)  
> > 0974-4d10-9689-e83da788c04a%[40googlegroups.com](http://40googlegroups.com)\>\>\>\>.  
> > \> \> For more options,  
> > visithttps://groups.google.\_\_\_\_com/groups/opt\_out  
> > \<[http://groups.google.com/\_\_groups/opt\_out](http://groups.google.com/__groups/opt_out) \<  
> > [http://groups.google.com/groups/opt\_out](http://groups.google.com/groups/opt_out)\>\>  
> > \<[http://groups.google.com/\_\_\_\_groups/opt\_out](http://groups.google.com/ ____ groups/opt_out) \<  
> > [http://groups.google.com/\_\_groups/opt\_out](http://groups.google.com/__groups/opt_out)\>  
> > \<[http://groups.google.com/\_\_groups/opt\_out](http://groups.google.com/__groups/opt_out) \<  
> > [http://groups.google.com/groups/opt\_out](http://groups.google.com/groups/opt_out)\>\>\>  
> > \<[https://groups.google.com/\_\_\_\_groups/opt\_out](https://groups.google.com/ ____ groups/opt_out) \<  
> > [https://groups.google.com/\_\_groups/opt\_out](https://groups.google.com/__groups/opt_out)\>  
> > \<[https://groups.google.com/\_\_groups/opt\_out](https://groups.google.com/__groups/opt_out) \<  
> > [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)\>\>  
> > \<[https://groups.google.com/\_\_\_\_groups/opt\_out](https://groups.google.com/ ____ groups/opt_out) \<  
> > [https://groups.google.com/\_\_groups/opt\_out](https://groups.google.com/__groups/opt_out)\>  
> > \<[https://groups.google.com/\_\_groups/opt\_out](https://groups.google.com/__groups/opt_out) \<  
> > [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)\>\>\>\>.
> > 
> > ```
> > >
> > > --
> > > Costin
> > >
> > > --
> > > You received this message because you are
> > 
> > ```
> > 
> > subscribed to the Google Groups "elasticsearch" group.  
> > \> To unsubscribe from this group and stop receiving  
> > emails from it, send an email to  
> > \>elasticsearc...@googlegroups.\_\_\_\_com \<mailto:  
> > elasticsearc...@\_\_googlegroups.com  
> > [mailto:elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com)\> \<javascript:\>.
> > 
> > ```
> > > To view this discussion on the web visit
> > 
> > >https://groups.google.com/d/ ____ msgid/elasticsearch/
> > 
> > ```
> > 
> > e29e342d-\_\_\_\_de74-4ed6-93d4-875fc728c5a5%\_\_\_\_40googlegroups.com \<  
> > [https://groups.google.com/d/\_\_msgid/elasticsearch/e29e342d-](https://groups.google.com/d/__msgid/elasticsearch/e29e342d-)  
> > \_\_de74-4ed6-93d4-875fc728c5a5%\_\_40googlegroups.com\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/e29e342d-
> > 
> > ```
> > 
> > \_\_de74-4ed6-93d4-875fc728c5a5%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/e29e342d-](https://groups.google.com/d/msgid/elasticsearch/e29e342d-)  
> > de74-4ed6-93d4-875fc728c5a5%[40googlegroups.com](http://40googlegroups.com)\>\>
> > 
> > ```
> > <https://groups.google.com/d/ ____ msgid/elasticsearch/
> > 
> > ```
> > 
> > e29e342d-\_\_\_\_de74-4ed6-93d4-875fc728c5a5%\_\_\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/\_\_msgid/elasticsearch/e29e342d-](https://groups.google.com/d/__msgid/elasticsearch/e29e342d-)  
> > \_\_de74-4ed6-93d4-875fc728c5a5%\_\_40googlegroups.com\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/e29e342d-
> > 
> > ```
> > 
> > \_\_de74-4ed6-93d4-875fc728c5a5%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/e29e342d-](https://groups.google.com/d/msgid/elasticsearch/e29e342d-)  
> > de74-4ed6-93d4-875fc728c5a5%[40googlegroups.com](http://40googlegroups.com)\>\>\>.
> > 
> > ```
> > > For more options, visithttps://groups.google.___
> > 
> > ```
> > 
> > \_com/groups/opt\_out  
> > \<[http://groups.google.com/\_\_groups/opt\_out](http://groups.google.com/__groups/opt_out) \<  
> > [http://groups.google.com/groups/opt\_out](http://groups.google.com/groups/opt_out)\>\>  
> > \<[https://groups.google.com/\_\_\_\_groups/opt\_out](https://groups.google.com/ ____ groups/opt_out) \<  
> > [https://groups.google.com/\_\_groups/opt\_out](https://groups.google.com/__groups/opt_out)\>
> > 
> > ```
> > <https://groups.google.com/__groups/opt_out <
> > 
> > ```
> > 
> > [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out)\>\>\>.
> > 
> > ```
> > --
> > Costin
> > 
> > --
> > You received this message because you are subscribed to
> > 
> > ```
> > 
> > the Google Groups "elasticsearch" group.  
> > To unsubscribe from this group and stop receiving emails  
> > from it, send an email to  
> > elasticsearch+unsubscribe@\_\_go\_\_oglegroups.com \<  
> > [http://googlegroups.com](http://googlegroups.com)\>  
> > \<mailto:elasticsearch%**[2Bunsubscribe@googlegroups.com](mailto:2Bunsubscribe@googlegroups.com) \<mailto:  
> > elasticsearch%252Bunsubscribe@googlegroups.com\>**\>  
> > \<[mailto:elasticsearch+\_\_\_\_unsubscribe@googlegroups.com](mailto:elasticsearch+ ____ unsubscribe@googlegroups.com)  
> > [mailto:elasticsearch%2B\_\_unsubscribe@googlegroups.com](mailto:elasticsearch%2B__unsubscribe@googlegroups.com) \<mailto:  
> > elasticsearch%**[2Bunsubscribe@googlegroups.com](mailto:2Bunsubscribe@googlegroups.com)  
> > [mailto:elasticsearch%252Bunsubscribe@googlegroups.com](mailto:elasticsearch%252Bunsubscribe@googlegroups.com)**\>\>.
> > 
> > ```
> > To view this discussion on the web visit
> > https://groups.google.com/d/ ____ msgid/elasticsearch/
> > 
> > ```
> > 
> > c1081bf2-\_\_\_\_117a-4af2-ba90-2c38a4572782%\_\_\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/\_\_msgid/elasticsearch/c1081bf2-](https://groups.google.com/d/__msgid/elasticsearch/c1081bf2-)  
> > \_\_117a-4af2-ba90-2c38a4572782%\_\_40googlegroups.com\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/c1081bf2-
> > 
> > ```
> > 
> > \_\_117a-4af2-ba90-2c38a4572782%\_\_40googlegroups.com  
> > \<[https://groups.google.com/d/msgid/elasticsearch/c1081bf2-](https://groups.google.com/d/msgid/elasticsearch/c1081bf2-)  
> > 117a-4af2-ba90-2c38a4572782%[40googlegroups.com](http://40googlegroups.com)\>\>
> > 
> > ```
> > <https://groups.google.com/d/ ____ msgid/elasticsearch/
> > 
> > ```
> > 
> > c1081bf2-\_\_\_\_117a-4af2-ba90-2c38a4572782%\__**[40googlegroups.com?utm](http://40googlegroups.com?utm)**_  
> > medium=\_\_email&utm\_source=\_\_footer  
> > \<[https://groups.google.com/d/\_\_msgid/elasticsearch/c1081bf2-](https://groups.google.com/d/__msgid/elasticsearch/c1081bf2-)  
> > \_\_117a-4af2-ba90-2c38a4572782%\__[40googlegroups.com?utm](http://40googlegroups.com?utm)_  
> > medium=\_\_email&utm\_source=footer\>
> > 
> > ```
> > <https://groups.google.com/d/__msgid/elasticsearch/c1081bf2-
> > 
> > ```
> > 
> > \_\_117a-4af2-ba90-2c38a4572782%\__[40googlegroups.com?utm](http://40googlegroups.com?utm)_  
> > medium=\_\_email&utm\_source=footer  
> > \<[https://groups.google.com/d/msgid/elasticsearch/c1081bf2-](https://groups.google.com/d/msgid/elasticsearch/c1081bf2-)  
> > 117a-4af2-ba90-2c38a4572782%[40googlegroups.com?utm\_medium=](http://40googlegroups.com?utm_medium=)  
> > email&utm\_source=footer\>\>\>.
> > 
> > ```
> > For more options, visit https://groups.google.com/d/__
> > 
> > ```
> > 
> > \_\_optout [https://groups.google.com/d/\_\_optout](https://groups.google.com/d/__optout)  
> > \<[https://groups.google.com/d/\_\_optout](https://groups.google.com/d/__optout) \<  
> > [https://groups.google.com/d/optout](https://groups.google.com/d/optout)\>\>.
> > 
> > ```
> > --
> > Costin
> > 
> > --
> > 
> > You received this message because you are subscribed to a
> > 
> > ```
> > 
> > topic in the Google Groups "elasticsearch" group.  
> > To unsubscribe from this topic, visit  
> > [https://groups.google.com/d/\_\_\_\_topic/elasticsearch/S-\_\_\_\_](https://groups.google.com/d/ ____topic/elasticsearch/S-____ )  
> > BrzwUHJbM/unsubscribe  
> > \<[https://groups.google.com/d/\_\_topic/elasticsearch/S-\_\_](https://groups.google.com/d/ __topic/elasticsearch/S-__ )  
> > BrzwUHJbM/unsubscribe\>  
> > \<[https://groups.google.com/d/\_\_topic/elasticsearch/S-\_\_](https://groups.google.com/d/ __topic/elasticsearch/S-__ )  
> > BrzwUHJbM/unsubscribe  
> > \<[https://groups.google.com/d/topic/elasticsearch/S-](https://groups.google.com/d/topic/elasticsearch/S-)  
> > BrzwUHJbM/unsubscribe\>\>.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GOd-wSu\_XQpc\_cjCcv\_ZgdiwEoJ6BT6VCAkL4ThvOHTAw%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GOd-wSu_XQpc_cjCcv_ZgdiwEoJ6BT6VCAkL4ThvOHTAw%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![costin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/costin/32/44950_2.png) [@costin](https://discuss.elastic.co/u/costin)\
**Post date:** [May 13, 2014, 7:39pm UTC](https://discuss.elastic.co/t/hadoop-getting-elasticsearch-hadoop-working-with-shark/15882/14 "2014-05-13T19:39:25Z")

</div>

Could you share your setup and configuration on a gist (the more info especially regarding the versions of the stack  
used helps)?  
Do you use just the input format or also the output format? To clarify - are you using Spark (Map/Reduce) or Shark (and  
the relevant  
Hive integration in es-hadoop)?

Cheers,

On 5/13/14 10:14 PM, Nick Pentreath wrote:

> Ok - well let me know when you're around.
> 
> The mapreduce inputformat works fine. I'm using it with Spark to access the ES data via ESInputFormat and run analytics  
> and machine learning jobs on that data, and the same \_ts field works and is the correct data (though it comes through as  
> org.apache.hadoop.io.Text, which I convert to Long or a DateTime as required).
> 
> Perhaps I'm missing it somewhere but is it possible to force a field to be a type? i.e. similar the es.field.mapping  
> could I tell it that it must parse the field as a string (since then I can take it and do whatever parsing / casting I  
> want).
> 
> I could just use the new Spark SQL module (which I'm seriously considering right now having explored it a bit in the  
> last few days), but some of the stuff we do requires a SQL Console and JDBC, so having Shark able to just pull in ES  
> data is definitely very useful...
> 
> On Tue, May 13, 2014 at 8:18 PM, Costin Leau \<[costin.leau@gmail.com](mailto:costin.leau@gmail.com) [mailto:costin.leau@gmail.com](mailto:costin.leau@gmail.com)\> wrote:
> 
> ```
> Hi Nick,
> 
> I'm glad to see you are making progress. This week I'm mainly on the road but maybe we can meet on the IRC next
> week, my invitation still stands :)
> Timestamp is relatively new type and doesn't handle timezones properly - it is backed by java.sq.Timestamp so it
> inherits a lot of its issues.
> For some reason the year in your date is rather off so it's worth checking the data read by es-hadoop before passing
> it to Hive (see [1]).
> I've had issues myself with it and it the moment the cluster is in a different timezone than the dataset itself
> things get buggy.
> Try using a UDF to do the conversion from the long to a timestamp - I've tried doing something similar in our
> conversion but since we don't know the timezones
> used, it's easy for things to get mixed.
> 
> Cheers,
> 
> [1] http://www.elasticsearch.org/ __guide/en/elasticsearch/hadoop/__ current/troubleshooting.html
> <http://www.elasticsearch.org/guide/en/elasticsearch/hadoop/current/troubleshooting.html>
> 
> On 5/13/14 8:25 PM, Nick Pentreath wrote:
> 
> Hi Costin
> 
> Sorry for the silence on this issue. This went a bit quiet.
> 
> But the good news is I've come back to it and managed to get it all working with the new shark 0.9.1 release and
> 2.0.0RC1. Actually if I used ADD JAR I got the same exception but when I just put the JAR into the shark lib/
> folder it
> worked fine (which seems to point to the classpath issue you mention).
> 
> However, I seem to have an issue with date <-> timestamp conversion.
> 
> I have a field in ES called "_ts" that has type "date" and the default format "dateOptionalTime". When I do a
> query that
> includes the timestamp it comes back NULL:
> 
> select ts from table ...
> (note I use a correct es.mapping.names to map the _ts field in ES to ts field in Hive/Shark that has timestamp
> type).
> 
> below is some of the debug-level output:
> 
> 14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the TIMESTAMP data type range so converted to null.
> Given data
> is :96997506-06-30 19:08:168:16.768
> 14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the TIMESTAMP data type range so converted to null.
> Given data
> is :96997605-06-28 19:08:168:16.768
> 14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the TIMESTAMP data type range so converted to null.
> Given data
> is :96997624-06-28 19:08:168:16.768
> 14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the TIMESTAMP data type range so converted to null.
> Given data
> is :96997629-06-28 19:08:168:16.768
> 14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the TIMESTAMP data type range so converted to null.
> Given data
> is :96997634-06-29 19:08:168:16.768
> NULL
> NULL
> NULL
> NULL
> NULL
> 
> The data that I index in the _ts field is timestamp in ms (long). It doesn't seem to be converted correctly but
> the data
> is correct (in ms at least) and I can query against it using date formats and date math in ES.
> 
> Example snippet from debug log from above:
> ,"_ts":1397130475607}}]}}"
> 
> Any ideas or am I doing something silly?
> 
> I do see that the Hive timestamp expects either seconds since epoch of a string-based format that has nanosecond
> granularity. Is this the issue with just ms long timestamp data?
> 
> Thanks
> Nick
> 
> On Thu, Mar 27, 2014 at 4:50 PM, Costin Leau <costin.leau@gmail.com <mailto:costin.leau@gmail.com>
> <mailto:costin.leau@gmail.com <mailto:costin.leau@gmail.com>>__> wrote:
> 
> Using the latest hive and hadoop is preferred as they contain various bug fixes.
> The error suggests a classpath issue - namely the same class is loaded twice for some reason and hence the
> casting
> fails.
> 
> Let's connect on IRC - give me a ping when you're available (user is costin).
> 
> Cheers,
> 
> On 3/27/14 4:29 PM, Nick Pentreath wrote:
> 
> Thanks for the response.
> 
> I tried latest Shark (cdh4 version of 0.9.1 here http://cloudera.rst.im/shark/ ) - this uses hadoop
> 1.0.4 and
> hive 0.11
> I believe, and build elasticsearch-hadoop from github master.
> 
> Still getting same error:
> org.elasticsearch.hadoop.hive. ____EsHiveInputFormat$__ EsHiveSplit cannot be cast to
> org.elasticsearch.hadoop.hive. ____EsHiveInputFormat$__ EsHiveSplit
> 
> Will using hive 0.11 / hadoop 1.0.4 vs hive 0.12 / hadoop 1.2.1 in es-hadoop master make a difference?
> 
> Anyone else actually got this working?
> 
> On Thu, Mar 20, 2014 at 2:44 PM, Costin Leau <costin.leau@gmail.com <mailto:costin.leau@gmail.com>
> <mailto:costin.leau@gmail.com <mailto:costin.leau@gmail.com>>
> <mailto:costin.leau@gmail.com <mailto:costin.leau@gmail.com> <mailto:costin.leau@gmail.com
> <mailto:costin.leau@gmail.com>> __>__ > wrote:
> 
> I recommend using master - there are several improvements done in this area. Also using the latest
> Shark
> (0.9.0) and
> Hive (0.12) will help.
> 
> On 3/20/14 12:00 PM, Nick Pentreath wrote:
> 
> Hi
> 
> I am struggling to get this working too. I'm just trying locally for now, running Shark 0.8.1,
> Hive
> 0.9.0 and ES
> 1.0.1
> with ES-hadoop 1.3.0.M2.
> 
> I managed to get a basic example working with WRITING into an index. But I'm really after
> READING and
> index.
> 
> I believe I have set everything up correctly, I've added the jar to Shark:
> ADD JAR /path/to/es-hadoop.jar;
> 
> created a table:
> CREATE EXTERNAL TABLE test_read (name string, price double)
> 
> STORED BY 'org.elasticsearch.hadoop. ______ hive.EsStorageHandler'
> 
> TBLPROPERTIES('es.resource' = 'test_index/test_type/_search? ______ q=*');
> 
> And then trying to 'SELECT * FROM test _read' gives me :
> 
> org.apache.spark. ______ SparkException: Job aborted: Task 3.0:0 failed more than 0 times;
> aborting job
> java.lang.ClassCastException:
> org.elasticsearch.hadoop.hive. ______EsHiveInputFormat$____ ESHiveSplit cannot
> be cast to
> org.elasticsearch.hadoop.hive. ______EsHiveInputFormat$____ ESHiveSplit
> 
> at
> org.apache.spark.scheduler. ______DAGScheduler$$anonfun$______ abortStage$1.apply( ______ DAGScheduler.scala:827)
> 
> at
> org.apache.spark.scheduler. ______DAGScheduler$$anonfun$______ abortStage$1.apply( ______ DAGScheduler.scala:825)
> 
> at scala.collection.mutable.______ResizableArray$class.foreach(______ResizableArray.scala:60)
> 
> at scala.collection.mutable.______ArrayBuffer.foreach(______ArrayBuffer.scala:47)
> 
> at org.apache.spark.scheduler.______DAGScheduler.abortStage(______DAGScheduler.scala:825)
> 
> at org.apache.spark.scheduler.______DAGScheduler.processEvent(______DAGScheduler.scala:440)
> 
> at org.apache.spark.scheduler. ______ DAGScheduler.org
> <http://org.apache.spark. __sch__ eduler.DAGScheduler.org <http://scheduler.DAGScheduler.org>
> <http://org.apache.spark.__scheduler.DAGScheduler.org
> <http://org.apache.spark.scheduler.DAGScheduler.org>>>$ ____apache$spark$__ scheduler$____DAGScheduler$$run(______DAGScheduler.scala:502)
> 
> at org.apache.spark.scheduler.______DAGScheduler$$anon$1.run(______DAGScheduler.scala:157)
> 
> FAILED: Execution Error, return code -101 from shark.execution.SparkTask
> 
> In fact I get the same error thrown when trying to READ from the table that I successfully
> WROTE to...
> 
> On Saturday, 22 February 2014 12:31:21 UTC+2, Costin Leau wrote:
> 
> Yeah, it might have been some sort of network configuration issue where services where
> running on
> different
> machines
> and
> localhost pointed to a different location.
> 
> Either way, I'm glad to hear things have are moving forward.
> 
> Cheers,
> 
> On 22/02/2014 1:06 AM, Max Lang wrote:
> > I managed to get it working on ec2 without issue this time. I'd say the biggest
> difference was
> that this
> time I set up a
> > dedicated ES machine. Is it possible that, because I was using a cluster with slaves,
> when I used
> "localhost" the slaves
> > couldn't find the ES instance running on the master? Or do all the requests go through
> the master?
> >
> >
> > On Wednesday, February 19, 2014 2:35:40 PM UTC-8, Costin Leau wrote:
> >
> > Hi,
> >
> > Setting logging in Hive/Hadoop can be tricky since the log4j needs to be picked up
> by the
> running JVM
> otherwise you
> > won't see anything.
> > Take a look at this link on how to tell Hive to use your logging settings [1].
> >
> > For the next release, we might introduce dedicated exceptions for the simple fact
> that some
> libraries, like Hive,
> > swallow the stack trace and it's unclear what the issue is which makes the exception
> (IllegalStateException) ambiguous.
> >
> > Let me know how it goes and whether you will encounter any issues with Shark. Or if
> you don't :)
> >
> > Thanks!
> >
> >
> 
> [1]https://cwiki.apache.org/ ______confluence/display/Hive/______ GettingStarted#GettingStarted- ______ ErrorLogs
> <https://cwiki.apache.org/ ____confluence/display/Hive/____ GettingStarted#GettingStarted- ____ ErrorLogs>
> <https://cwiki.apache.org/ ____confluence/display/Hive/____ GettingStarted#GettingStarted- ____ ErrorLogs
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs>>
> 
> <https://cwiki.apache.org/ ____confluence/display/Hive/____ GettingStarted#GettingStarted- ____ ErrorLogs
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs>
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs
> <https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs>>>
> 
> <https://cwiki.apache.org/ ______confluence/display/Hive/______ GettingStarted#GettingStarted- ______ ErrorLogs
> <https://cwiki.apache.org/ ____confluence/display/Hive/____ GettingStarted#GettingStarted- ____ ErrorLogs>
> <https://cwiki.apache.org/ ____confluence/display/Hive/____ GettingStarted#GettingStarted- ____ ErrorLogs
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs>>
> 
> <https://cwiki.apache.org/ ____confluence/display/Hive/____ GettingStarted#GettingStarted- ____ ErrorLogs
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs>
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs
> <https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs>>>>
> >
> 
> <https://cwiki.apache.org/ ______confluence/display/Hive/______ GettingStarted#GettingStarted- ______ ErrorLogs
> <https://cwiki.apache.org/ ____confluence/display/Hive/____ GettingStarted#GettingStarted- ____ ErrorLogs>
> <https://cwiki.apache.org/ ____confluence/display/Hive/____ GettingStarted#GettingStarted- ____ ErrorLogs
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs>>
> 
> <https://cwiki.apache.org/ ____confluence/display/Hive/____ GettingStarted#GettingStarted- ____ ErrorLogs
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs>
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs
> <https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs>>>
> 
> <https://cwiki.apache.org/ ______confluence/display/Hive/______ GettingStarted#GettingStarted- ______ ErrorLogs
> <https://cwiki.apache.org/ ____confluence/display/Hive/____ GettingStarted#GettingStarted- ____ ErrorLogs>
> <https://cwiki.apache.org/ ____confluence/display/Hive/____ GettingStarted#GettingStarted- ____ ErrorLogs
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs>>
> 
> <https://cwiki.apache.org/ ____confluence/display/Hive/____ GettingStarted#GettingStarted- ____ ErrorLogs
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs>
> <https://cwiki.apache.org/ __confluence/display/Hive/__ GettingStarted#GettingStarted-__ErrorLogs
> <https://cwiki.apache.org/confluence/display/Hive/GettingStarted#GettingStarted-ErrorLogs>>>>>
> >
> > On 20/02/2014 12:02 AM, Max Lang wrote:
> > > Hey Costin,
> > >
> > > Thanks for the swift reply. I abandoned EC2 to take that out of the equation and
> managed
> to get
> everything working
> > > locally using the latest version of everything (though I realized just now I'm
> still on
> hive 0.9).
> I'm guessing you're
> > > right about some port connection issue because I definitely had ES running on
> that machine.
> > >
> > > I changed hive-log4j.properties and added
> > > |
> > > #custom logging levels
> > > #log4j.logger.xxx=DEBUG
> > > log4j.logger.org <http://log4j.logger.org> <http://log4j.logger.org>
> <http://log4j.logger.org>. ______elasticsearch.hadoop.rest=______ TRACE
> > >log4j.logger.org. __elasticsea____ rch.hadoop.mr <http://elasticsea__rch.hadoop.mr>
> <http://elasticsearch.hadoop.__mr <http://elasticsearch.hadoop.mr>>
> <http://log4j.logger.org. __ela__ sticsearch.hadoop.mr <http://elasticsearch.hadoop.mr>
> <http://log4j.logger.org.__elasticsearch.hadoop.mr <http://log4j.logger.org.elasticsearch.hadoop.mr>>>
> <http://log4j.logger.org. __ela____ sticsearch.hadoop.mr <http://ela__sticsearch.hadoop.mr>
> <http://elasticsearch.hadoop.__mr <http://elasticsearch.hadoop.mr>>
> <http://log4j.logger.org. __ela__ sticsearch.hadoop.mr <http://elasticsearch.hadoop.mr>
> <http://log4j.logger.org.__elasticsearch.hadoop.mr <http://log4j.logger.org.elasticsearch.hadoop.mr>>>>
> <http://log4j.logger.org. __ela____ sticsearch.hadoop.mr <http://ela__sticsearch.hadoop.mr>
> <http://elasticsearch.hadoop.__mr <http://elasticsearch.hadoop.mr>>
> <http://log4j.logger.org. __ela__ sticsearch.hadoop.mr <http://elasticsearch.hadoop.mr>
> <http://log4j.logger.org.__elasticsearch.hadoop.mr <http://log4j.logger.org.elasticsearch.hadoop.mr>>>
> <http://log4j.logger.org. __ela____ sticsearch.hadoop.mr <http://ela__sticsearch.hadoop.mr>
> <http://elasticsearch.hadoop.__mr <http://elasticsearch.hadoop.mr>>
> <http://log4j.logger.org. __ela__ sticsearch.hadoop.mr <http://elasticsearch.hadoop.mr>
> <http://log4j.logger.org. __elasticsearch.hadoop.mr <http://log4j.logger.org.elasticsearch.hadoop.mr>>>>>=______ TRACE
> 
> > > |
> > >
> > > But I didn't see any trace logging. Hopefully I can get it working on EC2 without
> issue,
> but, for
> the future, is this
> > > the correct way to set TRACE logging?
> > >
> > > Oh and, for reference, I tried running without ES up and I got the following,
> exceptions:
> > >
> > > 2014-02-19 13:46:08,803 ERROR shark.SharkDriver (Logging.scala:logError(64)) -
> FAILED: Hive
> Internal Error:
> > > java.lang. ______ IllegalStateException(Cannot discover Elasticsearch version)
> > > java.lang. ______ IllegalStateException: Cannot discover Elasticsearch version
> > > at
> org.elasticsearch.hadoop.hive.______EsStorageHandler.init(______EsStorageHandler.java:101)
> > > at
> 
> org.elasticsearch.hadoop.hive. ______EsStorageHandler.______ configureOutputJobProperties( ______ EsStorageHandler.java:83)
> > > at
> 
> org.apache.hadoop.hive.ql. ______plan.PlanUtils.______ configureJobPropertiesForStora______geHandler(PlanUtils.java:____706)
> > > at
> 
> org.apache.hadoop.hive.ql. ______plan.PlanUtils.______ configureOutputJobPropertiesFo______rStorageHandler(PlanUtils.______java:675)
> > > at
> org.apache.hadoop.hive.ql. ______exec.FileSinkOperator.______ augmentPlan(FileSinkOperator. ______ java:764)
> > > at
> 
> org.apache.hadoop.hive.ql. ______parse.SemanticAnalyzer.______ putOpInsertMap( ______ SemanticAnalyzer.java:1518)
> > > at
> 
> org.apache.hadoop.hive.ql. ______parse.SemanticAnalyzer.______ genFileSinkPlan( ______ SemanticAnalyzer.java:4337)
> > > at
> 
> org.apache.hadoop.hive.ql. ______parse.SemanticAnalyzer.______ genPostGroupByBodyPlan( ______ SemanticAnalyzer.java:6207)
> > > at
> org.apache.hadoop.hive.ql. ______parse.SemanticAnalyzer.______ genBodyPlan(SemanticAnalyzer. ______ java:6138)
> > > at
> org.apache.hadoop.hive.ql. ______parse.SemanticAnalyzer.______ genPlan(SemanticAnalyzer.java: ______ 6764)
> > > at
> shark.parse. ______SharkSemanticAnalyzer.______ analyzeInternal( ______SharkSemanticAnalyzer.scala:______ 149)
> > > at
> 
> org.apache.hadoop.hive.ql. ______parse.BaseSemanticAnalyzer.______ analyze(BaseSemanticAnalyzer. ______ java:244)
> > > at shark.SharkDriver.compile( ______ SharkDriver.scala:215)
> > > at org.apache.hadoop.hive.ql.______Driver.compile(Driver.java:______336)
> > > at org.apache.hadoop.hive.ql. ______ Driver.run(Driver.java:895)
> > > at shark.SharkCliDriver.______processCmd(SharkCliDriver.______scala:324)
> > > at org.apache.hadoop.hive.cli.______CliDriver.processLine(______CliDriver.java:406)
> > > at shark.SharkCliDriver$.main( ______ SharkCliDriver.scala:232)
> > > at shark.SharkCliDriver.main( ______ SharkCliDriver.scala)
> 
> > > Caused by: java.io.IOException: Out of nodes and retries; caught exception
> > > at
> org.elasticsearch.hadoop.rest.______NetworkClient.execute(______NetworkClient.java:81)
> > > at org.elasticsearch.hadoop.rest.______RestClient.execute(____RestClient.__java:221)
> > > at org.elasticsearch.hadoop.rest.______RestClient.execute(____RestClient.__java:205)
> > > at org.elasticsearch.hadoop.rest.______RestClient.execute(____RestClient.__java:209)
> > > at org.elasticsearch.hadoop.rest.______RestClient.get(RestClient.______java:103)
> > > at
> org.elasticsearch.hadoop.rest.______RestClient.esVersion(______RestClient.java:274)
> > > at
> 
> org.elasticsearch.hadoop.rest. ______InitializationUtils.______ discoverEsVersion( ______ InitializationUtils.java:84)
> > > at
> org.elasticsearch.hadoop.hive.______EsStorageHandler.init(______EsStorageHandler.java:99)
> 
> > > ... 18 more
> > > Caused by: java.net.ConnectException: Connection refused
> > > at java.net.PlainSocketImpl. ______ socketConnect(Native Method)
> > > at java.net <http://java.net> <http://java.net>
> 
> <http://java.net>. ______AbstractPlainSocketImpl.______ doConnect( ______AbstractPlainSocketImpl.java:______ 339)
> > > at java.net <http://java.net> <http://java.net>
> 
> <http://java.net>. ______AbstractPlainSocketImpl.______ connectToAddress( ______AbstractPlainSocketImpl.java:______ 200)
> > > at java.net <http://java.net> <http://java.net>
> <http://java.net>. ______AbstractPlainSocketImpl.______ connect( ______AbstractPlainSocketImpl.java:______ 182)
> > > at java.net.SocksSocketImpl.______connect(SocksSocketImpl.java:______391)
> > > at java.net.Socket.connect( ______ Socket.java:579)
> > > at java.net.Socket.connect( ______ Socket.java:528)
> > > at java.net.Socket.<init>(Socket. ______ java:425)
> > > at java.net.Socket.<init>(Socket. ______ java:280)
> > > at
> 
> org.apache.commons.httpclient. ______protocol.______ DefaultProtocolSocketFactory.______createSocket(______DefaultProtocolSocketFactory. ______ java:80)
> > > at
> 
> org.apache.commons.httpclient. ______protocol.______ DefaultProtocolSocketFactory.______createSocket(______DefaultProtocolSocketFactory. ______ java:122)
> > > at
> org.apache.commons.httpclient.______HttpConnection.open(______HttpConnection.java:707)
> > > at
> 
> org.apache.commons.httpclient. ______HttpMethodDirector.______ executeWithRetry( ______ HttpMethodDirector.java:387)
> > > at
> 
> org.apache.commons.httpclient. ______HttpMethodDirector.______ executeMethod( ______ HttpMethodDirector.java:171)
> > > at
> org.apache.commons.httpclient.______HttpClient.executeMethod(______HttpClient.java:397)
> > > at
> org.apache.commons.httpclient.______HttpClient.executeMethod(______HttpClient.java:323)
> > > at
> 
> org.elasticsearch.hadoop.rest. ______commonshttp.______ CommonsHttpTransport.execute( ______CommonsHttpTransport.java:____ 160)
> > > at
> org.elasticsearch.hadoop.rest.______NetworkClient.execute(______NetworkClient.java:74)
> 
> > > ... 25 more
> > >
> > > Let me know if there's anything in particular you'd like me to try on EC2.
> > >
> > > (For posterity, the versions I used were: hadoop 2.2.0, hive 0.9.0, shark 8.1,
> spark 8.1,
> es-hadoop
> 1.3.0.M2, java
> > > 1.7.0_15, scala 2.9.3, elasticsearch 1.0.0)
> > >
> > > Thanks again,
> > > Max
> > >
> > > On Tuesday, February 18, 2014 10:16:38 PM UTC-8, Costin Leau wrote:
> > >
> > > The error indicates a network error - namely es-hadoop cannot connect to
> Elasticsearch
> on the
> default (localhost:9200)
> > > HTTP port. Can you double check whether that's indeed the case (using curl or
> even
> telnet on
> that port) - maybe the
> > > firewall prevents any connections to be made...
> > > Also you could try using the latest Hive, 0.12 and a more recent Hadoop such
> as 1.1.2
> or 1.2.1.
> > >
> > > Additionally, can you enable TRACE logging in your job on es-hadoop packages
> org.elasticsearch.hadoop.rest and
> > >org.elasticsearch.hadoop.mr <http://org.elasticsearch.hadoop.mr>
> <http://org.elasticsearch.__hadoop.mr <http://org.elasticsearch.hadoop.mr>>
> <http://org.elasticsearch. __ha__ doop.mr <http://hadoop.mr> <http://org.elasticsearch.__hadoop.mr
> <http://org.elasticsearch.hadoop.mr>>>
> <http://org.elasticsearch. __ha____ doop.mr <http://ha__doop.mr> <http://hadoop.mr>
> <http://org.elasticsearch. __ha__ doop.mr <http://hadoop.mr>
> <http://org.elasticsearch.__hadoop.mr <http://org.elasticsearch.hadoop.mr>>>>
> <http://org.elasticsearch. __ha____ doop.mr <http://ha__doop.mr> <http://hadoop.mr>
> <http://org.elasticsearch. __ha__ doop.mr <http://hadoop.mr>
> <http://org.elasticsearch.__hadoop.mr <http://org.elasticsearch.hadoop.mr>>>
> <http://org.elasticsearch. __ha____ doop.mr <http://ha__doop.mr> <http://hadoop.mr>
> <http://org.elasticsearch. __ha__ doop.mr <http://hadoop.mr>
> <http://org.elasticsearch.__hadoop.mr <http://org.elasticsearch.hadoop.mr>>>>>
> <http://org.elasticsearch. __ha____ doop.mr <http://ha__doop.mr> <http://hadoop.mr>
> <http://org.elasticsearch. __ha__ doop.mr <http://hadoop.mr>
> <http://org.elasticsearch.__hadoop.mr <http://org.elasticsearch.hadoop.mr>>>
> <http://org.elasticsearch. __ha____ doop.mr <http://ha__doop.mr> <http://hadoop.mr>
> <http://org.elasticsearch. __ha__ doop.mr <http://hadoop.mr>
> <http://org.elasticsearch.__hadoop.mr <http://org.elasticsearch.hadoop.mr>>>>
> 
> > <http://org.elasticsearch. __ha____ doop.mr <http://ha__doop.mr> <http://hadoop.mr>
> <http://org.elasticsearch. __ha__ doop.mr <http://hadoop.mr> <http://org.elasticsearch.__hadoop.mr
> <http://org.elasticsearch.hadoop.mr>>>
> <http://org.elasticsearch. __ha____ doop.mr <http://ha__doop.mr> <http://hadoop.mr>
> <http://org.elasticsearch. __ha__ doop.mr <http://hadoop.mr>
> <http://org.elasticsearch.__hadoop.mr <http://org.elasticsearch.hadoop.mr>>>>>> packages and report back ?
> 
> > >
> > > Thanks,
> > >
> > > On 19/02/2014 4:03 AM, Max Lang wrote:
> > > > I set everything up using this
> guide:https://github.com/ ______amplab/shark/wiki/Running-______ Shark-on-EC2
> <https://github.com/ ____amplab/shark/wiki/Running-____ Shark-on-EC2>
> <https://github.com/ __amplab/__ shark/wiki/Running- __Shark-on-__ EC2
> <https://github.com/ __amplab/shark/wiki/Running-__ Shark-on-EC2>>
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>>
> <https://github.com/amplab/ ______shark/wiki/Running-Shark-on-______ EC2
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2>
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>>
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>>>
> <https://github.com/amplab/ ______shark/wiki/Running-Shark-on-______ EC2
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2>
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>>
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>>
> <https://github.com/amplab/ ______shark/wiki/Running-Shark-on-______ EC2
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2>
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>>
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>>>>
> > > <https://github.com/amplab/ ______shark/wiki/Running-Shark-on-______ EC2
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2>
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>>
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>>
> <https://github.com/amplab/ ______shark/wiki/Running-Shark-on-______ EC2
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2>
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>>
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>>>
> > <https://github.com/amplab/ ______shark/wiki/Running-Shark-on-______ EC2
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2>
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>>
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>>
> <https://github.com/amplab/ ______shark/wiki/Running-Shark-on-______ EC2
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2>
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>>
> 
> <https://github.com/amplab/ ____shark/wiki/Running-Shark-on-____ EC2
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2>
> <https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2
> <https://github.com/amplab/shark/wiki/Running-Shark-on-EC2>>>>>> on an ec2 cluster. I've
> > > > copied the elasticsearch-hadoop jars into the hive lib directory and I have
> elasticsearch
> running on localhost:9200. I'm
> > > > running shark in a screen session with --service screenserver and
> connecting to it
> at the
> same time using shark -h
> > > > localhost.
> > > >
> > > > Unfortunately, when I attempt to write data into elasticsearch, it fails.
> Here's an
> example:
> > > >
> > > > |
> > > > [localhost:10000]shark>CREATE EXTERNAL TABLE wiki (id BIGINT,title
> STRING,last_modified
> STRING,xml STRING,text
> > > > STRING)ROW FORMAT DELIMITED FIELDS TERMINATED BY '\t'LOCATION
> 's3n://spark-data/wikipedia- ______ sample/';
> 
> > > > Timetaken (including network latency):0.159seconds
> > > > 14/02/1901:23:33INFO CliDriver:Timetaken (including network
> latency):0.159seconds
> > > >
> > > > [localhost:10000]shark>SELECT title FROM wiki LIMIT 1;
> > > > Alpokalja
> > > > Timetaken (including network latency):2.23seconds
> > > > 14/02/1901:23:48INFO CliDriver:Timetaken (including network
> latency):2.23seconds
> > > >
> > > > [localhost:10000]shark>CREATE EXTERNAL TABLE es_wiki (id BIGINT,title
> STRING,last_modified
> STRING,xml STRING,text
> > > > STRING)STORED BY
> 
> 'org.elasticsearch.hadoop. ______hive.EsStorageHandler'______ TBLPROPERTIES('es.resource'=' ______ wikipedia/article');
> 
> > > > Timetaken (including network latency):0.061seconds
> > > > 14/02/1901:33:51INFO CliDriver:Timetaken (including network
> latency):0.061seconds
> > > >
> > > > [localhost:10000]shark>INSERT OVERWRITE TABLE es_wiki SELECTw.id
> <http://w.id>,w.title,w.last _______ modified,w.xml,w.text FROM wiki w;
> > > > [HiveError]:Queryreturned non-zero
> code:9,cause:FAILED: ______ ExecutionError,returncode
> -101fromshark.execution. ______ SparkTask
> 
> > > > Timetaken (including network latency):3.575seconds
> > > > 14/02/1901:34:42INFO CliDriver:Timetaken (including network
> latency):3.575seconds
> > > > |
> > > >
> > > > *The stack trace looks like this:*
> > > >
> > > > org.apache.hadoop.hive.ql. ______ metadata.HiveException
> (org.apache.hadoop.hive.ql. ______ metadata.HiveException: java.io.IOException:
> 
> > > > Out of nodes and retries; caught exception)
> > > >
> > > >
> 
> org.apache.hadoop.hive.ql. ______exec.FileSinkOperator.______ processOp(FileSinkOperator.______java:602)shark.execution.______FileSinkOperator$$anonfun$______processPartition$1.apply(______FileSinkOperator.scala:84) ______shark.execution.______ FileSinkOperator$$anonfun$______processPartition$1.apply(______FileSinkOperator.scala:81) ______scala.collection.Iterator$______ class.foreach(Iterator.scala:______772)scala.collection.____Iterator$__$anon$19.foreach(____Iterator.__scala:399)shark.____execution. __FileSinkOperator.______ processPartition(______FileSinkOperator.scala:81)______shark.execution. ______FileSinkOperator$.writeFiles$______ 1(FileSinkOperator.scala:__207)____shark.execution. ______FileSinkOperator$$anonfun$______ executeProcessFileSinkPartitio______n$1.apply(__FileSinkOperator. ____scala:__ 211)shark.execution. ______FileSinkOperator$$anonfun$______ executeProcessFileSinkPartitio______n$1.apply(__FileSinkOperator. ____scala:__ 211)org.apache.spark. ______ scheduler.ResultTask.
> 
> ```

runTask(\_\_\_\_\_\_ResultTask.scala:107)org.\_\_\_\_\_\_apache.spark.scheduler.Ta

> ```
> sk.____run(Task.scala:53)org.__apache. ____spark.executor.__ Executor$ ____ Task
> 
> Runner$$anonfun$run$1. __apply$____ mcV$sp(Executor.scala:__215)____org.apac
> 
> he.spa
> 
> rk.dep
> >
> > loy.Sp
> > >
> > >
> 
> arkHadoopUtil.runAsUser(______SparkHadoopUtil.scala:50)org.______apache.spark.executor.______Executor$TaskRunner.run(______Executor.scala:182)java.util. ______concurrent.____ ThreadPoolExecutor.______runWorker(ThreadPoolExecutor.______java:1145)java.util. ______concurrent.ThreadPoolExecutor$______ Worker.run( ____ThreadPoolExecutor.__ java:615)____java.lang.Thread.run(____Thread.__java:744
> 
> >
> > >
> > > > I should be using Hive 0.9.0, shark 0.8.1, elasticsearch 1.0.0, Hadoop
> 1.0.4, and
> java 1.7.0_51
> > > > Based on my cursory look at the hadoop and elasticsearch-hadoop sources, it
> looks
> like hive
> is just rethrowing an
> > > > IOException it's getting from Spark, and elasticsearch-hadoop is just
> hitting those
> exceptions.
> > > > I suppose my questions are: Does this look like an issue with my
> ES/elasticsearch-hadoop
> config? And has anyone gotten
> > > > elasticsearch working with Spark/Shark?
> > > > Any ideas/insights are appreciated.
> > > > Thanks,Max
> > > >
> > > > --
> > > > You received this message because you are subscribed to the Google Groups
> "elasticsearch" group.
> > > > To unsubscribe from this group and stop receiving emails from it, send an
> email to
> > > >elasticsearc...@googlegroups. ______ com <mailto:elasticsearc...@
> <mailto:elasticsearc...@> __goog__ legroups.com <http://googlegroups.com>
> <mailto:elasticsearc...@__googlegroups.com <mailto:elasticsearc...@googlegroups.com>>> <javascript:>.
> 
> > > > To view this discussion on the web visit
> > >
> 
> >https://groups.google.com/d/ ______msgid/elasticsearch/__ 9486faff- ____3eaf-4344-8931-__ 3121bbc5d9c7% ______40googlegroups.com <https://groups.google.com/d/____ msgid/elasticsearch/9486faff- ____3eaf-4344-8931-3121bbc5d9c7%____ 40googlegroups.com> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____40googlegroups.com <https://groups.google.com/d/__ msgid/elasticsearch/9486faff- __3eaf-4344-8931-3121bbc5d9c7%__ 40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>>
> 
> <https://groups.google.com/d/ ______msgid/elasticsearch/__ 9486faff- ____3eaf-4344-8931-__ 3121bbc5d9c7% ______ 40googlegroups.com
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>>>
> >
> 
> <https://groups.google.com/d/ ______msgid/elasticsearch/__ 9486faff- ____3eaf-4344-8931-__ 3121bbc5d9c7% ______ 40googlegroups.com
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>>
> 
> <https://groups.google.com/d/ ______msgid/elasticsearch/__ 9486faff- ____3eaf-4344-8931-__ 3121bbc5d9c7% ______ 40googlegroups.com
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>>>>
> > >
> 
> <https://groups.google.com/d/ ______msgid/elasticsearch/__ 9486faff- ____3eaf-4344-8931-__ 3121bbc5d9c7% ______ 40googlegroups.com
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>>
> 
> <https://groups.google.com/d/ ______msgid/elasticsearch/__ 9486faff- ____3eaf-4344-8931-__ 3121bbc5d9c7% ______ 40googlegroups.com
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>>>
> >
> 
> <https://groups.google.com/d/ ______msgid/elasticsearch/__ 9486faff- ____3eaf-4344-8931-__ 3121bbc5d9c7% ______ 40googlegroups.com
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>>
> 
> <https://groups.google.com/d/ ______msgid/elasticsearch/__ 9486faff- ____3eaf-4344-8931-__ 3121bbc5d9c7% ______ 40googlegroups.com
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/9486faff-____ 3eaf-4344-8931-3121bbc5d9c7% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/9486faff-__ 3eaf-4344-8931-3121bbc5d9c7%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/9486faff-3eaf-4344-8931-3121bbc5d9c7%40googlegroups.com>>>>>>.
> > > > For more options, visithttps://groups.google. ______ com/groups/opt_out
> <http://groups.google.com/ ____groups/opt_out <http://groups.google.com/__ groups/opt_out>
> <http://groups.google.com/__groups/opt_out <http://groups.google.com/groups/opt_out>>>
> <http://groups.google.com/ ______groups/opt_out <http://groups.google.com/____ groups/opt_out>
> <http://groups.google.com/ ____groups/opt_out <http://groups.google.com/__ groups/opt_out>>
> <http://groups.google.com/ ____groups/opt_out <http://groups.google.com/__ groups/opt_out>
> <http://groups.google.com/__groups/opt_out <http://groups.google.com/groups/opt_out>>>>
> <http://groups.google.com/ ______groups/opt_out <http://groups.google.com/____ groups/opt_out>
> <http://groups.google.com/ ____groups/opt_out <http://groups.google.com/__ groups/opt_out>>
> <http://groups.google.com/ ____groups/opt_out <http://groups.google.com/__ groups/opt_out>
> <http://groups.google.com/__groups/opt_out <http://groups.google.com/groups/opt_out>>>
> 
> <http://groups.google.com/ ______ groups/opt_out
> <http://groups.google.com/ ____groups/opt_out> <http://groups.google.com/____ groups/opt_out
> <http://groups.google.com/__groups/opt_out>>
> <http://groups.google.com/ ____groups/opt_out <http://groups.google.com/__ groups/opt_out>
> <http://groups.google.com/__groups/opt_out <http://groups.google.com/groups/opt_out>>>>>
> <https://groups.google.com/ ______groups/opt_out <https://groups.google.com/____ groups/opt_out>
> <https://groups.google.com/ ____groups/opt_out <https://groups.google.com/__ groups/opt_out>>
> <https://groups.google.com/ ____groups/opt_out <https://groups.google.com/__ groups/opt_out>
> <https://groups.google.com/__groups/opt_out <https://groups.google.com/groups/opt_out>>>
> <https://groups.google.com/ ______ groups/opt_out
> <https://groups.google.com/ ____groups/opt_out> <https://groups.google.com/____ groups/opt_out
> <https://groups.google.com/__groups/opt_out>>
> <https://groups.google.com/ ____groups/opt_out <https://groups.google.com/__ groups/opt_out>
> <https://groups.google.com/__groups/opt_out <https://groups.google.com/groups/opt_out>>>>
> > <https://groups.google.com/ ______ groups/opt_out
> <https://groups.google.com/ ____groups/opt_out> <https://groups.google.com/____ groups/opt_out
> <https://groups.google.com/__groups/opt_out>>
> <https://groups.google.com/ ____groups/opt_out <https://groups.google.com/__ groups/opt_out>
> <https://groups.google.com/__groups/opt_out <https://groups.google.com/groups/opt_out>>>
> <https://groups.google.com/ ______groups/opt_out <https://groups.google.com/____ groups/opt_out>
> <https://groups.google.com/ ____groups/opt_out <https://groups.google.com/__ groups/opt_out>>
> <https://groups.google.com/ ____groups/opt_out <https://groups.google.com/__ groups/opt_out>
> <https://groups.google.com/__groups/opt_out <https://groups.google.com/groups/opt_out>>>>>>.
> 
> > >
> > > --
> > > Costin
> > >
> > > --
> > > You received this message because you are subscribed to the Google Groups
> "elasticsearch"
> group.
> > > To unsubscribe from this group and stop receiving emails from it, send an email to
> > >elasticsearc...@googlegroups. ______ com <mailto:elasticsearc...@
> <mailto:elasticsearc...@> __goog__ legroups.com <http://googlegroups.com>
> <mailto:elasticsearc...@__googlegroups.com <mailto:elasticsearc...@googlegroups.com>>> <javascript:>.
> 
> > > To view this discussion on the web visit
> >
> 
> >https://groups.google.com/d/ ______msgid/elasticsearch/__ 86187c3a- ____0974-4d10-9689-__ e83da788c04a% ______40googlegroups.com <https://groups.google.com/d/____ msgid/elasticsearch/86187c3a- ____0974-4d10-9689-e83da788c04a%____ 40googlegroups.com> <https://groups.google.com/d/ ____msgid/elasticsearch/86187c3a-____ 0974-4d10-9689-e83da788c04a% ____40googlegroups.com <https://groups.google.com/d/__ msgid/elasticsearch/86187c3a- __0974-4d10-9689-e83da788c04a%__ 40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/86187c3a-____ 0974-4d10-9689-e83da788c04a% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com>>>
> 
> <https://groups.google.com/d/ ______msgid/elasticsearch/__ 86187c3a- ____0974-4d10-9689-__ e83da788c04a% ______ 40googlegroups.com
> <https://groups.google.com/d/ ____msgid/elasticsearch/86187c3a-____ 0974-4d10-9689-e83da788c04a% ____ 40googlegroups.com>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/86187c3a-____ 0974-4d10-9689-e83da788c04a% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/86187c3a-____ 0974-4d10-9689-e83da788c04a% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com>>>>
> >
> 
> <https://groups.google.com/d/ ______msgid/elasticsearch/__ 86187c3a- ____0974-4d10-9689-__ e83da788c04a% ______ 40googlegroups.com
> <https://groups.google.com/d/ ____msgid/elasticsearch/86187c3a-____ 0974-4d10-9689-e83da788c04a% ____ 40googlegroups.com>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/86187c3a-____ 0974-4d10-9689-e83da788c04a% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/86187c3a-____ 0974-4d10-9689-e83da788c04a% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com>>>
> 
> <https://groups.google.com/d/ ______msgid/elasticsearch/__ 86187c3a- ____0974-4d10-9689-__ e83da788c04a% ______ 40googlegroups.com
> <https://groups.google.com/d/ ____msgid/elasticsearch/86187c3a-____ 0974-4d10-9689-e83da788c04a% ____ 40googlegroups.com>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/86187c3a-____ 0974-4d10-9689-e83da788c04a% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/86187c3a-____ 0974-4d10-9689-e83da788c04a% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/86187c3a-__ 0974-4d10-9689-e83da788c04a%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/86187c3a-0974-4d10-9689-e83da788c04a%40googlegroups.com>>>>>.
> > > For more options, visithttps://groups.google. ______ com/groups/opt_out
> <http://groups.google.com/ ____groups/opt_out <http://groups.google.com/__ groups/opt_out>
> <http://groups.google.com/__groups/opt_out <http://groups.google.com/groups/opt_out>>>
> <http://groups.google.com/ ______groups/opt_out <http://groups.google.com/____ groups/opt_out>
> <http://groups.google.com/ ____groups/opt_out <http://groups.google.com/__ groups/opt_out>>
> <http://groups.google.com/ ____groups/opt_out <http://groups.google.com/__ groups/opt_out>
> <http://groups.google.com/__groups/opt_out <http://groups.google.com/groups/opt_out>>>>
> <https://groups.google.com/ ______groups/opt_out <https://groups.google.com/____ groups/opt_out>
> <https://groups.google.com/ ____groups/opt_out <https://groups.google.com/__ groups/opt_out>>
> <https://groups.google.com/ ____groups/opt_out <https://groups.google.com/__ groups/opt_out>
> <https://groups.google.com/__groups/opt_out <https://groups.google.com/groups/opt_out>>>
> <https://groups.google.com/ ______ groups/opt_out
> <https://groups.google.com/ ____groups/opt_out> <https://groups.google.com/____ groups/opt_out
> <https://groups.google.com/__groups/opt_out>>
> <https://groups.google.com/ ____groups/opt_out <https://groups.google.com/__ groups/opt_out>
> <https://groups.google.com/__groups/opt_out <https://groups.google.com/groups/opt_out>>>>>.
> 
> >
> > --
> > Costin
> >
> > --
> > You received this message because you are subscribed to the Google Groups
> "elasticsearch" group.
> > To unsubscribe from this group and stop receiving emails from it, send an email to
> >elasticsearc...@googlegroups. ______ com <mailto:elasticsearc...@
> <mailto:elasticsearc...@> __goog__ legroups.com <http://googlegroups.com>
> <mailto:elasticsearc...@__googlegroups.com <mailto:elasticsearc...@googlegroups.com>>> <javascript:>.
> 
> > To view this discussion on the web visit
> 
> >https://groups.google.com/d/ ______msgid/elasticsearch/__ e29e342d- ____de74-4ed6-93d4-__ 875fc728c5a5% ______40googlegroups.com <https://groups.google.com/d/____ msgid/elasticsearch/e29e342d- ____de74-4ed6-93d4-875fc728c5a5%____ 40googlegroups.com> <https://groups.google.com/d/ ____msgid/elasticsearch/e29e342d-____ de74-4ed6-93d4-875fc728c5a5% ____40googlegroups.com <https://groups.google.com/d/__ msgid/elasticsearch/e29e342d- __de74-4ed6-93d4-875fc728c5a5%__ 40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/e29e342d-____ de74-4ed6-93d4-875fc728c5a5% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/e29e342d-__ de74-4ed6-93d4-875fc728c5a5%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/e29e342d-__ de74-4ed6-93d4-875fc728c5a5%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/e29e342d-de74-4ed6-93d4-875fc728c5a5%40googlegroups.com>>>
> 
> <https://groups.google.com/d/ ______msgid/elasticsearch/__ e29e342d- ____de74-4ed6-93d4-__ 875fc728c5a5% ______ 40googlegroups.com
> <https://groups.google.com/d/ ____msgid/elasticsearch/e29e342d-____ de74-4ed6-93d4-875fc728c5a5% ____ 40googlegroups.com>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/e29e342d-____ de74-4ed6-93d4-875fc728c5a5% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/e29e342d-__ de74-4ed6-93d4-875fc728c5a5%__40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/e29e342d-____ de74-4ed6-93d4-875fc728c5a5% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/e29e342d-__ de74-4ed6-93d4-875fc728c5a5%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/e29e342d-__ de74-4ed6-93d4-875fc728c5a5%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/e29e342d-de74-4ed6-93d4-875fc728c5a5%40googlegroups.com>>>>.
> 
> > For more options, visithttps://groups.google. ______ com/groups/opt_out
> <http://groups.google.com/ ____groups/opt_out <http://groups.google.com/__ groups/opt_out>
> <http://groups.google.com/__groups/opt_out <http://groups.google.com/groups/opt_out>>>
> <https://groups.google.com/ ______groups/opt_out <https://groups.google.com/____ groups/opt_out>
> <https://groups.google.com/ ____groups/opt_out <https://groups.google.com/__ groups/opt_out>>
> 
> <https://groups.google.com/ ____groups/opt_out <https://groups.google.com/__ groups/opt_out>
> <https://groups.google.com/__groups/opt_out <https://groups.google.com/groups/opt_out>>>>.
> 
> --
> Costin
> 
> --
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.
> To unsubscribe from this group and stop receiving emails from it, send an email to
> elasticsearch+unsubscribe@ __go____ oglegroups.com <http://go__oglegroups.com>
> <http://googlegroups.com>
> <mailto:elasticsearch% ____ 2Bunsubscribe@googlegroups.com
> <mailto:elasticsearch%25__2Bunsubscribe@googlegroups.com>
> <mailto:elasticsearch% __252Bunsubscribe@googlegroups.__ com
> <mailto:elasticsearch%25252Bunsubscribe@googlegroups.com>>__>
> <mailto:elasticsearch+ ______ unsubscribe@googlegroups.com
> <mailto:elasticsearch%2B ____ unsubscribe@googlegroups.com>
> <mailto:elasticsearch%2B ____ unsubscribe@googlegroups.com
> <mailto:elasticsearch%252B__unsubscribe@googlegroups.com>>
> <mailto:elasticsearch% ____2Bunsubscribe@googlegroups.com <mailto:elasticsearch%25__ 2Bunsubscribe@googlegroups.com>
> <mailto:elasticsearch% __252Bunsubscribe@googlegroups.__ com
> <mailto:elasticsearch%25252Bunsubscribe@googlegroups.com>>__>>.
> 
> To view this discussion on the web visit
> https://groups.google.com/d/ ______msgid/elasticsearch/__ c1081bf2- ____117a-4af2-ba90-__ 2c38a4572782% ______ 40googlegroups.com
> <https://groups.google.com/d/ ____msgid/elasticsearch/c1081bf2-____ 117a-4af2-ba90-2c38a4572782% ____ 40googlegroups.com>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/c1081bf2-____ 117a-4af2-ba90-2c38a4572782% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/c1081bf2-__ 117a-4af2-ba90-2c38a4572782%__40googlegroups.com>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/c1081bf2-____ 117a-4af2-ba90-2c38a4572782% ____ 40googlegroups.com
> <https://groups.google.com/d/ __msgid/elasticsearch/c1081bf2-__ 117a-4af2-ba90-2c38a4572782%__40googlegroups.com>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/c1081bf2-__ 117a-4af2-ba90-2c38a4572782%__40googlegroups.com
> <https://groups.google.com/d/msgid/elasticsearch/c1081bf2-117a-4af2-ba90-2c38a4572782%40googlegroups.com>>>
> 
> <https://groups.google.com/d/ ______msgid/elasticsearch/__ c1081bf2- ____117a-4af2-ba90-__ 2c38a4572782% ______40googlegroups.com?utm_____ medium= __email&utm_source=____ footer
> <https://groups.google.com/d/ ____msgid/elasticsearch/c1081bf2-____ 117a-4af2-ba90-2c38a4572782% ____40googlegroups.com?utm___ medium= __email&utm_source=__ footer>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/c1081bf2-____ 117a-4af2-ba90-2c38a4572782% ____40googlegroups.com?utm___ medium= __email&utm_source=__ footer
> <https://groups.google.com/d/ __msgid/elasticsearch/c1081bf2-__ 117a-4af2-ba90-2c38a4572782% __40googlegroups.com?utm_medium=__ email&utm_source=footer>>
> 
> <https://groups.google.com/d/ ____msgid/elasticsearch/c1081bf2-____ 117a-4af2-ba90-2c38a4572782% ____40googlegroups.com?utm___ medium= __email&utm_source=__ footer
> <https://groups.google.com/d/ __msgid/elasticsearch/c1081bf2-__ 117a-4af2-ba90-2c38a4572782% __40googlegroups.com?utm_medium=__ email&utm_source=footer>
> 
> <https://groups.google.com/d/ __msgid/elasticsearch/c1081bf2-__ 117a-4af2-ba90-2c38a4572782% __40googlegroups.com?utm_medium=__ email&utm_source=footer
> <https://groups.google.com/d/msgid/elasticsearch/c1081bf2-117a-4af2-ba90-2c38a4572782%40googlegroups.com?utm_medium=email&utm_source=footer>>>>.
> 
> For more options, visit https://groups.google.com/d/ ______ optout
> <https://groups.google.com/d/ ____optout> <https://groups.google.com/d/____ optout
> <https://groups.google.com/d/__optout>>
> <https://groups.google.com/d/ ____optout <https://groups.google.com/d/__ optout>
> <https://groups.google.com/d/__optout <https://groups.google.com/d/optout>>>.
> 
> --
> Costin
> 
> --
> 
> You received this message because you are subscribed to a topic in the Google Groups
> "elasticsearch" group.
> To unsubscribe from this topic, visit
> https://groups.google.com/d/ ______topic/elasticsearch/S-______ BrzwUHJbM/unsubscribe
> <https://groups.google.com/d/ ____topic/elasticsearch/S-____ BrzwUHJbM/unsubscribe>
> <https://groups.google.com/d/ ____topic/elasticsearch/S-____ BrzwUHJbM/unsubscribe
> <https://groups.google.com/d/ __topic/elasticsearch/S-__ BrzwUHJbM/unsubscribe>>
> <https://groups.google.com/d/ ____topic/elasticsearch/S-____ BrzwUHJbM/unsubscribe
> <https://groups.google.com/d/ __topic/elasticsearch/S-__ BrzwUHJbM/unsubscribe>
> <https://groups.google.com/d/ __topic/elasticsearch/S-__ BrzwUHJbM/unsubscribe
> <https://groups.google.com/d/topic/elasticsearch/S-BrzwUHJbM/unsubscribe>>>.
> 
> ```
> 
> --  
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an email to  
> [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com) [mailto:elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GOd-wSu\_XQpc\_cjCcv\_ZgdiwEoJ6BT6VCAkL4ThvOHTAw%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GOd-wSu_XQpc_cjCcv_ZgdiwEoJ6BT6VCAkL4ThvOHTAw%40mail.gmail.com)  
> [https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GOd-wSu\_XQpc\_cjCcv\_ZgdiwEoJ6BT6VCAkL4ThvOHTAw%40mail.gmail.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GOd-wSu_XQpc_cjCcv_ZgdiwEoJ6BT6VCAkL4ThvOHTAw%40mail.gmail.com?utm_medium=email&utm_source=footer).  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
Costin

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/537274ED.3010203%40gmail.com](https://groups.google.com/d/msgid/elasticsearch/537274ED.3010203%40gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Nick\_Pentreath](https://avatars.discourse-cdn.com/v4/letter/n/7ea924/32.png) [@Nick\_Pentreath](https://discuss.elastic.co/u/Nick_Pentreath)\
**Post date:** [May 16, 2014, 5:02pm UTC](https://discuss.elastic.co/t/hadoop-getting-elasticsearch-hadoop-working-with-shark/15882/15 "2014-05-16T17:02:08Z")

</div>

Hi Costin

Here is some info on setup, versions, and the simplified version of code /  
data: [Elasticsearch / Spark / Shark setup · GitHub](https://gist.github.com/MLnick/bb7c6e87f5c53be2cce4)

I am using both Spark (the ESInputFormat directly) and Shark (ES hive jar  
bundle). Spark works fine but does return Text (as opposed to Long or  
whatever), while Shark returns NULLS (with the debug info as per below).

Hopefully this helps, ping me if you need more info ?

N

On Tue, May 13, 2014 at 9:39 PM, Costin Leau [costin.leau@gmail.com](mailto:costin.leau@gmail.com) wrote:

> Could you share your setup and configuration on a gist (the more info  
> especially regarding the versions of the stack used helps)?  
> Do you use just the input format or also the output format? To clarify -  
> are you using Spark (Map/Reduce) or Shark (and the relevant  
> Hive integration in es-hadoop)?
> 
> Cheers,
> 
> On 5/13/14 10:14 PM, Nick Pentreath wrote:
> 
> > Ok - well let me know when you're around.
> > 
> > The mapreduce inputformat works fine. I'm using it with Spark to access  
> > the ES data via ESInputFormat and run analytics  
> > and machine learning jobs on that data, and the same \_ts field works and  
> > is the correct data (though it comes through as  
> > org.apache.hadoop.io.Text, which I convert to Long or a DateTime as  
> > required).
> > 
> > Perhaps I'm missing it somewhere but is it possible to force a field to  
> > be a type? i.e. similar the es.field.mapping  
> > could I tell it that it must parse the field as a string (since then I  
> > can take it and do whatever parsing / casting I  
> > want).
> > 
> > I could just use the new Spark SQL module (which I'm seriously  
> > considering right now having explored it a bit in the  
> > last few days), but some of the stuff we do requires a SQL Console and  
> > JDBC, so having Shark able to just pull in ES  
> > data is definitely very useful...
> > 
> > On Tue, May 13, 2014 at 8:18 PM, Costin Leau \<[costin.leau@gmail.com](mailto:costin.leau@gmail.com)\<mailto:  
> > [costin.leau@gmail.com](mailto:costin.leau@gmail.com)\>\> wrote:
> > 
> > ```
> > Hi Nick,
> > 
> > I'm glad to see you are making progress. This week I'm mainly on the
> > 
> > ```
> > 
> > road but maybe we can meet on the IRC next  
> > week, my invitation still stands 🙂  
> > Timestamp is relatively new type and doesn't handle timezones  
> > properly - it is backed by java.sq.Timestamp so it  
> > inherits a lot of its issues.  
> > For some reason the year in your date is rather off so it's worth  
> > checking the data read by es-hadoop before passing  
> > it to Hive (see [1]).  
> > I've had issues myself with it and it the moment the cluster is in a  
> > different timezone than the dataset itself  
> > things get buggy.  
> > Try using a UDF to do the conversion from the long to a timestamp -  
> > I've tried doing something similar in our  
> > conversion but since we don't know the timezones  
> > used, it's easy for things to get mixed.
> > 
> > ```
> > Cheers,
> > 
> > [1] http://www.elasticsearch.org/__guide/en/elasticsearch/
> > 
> > ```
> > 
> > hadoop/\_\_current/troubleshooting.html
> > 
> > ```
> > <http://www.elasticsearch.org/guide/en/elasticsearch/hadoop/
> > 
> > ```
> > 
> > current/troubleshooting.html\>
> > 
> > ```
> > On 5/13/14 8:25 PM, Nick Pentreath wrote:
> > 
> > Hi Costin
> > 
> > Sorry for the silence on this issue. This went a bit quiet.
> > 
> > But the good news is I've come back to it and managed to get it
> > 
> > ```
> > 
> > all working with the new shark 0.9.1 release and  
> > 2.0.0RC1. Actually if I used ADD JAR I got the same exception but  
> > when I just put the JAR into the shark lib/  
> > folder it  
> > worked fine (which seems to point to the classpath issue you  
> > mention).
> > 
> > ```
> > However, I seem to have an issue with date <-> timestamp
> > 
> > ```
> > 
> > conversion.
> > 
> > ```
> > I have a field in ES called "_ts" that has type "date" and the
> > 
> > ```
> > 
> > default format "dateOptionalTime". When I do a  
> > query that  
> > includes the timestamp it comes back NULL:
> > 
> > ```
> > select ts from table ...
> > (note I use a correct es.mapping.names to map the _ts field in ES
> > 
> > ```
> > 
> > to ts field in Hive/Shark that has timestamp  
> > type).
> > 
> > ```
> > below is some of the debug-level output:
> > 
> > 14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the
> > 
> > ```
> > 
> > TIMESTAMP data type range so converted to null.  
> > Given data  
> > is :96997506-06-30 19:08:168:16.768  
> > 14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the  
> > TIMESTAMP data type range so converted to null.  
> > Given data  
> > is :96997605-06-28 19:08:168:16.768  
> > 14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the  
> > TIMESTAMP data type range so converted to null.  
> > Given data  
> > is :96997624-06-28 19:08:168:16.768  
> > 14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the  
> > TIMESTAMP data type range so converted to null.  
> > Given data  
> > is :96997629-06-28 19:08:168:16.768  
> > 14/05/13 19:19:47 DEBUG lazy.LazyPrimitive: Data not in the  
> > TIMESTAMP data type range so converted to null.  
> > Given data  
> > is :96997634-06-29 19:08:168:16.768  
> > NULL  
> > NULL  
> > NULL  
> > NULL  
> > NULL
> > 
> > ```
> > The data that I index in the _ts field is timestamp in ms (long).
> > 
> > ```
> > 
> > It doesn't seem to be converted correctly but  
> > the data  
> > is correct (in ms at least) and I can query against it using date  
> > formats and date math in ES.
> > 
> > ```
> > Example snippet from debug log from above:
> > ,"_ts":1397130475607}}]}}"
> > 
> > Any ideas or am I doing something silly?
> > 
> > I do see that the Hive timestamp expects either seconds since
> > 
> > ```
> > 
> > epoch of a string-based format that has nanosecond  
> > granularity. Is this the issue with just ms long timestamp data?
> > 
> > ```
> > Thanks
> > Nick
> > 
> > On Thu, Mar 27, 2014 at 4:50 PM, Costin Leau <
> > 
> > ```
> > 
> > [costin.leau@gmail.com](mailto:costin.leau@gmail.com) [mailto:costin.leau@gmail.com](mailto:costin.leau@gmail.com)  
> > \<[mailto:costin.leau@gmail.com](mailto:costin.leau@gmail.com) [mailto:costin.leau@gmail.com](mailto:costin.leau@gmail.com)\>\_\_\>  
> > wrote:
> > 
> > ```
> > Using the latest hive and hadoop is preferred as they
> > 
> > ```
> > 
> > contain various bug fixes.  
> > The error suggests a classpath issue - namely the same class  
> > is loaded twice for some reason and hence the  
> > casting  
> > fails.
> > 
> > ```
> > Let's connect on IRC - give me a ping when you're available
> > 
> > ```
> > 
> > (user is costin).
> > 
> > ```
> > Cheers,
> > 
> > On 3/27/14 4:29 PM, Nick Pentreath wrote:
> > 
> > Thanks for the response.
> > 
> > I tried latest Shark (cdh4 version of 0.9.1 here
> > 
> > ```
> > 
> > [http://cloudera.rst.im/shark/](http://cloudera.rst.im/shark/) ) - this uses hadoop  
> > 1.0.4 and  
> > hive 0.11  
> > I believe, and build elasticsearch-hadoop from github  
> > master.
> > 
> > ```
> > Still getting same error:
> > org.elasticsearch.hadoop.hive. ____EsHiveInputFormat$__ EsHiveSplit
> > 
> > ```
> > 
> > cannot be cast to  
> > org.elasticsearch.hadoop.hive.\_\_ **EsHiveInputFormat$**  
> > EsHiveSplit
> > 
> > ```
> > Will using hive 0.11 / hadoop 1.0.4 vs hive 0.12 /
> > 
> > ```
> > 
> > hadoop 1.2.1 in es-hadoop master make a difference?
> > 
> > ```
> > Anyone else actually got this working?
> > 
> > On Thu, Mar 20, 2014 at 2:44 PM, Costin Leau <
> > 
> > ```
> > 
> > [costin.leau@gmail.com](mailto:costin.leau@gmail.com) [mailto:costin.leau@gmail.com](mailto:costin.leau@gmail.com)  
> > \<[mailto:costin.leau@gmail.com](mailto:costin.leau@gmail.com) [mailto:costin.leau@gmail.com](mailto:costin.leau@gmail.com)\>  
> > \<[mailto:costin.leau@gmail.com](mailto:costin.leau@gmail.com) \<mailto:  
> > [costin.leau@gmail.com](mailto:costin.leau@gmail.com)\> \<[mailto:costin.leau@gmail.com](mailto:costin.leau@gmail.com)  
> > [mailto:costin.leau@gmail.com](mailto:costin.leau@gmail.com)\> **\>** \> wrote:
> > 
> > ```
> > I recommend using master - there are several
> > 
> > ```
> > 
> > improvements done in this area. Also using the latest  
> > Shark  
> > (0.9.0) and  
> > Hive (0.12) will help.
> > 
> > ```
> > On 3/20/14 12:00 PM, Nick Pentreath wrote:
> > 
> > Hi
> > 
> > I am struggling to get this working too. I'm
> > 
> > ```
> > 
> > just trying locally for now, running Shark 0.8.1,  
> > Hive  
> > 0.9.0 and ES  
> > 1.0.1  
> > with ES-hadoop 1.3.0.M2.
> > 
> > ```
> > I managed to get a basic example working with
> > 
> > ```
> > 
> > WRITING into an index. But I'm really after  
> > READING and  
> > index.
> > 
> > ```
> > I believe I have set everything up correctly,
> > 
> > ```
> > 
> > I've added the jar to Shark:  
> > ADD JAR /path/to/es-hadoop.jar;
> > 
> > ```
> > created a table:
> > CREATE EXTERNAL TABLE test_read (name string,
> > 
> > ```
> > 
> > price double)
> > 
> > ```
> > STORED BY 'org.elasticsearch.hadoop. ____
> > 
> > ```
> > 
> > \_\_hive.EsStorageHandler'
> > 
> > ```
> > TBLPROPERTIES('es.resource' =
> > 
> > ```
> > 
> > 'test\_index/test\_type/\_search?\_\_\_\_\_\_q=\*');
> > 
> > ```
> > And then trying to 'SELECT * FROM test _read'
> > 
> > ```
> > 
> > gives me :
> > 
> > ```
> > org.apache.spark. ______ SparkException: Job
> > 
> > ```
> > 
> > aborted: Task 3.0:0 failed more than 0 times;  
> > aborting job  
> > java.lang.ClassCastException:  
> > org.elasticsearch.hadoop.hive.\_\_\_\_\_\_EsHiveInputFormat$\_\_\_\_ESHiveSplit  
> > cannot  
> > be cast to  
> > org.elasticsearch.hadoop.hive.  
> > \_\_\_\_\_\_EsHiveInputFormat$\_\_\_\_ESHiveSplit
> > 
> > ```
> > at
> > org.apache.spark.scheduler. ______DAGScheduler$$anonfun$_____
> > 
> > ```
> > 
> > \_abortStage$1.apply(\_\_\_\_\_\_DAGScheduler.scala:827)
> > 
> > ```
> > at
> > org.apache.spark.scheduler. ______DAGScheduler$$anonfun$_____
> > 
> > ```
> > 
> > \_abortStage$1.apply(\_\_\_\_\_\_DAGScheduler.scala:825)
> > 
> > ```
> > at scala.collection.mutable. _____
> > 
> > ```
> > 
> > \_ResizableArray$class.foreach(\_\_\_\_\_\_ResizableArray.scala:60)
> > 
> > ```
> > at scala.collection.mutable. _____
> > 
> > ```
> > 
> > \_ArrayBuffer.foreach(\_\_\_\_\_\_ArrayBuffer.scala:47)
> > 
> > ```
> > at org.apache.spark.scheduler.___
> > 
> > ```
> > 
> > \_\_\_DAGScheduler.abortStage(\_\_\_\_\_\_DAGScheduler.scala:825)
> > 
> > ```
> > at org.apache.spark.scheduler.___
> > 
> > ```
> > 
> > \_\_\_DAGScheduler.processEvent(\_\_\_\_\_\_DAGScheduler.scala:440)
> > 
> > ```
> > at org.apache.spark.scheduler.___
> > 
> > ```
> > 
> > \_\_\_DAGScheduler.org  
> > \<[http://org.apache.spark](http://org.apache.spark).\_\_sch  
> > \_\_eduler.DAGScheduler.org [http://scheduler.DAGScheduler.org](http://scheduler.DAGScheduler.org)  
> > \<[http://org.apache.spark](http://org.apache.spark).\_\_scheduler.DAGScheduler.org  
> > [http://org.apache.spark.scheduler.DAGScheduler.org](http://org.apache.spark.scheduler.DAGScheduler.org)\>\>$  
> > \_\_\_\_apache$spark$**scheduler$DAGScheduler$$run(**  
> > DAGScheduler.scala:502)
> > 
> > ```
> > at org.apache.spark.scheduler.___
> > 
> > ```
> > 
> > \_\_\_DAGScheduler$$anon$1.run(\_\_\_\_\_\_DAGScheduler.scala:157)
> > 
> > ```
> > FAILED: Execution Error, return code -101 from
> > 
> > ```
> > 
> > shark.execution.SparkTask
> > 
> > ```
> > In fact I get the same error thrown when trying
> > 
> > ```
> > 
> > to READ from the table that I successfully  
> > WROTE to...
> > 
> > ```
> > On Saturday, 22 February 2014 12:31:21 UTC+2,
> > 
> > ```
> > 
> > Costin Leau wrote:
> > 
> > ```
> > Yeah, it might have been some sort of
> > 
> > ```
> > 
> > network configuration issue where services where  
> > running on  
> > different  
> > machines  
> > and  
> > localhost pointed to a different location.
> > 
> > ```
> > Either way, I'm glad to hear things have
> > 
> > ```
> > 
> > are moving forward.
> > 
> > ```
> > Cheers,
> > 
> > On 22/02/2014 1:06 AM, Max Lang wrote:
> > > I managed to get it working on ec2
> > 
> > ```
> > 
> > without issue this time. I'd say the biggest  
> > difference was  
> > that this  
> > time I set up a  
> > \> dedicated ES machine. Is it possible  
> > that, because I was using a cluster with slaves,  
> > when I used  
> > "localhost" the slaves  
> > \> couldn't find the ES instance running on  
> > the master? Or do all the requests go through  
> > the master?  
> > \>  
> > \>  
> > \> On Wednesday, February 19, 2014 2:35:40  
> > PM UTC-8, Costin Leau wrote:  
> > \>  
> > \> Hi,  
> > \>  
> > \> Setting logging in Hive/Hadoop can  
> > be tricky since the log4j needs to be picked up  
> > by the  
> > running JVM  
> > otherwise you  
> > \> won't see anything.  
> > \> Take a look at this link on how to  
> > tell Hive to use your logging settings [1].  
> > \>  
> > \> For the next release, we might  
> > introduce dedicated exceptions for the simple fact  
> > that some  
> > libraries, like Hive,  
> > \> swallow the stack trace and it's  
> > unclear what the issue is which makes the exception  
> > (IllegalStateException) ambiguous.  
> > \>  
> > \> Let me know how it goes and whether  
> > you will encounter any issues with Shark. Or if  
> > you don't 🙂  
> > \>  
> > \> Thanks!  
> > \>  
> > \>
> > 
> > ```
> > [1]https://cwiki.apache.org/ ______confluence/display/Hive/__
> > 
> > ```
> > 
> > \_\_\_\_GettingStarted#GettingStarted-\_\_\_\_\_\_ErrorLogs  
> > \<[https://cwiki.apache.org/\_\_\_\_confluence/display/Hive/\_\_\_\_](https://cwiki.apache.org/ ____confluence/display/Hive/____ )  
> > GettingStarted#GettingStarted-\_\_\_\_ErrorLogs\>
> > 
> > ```
> > <https://cwiki.apache.org/ ____
> > 
> > ```
> > 
> > confluence/display/Hive/\_\_\_\_GettingStarted#GettingStarted-\_\_\_\_ErrorLogs  
> > \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-\_\_ErrorLogs\>\>
> > 
> > ```
> > <https://cwiki.apache.org/ ____confluence/display/Hive/____
> > 
> > ```
> > 
> > GettingStarted#GettingStarted-\_\_\_\_ErrorLogs  
> > \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-\_\_ErrorLogs\>  
> > \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-\_\_ErrorLogs  
> > \<[Home - Apache Hive - Apache Software Foundation](https://cwiki.apache.org/confluence/display/Hive/)  
> > GettingStarted#GettingStarted-ErrorLogs\>\>\>
> > 
> > ```
> > <https://cwiki.apache.org/ ______confluence/display/Hive/____
> > 
> > ```
> > 
> > \_\_GettingStarted#GettingStarted-\_\_\_\_\_\_ErrorLogs  
> > \<[https://cwiki.apache.org/\_\_\_\_confluence/display/Hive/\_\_\_\_](https://cwiki.apache.org/ ____confluence/display/Hive/____ )  
> > GettingStarted#GettingStarted-\_\_\_\_ErrorLogs\>
> > 
> > ```
> > <https://cwiki.apache.org/ ____
> > 
> > ```
> > 
> > confluence/display/Hive/\_\_\_\_GettingStarted#GettingStarted-\_\_\_\_ErrorLogs  
> > \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-\_\_ErrorLogs\>\>
> > 
> > ```
> > <https://cwiki.apache.org/ ____confluence/display/Hive/____
> > 
> > ```
> > 
> > GettingStarted#GettingStarted-\_\_\_\_ErrorLogs  
> > \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-\_\_ErrorLogs\>  
> > \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-\_\_ErrorLogs  
> > \<[Home - Apache Hive - Apache Software Foundation](https://cwiki.apache.org/confluence/display/Hive/)  
> > GettingStarted#GettingStarted-ErrorLogs\>\>\>\>  
> > \>
> > 
> > ```
> > <https://cwiki.apache.org/ ______confluence/display/Hive/____
> > 
> > ```
> > 
> > \_\_GettingStarted#GettingStarted-\_\_\_\_\_\_ErrorLogs  
> > \<[https://cwiki.apache.org/\_\_\_\_confluence/display/Hive/\_\_\_\_](https://cwiki.apache.org/ ____confluence/display/Hive/____ )  
> > GettingStarted#GettingStarted-\_\_\_\_ErrorLogs\>
> > 
> > ```
> > <https://cwiki.apache.org/ ____
> > 
> > ```
> > 
> > confluence/display/Hive/\_\_\_\_GettingStarted#GettingStarted-\_\_\_\_ErrorLogs  
> > \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-\_\_ErrorLogs\>\>
> > 
> > ```
> > <https://cwiki.apache.org/ ____confluence/display/Hive/____
> > 
> > ```
> > 
> > GettingStarted#GettingStarted-\_\_\_\_ErrorLogs  
> > \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-\_\_ErrorLogs\>  
> > \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-\_\_ErrorLogs  
> > \<[Home - Apache Hive - Apache Software Foundation](https://cwiki.apache.org/confluence/display/Hive/)  
> > GettingStarted#GettingStarted-ErrorLogs\>\>\>
> > 
> > ```
> > <https://cwiki.apache.org/ ______confluence/display/Hive/____
> > 
> > ```
> > 
> > \_\_GettingStarted#GettingStarted-\_\_\_\_\_\_ErrorLogs  
> > \<[https://cwiki.apache.org/\_\_\_\_confluence/display/Hive/\_\_\_\_](https://cwiki.apache.org/ ____confluence/display/Hive/____ )  
> > GettingStarted#GettingStarted-\_\_\_\_ErrorLogs\>  
> > \<[https://cwiki.apache.org/\_\_\_\_](https://cwiki.apache.org/ ____ )  
> > confluence/display/Hive/\_\_\_\_GettingStarted#GettingStarted-\_\_\_\_ErrorLogs  
> > \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-\_\_ErrorLogs\>\>
> > 
> > ```
> > <https://cwiki.apache.org/ ____confluence/display/Hive/____
> > 
> > ```
> > 
> > GettingStarted#GettingStarted-\_\_\_\_ErrorLogs  
> > \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-**ErrorLogs\>  
> > \<[https://cwiki.apache.org/\_\_confluence/display/Hive/\_\_](https://cwiki.apache.org/ __confluence/display/Hive/__ )  
> > GettingStarted#GettingStarted-ErrorLogs  
> > \<[Home - Apache Hive - Apache Software Foundation](https://cwiki.apache.org/confluence/display/Hive/)  
> > GettingStarted#GettingStarted-ErrorLogs\>\>\>\>\>  
> > \>  
> > \> On 20/02/2014 12:02 AM, Max Lang  
> > wrote:  
> > \> \> Hey Costin,  
> > \> \>  
> > \> \> Thanks for the swift reply. I  
> > abandoned EC2 to take that out of the equation and  
> > managed  
> > to get  
> > everything working  
> > \> \> locally using the latest version  
> > of everything (though I realized just now I'm  
> > still on  
> > hive 0.9).  
> > I'm guessing you're  
> > \> \> right about some port connection  
> > issue because I definitely had ES running on  
> > that machine.  
> > \> \>  
> > \> \> I changed hive-log4j.properties  
> > and added  
> > \> \> |  
> > \> \> #custom logging levels  
> > \> \> #log4j.logger.xxx=DEBUG  
> > \> \> [log4j.logger.org](http://log4j.logger.org) \<  
> > [http://log4j.logger.org](http://log4j.logger.org)\> [http://log4j.logger.org](http://log4j.logger.org)  
> > [http://log4j.logger.org](http://log4j.logger.org).**  
> > \_\_elasticsearch.hadoop.rest=\_\_\_\_\_\_TRACE  
> > \> \>[log4j.logger.org](http://log4j.logger.org).\_\_elasticsea  
> > \_\_\_\_rch.hadoop.mr [http://elasticsea\_\_rch.hadoop.mr](http://elasticsea__rch.hadoop.mr)  
> > \<[http://elasticsearch.hadoop](http://elasticsearch.hadoop).\_\_mr \<[http://elasticsearch.hadoop.mr](http://elasticsearch.hadoop.mr)
> > 
> > > >
> > 
> > ```
> > <http://log4j.logger.org. __ela__ sticsearch.hadoop.mr <
> > 
> > ```
> > 
> > [http://elasticsearch.hadoop.mr](http://elasticsearch.hadoop.mr)\>  
> > \<[http://log4j.logger.org](http://log4j.logger.org).\_\_elasticsearch.hadoop.mr \<  
> > [http://log4j.logger.org.elasticsearch.hadoop.mr](http://log4j.logger.org.elasticsearch.hadoop.mr)\>\>\>  
> > \<[http://log4j.logger.org](http://log4j.logger.org).\_\_ela  
> > \_\_\_\_sticsearch.hadoop.mr [http://ela\_\_sticsearch.hadoop.mr](http://ela__sticsearch.hadoop.mr)  
> > \<[http://elasticsearch.hadoop](http://elasticsearch.hadoop).\_\_mr \<[http://elasticsearch.hadoop.mr](http://elasticsearch.hadoop.mr)
> > 
> > > >
> > 
> > ```
> > <http://log4j.logger.org. __ela__ sticsearch.hadoop.mr <
> > 
> > ```
> > 
> > [http://elasticsearch.hadoop.mr](http://elasticsearch.hadoop.mr)\>  
> > \<[http://log4j.logger.org](http://log4j.logger.org).\_\_elasticsearch.hadoop.mr \<  
> > [http://log4j.logger.org.elasticsearch.hadoop.mr](http://log4j.logger.org.elasticsearch.hadoop.mr)\>\>\>\>  
> > \<[http://log4j.logger.org](http://log4j.logger.org).\_\_ela  
> > \_\_\_\_sticsearch.hadoop.mr [http://ela\_\_sticsearch.hadoop.mr](http://ela__sticsearch.hadoop.mr)  
> > \<[http://elasticsearch.hadoop](http://elasticsearch.hadoop).\_\_mr \<[http://elasticsearch.hadoop.mr](http://elasticsearch.hadoop.mr)
> > 
> > > >
> > 
> > ```
> > <http://log4j.logger.org. __ela__ sticsearch.hadoop.mr <
> > 
> > ```
> > 
> > [http://elasticsearch.hadoop.mr](http://elasticsearch.hadoop.mr)\>  
> > \<[http://log4j.logger.org](http://log4j.logger.org).\_\_elasticsearch.hadoop.mr \<  
> > [http://log4j.logger.org.elasticsearch.hadoop.mr](http://log4j.logger.org.elasticsearch.hadoop.mr)\>\>\>  
> > \<[http://log4j.logger.org](http://log4j.logger.org).\_\_ela  
> > \_\_\_\_sticsearch.hadoop.mr [http://ela\_\_sticsearch.hadoop.mr](http://ela__sticsearch.hadoop.mr)  
> > \<[http://elasticsearch.hadoop](http://elasticsearch.hadoop).\_\_mr \<[http://elasticsearch.hadoop.mr](http://elasticsearch.hadoop.mr)
> > 
> > > >
> > 
> > ```
> > <http://log4j.logger.org. __ela__ sticsearch.hadoop.mr <
> > 
> > ```
> > 
> > [http://elasticsearch.hadoop.mr](http://elasticsearch.hadoop.mr)\>  
> > \<[http://log4j.logger.org](http://log4j.logger.org).\_\_elasticsearch.hadoop.mr \<  
> > [http://log4j.logger.org.elasticsearch.hadoop.mr](http://log4j.logger.org.elasticsearch.hadoop.mr)\>\>\>\>\>=\_\_\_\_\_\_TRACE
> > 
> > ```
> > > > |
> > > >
> > > > But I didn't see any trace
> > 
> > ```
> > 
> > logging. Hopefully I can get it working on EC2 without  
> > issue,  
> > but, for  
> > the future, is this  
> > \> \> the correct way to set TRACE  
> > logging?  
> > \> \>  
> > \> \> Oh and, for reference, I tried  
> > running without ES up and I got the following,  
> > exceptions:  
> > \> \>  
> > \> \> 2014-02-19 13:46:08,803 ERROR  
> > shark.SharkDriver (Logging.scala:logError(64)) -  
> > FAILED: Hive  
> > Internal Error:  
> > \> \> java.lang.\_\_\_\_\_\_IllegalStateException(Cannot  
> > discover Elasticsearch version)  
> > \> \> java.lang.\_\_\_\_\_\_IllegalStateException:  
> > Cannot discover Elasticsearch version  
> > \> \> at  
> > org.elasticsearch.hadoop.hive.\_\_\_\_**EsStorageHandler.init(**  
> > \_\_\_\_EsStorageHandler.java:101)  
> > \> \> at
> > 
> > ```
> > org.elasticsearch.hadoop.hive. ______EsStorageHandler.______
> > 
> > ```
> > 
> > configureOutputJobProperties(\_\_\_\_\_\_EsStorageHandler.java:83)  
> > \> \> at
> > 
> > ```
> > org.apache.hadoop.hive.ql. ______plan.PlanUtils.______
> > 
> > ```
> > 
> > configureJobPropertiesForStora\_\_\_\_\_\_geHandler(PlanUtils.java:\_\_\_\_706)  
> > \> \> at
> > 
> > ```
> > org.apache.hadoop.hive.ql. ______plan.PlanUtils.______
> > 
> > ```
> > 
> > configureOutputJobPropertiesFo\_\_\_\_\_\_rStorageHandler(  
> > PlanUtils.\_\_**java:675)  
> > \> \> at  
> > org.apache.hadoop.hive.ql.**  
> > \_\_exec.FileSinkOperator.\_\_\_\_\__augmentPlan(FileSinkOperator._  
> > \_\_\_\_\_java:764)  
> > \> \> at
> > 
> > ```
> > org.apache.hadoop.hive.ql. ______parse.SemanticAnalyzer._____
> > 
> > ```
> > 
> > \_putOpInsertMap(\_\_\_\_\_\_SemanticAnalyzer.java:1518)  
> > \> \> at
> > 
> > ```
> > org.apache.hadoop.hive.ql. ______parse.SemanticAnalyzer._____
> > 
> > ```
> > 
> > \_genFileSinkPlan(\_\_\_\_\_\_SemanticAnalyzer.java:4337)  
> > \> \> at
> > 
> > ```
> > org.apache.hadoop.hive.ql. ______parse.SemanticAnalyzer._____
> > 
> > ```
> > 
> > \_genPostGroupByBodyPlan(\_\_**SemanticAnalyzer.java:6207)  
> > \> \> at  
> > org.apache.hadoop.hive.ql.**  
> > \_\_parse.SemanticAnalyzer.\_\_\_\_\_\_genBodyPlan(SemanticAnalyzer.  
> > \_\_**java:6138)  
> > \> \> at  
> > org.apache.hadoop.hive.ql.**  
> > \_\_parse.SemanticAnalyzer.\_\_\_\_\_\_genPlan(SemanticAnalyzer.java:\_\_\_\_\_\_6764)  
> > \> \> at  
> > shark.parse. **SharkSemanticAnalyzer.**  
> > analyzeInternal(\_\_\_\_\_\_SharkSemanticAnalyzer.scala:\_\_\_\_\_\_149)  
> > \> \> at
> > 
> > ```
> > org.apache.hadoop.hive.ql. ______ parse.BaseSemanticAnalyzer._
> > 
> > ```
> > 
> > \_\_\_\_\_analyze(BaseSemanticAnalyzer.**java:244)  
> > \> \> at shark.SharkDriver.compile(  
> > SharkDriver.scala:215)  
> > \> \> at org.apache.hadoop.hive.ql.**  
> > **Driver.compile(Driver.java:336)  
> > \> \> at org.apache.hadoop.hive.ql.  
> > Driver.run(Driver.java:895)  
> > \> \> at shark.SharkCliDriver.**  
> > processCmd(SharkCliDriver.\_\_\__**scala:324)  
> > \> \> at org.apache.hadoop.hive.cli.**_  
> > \_\_\_CliDriver.processLine(\_\__**CliDriver.java:406)  
> > \> \> at shark.SharkCliDriver$.main(**  
> > **SharkCliDriver.scala:232)  
> > \> \> at shark.SharkCliDriver.main(**_  
> > \_\_SharkCliDriver.scala)
> > 
> > ```
> > > > Caused by: java.io.IOException:
> > 
> > ```
> > 
> > Out of nodes and retries; caught exception  
> > \> \> at  
> > org.elasticsearch.hadoop.rest.\_\_\_\_**NetworkClient.execute(**  
> > \_\_\_\_NetworkClient.java:81)  
> > \> \> at org.elasticsearch.hadoop.rest.  
> > \_\_\_\_\_\_RestClient.execute(\_\_\_\_RestClient.\_\_java:221)  
> > \> \> at org.elasticsearch.hadoop.rest.  
> > \_\_\_\_\_\_RestClient.execute(\_\_\_\_RestClient.\_\_java:205)  
> > \> \> at org.elasticsearch.hadoop.rest.  
> > \_\_\_\_\_\_RestClient.execute(\_\_\_\_RestClient.\_\_java:209)  
> > \> \> at org.elasticsearch.hadoop.rest.  
> > \_\_\_\_\_\_RestClient.get(RestClient.\_\_\_\_\_\_java:103)  
> > \> \> at  
> > org.elasticsearch.hadoop.rest.\_\_\__**RestClient.esVersion(**_  
> > \_\_\_RestClient.java:274)  
> > \> \> at
> > 
> > ```
> > org.elasticsearch.hadoop.rest. ______InitializationUtils.____
> > 
> > ```
> > 
> > \_\_discoverEsVersion(\_\_\_\_\_\_InitializationUtils.java:84)  
> > \> \> at  
> > org.elasticsearch.hadoop.hive.\_\_\_\_**EsStorageHandler.init(**  
> > \_\_\_\_EsStorageHandler.java:99)
> > 
> > ```
> > > > ... 18 more
> > > > Caused by:
> > 
> > ```
> > 
> > java.net.ConnectException: Connection refused  
> > \> \> at java.net.PlainSocketImpl.\_\_\_\_\_\_socketConnect(Native  
> > Method)
> > 
> > ```
> > > > at java.net <http://java.net> <
> > 
> > ```
> > 
> > [http://java.net](http://java.net)\>
> > 
> > ```
> > <http://java.net>. ______AbstractPlainSocketImpl.______
> > 
> > ```
> > 
> > doConnect(\_\_\_\_\_\_AbstractPlainSocketImpl.java:\_\_\_\_\_\_339)
> > 
> > ```
> > > > at java.net <http://java.net> <
> > 
> > ```
> > 
> > [http://java.net](http://java.net)\>
> > 
> > ```
> > <http://java.net>. ______AbstractPlainSocketImpl.______
> > 
> > ```
> > 
> > connectToAddress(\_\_\_\_\_\_AbstractPlainSocketImpl.java:\_\_\_\_\_\_200)
> > 
> > ```
> > > > at java.net <http://java.net> <
> > 
> > ```
> > 
> > [http://java.net](http://java.net)\>  
> > [http://java.net](http://java.net). **AbstractPlainSocketImpl.**  
> > connect(\_\_**AbstractPlainSocketImpl.java:_182)  
> > \> \> at java.net.SocksSocketImpl.  
> > connect(SocksSocketImpl.java:391)  
> > \> \> at java.net.Socket.connect(  
> > Socket.java:579)  
> > \> \> at java.net.Socket.connect(_**  
> > Socket.java:528)  
> > \> \> at java.net.Socket.(Socket.  
> > \_\_\_\_\_\_java:425)  
> > \> \> at java.net.Socket.(Socket.  
> > \_\_\_\_\_\_java:280)  
> > \> \> at
> > 
> > ```
> > org.apache.commons.httpclient. ______protocol.______
> > 
> > ```
> > 
> > DefaultProtocolSocketFactory.**createSocket(**  
> > DefaultProtocolSocketFactory.\_\_\_\_\_\_java:80)  
> > \> \> at
> > 
> > ```
> > org.apache.commons.httpclient. ______protocol.______
> > 
> > ```
> > 
> > DefaultProtocolSocketFactory.**createSocket(**  
> > DefaultProtocolSocketFactory.\_\_\_\_\_\_java:122)  
> > \> \> at  
> > org.apache.commons.httpclient.\_\_**HttpConnection.open(**  
> > \_\_HttpConnection.java:707)  
> > \> \> at
> > 
> > ```
> > org.apache.commons.httpclient. ______HttpMethodDirector._____
> > 
> > ```
> > 
> > \_executeWithRetry(\_\_\_\_\_\_HttpMethodDirector.java:387)  
> > \> \> at
> > 
> > ```
> > org.apache.commons.httpclient. ______HttpMethodDirector._____
> > 
> > ```
> > 
> > \_executeMethod(\_\_\_\_\_\_HttpMethodDirector.java:171)  
> > \> \> at  
> > org.apache.commons.httpclient.\_\_\_\_\_\_HttpClient.  
> > executeMethod(\_\_\_\_\_\_HttpClient.java:397)  
> > \> \> at  
> > org.apache.commons.httpclient.\_\_\_\_\_\_HttpClient.  
> > executeMethod(\_\_\_\_\_\_HttpClient.java:323)  
> > \> \> at
> > 
> > ```
> > org.elasticsearch.hadoop.rest. ______commonshttp.______
> > 
> > ```
> > 
> > CommonsHttpTransport.execute(\_\_\_\_\_\_CommonsHttpTransport.java:\_\_\_\_160)  
> > \> \> at  
> > org.elasticsearch.hadoop.rest.\_\_\_\_**NetworkClient.execute(**  
> > \_\_\_\_NetworkClient.java:74)
> > 
> > ```
> > > > ... 25 more
> > > >
> > > > Let me know if there's anything in
> > 
> > ```
> > 
> > particular you'd like me to try on EC2.  
> > \> \>  
> > \> \> (For posterity, the versions I  
> > used were: hadoop 2.2.0, hive 0.9.0, shark 8.1,  
> > spark 8.1,  
> > es-hadoop  
> > 1.3.0.M2, java  
> > \> \> 1.7.0\_15, scala 2.9.3,  
> > elasticsearch 1.0.0)  
> > \> \>  
> > \> \> Thanks again,  
> > \> \> Max  
> > \> \>  
> > \> \> On Tuesday, February 18, 2014  
> > 10:16:38 PM UTC-8, Costin Leau wrote:  
> > \> \>  
> > \> \> The error indicates a network  
> > error - namely es-hadoop cannot connect to  
> > Elasticsearch  
> > on the  
> > default (localhost:9200)  
> > \> \> HTTP port. Can you double  
> > check whether that's indeed the case (using curl or  
> > even  
> > telnet on  
> > that port) - maybe the  
> > \> \> firewall prevents any  
> > connections to be made...  
> > \> \> Also you could try using the  
> > latest Hive, 0.12 and a more recent Hadoop such  
> > as 1.1.2  
> > or 1.2.1.  
> > \> \>  
> > \> \> Additionally, can you enable  
> > TRACE logging in your job on es-hadoop packages  
> > org.elasticsearch.hadoop.rest and  
> > \> \>org.elasticsearch.hadoop.mr \<  
> > [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)\>  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_hadoop.mr \<[http://org.elasticsearch](http://org.elasticsearch).  
> > hadoop.mr\>\>  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_ha\_\_doop.mr \<  
> > [http://hadoop.mr](http://hadoop.mr)\> \<[http://org.elasticsearch](http://org.elasticsearch).\_\_hadoop.mr  
> > [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)\>\>  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_ha\_\_\_\_doop.mr \<  
> > http://ha\_\_doop.mr\> [http://hadoop.mr](http://hadoop.mr)
> > 
> > ```
> > <http://org.elasticsearch. __ha__ doop.mr <http://hadoop.mr>
> > <http://org.elasticsearch.__hadoop.mr <
> > 
> > ```
> > 
> > [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)\>\>\>\>  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_ha\_\_\_\_doop.mr \<  
> > http://ha\_\_doop.mr\> [http://hadoop.mr](http://hadoop.mr)
> > 
> > ```
> > <http://org.elasticsearch. __ha__ doop.mr <http://hadoop.mr>
> > <http://org.elasticsearch.__hadoop.mr <
> > 
> > ```
> > 
> > [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)\>\>\>  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_ha\_\_\_\_doop.mr\<  
> > http://ha\_\_doop.mr\> [http://hadoop.mr](http://hadoop.mr)
> > 
> > ```
> > <http://org.elasticsearch. __ha__ doop.mr <http://hadoop.mr>
> > <http://org.elasticsearch.__hadoop.mr <
> > 
> > ```
> > 
> > [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)\>\>\>\>\>  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_ha\_\_\_\_doop.mr \<  
> > http://ha\_\_doop.mr\> [http://hadoop.mr](http://hadoop.mr)
> > 
> > ```
> > <http://org.elasticsearch. __ha__ doop.mr <http://hadoop.mr>
> > <http://org.elasticsearch.__hadoop.mr <
> > 
> > ```
> > 
> > [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)\>\>\>  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_ha\_\_\_\_doop.mr [http://ha\_\_doop.mr](http://ha__doop.mr) \<  
> > [http://hadoop.mr](http://hadoop.mr)\>
> > 
> > ```
> > <http://org.elasticsearch. __ha__ doop.mr <
> > 
> > ```
> > 
> > [http://hadoop.mr](http://hadoop.mr)\>  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_hadoop.mr \<[http://org.elasticsearch](http://org.elasticsearch).  
> > hadoop.mr\>\>\>\>
> > 
> > ```
> > > <http://org.elasticsearch.__ha
> > 
> > ```
> > 
> > \_\_\_\_doop.mr [http://ha\_\_doop.mr](http://ha__doop.mr) [http://hadoop.mr](http://hadoop.mr)
> > 
> > ```
> > <http://org.elasticsearch. __ha__ doop.mr <
> > 
> > ```
> > 
> > [http://hadoop.mr](http://hadoop.mr)\> \<[http://org.elasticsearch](http://org.elasticsearch).\_\_hadoop.mr  
> > [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)\>\>  
> > \<[http://org.elasticsearch](http://org.elasticsearch).\_\_ha\_\_\_\_doop.mr \<  
> > http://ha\_\_doop.mr\> [http://hadoop.mr](http://hadoop.mr)
> > 
> > ```
> > <http://org.elasticsearch. __ha__ doop.mr <http://hadoop.mr>
> > <http://org.elasticsearch.__hadoop.mr <
> > 
> > ```
> > 
> > [http://org.elasticsearch.hadoop.mr](http://org.elasticsearch.hadoop.mr)\>\>\>\>\>\> packages and report back ?
> > 
> > ```
> > > >
> > > > Thanks,
> > > >
> > > > On 19/02/2014 4:03 AM, Max
> > 
> > ```
> > 
> > Lang wrote:  
> > \> \> \> I set everything up using  
> > this  
> > guide:[https://github.com/\_\_\_\_\_](https://github.com/ _____ )  
> > \_amplab/shark/wiki/Running-\_\_\_\_\_\_Shark-on-EC2  
> > \<[https://github.com/\_\_\_\_amplab/shark/wiki/Running-\_\_\_\_](https://github.com/ ____amplab/shark/wiki/Running-____ )  
> > Shark-on-EC2\>  
> > \<[https://github.com/\_\_amplab/\_](https://github.com/__amplab/_)  
> > \_shark/wiki/Running-\_\_Shark-on-\_\_EC2  
> > [https://github.com/\_\_amplab/shark/wiki/Running-\_\_Shark-on-EC2](https://github.com/ __amplab/shark/wiki/Running-__ Shark-on-EC2)\>
> > 
> > ```
> > <https://github.com/amplab/___
> > 
> > ```
> > 
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)  
> > \<[https://github.com/amplab/\_\_](https://github.com/amplab/__)  
> > shark/wiki/Running-Shark-on-\_\_EC2  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>\>  
> > \<[https://github.com/amplab/\_\_\_](https://github.com/amplab/___)  
> > \_\_\_shark/wiki/Running-Shark-on-\_\_\_\_\_\_EC2  
> > \<[https://github.com/amplab/\_\_\_\_shark/wiki/Running-Shark-on-\_](https://github.com/amplab/ ____ shark/wiki/Running-Shark-on-_)  
> > \_\_\_EC2\>
> > 
> > ```
> > <https://github.com/amplab/___
> > 
> > ```
> > 
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)\>  
> > \<[https://github.com/amplab/\_\_\_](https://github.com/amplab/___)  
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)  
> > \<[https://github.com/amplab/\_\_](https://github.com/amplab/__)  
> > shark/wiki/Running-Shark-on-\_\_EC2  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>\>\>  
> > \<[https://github.com/amplab/\_\_\_](https://github.com/amplab/___)  
> > \_\_\_shark/wiki/Running-Shark-on-\_\_\_\_\_\_EC2  
> > \<[https://github.com/amplab/\_\_\_\_shark/wiki/Running-Shark-on-\_](https://github.com/amplab/ ____ shark/wiki/Running-Shark-on-_)  
> > \_\_\_EC2\>
> > 
> > ```
> > <https://github.com/amplab/___
> > 
> > ```
> > 
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)\>  
> > \<[https://github.com/amplab/\_\_\_](https://github.com/amplab/___)  
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)  
> > \<[https://github.com/amplab/\_\_](https://github.com/amplab/__)  
> > shark/wiki/Running-Shark-on-\_\_EC2  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>\>  
> > \<[https://github.com/amplab/\_\_\_](https://github.com/amplab/___)  
> > \_\_\_shark/wiki/Running-Shark-on-\_\_\_\_\_\_EC2  
> > \<[https://github.com/amplab/\_\_\_\_shark/wiki/Running-Shark-on-\_](https://github.com/amplab/ ____ shark/wiki/Running-Shark-on-_)  
> > \_\_\_EC2\>
> > 
> > ```
> > <https://github.com/amplab/___
> > 
> > ```
> > 
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)\>  
> > \<[https://github.com/amplab/\_\_\_](https://github.com/amplab/___)  
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)  
> > \<[https://github.com/amplab/\_\_](https://github.com/amplab/__)  
> > shark/wiki/Running-Shark-on-\_\_EC2  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>\>\>\>  
> > \> \> \<[https://github.com/amplab/\_\_\_](https://github.com/amplab/___)  
> > \_\_\_shark/wiki/Running-Shark-on-\_\_\_\_\_\_EC2  
> > \<[https://github.com/amplab/\_\_\_\_shark/wiki/Running-Shark-on-\_](https://github.com/amplab/ ____ shark/wiki/Running-Shark-on-_)  
> > \_\_\_EC2\>
> > 
> > ```
> > <https://github.com/amplab/___
> > 
> > ```
> > 
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)\>  
> > \<[https://github.com/amplab/\_\_\_](https://github.com/amplab/___)  
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)  
> > \<[https://github.com/amplab/\_\_](https://github.com/amplab/__)  
> > shark/wiki/Running-Shark-on-\_\_EC2  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>\>  
> > \<[https://github.com/amplab/\_\_\_](https://github.com/amplab/___)  
> > \_\_\_shark/wiki/Running-Shark-on-\_\_\_\_\_\_EC2  
> > \<[https://github.com/amplab/\_\_\_\_shark/wiki/Running-Shark-on-\_](https://github.com/amplab/ ____ shark/wiki/Running-Shark-on-_)  
> > \_\_\_EC2\>
> > 
> > ```
> > <https://github.com/amplab/___
> > 
> > ```
> > 
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)\>  
> > \<[https://github.com/amplab/\_\_\_](https://github.com/amplab/___)  
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)  
> > \<[https://github.com/amplab/\_\_](https://github.com/amplab/__)  
> > shark/wiki/Running-Shark-on-\_\_EC2  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>\>\>  
> > \> \<[https://github.com/amplab/\_\_\_](https://github.com/amplab/___)  
> > \_\_\_shark/wiki/Running-Shark-on-\_\_\_\_\_\_EC2  
> > \<[https://github.com/amplab/\_\_\_\_shark/wiki/Running-Shark-on-\_](https://github.com/amplab/ ____ shark/wiki/Running-Shark-on-_)  
> > \_\_\_EC2\>
> > 
> > ```
> > <https://github.com/amplab/___
> > 
> > ```
> > 
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)\>  
> > \<[https://github.com/amplab/\_\_\_](https://github.com/amplab/___)  
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)  
> > \<[https://github.com/amplab/\_\_](https://github.com/amplab/__)  
> > shark/wiki/Running-Shark-on-\_\_EC2  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>\>  
> > \<[https://github.com/amplab/\_\_\_](https://github.com/amplab/___)  
> > \_\_\_shark/wiki/Running-Shark-on-\_\_\_\_\_\_EC2  
> > \<[https://github.com/amplab/\_\_\_\_shark/wiki/Running-Shark-on-\_](https://github.com/amplab/ ____ shark/wiki/Running-Shark-on-_)  
> > \_\_\_EC2\>  
> > \<[https://github.com/amplab/\_\_\_](https://github.com/amplab/___)  
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)\>
> > 
> > ```
> > <https://github.com/amplab/___
> > 
> > ```
> > 
> > \_shark/wiki/Running-Shark-on-\_\_\_\_EC2  
> > [https://github.com/amplab/\_\_shark/wiki/Running-Shark-on-\_\_EC2](https://github.com/amplab/ __shark/wiki/Running-Shark-on-__ EC2)  
> > \<[https://github.com/amplab/\_\_](https://github.com/amplab/__)  
> > shark/wiki/Running-Shark-on-\_\_EC2  
> > [https://github.com/amplab/shark/wiki/Running-Shark-on-EC2](https://github.com/amplab/shark/wiki/Running-Shark-on-EC2)\>\>\>\>\>  
> > on an ec2 cluster. I've  
> > \> \> \> copied the  
> > elasticsearch-hadoop jars into the hive lib directory and I have  
> > elasticsearch  
> > running on localhost:9200. I'm  
> > \> \> \> running shark in a screen  
> > session with --service screenserver and  
> > connecting to it  
> > at the  
> > same time using shark -h  
> > \> \> \> localhost.  
> > \> \> \>  
> > \> \> \> Unfortunately, when I  
> > attempt to write data into elasticsearch, it fails.  
> > Here's an  
> > example:  
> > \> \> \>  
> > \> \> \> |  
> > \> \> \>  
> > [localhost:10000]shark\>CREATE EXTERNAL TABLE wiki (id BIGINT,title  
> > STRING,last\_modified  
> > STRING,xml STRING,text  
> > \> \> \> STRING)ROW FORMAT DELIMITED  
> > FIELDS TERMINATED BY '\t'LOCATION  
> > 's3n://spark-data/wikipedia-\_\_\_\_\_\_sample/';
> > 
> > ```
> > > > > Timetaken (including network
> > 
> > ```
> > 
> > latency):0.159seconds  
> > \> \> \> 14/02/1901:23:33INFO  
> > CliDriver:Timetaken (including network  
> > latency):0.159seconds  
> > \> \> \>  
> > \> \> \>  
> > [localhost:10000]shark\>SELECT title FROM wiki LIMIT 1;  
> > \> \> \> Alpokalja  
> > \> \> \> Timetaken (including network  
> > latency):2.23seconds  
> > \> \> \> 14/02/1901:23:48INFO  
> > CliDriver:Timetaken (including network  
> > latency):2.23seconds  
> > \> \> \>  
> > \> \> \>  
> > [localhost:10000]shark\>CREATE EXTERNAL TABLE es\_wiki (id BIGINT,title  
> > STRING,last\_modified  
> > STRING,xml STRING,text  
> > \> \> \> STRING)STORED BY
> > 
> > ```
> > 'org.elasticsearch.hadoop. ______hive.EsStorageHandler'______
> > 
> > ```
> > 
> > TBLPROPERTIES('es.resource'='\_\_\_\_\_\_wikipedia/article');
> > 
> > ```
> > > > > Timetaken (including network
> > 
> > ```
> > 
> > latency):0.061seconds  
> > \> \> \> 14/02/1901:33:51INFO  
> > CliDriver:Timetaken (including network  
> > latency):0.061seconds  
> > \> \> \>  
> > \> \> \>  
> > [localhost:10000]shark\>INSERT OVERWRITE TABLE es\_wiki SELECTw.id  
> > [http://w.id](http://w.id),w.title,w.last\_\_\_\_\_\_\_modified,w.xml,w.text  
> > FROM wiki w;  
> > \> \> \> [HiveError]:Queryreturned  
> > non-zero  
> > code:9,cause:FAILED:\_\_\_\_\_\_ExecutionError,returncode  
> > -101fromshark.execution.\_\_\_\_\_\_SparkTask
> > 
> > ```
> > > > > Timetaken (including network
> > 
> > ```
> > 
> > latency):3.575seconds  
> > \> \> \> 14/02/1901:34:42INFO  
> > CliDriver:Timetaken (including network  
> > latency):3.575seconds  
> > \> \> \> |  
> > \> \> \>  
> > \> \> \> _The stack trace looks like  
> > this:_  
> > \> \> \>  
> > \> \> \>  
> > org.apache.hadoop.hive.ql.\_\_\_\_\_\_metadata.HiveException  
> > (org.apache.hadoop.hive.ql.\_\_\_\_\_\_metadata.HiveException:  
> > java.io.IOException:
> > 
> > ```
> > > > > Out of nodes and retries;
> > 
> > ```
> > 
> > caught exception)  
> > \> \> \>  
> > \> \> \>
> > 
> > ```
> > org.apache.hadoop.hive.ql. ______exec.FileSinkOperator.______
> > 
> > ```
> > 
> > processOp(FileSinkOperator.\_\_\_\_**java:602)shark.execution.**  
> > \_\_\_\_FileSinkOperator$$anonfun$\_\_\_\_\_\_processPartition$1.  
> > apply(\_\_\_\_\_\_FileSinkOperator.scala:84) **shark.execution.**  
> > FileSinkOperator$$anonfun$\_\_\__**processPartition$1.apply(**_  
> > \_\_\_FileSinkOperator.scala:81)\_\_\_\_\_\_scala.collection.  
> > Iterator$\_\_\_\_\_\_class.foreach(Iterator.scala:**772)  
> > scala.collection.\_\_\_\_Iterator$$anon$19.foreach(**  
> > Iterator.\_\_scala:399)shark.\_\_\_\_execution.\_\_FileSinkOperator.  
> > \_\_\_\_\_\_processPartition(\_\_**FileSinkOperator.scala:81)**  
> > \_\_shark.execution.\_\_\_\_\__FileSinkOperator$.writeFiles$_  
> > \_\_\_\_\_1(FileSinkOperator.scala:\_\_207) **shark.execution.**  
> > \_\_FileSinkOperator$$anonfun$\_\_\_\_\_\_executeProcessFileSinkPartitio  
> > \_\_\_\_\_\_n$1.apply(\_\_FileSinkOperator.**scala:211)shark.execution.**  
> > FileSinkOperator$$anonfun$\_\_\_\_\_\_executeProcessFileSinkPartitio  
> > \_\_\_\_\_\_n$1.apply(\_\_FileSinkOperator.\_\_ **scala:**  
> > 211)org.apache.spark.\_\_\_\_\_\_scheduler.ResultTask.
> 
> runTask(\_\_\_\_\_\_ResultTask.scala:107)org.\_\_\_\_\_\_apache.spark.scheduler.Ta
> 
> > ```
> > sk.____run(Task.scala:53)org.__apache. ____spark.executor.__
> > 
> > ```
> > 
> > Executor$\_\_\_\_Task
> > 
> > ```
> > Runner$$anonfun$run$1. __apply$____ mcV$sp(Executor.scala:__
> > 
> > ```
> > 
> > 215)\_\_\_\_org.apac
> > 
> > ```
> > he.spa
> > 
> > rk.dep
> > >
> > > loy.Sp
> > > >
> > > >
> > 
> > arkHadoopUtil.runAsUser( ______ SparkHadoopUtil.scala:50)org._
> > 
> > ```
> > 
> > \_\_\_\_\_apache.spark.executor.\_\_\__**Executor$TaskRunner.run(**_  
> > \_\_\_Executor.scala:182)java.util.\_\_ **concurrent.**  
> > ThreadPoolExecutor.\_\_\_\_\__runWorker(ThreadPoolExecutor._  
> > \_\_\_\_\_java:1145)java.util. **concurrent.ThreadPoolExecutor$**  
> > Worker.run(\_\_\_\_ThreadPoolExecutor.\_\_java:615)  
> > \_\_\_\_java.lang.Thread.run(\_\_\_\_Thread.\_\_java:744
> > 
> > ```
> > >
> > > >
> > > > > I should be using Hive
> > 
> > ```
> > 
> > 0.9.0, shark 0.8.1, elasticsearch 1.0.0, Hadoop  
> > 1.0.4, and  
> > java 1.7.0\_51  
> > \> \> \> Based on my cursory look at  
> > the hadoop and elasticsearch-hadoop sources, it  
> > looks  
> > like hive  
> > is just rethrowing an  
> > \> \> \> IOException it's getting  
> > from Spark, and elasticsearch-hadoop is just  
> > hitting those  
> > exceptions.  
> > \> \> \> I suppose my questions are:  
> > Does this look like an issue with my  
> > ES/elasticsearch-hadoop  
> > config? And has anyone gotten  
> > \> \> \> elasticsearch working with  
> > Spark/Shark?  
> > \> \> \> Any ideas/insights are  
> > appreciated.  
> > \> \> \> Thanks,Max  
> > \> \> \>  
> > \> \> \> --  
> > \> \> \> You received this message  
> > because you are subscribed to the Google Groups  
> > "elasticsearch" group.  
> > \> \> \> To unsubscribe from this  
> > group and stop receiving emails from it, send an  
> > email to  
> > \> \> \>elasticsearc...@googlegroups.\_\_\_\_\_\_com  
> > \<mailto:elasticsearc...@  
> > [mailto:elasticsearc...@](mailto:elasticsearc...@)\_\_goog\_\_legroups.com \<  
> > [http://googlegroups.com](http://googlegroups.com)\>
> > 
> > ```
> > <mailto:elasticsearc...@__googlegroups.com <mailto:
> > 
> > ```
> > 
> > [elasticsearc...@googlegroups.com](mailto:elasticsearc...@googlegroups.com)\>\>\> \<javascript:\>.
> > 
> > ```
> > > > > To view this discussion on
> > 
> > ```
> > 
> > the web visit  
> > \> \>
> > 
> > ```
> > >
> > 
> > ```
> 
> ...
> 
> [Message clipped]

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GOCjCvN2z0uijL6\_G3qF5ki4afuMTzBLy%3D%2Bj5W6qqSuSQ%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CALD%2B6GOCjCvN2z0uijL6_G3qF5ki4afuMTzBLy%3D%2Bj5W6qqSuSQ%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:28am UTC](https://discuss.elastic.co/t/hadoop-getting-elasticsearch-hadoop-working-with-shark/15882/16 "2017-07-06T01:28:50Z")

</div>


