# JDBC river doesn't start after delete & recreate

**URL:** <https://discuss.elastic.co/t/jdbc-river-doesnt-start-after-delete-recreate/14755>\
**Category:** Elasticsearch\
**Created:** [December 6, 2013, 6:28pm UTC](https://discuss.elastic.co/t/jdbc-river-doesnt-start-after-delete-recreate/14755 "2013-12-06T18:28:59Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![Justin\_Doles](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/justin_doles/32/40730_2.png) [@Justin\_Doles](https://discuss.elastic.co/u/Justin_Doles)\
**Post date:** [December 6, 2013, 6:28pm UTC](https://discuss.elastic.co/t/jdbc-river-doesnt-start-after-delete-recreate/14755/1 "2013-12-06T18:28:59Z")

</div>

I'm having an issue with JDBC rivers. I'm running ES 0.90.7 and JDBC river  
2.2.3. I have a 3 node cluster: 1 w/no data (Windows) and 2 w/ data  
(Linux).

I can create a simple river initially.

{  
"type" : "jdbc",  
"jdbc" : {  
"strategy" : "oneshot",  
"driver" : "com.mysql.jdbc.Driver",  
"url" : "jdbc:mysql://192.168.1.1:6033/test",  
"user" : "test\_account",  
"password" : "test\_password",  
"sql" : "SELECT `orders`.`id` AS `_id`, `orders`.`name`,  
`orders`.`description`, BinaryToGuid(`orders`.`guid`), `orders`.`number`  
FROM `test`.`orders`;"  
},  
"index" : {  
"index" : "orders",  
"type" : "order"  
}  
}

This works for the initial load. If I stop the river while it's running by  
deleting it, I cannot start another river with the same parameters until  
stop ES on the data node that was processing this river. Once I do that,  
the other node starts the river. Is there something I'm not understanding?  
I don't see any errors in the logs.

Thanks.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/7614d944-d89a-463e-8c46-6e07134f230f%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/7614d944-d89a-463e-8c46-6e07134f230f%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [December 6, 2013, 7:19pm UTC](https://discuss.elastic.co/t/jdbc-river-doesnt-start-after-delete-recreate/14755/2 "2013-12-06T19:19:09Z")

</div>

Not sure if this is related to rivers in general.

The JDBC river runs in a separate thread and writes state info at each  
cycle into the private river index. Deleting a river while the river is  
running may not remove the river resources completely, or it may hang doing  
this. This could explain why the cluster is thinking it should restart the  
river instance at the other node. At least, I have not tested these  
situations.

Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoGKk6KAy-ZX0xWrc%2BQXnQ\_4dSZT0%2BBebo-iOS8vX8CqSw%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoGKk6KAy-ZX0xWrc%2BQXnQ_4dSZT0%2BBebo-iOS8vX8CqSw%40mail.gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Justin\_Doles](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/justin_doles/32/40730_2.png) [@Justin\_Doles](https://discuss.elastic.co/u/Justin_Doles)\
**Post date:** [December 6, 2013, 8:21pm UTC](https://discuss.elastic.co/t/jdbc-river-doesnt-start-after-delete-recreate/14755/3 "2013-12-06T20:21:51Z")

</div>

You may be right. I'm far from an expert in ES. If I delete the first  
river (my\_river\_1) while it's running (delete is successful) then create a  
new river (my\_river\_2) with the same parameters, the second river won't  
start until I stop the ES node that was processing the first river.

I've tried waiting a few minutes between, but it doesn't seem to matter.

Justin

On Friday, December 6, 2013 2:19:09 PM UTC-5, Jörg Prante wrote:

> Not sure if this is related to rivers in general.
> 
> The JDBC river runs in a separate thread and writes state info at each  
> cycle into the private river index. Deleting a river while the river is  
> running may not remove the river resources completely, or it may hang doing  
> this. This could explain why the cluster is thinking it should restart the  
> river instance at the other node. At least, I have not tested these  
> situations.
> 
> Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/10301994-02a9-492e-ab4d-2bf37235e115%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/10301994-02a9-492e-ab4d-2bf37235e115%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Justin\_Doles](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/justin_doles/32/40730_2.png) [@Justin\_Doles](https://discuss.elastic.co/u/Justin_Doles)\
**Post date:** [December 7, 2013, 2:38am UTC](https://discuss.elastic.co/t/jdbc-river-doesnt-start-after-delete-recreate/14755/4 "2013-12-07T02:38:41Z")

</div>

I did some more digging and you're right about it being a river issue (as  
far as I can tell). It looks like the rivers all get assigned to a single  
node . If I delete a river midway, it won't run any additional rivers that  
are created. But once that node is shutdown, the other rivers I created  
after the delete get assigned to another node and begin to process.

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

> _Rivers are singletons within the cluster. They get allocated  
> automatically to one of the nodes and run. If that node fails, a river will  
> be automatically allocated to another node._

I can't tell if this is a bug or not though. Or maybe there's a time I  
need to wait?

Justin

On Friday, December 6, 2013 3:21:51 PM UTC-5, Justin Doles wrote:

> You may be right. I'm far from an expert in ES. If I delete the first  
> river (my\_river\_1) while it's running (delete is successful) then create a  
> new river (my\_river\_2) with the same parameters, the second river won't  
> start until I stop the ES node that was processing the first river.
> 
> I've tried waiting a few minutes between, but it doesn't seem to matter.
> 
> Justin
> 
> On Friday, December 6, 2013 2:19:09 PM UTC-5, Jörg Prante wrote:
> 
> > Not sure if this is related to rivers in general.
> > 
> > The JDBC river runs in a separate thread and writes state info at each  
> > cycle into the private river index. Deleting a river while the river is  
> > running may not remove the river resources completely, or it may hang doing  
> > this. This could explain why the cluster is thinking it should restart the  
> > river instance at the other node. At least, I have not tested these  
> > situations.
> > 
> > Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/41c5fa8e-f5d9-4531-8cbf-761b0f2b9d1e%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/41c5fa8e-f5d9-4531-8cbf-761b0f2b9d1e%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [December 7, 2013, 11:10am UTC](https://discuss.elastic.co/t/jdbc-river-doesnt-start-after-delete-recreate/14755/5 "2013-12-07T11:10:50Z")

</div>

I think the concept of river is broken. For example:

- it is assumed that river instances shall always run. If a node fails, all  
the river instances on that node are started again on other nodes. The idea  
is to run river contiuously without interruption so no data gets lost.

- the river cluster service does not watch what river instances are  
currently doing and what the river instance state is since the river  
instance state is private to the river.

- if a river instance is deleted, the river cluster service must know this  
instance is permanently removed. But what is permanently if you can  
recreate a river instance under the same name after a deletion?

There were discussions that rivers may be deprecated in favor of message  
queues like logstash.

I think it would be a good idea to improve the river concept to a truly  
distributed design.

For example:

- rivers should be aware of many river instances in parallel so they could  
share the work by dividing the workload

- a river instance should always be distributed to many nodes, and by river  
instance creation, a plan of execution is announced to all river instances

- river instances should (similar to web crawlers) receive a list of URLs  
of sources they can process in parallel. The URLs carry schemes for custom  
URL handlers (like twitter://, wikipedia://, jdbc:// etc.) Dispatching the  
URLs would be a central task at river initiation phase, probably of the ES  
master node, or the node that receives a river creation request. The state  
of each (active) URL should be available in the cluster state

- and, river instances should be identifiable by the cluster service by an  
ID, and should respond with a state message if they are asked for a report.  
Also, a river instance should be able to receive stop signals and react in  
a predictable way (finishing the URL queue, finishing current URL then  
abort the URL queue, or abort immediately)

- river instances should be able to shutdown automatically if the list of  
URLs they received is done and delete themselves from the active river  
instance list in the cluster state

- plan of execution could also be defined by a cron-like request

- nodes should be configurable if they can run river instances or not

- the number of river instances could also be a parameter in a river  
creation request. So if the number of URLs to be processed exceed the  
available river nodes, they would have to be executed in a queue

- by providing a standard bulk indexing procedure in a new generic river  
framework common to all rivers, writing custom code for rivers would reduce  
to the mere task of handling a single URL for fetching data and construct  
JSON documents in a stream-like manner, maybe with something like JSON-Path  
keys for inserting values.

So many wishes.... sorry for that. But it's christmas time 😉

Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoHDfOLg6ETjVxBdD-kOq1UfKA9Rt9qzWuJJ33VeS4OW\_Q%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoHDfOLg6ETjVxBdD-kOq1UfKA9Rt9qzWuJJ33VeS4OW_Q%40mail.gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Gabe\_Gorelick\_Feldma](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gabe_gorelick_feldma/32/1423_2.png) [@Gabe\_Gorelick\_Feldma](https://discuss.elastic.co/u/Gabe_Gorelick_Feldma)\
**Post date:** [December 7, 2013, 3:58pm UTC](https://discuss.elastic.co/t/jdbc-river-doesnt-start-after-delete-recreate/14755/6 "2013-12-07T15:58:14Z")

</div>

Given that rivers in general seem to be flawed, is there an easy way  
clients can tell whether a JDBC river job is running (so they don't try to  
delete it)? Maybe a field in the internal JDBC river document? I haven't  
seen any documentation on the structure of that doc.

On Saturday, December 7, 2013 6:10:50 AM UTC-5, Jörg Prante wrote:

> I think the concept of river is broken. For example:
> 
> - it is assumed that river instances shall always run. If a node fails,  
> all the river instances on that node are started again on other nodes. The  
> idea is to run river contiuously without interruption so no data gets lost.
> 
> - the river cluster service does not watch what river instances are  
> currently doing and what the river instance state is since the river  
> instance state is private to the river.
> 
> - if a river instance is deleted, the river cluster service must know this  
> instance is permanently removed. But what is permanently if you can  
> recreate a river instance under the same name after a deletion?
> 
> There were discussions that rivers may be deprecated in favor of message  
> queues like logstash.
> 
> I think it would be a good idea to improve the river concept to a truly  
> distributed design.
> 
> For example:
> 
> - rivers should be aware of many river instances in parallel so they could  
> share the work by dividing the workload
> 
> - a river instance should always be distributed to many nodes, and by  
> river instance creation, a plan of execution is announced to all river  
> instances
> 
> - river instances should (similar to web crawlers) receive a list of URLs  
> of sources they can process in parallel. The URLs carry schemes for custom  
> URL handlers (like twitter://, wikipedia://, jdbc:// etc.) Dispatching the  
> URLs would be a central task at river initiation phase, probably of the ES  
> master node, or the node that receives a river creation request. The state  
> of each (active) URL should be available in the cluster state
> 
> - and, river instances should be identifiable by the cluster service by an  
> ID, and should respond with a state message if they are asked for a report.  
> Also, a river instance should be able to receive stop signals and react in  
> a predictable way (finishing the URL queue, finishing current URL then  
> abort the URL queue, or abort immediately)
> 
> - river instances should be able to shutdown automatically if the list of  
> URLs they received is done and delete themselves from the active river  
> instance list in the cluster state
> 
> - plan of execution could also be defined by a cron-like request
> 
> - nodes should be configurable if they can run river instances or not
> 
> - the number of river instances could also be a parameter in a river  
> creation request. So if the number of URLs to be processed exceed the  
> available river nodes, they would have to be executed in a queue
> 
> - by providing a standard bulk indexing procedure in a new generic river  
> framework common to all rivers, writing custom code for rivers would reduce  
> to the mere task of handling a single URL for fetching data and construct  
> JSON documents in a stream-like manner, maybe with something like JSON-Path  
> keys for inserting values.
> 
> So many wishes.... sorry for that. But it's christmas time 😉
> 
> Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/897645ca-013c-40d7-9f4c-102f2fac9912%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/897645ca-013c-40d7-9f4c-102f2fac9912%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [December 7, 2013, 5:36pm UTC](https://discuss.elastic.co/t/jdbc-river-doesnt-start-after-delete-recreate/14755/7 "2013-12-07T17:36:06Z")

</div>

Finding out if a JDBC river job runs has to be implemented, it is not  
present yet.

Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoHqysv5diJw\_pKzk\_CRz%2BFhFbhdgnNwCFhuEy6ua0VRHg%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoHqysv5diJw_pKzk_CRz%2BFhFbhdgnNwCFhuEy6ua0VRHg%40mail.gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Gabe\_Gorelick\_Feldma](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/gabe_gorelick_feldma/32/1423_2.png) [@Gabe\_Gorelick\_Feldma](https://discuss.elastic.co/u/Gabe_Gorelick_Feldma)\
**Post date:** [December 7, 2013, 7:56pm UTC](https://discuss.elastic.co/t/jdbc-river-doesnt-start-after-delete-recreate/14755/8 "2013-12-07T19:56:44Z")

</div>

How hard would it be to implement? I'm happy to help if someone points me  
in the right direction.

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/c8667406-1bee-47c8-82fa-1f02e1a79d10%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/c8667406-1bee-47c8-82fa-1f02e1a79d10%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [December 7, 2013, 9:45pm UTC](https://discuss.elastic.co/t/jdbc-river-doesnt-start-after-delete-recreate/14755/9 "2013-12-07T21:45:10Z")

</div>

It's very easy, I added an issue. An activity flag can be added to the  
river state document.

> <https://github.com/jprante/elasticsearch-jdbc/issues/146>

Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoHZk6sPZF\_GN0%2BEpAUrmfTxcZA%2B%2BCz2v-nUrwmjoHt1hg%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoHZk6sPZF_GN0%2BEpAUrmfTxcZA%2B%2BCz2v-nUrwmjoHt1hg%40mail.gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Justin\_Doles](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/justin_doles/32/40730_2.png) [@Justin\_Doles](https://discuss.elastic.co/u/Justin_Doles)\
**Post date:** [December 9, 2013, 4:35pm UTC](https://discuss.elastic.co/t/jdbc-river-doesnt-start-after-delete-recreate/14755/10 "2013-12-09T16:35:25Z")

</div>

I did find a way to prevent nodes from running rivers.  
node.river: false|true

I set this to true on both my data nodes and false on my master node. So  
far so good.

I also noticed that multiple rivers ran on different nodes. I'm not  
certain if this was a side effect of that setting or a coincidence. I'll  
be doing more testing in the next couple weeks.

All your ideas have merit. There is definitely room to improve rivers.  
Not always running and a reliable status would be huge.

Justin

On Saturday, December 7, 2013 6:10:50 AM UTC-5, Jörg Prante wrote:

> I think the concept of river is broken. For example:
> 
> - it is assumed that river instances shall always run. If a node fails,  
> all the river instances on that node are started again on other nodes. The  
> idea is to run river contiuously without interruption so no data gets lost.
> 
> - the river cluster service does not watch what river instances are  
> currently doing and what the river instance state is since the river  
> instance state is private to the river.
> 
> - if a river instance is deleted, the river cluster service must know this  
> instance is permanently removed. But what is permanently if you can  
> recreate a river instance under the same name after a deletion?
> 
> There were discussions that rivers may be deprecated in favor of message  
> queues like logstash.
> 
> I think it would be a good idea to improve the river concept to a truly  
> distributed design.
> 
> For example:
> 
> - rivers should be aware of many river instances in parallel so they could  
> share the work by dividing the workload
> 
> - a river instance should always be distributed to many nodes, and by  
> river instance creation, a plan of execution is announced to all river  
> instances
> 
> - river instances should (similar to web crawlers) receive a list of URLs  
> of sources they can process in parallel. The URLs carry schemes for custom  
> URL handlers (like twitter://, wikipedia://, jdbc:// etc.) Dispatching the  
> URLs would be a central task at river initiation phase, probably of the ES  
> master node, or the node that receives a river creation request. The state  
> of each (active) URL should be available in the cluster state
> 
> - and, river instances should be identifiable by the cluster service by an  
> ID, and should respond with a state message if they are asked for a report.  
> Also, a river instance should be able to receive stop signals and react in  
> a predictable way (finishing the URL queue, finishing current URL then  
> abort the URL queue, or abort immediately)
> 
> - river instances should be able to shutdown automatically if the list of  
> URLs they received is done and delete themselves from the active river  
> instance list in the cluster state
> 
> - plan of execution could also be defined by a cron-like request
> 
> - nodes should be configurable if they can run river instances or not
> 
> - the number of river instances could also be a parameter in a river  
> creation request. So if the number of URLs to be processed exceed the  
> available river nodes, they would have to be executed in a queue
> 
> - by providing a standard bulk indexing procedure in a new generic river  
> framework common to all rivers, writing custom code for rivers would reduce  
> to the mere task of handling a single URL for fetching data and construct  
> JSON documents in a stream-like manner, maybe with something like JSON-Path  
> keys for inserting values.
> 
> So many wishes.... sorry for that. But it's christmas time 😉
> 
> Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/d57f4561-9ee0-4674-9a8c-56cf4afb21ba%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/d57f4561-9ee0-4674-9a8c-56cf4afb21ba%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Justin\_Doles](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/justin_doles/32/40730_2.png) [@Justin\_Doles](https://discuss.elastic.co/u/Justin_Doles)\
**Post date:** [December 9, 2013, 4:41pm UTC](https://discuss.elastic.co/t/jdbc-river-doesnt-start-after-delete-recreate/14755/11 "2013-12-09T16:41:48Z")

</div>

I swore I saw node.river accepted true|false, but it's _none_ or a comma  
separated list.

> **[Elasticsearch Platform — Find real-time answers at scale](https://www.elastic.co)**
>
> Power insights and outcomes with the Elasticsearch Platform and AI. See into your data and find answers that matter with enterprise solutions designed to help you build, observe, and protect. Try Elasticsearch free today.

On Monday, December 9, 2013 11:35:25 AM UTC-5, Justin Doles wrote:

> I did find a way to prevent nodes from running rivers.  
> node.river: false|true
> 
> I set this to true on both my data nodes and false on my master node. So  
> far so good.
> 
> I also noticed that multiple rivers ran on different nodes. I'm not  
> certain if this was a side effect of that setting or a coincidence. I'll  
> be doing more testing in the next couple weeks.
> 
> All your ideas have merit. There is definitely room to improve rivers.  
> Not always running and a reliable status would be huge.
> 
> Justin
> 
> On Saturday, December 7, 2013 6:10:50 AM UTC-5, Jörg Prante wrote:
> 
> > I think the concept of river is broken. For example:
> > 
> > - it is assumed that river instances shall always run. If a node fails,  
> > all the river instances on that node are started again on other nodes. The  
> > idea is to run river contiuously without interruption so no data gets lost.
> > 
> > - the river cluster service does not watch what river instances are  
> > currently doing and what the river instance state is since the river  
> > instance state is private to the river.
> > 
> > - if a river instance is deleted, the river cluster service must know  
> > this instance is permanently removed. But what is permanently if you can  
> > recreate a river instance under the same name after a deletion?
> > 
> > There were discussions that rivers may be deprecated in favor of message  
> > queues like logstash.
> > 
> > I think it would be a good idea to improve the river concept to a truly  
> > distributed design.
> > 
> > For example:
> > 
> > - rivers should be aware of many river instances in parallel so they  
> > could share the work by dividing the workload
> > 
> > - a river instance should always be distributed to many nodes, and by  
> > river instance creation, a plan of execution is announced to all river  
> > instances
> > 
> > - river instances should (similar to web crawlers) receive a list of URLs  
> > of sources they can process in parallel. The URLs carry schemes for custom  
> > URL handlers (like twitter://, wikipedia://, jdbc:// etc.) Dispatching the  
> > URLs would be a central task at river initiation phase, probably of the ES  
> > master node, or the node that receives a river creation request. The state  
> > of each (active) URL should be available in the cluster state
> > 
> > - and, river instances should be identifiable by the cluster service by  
> > an ID, and should respond with a state message if they are asked for a  
> > report. Also, a river instance should be able to receive stop signals and  
> > react in a predictable way (finishing the URL queue, finishing current URL  
> > then abort the URL queue, or abort immediately)
> > 
> > - river instances should be able to shutdown automatically if the list of  
> > URLs they received is done and delete themselves from the active river  
> > instance list in the cluster state
> > 
> > - plan of execution could also be defined by a cron-like request
> > 
> > - nodes should be configurable if they can run river instances or not
> > 
> > - the number of river instances could also be a parameter in a river  
> > creation request. So if the number of URLs to be processed exceed the  
> > available river nodes, they would have to be executed in a queue
> > 
> > - by providing a standard bulk indexing procedure in a new generic river  
> > framework common to all rivers, writing custom code for rivers would reduce  
> > to the mere task of handling a single URL for fetching data and construct  
> > JSON documents in a stream-like manner, maybe with something like JSON-Path  
> > keys for inserting values.
> > 
> > So many wishes.... sorry for that. But it's christmas time 😉
> > 
> > Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/509589f4-686a-4034-a122-4eb8b1fb75c9%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/509589f4-686a-4034-a122-4eb8b1fb75c9%40googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [December 9, 2013, 5:05pm UTC](https://discuss.elastic.co/t/jdbc-river-doesnt-start-after-delete-recreate/14755/12 "2013-12-09T17:05:17Z")

</div>

Yes, it's "_none_". False/true is not recognized. This peculiarity could be  
easily fixed, the code is in  
[https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/river/cluster/RiverNodeHelper.java#L40](https://github.com/elasticsearch/elasticsearch/blob/master/src/main/java/org/elasticsearch/river/cluster/RiverNodeHelper.java#L40)

Jörg

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAKdsXoEBdNsxfs2tWtzHTzC-BdsjKAch2SRKiiw9MGTphBP0wg%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAKdsXoEBdNsxfs2tWtzHTzC-BdsjKAch2SRKiiw9MGTphBP0wg%40mail.gmail.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:02am UTC](https://discuss.elastic.co/t/jdbc-river-doesnt-start-after-delete-recreate/14755/13 "2017-07-06T02:02:25Z")

</div>


