# MySQL River?

**URL:** <https://discuss.elastic.co/t/mysql-river/6072>\
**Category:** Elasticsearch\
**Created:** [December 5, 2011, 10:46pm UTC](https://discuss.elastic.co/t/mysql-river/6072 "2011-12-05T22:46:14Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [December 5, 2011, 10:46pm UTC](https://discuss.elastic.co/t/mysql-river/6072/1 "2011-12-05T22:46:14Z")

</div>

Hello,

How do people do MySQL (or any other RDBMS) indexing in bulk an  
incrementally? Is there something like Solr's DataImportHandler in  
ES?  
I know there are Rivers and I _assume_ it makes sense to implement a  
tool for indexing database content as a River, but I don't see a MySQL  
River.

Or are people writing standalone and external indexer apps that push  
data from DB to ES? (and thus probably creating a SPOF in their  
indexing pipeline)

## Thanks, Otis

Sematext is Hiring World-Wide -- [http://sematext.com/about/jobs.html](http://sematext.com/about/jobs.html)

---

<div class="post-metadata">

**Author:** ![Damien\_Hardy](https://avatars.discourse-cdn.com/v4/letter/d/779978/32.png) [@Damien\_Hardy](https://discuss.elastic.co/u/Damien_Hardy)\
**Post date:** [December 6, 2011, 8:08am UTC](https://discuss.elastic.co/t/mysql-river/6072/2 "2011-12-06T08:08:06Z")

</div>

Hello,

Maybe you can look after MySQL trigger to fire a POST request to  
elasticsearch base on UDF like : [http://code.google.com/p/mysql-udf-http/](http://code.google.com/p/mysql-udf-http/)

Cheers,

--  
Damien

2011/12/5 Otis Gospodnetic [otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com)

> Hello,
> 
> How do people do MySQL (or any other RDBMS) indexing in bulk an  
> incrementally? Is there something like Solr's DataImportHandler in  
> ES?  
> I know there are Rivers and I _assume_ it makes sense to implement a  
> tool for indexing database content as a River, but I don't see a MySQL  
> River.
> 
> Or are people writing standalone and external indexer apps that push  
> data from DB to ES? (and thus probably creating a SPOF in their  
> indexing pipeline)
> 
> ## Thanks, Otis
> 
> Sematext is Hiring World-Wide -- [Jobs - Sematext](http://sematext.com/about/jobs.html)

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [December 6, 2011, 5:51pm UTC](https://discuss.elastic.co/t/mysql-river/6072/3 "2011-12-06T17:51:00Z")

</div>

Hi,

OK, so this is still external, except it's a push instead of a pull  
(like in Solr's DIH).  
So no MySQL River.

Is there a technical reason why one could _not_ implement a MySQL  
River, or is it just that nobody had this MySQL==\>ES itch yet?

## Thanks, Otis

Sematext is Hiring World-Wide -- [Jobs](http://sematext.com/about/jobs.html)

On Dec 6, 3:08 am, Damien Hardy [damienhardy....@gmail.com](mailto:damienhardy....@gmail.com) wrote:

> Hello,
> 
> Maybe you can look after MySQL trigger to fire a POST request to  
> elasticsearch base on UDF like :[Google Code Archive - Long-term storage for Google Code Project Hosting.](http://code.google.com/p/mysql-udf-http/)
> 
> Cheers,
> 
> --  
> Damien
> 
> 2011/12/5 Otis Gospodnetic [otis.gospodne...@gmail.com](mailto:otis.gospodne...@gmail.com)
> 
> > Hello,
> 
> > How do people do MySQL (or any other RDBMS) indexing in bulk an  
> > incrementally? Is there something like Solr's DataImportHandler in  
> > ES?  
> > I know there are Rivers and I _assume_ it makes sense to implement a  
> > tool for indexing database content as a River, but I don't see a MySQL  
> > River.
> 
> > Or are people writing standalone and external indexer apps that push  
> > data from DB to ES? (and thus probably creating a SPOF in their  
> > indexing pipeline)
> 
> > ## Thanks, Otis
> > 
> > Sematext is Hiring World-Wide --[Jobs](http://sematext.com/about/jobs.html)

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 6, 2011, 6:49pm UTC](https://discuss.elastic.co/t/mysql-river/6072/4 "2011-12-06T18:49:55Z")

</div>

In my project (with postgresql), I choose to push documents in ES each time I create, update or delete an Hibernate entity.  
That's why I don't use a river by now.

Problem is when I need to reindex all my documents from my database.  
At the present time, I create a "batch" for that but I'm thinking of using Talend for that.

That's not a river, for sure !

David 😉  
@dadoonet

Le 6 déc. 2011 à 18:51, Otis Gospodnetic [otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com) a écrit :

> Hi,
> 
> OK, so this is still external, except it's a push instead of a pull  
> (like in Solr's DIH).  
> So no MySQL River.
> 
> Is there a technical reason why one could _not_ implement a MySQL  
> River, or is it just that nobody had this MySQL==\>ES itch yet?
> 
> ## Thanks, Otis
> 
> Sematext is Hiring World-Wide -- [Jobs - Sematext](http://sematext.com/about/jobs.html)
> 
> On Dec 6, 3:08 am, Damien Hardy [damienhardy....@gmail.com](mailto:damienhardy....@gmail.com) wrote:
> 
> > Hello,
> > 
> > Maybe you can look after MySQL trigger to fire a POST request to  
> > elasticsearch base on UDF like :[http://code.google.com/p/mysql-udf-http/](http://code.google.com/p/mysql-udf-http/)
> > 
> > Cheers,
> > 
> > --  
> > Damien
> > 
> > 2011/12/5 Otis Gospodnetic [otis.gospodne...@gmail.com](mailto:otis.gospodne...@gmail.com)
> > 
> > > Hello,
> > 
> > > How do people do MySQL (or any other RDBMS) indexing in bulk an  
> > > incrementally? Is there something like Solr's DataImportHandler in  
> > > ES?  
> > > I know there are Rivers and I _assume_ it makes sense to implement a  
> > > tool for indexing database content as a River, but I don't see a MySQL  
> > > River.
> > 
> > > Or are people writing standalone and external indexer apps that push  
> > > data from DB to ES? (and thus probably creating a SPOF in their  
> > > indexing pipeline)
> > 
> > > ## Thanks, Otis
> > > 
> > > Sematext is Hiring World-Wide --[Jobs - Sematext](http://sematext.com/about/jobs.html)

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [December 6, 2011, 7:48pm UTC](https://discuss.elastic.co/t/mysql-river/6072/5 "2011-12-06T19:48:54Z")

</div>

The problem with a generic database river is handling deletes.

On Tue, Dec 6, 2011 at 8:49 PM, David Pilato [david@pilato.fr](mailto:david@pilato.fr) wrote:

> In my project (with postgresql), I choose to push documents in ES each  
> time I create, update or delete an Hibernate entity.  
> That's why I don't use a river by now.
> 
> Problem is when I need to reindex all my documents from my database.  
> At the present time, I create a "batch" for that but I'm thinking of using  
> Talend for that.
> 
> That's not a river, for sure !
> 
> David 😉  
> @dadoonet
> 
> Le 6 déc. 2011 à 18:51, Otis Gospodnetic [otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com) a  
> écrit :
> 
> > Hi,
> > 
> > OK, so this is still external, except it's a push instead of a pull  
> > (like in Solr's DIH).  
> > So no MySQL River.
> > 
> > Is there a technical reason why one could _not_ implement a MySQL  
> > River, or is it just that nobody had this MySQL==\>ES itch yet?
> > 
> > ## Thanks, Otis
> > 
> > Sematext is Hiring World-Wide -- [Jobs - Sematext](http://sematext.com/about/jobs.html)
> > 
> > On Dec 6, 3:08 am, Damien Hardy [damienhardy....@gmail.com](mailto:damienhardy....@gmail.com) wrote:
> > 
> > > Hello,
> > > 
> > > Maybe you can look after MySQL trigger to fire a POST request to  
> > > elasticsearch base on UDF like :  
> > > [http://code.google.com/p/mysql-udf-http/](http://code.google.com/p/mysql-udf-http/)
> > > 
> > > Cheers,
> > > 
> > > --  
> > > Damien
> > > 
> > > 2011/12/5 Otis Gospodnetic [otis.gospodne...@gmail.com](mailto:otis.gospodne...@gmail.com)
> > > 
> > > > Hello,
> > > 
> > > > How do people do MySQL (or any other RDBMS) indexing in bulk an  
> > > > incrementally? Is there something like Solr's DataImportHandler in  
> > > > ES?  
> > > > I know there are Rivers and I _assume_ it makes sense to implement a  
> > > > tool for indexing database content as a River, but I don't see a MySQL  
> > > > River.
> > > 
> > > > Or are people writing standalone and external indexer apps that push  
> > > > data from DB to ES? (and thus probably creating a SPOF in their  
> > > > indexing pipeline)
> > > 
> > > > ## Thanks, Otis
> > > > 
> > > > Sematext is Hiring World-Wide --[Jobs - Sematext](http://sematext.com/about/jobs.html)

---

<div class="post-metadata">

**Author:** ![otisg](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/otisg/32/492_2.png) [@otisg](https://discuss.elastic.co/u/otisg)\
**Post date:** [December 7, 2011, 5:21am UTC](https://discuss.elastic.co/t/mysql-river/6072/6 "2011-12-07T05:21:52Z")

</div>

Hello,

On Dec 6, 2:48 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:

> The problem with a generic database river is handling deletes.

Solr's DIH relies on DB doing soft deletes. Could an ES RDBMS/MySQL  
River not do the same thing?

## Thanks, Otis

Sematext is Hiring World-Wide -- [Jobs](http://sematext.com/about/jobs.html)

> On Tue, Dec 6, 2011 at 8:49 PM, David Pilato [da...@pilato.fr](mailto:da...@pilato.fr) wrote:
> 
> > In my project (with postgresql), I choose to push documents in ES each  
> > time I create, update or delete an Hibernate entity.  
> > That's why I don't use a river by now.
> 
> > Problem is when I need to reindex all my documents from my database.  
> > At the present time, I create a "batch" for that but I'm thinking of using  
> > Talend for that.
> 
> > That's not a river, for sure !
> 
> > David 😉  
> > @dadoonet
> 
> > Le 6 déc. 2011 à 18:51, Otis Gospodnetic [otis.gospodne...@gmail.com](mailto:otis.gospodne...@gmail.com) a  
> > écrit :
> 
> > > Hi,
> 
> > > OK, so this is still external, except it's a push instead of a pull  
> > > (like in Solr's DIH).  
> > > So no MySQL River.
> 
> > > Is there a technical reason why one could _not_ implement a MySQL  
> > > River, or is it just that nobody had this MySQL==\>ES itch yet?
> 
> > > ## Thanks, Otis
> > > 
> > > Sematext is Hiring World-Wide --[Jobs](http://sematext.com/about/jobs.html)
> 
> > > On Dec 6, 3:08 am, Damien Hardy [damienhardy....@gmail.com](mailto:damienhardy....@gmail.com) wrote:
> > > 
> > > > Hello,
> 
> > > > Maybe you can look after MySQL trigger to fire a POST request to  
> > > > elasticsearch base on UDF like :  
> > > > [Google Code Archive - Long-term storage for Google Code Project Hosting.](http://code.google.com/p/mysql-udf-http/)
> 
> > > > Cheers,
> 
> > > > --  
> > > > Damien
> 
> > > > 2011/12/5 Otis Gospodnetic [otis.gospodne...@gmail.com](mailto:otis.gospodne...@gmail.com)
> 
> > > > > Hello,
> 
> > > > > How do people do MySQL (or any other RDBMS) indexing in bulk an  
> > > > > incrementally? Is there something like Solr's DataImportHandler in  
> > > > > ES?  
> > > > > I know there are Rivers and I _assume_ it makes sense to implement a  
> > > > > tool for indexing database content as a River, but I don't see a MySQL  
> > > > > River.
> 
> > > > > Or are people writing standalone and external indexer apps that push  
> > > > > data from DB to ES? (and thus probably creating a SPOF in their  
> > > > > indexing pipeline)
> 
> > > > > ## Thanks, Otis
> > > > > 
> > > > > Sematext is Hiring World-Wide --[Jobs](http://sematext.com/about/jobs.html)

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [December 7, 2011, 3:30pm UTC](https://discuss.elastic.co/t/mysql-river/6072/7 "2011-12-07T15:30:45Z")

</div>

By soft deletes you mean marking rows as deleted? Then its really up to the  
application you write to work like that. I personally not a fan of "xml  
configuration" to invent a (programming) language to query database and  
index. If one is needed, one can write something that does that easily and  
cron it (or write it as a river). Most times, the indexing process is more  
complex than just index a single table into a document, it involves joins  
and possibly more criteria that are best expressed (and maintained) in  
your favorite lang.

On Wed, Dec 7, 2011 at 7:21 AM, Otis Gospodnetic \<[otis.gospodnetic@gmail.com](mailto:otis.gospodnetic@gmail.com)

> wrote:

> Hello,
> 
> On Dec 6, 2:48 pm, Shay Banon [kim...@gmail.com](mailto:kim...@gmail.com) wrote:
> 
> > The problem with a generic database river is handling deletes.
> 
> Solr's DIH relies on DB doing soft deletes. Could an ES RDBMS/MySQL  
> River not do the same thing?
> 
> ## Thanks, Otis
> 
> Sematext is Hiring World-Wide -- [Jobs - Sematext](http://sematext.com/about/jobs.html)
> 
> > On Tue, Dec 6, 2011 at 8:49 PM, David Pilato [da...@pilato.fr](mailto:da...@pilato.fr) wrote:
> > 
> > > In my project (with postgresql), I choose to push documents in ES each  
> > > time I create, update or delete an Hibernate entity.  
> > > That's why I don't use a river by now.
> > 
> > > Problem is when I need to reindex all my documents from my database.  
> > > At the present time, I create a "batch" for that but I'm thinking of  
> > > using  
> > > Talend for that.
> > 
> > > That's not a river, for sure !
> > 
> > > David 😉  
> > > @dadoonet
> > 
> > > Le 6 déc. 2011 à 18:51, Otis Gospodnetic [otis.gospodne...@gmail.com](mailto:otis.gospodne...@gmail.com)  
> > > a  
> > > écrit :
> > 
> > > > Hi,
> > 
> > > > OK, so this is still external, except it's a push instead of a pull  
> > > > (like in Solr's DIH).  
> > > > So no MySQL River.
> > 
> > > > Is there a technical reason why one could _not_ implement a MySQL  
> > > > River, or is it just that nobody had this MySQL==\>ES itch yet?
> > 
> > > > ## Thanks, Otis
> > > > 
> > > > Sematext is Hiring World-Wide --[Jobs - Sematext](http://sematext.com/about/jobs.html)
> > 
> > > > On Dec 6, 3:08 am, Damien Hardy [damienhardy....@gmail.com](mailto:damienhardy....@gmail.com) wrote:
> > > > 
> > > > > Hello,
> > 
> > > > > Maybe you can look after MySQL trigger to fire a POST request to  
> > > > > elasticsearch base on UDF like :  
> > > > > [http://code.google.com/p/mysql-udf-http/](http://code.google.com/p/mysql-udf-http/)
> > 
> > > > > Cheers,
> > 
> > > > > --  
> > > > > Damien
> > 
> > > > > 2011/12/5 Otis Gospodnetic [otis.gospodne...@gmail.com](mailto:otis.gospodne...@gmail.com)
> > 
> > > > > > Hello,
> > 
> > > > > > How do people do MySQL (or any other RDBMS) indexing in bulk an  
> > > > > > incrementally? Is there something like Solr's DataImportHandler in  
> > > > > > ES?  
> > > > > > I know there are Rivers and I _assume_ it makes sense to implement  
> > > > > > a  
> > > > > > tool for indexing database content as a River, but I don't see a  
> > > > > > MySQL  
> > > > > > River.
> > 
> > > > > > Or are people writing standalone and external indexer apps that  
> > > > > > push  
> > > > > > data from DB to ES? (and thus probably creating a SPOF in their  
> > > > > > indexing pipeline)
> > 
> > > > > > ## Thanks, Otis
> > > > > > 
> > > > > > Sematext is Hiring World-Wide --  
> > > > > > [Jobs - Sematext](http://sematext.com/about/jobs.html)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 3:46am UTC](https://discuss.elastic.co/t/mysql-river/6072/8 "2017-07-06T03:46:13Z")

</div>


