# Synchronise websites and Elasticsearch

**URL:** <https://discuss.elastic.co/t/synchronise-websites-and-elasticsearch/60710>\
**Category:** Elasticsearch\
**Created:** [September 16, 2016, 1:07pm UTC](https://discuss.elastic.co/t/synchronise-websites-and-elasticsearch/60710 "2016-09-16T13:07:10Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![adrianolimit](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/adrianolimit/32/11930_2.png) [@adrianolimit](https://discuss.elastic.co/u/adrianolimit)\
**Post date:** [September 16, 2016, 1:07pm UTC](https://discuss.elastic.co/t/synchronise-websites-and-elasticsearch/60710/1 "2016-09-16T13:07:10Z")

</div>

Hello everyone,

Do you have any idea how can I process to synchronise my websites (built with Wordpress, Drupal, Joomla...) with my Elasticsearch?

Thank you in advance.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [September 16, 2016, 1:20pm UTC](https://discuss.elastic.co/t/synchronise-websites-and-elasticsearch/60710/2 "2016-09-16T13:20:42Z")

</div>

You can may be try to find a webcrawler but IMO it would be too much generic.  
I'd use dedicated connectors.

For example, here is an article for Drupal which can help you: [http://redcrackle.com/blog/configuring-drupal-elasticsearch-facet-search-functionality](http://redcrackle.com/blog/configuring-drupal-elasticsearch-facet-search-functionality)

I hope this helps.

---

<div class="post-metadata">

**Author:** ![adrianolimit](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/adrianolimit/32/11930_2.png) [@adrianolimit](https://discuss.elastic.co/u/adrianolimit)\
**Post date:** [September 16, 2016, 9:49pm UTC](https://discuss.elastic.co/t/synchronise-websites-and-elasticsearch/60710/3 "2016-09-16T21:49:12Z")

</div>

> [@dadoonet](#):
>
> IMO

Thank you @dadoonet for your answer, I will read this article. Do you think that is better to use webcrawler/dedicated connector or do something directly in the database (trigger/transaction/logs)?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [September 17, 2016, 7:29am UTC](https://discuss.elastic.co/t/synchronise-websites-and-elasticsearch/60710/4 "2016-09-17T07:29:52Z")

</div>

I always prefer sending the data to elasticsearch within the same "transaction" which saves your data to the database.  
I wrote an article about it: [http://david.pilato.fr/blog/2015/05/09/advanced-search-for-your-legacy-application/](http://david.pilato.fr/blog/2015/05/09/advanced-search-for-your-legacy-application/)

Another approach could be to reindex all your system every night in another index and then switch the alias but it's far away from real time. I mean that it works well if you don't care about updates in your DB during the day.

The closer you are to the application which is generating the data, the better.  
So if you are using Drupal and have a connector for that, you should use it.  
Same for other systems.

If you can't do that, because there is no way to extend the application, then yes you can use [logstash](https://www.elastic.co/guide/en/logstash/current/plugins-inputs-jdbc.html) or [elasticsearch-jdbc](https://github.com/jprante/elasticsearch-jdbc) for that. Note that dealing with updates and deletes could be hard.

HTH

---

<div class="post-metadata">

**Author:** ![adrianolimit](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/adrianolimit/32/11930_2.png) [@adrianolimit](https://discuss.elastic.co/u/adrianolimit)\
**Post date:** [September 20, 2016, 8:55am UTC](https://discuss.elastic.co/t/synchronise-websites-and-elasticsearch/60710/5 "2016-09-20T08:55:54Z")

</div>

@dadoonet Your article is very interesting 🙂

I would like to set up this process :

1. Run logstash each \* minutes with jdbc plugin in order to do "select \* from ..." and store the data in the index\_date
2. When the indexation is finished, I would like to add an alias to my index like "index\_current" where my application will do searches.
3. Next time that logstash will run, I will reproduce the process and create new index\_date and when the indexation will be finished I will add an alias "index\_current" to my alias and delete the oldest.

Do you think it is a good idea?

I have one question : How can I know when Logstash has finished to collect the data?

Thank you in advance.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [September 20, 2016, 12:26pm UTC](https://discuss.elastic.co/t/synchronise-websites-and-elasticsearch/60710/6 "2016-09-20T12:26:53Z")

</div>

> Do you think it is a good idea?

Yes.

> How can I know when Logstash has finished to collect the data?

I think that Logstash will exit after the end of the job.

Look at the documentation: [Jdbc input plugin | Logstash Reference [8.11] | Elastic](https://www.elastic.co/guide/en/logstash/current/plugins-inputs-jdbc.html)

> You can periodically schedule ingestion using a cron syntax (see schedule setting) or run the query one time to load data into Logstash.

And [Jdbc input plugin | Logstash Reference [8.11] | Elastic](https://www.elastic.co/guide/en/logstash/current/plugins-inputs-jdbc.html#plugins-inputs-jdbc-schedule)

So if you don't set `schedule` your logstash job will end after having processed all the data.

---

<div class="post-metadata">

**Author:** ![adrianolimit](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/adrianolimit/32/11930_2.png) [@adrianolimit](https://discuss.elastic.co/u/adrianolimit)\
**Post date:** [September 20, 2016, 1:36pm UTC](https://discuss.elastic.co/t/synchronise-websites-and-elasticsearch/60710/7 "2016-09-20T13:36:04Z")

</div>

Thank you again @dadoonet for your response, I checked and it's true, Logstash exits at the end of the job 🙂

---

<div class="post-metadata">

**Author:** ![adrianolimit](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/adrianolimit/32/11930_2.png) [@adrianolimit](https://discuss.elastic.co/u/adrianolimit)\
**Post date:** [September 21, 2016, 11:20pm UTC](https://discuss.elastic.co/t/synchronise-websites-and-elasticsearch/60710/8 "2016-09-21T23:20:43Z")

</div>

I think that is one of the best solutions in order to have zero downtime, the other solution as you have mentioned would be to execute Elasticsearch's commands directly from the application at the same time than the database's transactions (Mysql, Oracle...).

Thank you again for your advice @dadoonet

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:18pm UTC](https://discuss.elastic.co/t/synchronise-websites-and-elasticsearch/60710/9 "2017-07-05T22:18:21Z")

</div>


