# Logstash JDBC Input Plugin for streaming data

**URL:** <https://discuss.elastic.co/t/logstash-jdbc-input-plugin-for-streaming-data/28033>\
**Category:** Logstash\
**Created:** [August 25, 2015, 3:56pm UTC](https://discuss.elastic.co/t/logstash-jdbc-input-plugin-for-streaming-data/28033 "2015-08-25T15:56:45Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Navneet\_Mathpal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/navneet_mathpal/32/3677_2.png) [@Navneet\_Mathpal](https://discuss.elastic.co/u/Navneet_Mathpal)\
**Post date:** [August 25, 2015, 3:56pm UTC](https://discuss.elastic.co/t/logstash-jdbc-input-plugin-for-streaming-data/28033/1 "2015-08-25T15:56:45Z")

</div>

Can we use logstash-jdbc-input plugin for streaming data , ex: if I have data available in my database and I run JDBC input plugin , it will index the data into es, but if after some time more data comes of database , Is jdbc input plugin is able to index that data without restarting the logstash ?

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [August 25, 2015, 5:42pm UTC](https://discuss.elastic.co/t/logstash-jdbc-input-plugin-for-streaming-data/28033/2 "2015-08-25T17:42:59Z")

</div>

Yes, you can run it periodically and have it pick up new data. See the [State section of the documentation](https://www.elastic.co/guide/en/logstash/current/plugins-inputs-jdbc.html#_state).

---

<div class="post-metadata">

**Author:** ![Navneet\_Mathpal](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/navneet_mathpal/32/3677_2.png) [@Navneet\_Mathpal](https://discuss.elastic.co/u/Navneet_Mathpal)\
**Post date:** [August 26, 2015, 5:01am UTC](https://discuss.elastic.co/t/logstash-jdbc-input-plugin-for-streaming-data/28033/3 "2015-08-26T05:01:52Z")

</div>

@magnusbaeck I have a concern is that , If I have a very large data set , and once it indexes the data into es and if I schedule it again , will it again try to index the whole data set ?  
because if it does so , it would be an overburden on application , isn't ?  
(I know we can handle the duplicate rows but as a performance point of view how feasible it would be ?)

Thanks 😃

---

<div class="post-metadata">

**Author:** ![magnusbaeck](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/magnusbaeck/32/44943_2.png) [@magnusbaeck](https://discuss.elastic.co/u/magnusbaeck)\
**Post date:** [August 26, 2015, 5:46am UTC](https://discuss.elastic.co/t/logstash-jdbc-input-plugin-for-streaming-data/28033/4 "2015-08-26T05:46:40Z")

</div>

As the documentation I linked to tries to explain, your query is served with a parameter that contains the timestamp when the query was run the last time. You can use that to select only the rows that have been updated since the last run. Duplicates in the Elasticsearch output can be avoided by setting the document id to e.g. the primary key from the source database.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 5:30am UTC](https://discuss.elastic.co/t/logstash-jdbc-input-plugin-for-streaming-data/28033/5 "2017-07-06T05:30:54Z")

</div>


