# Continuous data read using Elasticsearch input plugin in logstash

**URL:** <https://discuss.elastic.co/t/continuous-data-read-using-elasticsearch-input-plugin-in-logstash/365198>\
**Category:** Logstash\
**Created:** [August 20, 2024, 12:30pm UTC](https://discuss.elastic.co/t/continuous-data-read-using-elasticsearch-input-plugin-in-logstash/365198 "2024-08-20T12:30:11Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![venkatkumar229](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/venkatkumar229/32/104663_2.png) [@venkatkumar229](https://discuss.elastic.co/u/venkatkumar229)\
**Post date:** [August 20, 2024, 12:30pm UTC](https://discuss.elastic.co/t/continuous-data-read-using-elasticsearch-input-plugin-in-logstash/365198/1 "2024-08-20T12:30:11Z")

</div>

Hi Team,

I have a requirement where i need to read the data from an elasticsearch index(datastream) and perform "split" operation to split the array records into individual documents and store into a destination index(datastream) in elasticsearch.

I have created a logstash pipeline which will run for every hour and read the data from last 1hour. But i have observed that sometimes i am seeing some data missing without any errors in logstash and no data in dead letter queue or in persistent queue. Is there any way that we can track the documents processed and make sure that the no documents got missed while reading/processing the data.

my Pipeline:

```auto
input {
       elasticsearch {
        cloud_id => "xxxxxxxxxx"
        index => "orders-data"
        query => '{"query": {"range": {"event.ingest": {"gte": "now-1h"}}}}'
        schedule => "0 * * * *"
  }
}
 
 
filter {
    split {
        field => "[orders]"
    }
 
}
 
output {
            elasticsearch {
        cloud_id => "xxxx"
        index => "splitted-orders"
        ssl => true
        action => "create"
    }
}

```

When i am working on indexes i used to update the source index by adding a field called processed = true when logstash processed the document. But now i am unable to updated the source index as it is a datastream.

```auto
input {
       elasticsearch {
        cloud_id => "xxxxxxxxxx"
        index => "orders-data"
        query => '{"query": {"range": {"event.ingest": {"gte": "now-1h"}}}}'
        schedule => "0 * * * *"
		docinfo => true
        docinfo_target => "[@metadata][doc]"
  }
}
 
 
filter {
	mutate{
		add_field => {"[processed]" => true}
	}
    split {
        field => "[orders]"
    }
 
}
 
output {
    elasticsearch {
        cloud_id => "xxxx"
        index => "orders-data"
        ssl => true
        action => "update"
		document_id => "%{[@metadata][doc][_id]}"
    }
    elasticsearch {
        cloud_id => "xxxx"
        index => "splitted-orders"
        ssl => true
        action => "create"
    }
}

```

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [August 20, 2024, 12:43pm UTC](https://discuss.elastic.co/t/continuous-data-read-using-elasticsearch-input-plugin-in-logstash/365198/2 "2024-08-20T12:43:42Z")

</div>

> [@venkatkumar229](#):
>
> I have created a logstash pipeline which will run for every hour and read the data from last 1hour. But i have observed that sometimes i am seeing some data missing without any errors in logstash and no data in dead letter queue or in persistent queue.

How did you identify the missing documents? Have you checked their `event.ingest` time to see if there is some pattern?

Your issue may be related to the fact that your query will run at the minute 0 of every hour, but you are looking for `now-1h`, this may lead to gaps because there is no guarantee that the query will run at the exact time, it may have some delay, and also you may have some documents added in the end of the last hour that were still not searchable.

Try to increase the `now-1h` to something like `now-70m` or `now-80m`.

Your look back time needs to be larger than your schedule.

---

<div class="post-metadata">

**Author:** ![venkatkumar229](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/venkatkumar229/32/104663_2.png) [@venkatkumar229](https://discuss.elastic.co/u/venkatkumar229)\
**Post date:** [August 21, 2024, 11:15am UTC](https://discuss.elastic.co/t/continuous-data-read-using-elasticsearch-input-plugin-in-logstash/365198/3 "2024-08-21T11:15:45Z")

</div>

Hi @leandrojmp ,

If i increase the now-1h to now-70m i may i get the duplicate records right?

Regards,  
Praveen Kumar
