# Logstash Elasticsearch Input - Query in @timestamp order

**URL:** <https://discuss.elastic.co/t/logstash-elasticsearch-input-query-in-timestamp-order/210607>\
**Category:** Logstash\
**Created:** [December 5, 2019, 12:18am UTC](https://discuss.elastic.co/t/logstash-elasticsearch-input-query-in-timestamp-order/210607 "2019-12-05T00:18:20Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![Paul\_Ainslie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/paul_ainslie/32/55031_2.png) [@Paul\_Ainslie](https://discuss.elastic.co/u/Paul_Ainslie)\
**Post date:** [December 5, 2019, 12:18am UTC](https://discuss.elastic.co/t/logstash-elasticsearch-input-query-in-timestamp-order/210607/1 "2019-12-05T00:18:20Z")

</div>

I'm querying an index with Nginx access records, running it thru the Aggregate filter plugin and writing it to a file. However, the elasticsearch input plugin is not reading the records in @timestamp order - they are written in a random order.

Furthermore, the logstash job has been running for over 4 days on an index that has 17 million records and 11.2 GB of data. I wonder if it's timing out and starting from the beginning again.

Any suggestions given the following config?

```auto
    input {
        elasticsearch {
          hosts => ["https://m-elsc-006.on4cdn.net:9200"]
          index => "mediaserver-2019.11.11-000020"
          user => "elastic"
          password => " ******************"
          size => 5000
          scroll => "5m"
          query => '{ "sort": ["@timestamp"] }'
          docinfo => true
          docinfo_fields => ["_type", "_id"]
        }
    }

```

Of note, I always get this error (even without the aggregate filter)

```auto
[WARN][org.logstash.instrument.metrics.gauge.LazyDelegatingGauge][mediaserver_access_TEST] A gauge metric of an unknown type (org.jruby.RubyArray) has been create for key: cluster_uuids. This may result in invalid serialization. It is recommended to log an issue to the responsible developer/development team.

```

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [December 5, 2019, 12:50am UTC](https://discuss.elastic.co/t/logstash-elasticsearch-input-query-in-timestamp-order/210607/2 "2019-12-05T00:50:55Z")

</div>

If pipeline.java\_execution is enabled (which became the default in v7.0) then logstash [will](https://github.com/elastic/logstash/issues/10938) re-order events even with pipeline.workers set to 1.

If you have pipeline.workers set to 1 and java\_execution disabled then you may have an elasticsearch question rather than a logstash question.

That WARN is completely normal in recent versions.

---

<div class="post-metadata">

**Author:** ![Paul\_Ainslie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/paul_ainslie/32/55031_2.png) [@Paul\_Ainslie](https://discuss.elastic.co/u/Paul_Ainslie)\
**Post date:** [December 5, 2019, 4:17am UTC](https://discuss.elastic.co/t/logstash-elasticsearch-input-query-in-timestamp-order/210607/3 "2019-12-05T04:17:22Z")

</div>

I'm running Logstash 7.4 (and reading from Elasticsearch 7.4). I've now set the following in logstash.yml and it's the same problem. 😒 It's processing 5 days worth of log files... it gets thru 5 days of records in @timstamp order (~33,000 records out of 17.1 million) , then loops back to the beginning and processes another 5 days/~30,000 records, etc.

```auto
pipeline.workers: 1
pipeline.java_execution: false

```

---

<div class="post-metadata">

**Author:** ![BeMoore](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/bemoore/32/58724_2.png) [@BeMoore](https://discuss.elastic.co/u/BeMoore)\
**Post date:** [December 5, 2019, 10:46am UTC](https://discuss.elastic.co/t/logstash-elasticsearch-input-query-in-timestamp-order/210607/4 "2019-12-05T10:46:45Z")

</div>

Paul,

It sounds like you need to use a sincedb to maintain track of where your logstash is getting to in its reading of the logfiles.

read this

[https://www.elastic.co/guide/en/logstash/current/plugins-inputs-file.html#\_tracking\_of\_current\_position\_in\_watched\_files](https://www.elastic.co/guide/en/logstash/current/plugins-inputs-file.html#_tracking_of_current_position_in_watched_files)

i would create a directory inside your /etc/logstash called sincedb  
then add the following to your input section of your logstash .conf in conf.d folder

`sincedb_path => "/etc/logstash/sincedb/myfilesincedb_name.dat"`

this way, it should only read to the last recorded point ( in that dat file you'll see a timestamp ) it won't re-read this whole file, but only events after the time stamp.

So another trick for ingestion from beginning is to also set in the input section of your conf file 🙂  
`start_position => "beginning"`

so that is knows to start at the top and work its ways down to the end. once its completed and change the "beginning" to "end" and it will only ever start from the end of the log file. You can also mix and match start\_position and sincedb in the same conf, it just depends on what your aim is. Experiment and have fun!

---

<div class="post-metadata">

**Author:** ![Paul\_Ainslie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/paul_ainslie/32/55031_2.png) [@Paul\_Ainslie](https://discuss.elastic.co/u/Paul_Ainslie)\
**Post date:** [December 5, 2019, 1:36pm UTC](https://discuss.elastic.co/t/logstash-elasticsearch-input-query-in-timestamp-order/210607/5 "2019-12-05T13:36:49Z")

</div>

The `synced_path` option doesn't work with the elasticsearch input plugin, I get [ERROR]..."Something is wrong with your configuration."

I'm now trying it without the aggregate filter plugin (see below) - still the same problem.

```auto
input {
    elasticsearch {
      hosts => ["https://m-elsc-006.on4cdn.net:9200"]
      index => "mediaserver-2019.11.11-000020"
      user => "elastic"
      password => " **************"
      size => 5000
      scroll => "10m"
      query => '{ "sort": ["@timestamp"] }'
      docinfo => true
# docinfo_fields => ["_type", "_id"]
    }
}

output {
    file {
        path => "/var/log/logstash/2019.11.11-000020_TEST.json"
        codec => "json_lines"
    }
}

```

---

<div class="post-metadata">

**Author:** ![Paul\_Ainslie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/paul_ainslie/32/55031_2.png) [@Paul\_Ainslie](https://discuss.elastic.co/u/Paul_Ainslie)\
**Post date:** [December 5, 2019, 10:53pm UTC](https://discuss.elastic.co/t/logstash-elasticsearch-input-query-in-timestamp-order/210607/6 "2019-12-05T22:53:52Z")

</div>

@Badger You mentioned that, "...you may have an elasticsearch question rather than a logstash question." Is there any way for me to test and confirm that this is an Elasticsearch issue?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [December 5, 2019, 10:57pm UTC](https://discuss.elastic.co/t/logstash-elasticsearch-input-query-in-timestamp-order/210607/7 "2019-12-05T22:57:58Z")

</div>

Is the pipeline getting restarted?

---

<div class="post-metadata">

**Author:** ![Paul\_Ainslie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/paul_ainslie/32/55031_2.png) [@Paul\_Ainslie](https://discuss.elastic.co/u/Paul_Ainslie)\
**Post date:** [December 5, 2019, 11:09pm UTC](https://discuss.elastic.co/t/logstash-elasticsearch-input-query-in-timestamp-order/210607/8 "2019-12-05T23:09:36Z")

</div>

The logstash-plain.log file doesn't show any issues. But it's only reading a small percentage of the index and going back to the beginning to read more records. Not sure if that's what you're asking.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [December 6, 2019, 1:18am UTC](https://discuss.elastic.co/t/logstash-elasticsearch-input-query-in-timestamp-order/210607/9 "2019-12-06T01:18:28Z")

</div>

If the pipeline restarted it would re-run the same query and start fetching the same set of records. That suggests to me that it is being restarted. I would expect that to get logged.

When I said you might have an elasticsearch question I meant that the query might be wrong.

---

<div class="post-metadata">

**Author:** ![Paul\_Ainslie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/paul_ainslie/32/55031_2.png) [@Paul\_Ainslie](https://discuss.elastic.co/u/Paul_Ainslie)\
**Post date:** [December 6, 2019, 1:46am UTC](https://discuss.elastic.co/t/logstash-elasticsearch-input-query-in-timestamp-order/210607/10 "2019-12-06T01:46:37Z")

</div>

My bad, yes, the pipeline is getting restarted (I presume when it's completed). I'm running it as a service, is that my problem?

These messages keep happening every hour or so.

```auto
[INFO][logstash.pipeline][mediaserver_access_TEST] Pipeline has terminated {:pipeline_id=>"mediaserver_access_TEST", :thread=>"#<Thread:0x262149ab run>"}
[INFO][logstash.runner] Logstash shut down.

```

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [December 6, 2019, 1:38pm UTC](https://discuss.elastic.co/t/logstash-elasticsearch-input-query-in-timestamp-order/210607/11 "2019-12-06T13:38:07Z")

</div>

If logstash restarts the elasticsearch input will run the same query again and fetch the same data.

---

<div class="post-metadata">

**Author:** ![Paul\_Ainslie](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/paul_ainslie/32/55031_2.png) [@Paul\_Ainslie](https://discuss.elastic.co/u/Paul_Ainslie)\
**Post date:** [December 7, 2019, 9:33pm UTC](https://discuss.elastic.co/t/logstash-elasticsearch-input-query-in-timestamp-order/210607/12 "2019-12-07T21:33:15Z")

</div>

The only way I could solve this was by running Logstash on the command line (vs. as a service). I ran it like this...

```auto
sudo nohup /usr/share/logstash/bin/logstash --path.settings /etc/logstash/ > ~/logstash_oneshot.log &

```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 4, 2020, 9:33pm UTC](https://discuss.elastic.co/t/logstash-elasticsearch-input-query-in-timestamp-order/210607/13 "2020-01-04T21:33:17Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
