# Elastic Cloud Persistent Queue?

**URL:** <https://discuss.elastic.co/t/elastic-cloud-persistent-queue/345066>\
**Category:** Elasticsearch\
**Created:** [October 16, 2023, 6:03am UTC](https://discuss.elastic.co/t/elastic-cloud-persistent-queue/345066 "2023-10-16T06:03:33Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![wwalker](https://avatars.discourse-cdn.com/v4/letter/w/43a26b/32.png) [@wwalker](https://discuss.elastic.co/u/wwalker)\
**Post date:** [October 16, 2023, 6:03am UTC](https://discuss.elastic.co/t/elastic-cloud-persistent-queue/345066/1 "2023-10-16T06:03:33Z")

</div>

I am ingesting logs from an on-prem logstash to Elastic Cloud. My Logstash instance has persistent queue enabled. I ingested a large set of data, about 50 million events from my on-prem Elasticsearch instance using the Elasticsearch input, and saw the Logstash queue fill and then empty. My Logstash queue has been empty for five hours now and when I search for the logs that I ingested, the number of events on-prem do not match what's in the cloud. When I refresh the cloud instance, the number is slowly going up. Is there some kind of queue or cache of some sort in Elastic Cloud?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 16, 2023, 6:21am UTC](https://discuss.elastic.co/t/elastic-cloud-persistent-queue/345066/2 "2023-10-16T06:21:11Z")

</div>

> [@wwalker](#):
>
> My Logstash queue has been empty for five hours now and when I search for the logs that I ingested, the number of events on-prem do not match what's in the cloud.

How do you check/verify this? Are you using the [cat indices API](https://www.elastic.co/guide/en/elasticsearch/reference/8.10/cat-indices.html)?

> [@wwalker](#):
>
> When I refresh the cloud instance, the number is slowly going up. Is there some kind of queue or cache of some sort in Elastic Cloud?

No, there is no built in queue. Depending on how you are verifying the document count, one reason for data to show up over time is that the timesatmp set for indexed data is incorrect. All timestamps indexed into Elasticsearch are in UTC timezone so ingesting local timezone timestamps without timezone specified may make data appear to show up over time as they future dated.

It would help if you could share a document that showed up late as well as your Logstash pipeline.

---

<div class="post-metadata">

**Author:** ![wwalker](https://avatars.discourse-cdn.com/v4/letter/w/43a26b/32.png) [@wwalker](https://discuss.elastic.co/u/wwalker)\
**Post date:** [October 16, 2023, 6:34am UTC](https://discuss.elastic.co/t/elastic-cloud-persistent-queue/345066/3 "2023-10-16T06:34:59Z")

</div>

I configured logstash to pull all documents in a specific index and configured the input to also add the old index name to a new field in each event. I then pull up Kibana's Discover and filter on-prem and Elastic Cloud by that index name to get the total number.

I'm checking my Logstash persistent queue using Kibana's monitoring page. However, I just looked at the pipeline metrics and it shows it's still pulling in documents on the Elasticsearch input...at a significantly slower rate. Initial ingest rate was about 4,100 events/second but it tailed off to about 900 events/second....wonder why it's doing that...

![image](https://us1.discourse-cdn.com/elastic/original/3X/3/e/3ea9fe5775d1b5f8c752f934ee1b23d71dd568eb.png)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 16, 2023, 7:13am UTC](https://discuss.elastic.co/t/elastic-cloud-persistent-queue/345066/4 "2023-10-16T07:13:12Z")

</div>

Are you specifying the document ID in your Elasticsearch output? If you do, each insert will be treated as a potential update, which often slows down the indexing throughput as the dstimation index size grows.

---

<div class="post-metadata">

**Author:** ![wwalker](https://avatars.discourse-cdn.com/v4/letter/w/43a26b/32.png) [@wwalker](https://discuss.elastic.co/u/wwalker)\
**Post date:** [October 16, 2023, 1:21pm UTC](https://discuss.elastic.co/t/elastic-cloud-persistent-queue/345066/5 "2023-10-16T13:21:13Z")

</div>

> [@Christian\_Dahlqvist](#):
>
> Are you specifying the document ID in your Elasticsearch output?

At the moment, I am not.

---

<div class="post-metadata">

**Author:** ![wwalker](https://avatars.discourse-cdn.com/v4/letter/w/43a26b/32.png) [@wwalker](https://discuss.elastic.co/u/wwalker)\
**Post date:** [October 23, 2023, 4:00pm UTC](https://discuss.elastic.co/t/elastic-cloud-persistent-queue/345066/6 "2023-10-23T16:00:53Z")

</div>

I managed to flatten the curve and increase throughput to where my cloud instance is now the bottleneck. I currently have 8 vCPU assigned to the VM with 32 GB of RAM. I upped `pipeline.workers` to 32, set the Elasticsearch `size` to 7500, and `scroll` to 8m.

Initial ingest hits 9,400 e/s with a final rate of 3,400 e/s. 2.29 and 3.7 times higher respectively...can't complain about those gains. 45 million documents processed by Logstash in about 2 hours.

![image](https://us1.discourse-cdn.com/elastic/original/3X/c/c/cc1b10c3df8e980b8c280c2a4681ba1f4ad2c634.png)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 20, 2023, 4:01pm UTC](https://discuss.elastic.co/t/elastic-cloud-persistent-queue/345066/7 "2023-11-20T16:01:03Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
