# AWS/CloudTrail Integration for Elastic Agent

**URL:** <https://discuss.elastic.co/t/aws-cloudtrail-integration-for-elastic-agent/371766>\
**Category:** Elastic Observability\
**Created:** [December 10, 2024, 2:18pm UTC](https://discuss.elastic.co/t/aws-cloudtrail-integration-for-elastic-agent/371766 "2024-12-10T14:18:59Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![jmello31](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jmello31/32/139875_2.png) [@jmello31](https://discuss.elastic.co/u/jmello31)\
**Post date:** [December 10, 2024, 2:18pm UTC](https://discuss.elastic.co/t/aws-cloudtrail-integration-for-elastic-agent/371766/1 "2024-12-10T14:18:59Z")

</div>

Hello! I am using the default overarching AWS integration for Elastic Agent in order to collect CloudTrail logs from S3. I am successfully collecting logs, but I have a HUGE dataset and it is still processing logs from September even after I enabled it yesterday.

Does the integration read from when the trail first started and work towards real time? Is there a way to fix it so it only cares about new events from X start date? Would increasing CPU/RAM requests on my ingest nodes make it go faster?

TLDR: I have the CloudTrail integration within the AWS integration working successfully, but it will be stuck for days catching up to get to real time.

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [December 10, 2024, 3:05pm UTC](https://discuss.elastic.co/t/aws-cloudtrail-integration-for-elastic-agent/371766/2 "2024-12-10T15:05:57Z")

</div>

> [@jmello31](#):
>
> Does the integration read from when the trail first started and work towards real time? Is there a way to fix it so it only cares about new events from X start date?

I'm assuming you configured it to do polling on the s3 bucket with your cloudtrail logs, right?

If so, it will process all files in the bucket, depending on the number of files this can take a really long time as polling from s3 can be pretty slow.

> [@jmello31](#):
>
> Would increasing CPU/RAM requests on my ingest nodes make it go faster?

I don't think so, you can try to increase the number of workers in the integration to see if this improves.

 ![Screenshot from 2024-12-10 11-45-51](https://us1.discourse-cdn.com/elastic/original/3X/8/d/8d6d566767dad0e9b472b01da0c5b1f017ffa398.png)

Another thing is, depending on the amount of events you get in your Cloudtrail, you will never be able to have it in realtime using polling mode as it does not scale, the recommendation is to use SQS notifications, this allows you to have multiple agents consuming the data in the case of a high rate cloudtrail bucket.

But SQS notifications only works from the moment they were configured, it does not work for data that it is already in the bucket.

---

<div class="post-metadata">

**Author:** ![jmello31](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jmello31/32/139875_2.png) [@jmello31](https://discuss.elastic.co/u/jmello31)\
**Post date:** [December 10, 2024, 3:09pm UTC](https://discuss.elastic.co/t/aws-cloudtrail-integration-for-elastic-agent/371766/3 "2024-12-10T15:09:03Z")

</div>

This was very helpful, thank you!

---

<div class="post-metadata">

**Author:** ![jmello31](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jmello31/32/139875_2.png) [@jmello31](https://discuss.elastic.co/u/jmello31)\
**Post date:** [December 10, 2024, 3:15pm UTC](https://discuss.elastic.co/t/aws-cloudtrail-integration-for-elastic-agent/371766/4 "2024-12-10T15:15:47Z")

</div>

Should I use the dedicated CloudTrail Integration you think? Would I just have to enable event notification to SQS in my Cloudtrail bucket and then provide the queue name to the integration?

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [December 10, 2024, 3:30pm UTC](https://discuss.elastic.co/t/aws-cloudtrail-integration-for-elastic-agent/371766/5 "2024-12-10T15:30:20Z")

</div>

> [@jmello31](#):
>
> Should I use the dedicated CloudTrail Integration you think?

It is the same integration, the difference is if you add the _AWS Cloudtrail_ integration, it will show you only the cloudtrail settings to configure, after you install it if you go to edit you will see all other AWS integrations as disabled.

> [@jmello31](#):
>
> Would I just have to enable event notification to SQS in my Cloudtrail bucket and then provide the queue name to the integration?

You need to configure notifications from your cloudtrail s3 bucket to a sqs queue, and then use this queue in the configuration in Elastic Agent.

But this works for objects created after the notification was configured.

You cannot have both configured, so I recommend that first you wait for it to process the old data, if you do not care for old data, then just disable it and configure to use the SQS notifications.

---

<div class="post-metadata">

**Author:** ![jmello31](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jmello31/32/139875_2.png) [@jmello31](https://discuss.elastic.co/u/jmello31)\
**Post date:** [December 10, 2024, 3:33pm UTC](https://discuss.elastic.co/t/aws-cloudtrail-integration-for-elastic-agent/371766/6 "2024-12-10T15:33:12Z")

</div>

Perfect, thank you so much! The main goal is to ingest Cloudtrail logs in real time as much as possible. It definitely looks like the route you mentioned with SQS is the best way to achieve that!

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [December 10, 2024, 3:34pm UTC](https://discuss.elastic.co/t/aws-cloudtrail-integration-for-elastic-agent/371766/7 "2024-12-10T15:34:44Z")

</div>

Yeah, using SQS notifications is the best approach as you can scale the number of agents.

I think this is the AWS documention: [Walkthrough: Configuring a bucket for notifications (SNS topic or SQS queue) - Amazon Simple Storage Service](https://docs.aws.amazon.com/AmazonS3/latest/userguide/ways-to-add-notification-config-to-bucket.html#step1-create-sqs-queue-for-notification)

---

<div class="post-metadata">

**Author:** ![strawgate](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/strawgate/32/131008_2.png) [@strawgate](https://discuss.elastic.co/u/strawgate)\
**Post date:** [December 11, 2024, 12:38am UTC](https://discuss.elastic.co/t/aws-cloudtrail-integration-for-elastic-agent/371766/8 "2024-12-11T00:38:21Z")

</div>

If you are using elastic cloud you could also explore sending data using firehose, see [Amazon Kinesis Data Firehose overview | Amazon Kinesis Data Firehose Ingest Guide | Elastic](https://www.elastic.co/guide/en/kinesis/master/aws-firehose.html)

This is a great option as it really simplifies the architecture.

Otherwise, if you're not in cloud I would echo the previous recommendation and strongly encourage using SQS + S3. To improve throughput I would recommend following the recommendations here [Get the most from Elastic Agent with Amazon S3 and SQS | Elastic Blog](https://www.elastic.co/blog/elastic-agent-amazon-s3-sqs) which recommend using the "throughput" performance preset (see more info here: [Using Elastic Agent Performance Presets in 8.12 | Elastic Blog](https://www.elastic.co/blog/using-elastic-agent-performance-presets-in-8-12)) under fleet \> settings \> output
