Skip to main content

The logging feedback loop that quadrupled an AWS bill


A client’s AWS bill went from roughly $1.7K in one month to $7.8K the next. Nothing in the application had changed. Three security and logging features, each reasonable on its own, had been connected so that each one’s output became another’s input. Logging needs to be designed as a system, because a single console change can turn it into a loop that pays for itself to grow.

The three ingredients

This account had recently had its security tooling built out quickly. By early September it had two of the three pieces in place:

  • GuardDuty Malware Protection for S3 on the CloudTrail log bucket. This feature scans every new object in a bucket, and it is billed per object.
  • A CloudTrail trail recording S3 data events for every bucket, with no bucket filter, and streaming those events to CloudWatch Logs as well. The third piece arrived over about five minutes one night, when S3 server access logging was turned on for 14 buckets by hand in the console. Every one of them was pointed at the CloudTrail log bucket, and that included the CloudTrail bucket itself.

How the loop worked

There were two loops, and the malware scanner sat in both.

Loop A. GuardDuty scans a new log file, which means reading it and writing a tag with the result. Access logging on the bucket records those requests and delivers a new access-log file into the same bucket. That file is new, so GuardDuty scans it, and the cycle repeats.

Loop B. The trail records the same reads, tag writes and log deliveries as S3 data events, and writes them back into the bucket as more log files. Those get scanned too, and each scan generates more data events.

AWS’s own documentation says the destination bucket for access logs should not have access logging enabled, and that delivering logs to the source bucket causes an infinite loop. The malware plan turned that loop into a large bill, because every turn of it was a billable scan.

The evidence of a closed loop was clear. In a two-hour sample on the second day, GuardDuty tag writes and access-log deliveries on the CloudTrail bucket ran almost one-for-one, at about 57,000 an hour each. Application traffic does not look like that.

What it cost

Line itemMonth beforeLoop month
GuardDuty malware scanshundreds of thousands, tens of dollars~27 million, over $5,000
CloudTrail data events recorded~2.6 million~110 million
CloudTrail bucket objectsunder 1 millionover 22 million
Total monthly bill~$1.7K~$7.8K

The loop ran for about three weeks at roughly $250–290 a day. Smaller costs followed the same curve: S3 request charges, CloudWatch Logs ingestion of the extra data events, EventBridge events for each new object, and a support plan billed as a percentage of total spend. If it had been left running, the estimate was about $8,700 a month.

Why nobody caught it for three weeks

The signals were there from the first full day. Daily scans jumped from about 12,000 to about 790,000 and kept climbing. A budget forecast alert should have fired the next day, and the actual-spend alerts within a week. All of them went to a single mailbox with no named owner, and nobody acted on them. Cost anomaly detection wasn’t set up until the evening the loop was finally addressed.

The first remediation attempt also shows how this goes wrong. The malware plan on the log bucket was deleted, which stopped the scanning. A new plan was then created on the same bucket 42 seconds later, which restarted the loop until it was removed again. Removing the scanner treats the symptom. The feedback path is the self-targeted logging plus the unfiltered data-event selector, and that has to go first, or anything per-object attached to the bucket later can start the loop again.

How to check your own account

You don’t need special tooling.

  • Cost Explorer, grouped by service and then by usage type, daily. Look for MalwareProtectionS3ScanRequest and DataEventsRecorded lines. A per-service jump from a few dollars a day to hundreds is unmistakable once you look at it daily rather than monthly.
  • GuardDuty usage statistics. Look at which buckets and features account for the cost, and check for malware plans on buckets that only receive logs.
  • Bucket logging targets. List every bucket’s logging configuration and flag any bucket that logs into itself, or into a bucket that something else scans or triggers on.
for b in $(aws s3api list-buckets --query 'Buckets[].Name' --output text); do
  t=$(aws s3api get-bucket-logging --bucket "$b" \
      --query 'LoggingEnabled.TargetBucket' --output text)
  [ "$t" = "$b" ] && echo "SELF-LOGGING: $b"
  [ "$t" != "None" ] && echo "$b -> $t"
done
  • Object counts. CloudWatch’s daily NumberOfObjects metric for your log buckets. Ours went from under a million objects to over 22 million in about two and a half weeks.

Designing logging so it can’t loop

  • Use a dedicated access-log bucket that has its own access logging off, no malware plan, a per-source prefix, an expiration rule, and a bucket policy that only lets the S3 logging service write to it.
  • Never send access logs into the CloudTrail bucket. Keep log types separate, so lifecycle rules, queries and permissions can be applied to each.
  • Exclude log buckets from CloudTrail data events. Add resources.ARN NotStartsWith conditions for your log buckets to the advanced event selector. Better still, record data events only for buckets that hold application data.
  • Put Malware Protection for S3 only on buckets that take uploads from untrusted sources, and scope it to the upload prefixes. Gzipped JSON from AWS services doesn’t need scanning. In this account, the buckets that actually received user uploads had no protection at all until the incident.
  • Manage logging, trail selectors and GuardDuty plans as infrastructure-as-code with peer review. Add an AWS Config rule or a pre-deployment check that flags any bucket whose logging target is itself or the CloudTrail bucket.
  • Route budgets and anomaly detection to people who will act. That means a named owner and a team channel, a one-business-day response expectation, a low anomaly threshold, and a per-service budget for GuardDuty.
  • Check the result the next day. After any logging or security-service change, look at that service’s line in Cost Explorer and at the target bucket’s object count.

A note on evidence

We did the investigation read-only. We captured every finding to an evidence pack with a SHA-256 manifest, used CloudTrail’s signed digest files to confirm the sampled log files were unaltered originals, and wrote a script that recomputes each figure in the report from the pack. That made the findings checkable by the client, and gave a solid basis for asking AWS for a billing adjustment. One honest footnote: our own reads of the log bucket triggered a GuardDuty anomalous-behaviour finding, which we documented and attributed to the audit. Measured, not assumed.

How we can help

If your AWS bill has moved and nobody can say why, we can trace it to the configuration change behind it, verify the finding, and help you put guardrails in place so it doesn’t happen again. See our services for how we approach platform reliability and FinOps work.

Back to blog

Not sure where to start? Start with an assessment.

A senior review of your app, cloud estate, or AI platform, scoped and quoted before work starts, that ends in a prioritized plan, so you decide what to fix and when.

Talk to an engineer