A Cost-Anomaly Job Can Run Green and Check Nothing: CUR Lag and a date == today Gate

Our cost-anomaly sweep runs every day at 08:00 UTC. It reads daily cost rows built from AWS billing data, and it only evaluated a service if that service had a row dated today.
Billing data lags real usage. At 08:00 UTC a row dated today almost never exists. So the sweep evaluated zero services on every run, with no error.
This is the full story behind a correction we added to one of our own posts, and the three things we would tell anyone building on Cost and Usage Report (CUR) data.
The honest link first
In August we published Catching a Cost Spike the Same Day It Happens. It walked through the z-score math behind our anomaly detector. The math was right. The promise in the title was not.
On September 21 we added a correction note to the top of that post. It says, in short: when the post went out, the scheduled sweep did not catch spikes the same day, because it only looked at rows dated today, and those rows almost never existed yet. It also says a second bug stopped the sweep from resolving Shield-tier users at all.
That note is two sentences long. This post is the rest of it.
Bug 1: a date == today gate on lagged data
Here is the shape of the old code, trimmed:
today = datetime.now(timezone.utc).date().isoformat()
for service, daily_costs in service_costs.items():
today_cost = next(
(d['cost'] for d in daily_costs if d['date'] == today),
None
)
if today_cost is None:
continue
# ... z-score against the trailing history ...
Read the continue line with lagged data in mind. At 08:00 UTC the newest cost row for an account is usually yesterday or the day before. No row matches today. Every service hits continue. The loop finishes, the function returns an empty list, and the Lambda exits cleanly.
Nothing about that run looks wrong from the outside. No exception, no timeout, no error log. An empty list of anomalies is also exactly what a quiet, healthy account produces. The job was indistinguishable from a job that checked everything and found nothing.
That is the trap. A scheduled job's success status tells you it finished. It says nothing about how much it covered.
Bug 2: a lookup on an attribute no row ever had
The second bug sat one step earlier. The sweep resolved each Shield-tier user through a cognito_user_sub attribute. No row in our users table has ever had that attribute: the table's key is id, which holds the same Cognito identifier.
So every user's ID resolved to None, and the sweep could not process a single Shield user.
Our tests didn't catch it because the test fixtures gave fake users a cognito_user_sub field. The fixtures agreed with the code, so the tests passed. The fix changed the lookup to id, and the tests now use users shaped like real rows (an id, no cognito_user_sub). Each new or updated test was confirmed to fail against the pre-fix handler before we trusted it.
Two bugs, stacked. Either one alone was enough to make the sweep evaluate nothing.
The fix: evaluate the newest complete day, and say so when it's too old
We fixed both on September 15 and 16, 2026. The date fix landed in both places that run detection, the scheduled sweep and the on-demand path, and it changed four things.
1. Evaluate the newest date that has data. Instead of today, the sweep picks the newest date with a cost row for the account, across all of its services, and evaluates that day. Detection now runs behind real usage by the billing lag, typically a day or two, not on the same day.
2. Skip accounts whose data is too old, loudly. If an account's newest cost row is more than 3 days old, the sweep skips it instead of evaluating a stale day. Two things record the skip:
- A log line at WARNING. Not INFO, because INFO is dropped in our API Lambda, which runs the same logic when someone triggers detection on demand. We learned that one separately, and it applies here too: a log you never see is not a signal.
- A CloudWatch metric,
CloudWise/Anomaly/StaleCostData, with anEnvironmentdimension only. The commit priced it at about $0.30 a month per environment, about $0.60 across staging and production.
3. Alert once per evaluated day. Evaluating a lagged day means two consecutive daily runs can look at the same date. A separate dedupe keyed on (account, service, evaluated_date) makes sure that day alerts exactly once.
4. Keep newer rows out of the baseline. The baseline now only uses rows before the evaluated day (< target_date, not != target_date). Otherwise a newer, partial CUR row could slip into the history and manufacture a spike.
The z-score thresholds and the $5 floor from the original post are unchanged.
What we would tell anyone building on CUR
Know your data's lag before you write "today". Billing data is not a live feed. Pick the evaluation date from the data you actually have, not from the clock. If your code has date == today anywhere near cost data, check what time the job runs and what the newest row is at that moment.
Count what you evaluated. The number that would have exposed this bug on day one was not "anomalies found". It was "services evaluated". A run that evaluates zero services is not a quiet day. It is a broken job, and it should look different from one.
Alarm on zero coverage, not only on errors. An error alarm would not have caught this, because there was no error. Coverage needs its own signal.
To be straight about where we are with that last one: today the sweep emits the StaleCostData metric, and it logs the skip at WARNING. There is no alarm on the metric yet, and the sweep does not yet publish a count of services evaluated per run. That is the gap still open, and it's the one we would close first in anyone else's system too.
Make test data look like production data. The Shield bug survived because the fixtures had a field no real row has. A fixture shaped from a real row, with the real keys and nothing extra, would have failed on the first run.
Why this matters for a cost tool
A cost tool that errors is annoying. A cost tool that runs, reports success, and checks nothing is worse, because it hands you confidence you didn't earn. The fix here was small. Finding it required asking a question the job's status could never answer: how many things did you actually look at?
If you want to see what a read-only pass over your own account finds, CloudWise runs 187 waste checks across 40+ AWS services, and it takes about five minutes.
Free, read-only AWS waste scan: cloudcostwise.io. Five minutes, no card.
Stop wasting money on AWS
CloudWise monitors 45 AWS services and finds waste automatically. Free forever.
Start Free Scan →