Case study
Problem
Pricing alerts land in one Slack channel, and each needs a manual dig across Airflow, Datadog, CloudWatch, SQS and GitHub. Most turn out to be the same few things — a transient failure, a pipeline waiting on another, a deploy that didn't finish, a task out of memory — and the on-call re-derives the diagnosis every time.
What it does
An agent in Claude Code or Codex, run hourly (or by hand) on the on-call's laptop, sweeps every unresolved pricing alert. For each, one call gathers the core evidence, with log searches as follow-ups; the agent matches it against a decision table written from real alert history and acts: it restarts the task, reruns failed CI jobs, opens a draft PR, proposes dead-letter queue changes, or escalates with the evidence. Restarts and reruns never wait for anyone; a person is needed only to merge a PR or approve a queue change. The guards — scope, caps, dry runs, approvals, where it may post — are enforced by a plain-Python core behind a thin MCP server, so the agent can't work around a refusal.
System
How an alert gets handled — and what runs underneath. Pick a problem below, or click Investigate to open it.
One sweep · ⊕ opens a step
An alert fires and the hourly sweep picks it up. Here: a pricing task failed in Airflow.
Architecture
- 01An alert fires: a failed pricing task, a pricing alert, or a failed CI run
- 02Hourly sweep collects every unresolved alert
- 03Investigate: the core evidence, then follow-up log searches if needed
- 04Decide against a decision table built from alert history
- 05Act through guarded tools — restart, rerun, draft PR, queue proposal, or escalate
- 06Reply in the alert's own thread; the next sweep picks up a merge or an approval
Stack
- AgentClaude Code / Codex running an on-call skill
- ToolsMCP server over a plain-Python core (14 tools)
- GuardsRulebook + checks in code: scope, caps, dry runs, approvals
- ReadsAirflow, Datadog, CloudWatch, SQS + Postgres, GitHub, MLflow
- AlertsRootly + Slack threads
- RuntimeOn-call laptop over Twingate, credentials from the on-call's AWS login
Outcomes
- Merged and live on the pricing on-call, with writes on
- Restarts and CI reruns happen without waiting for anyone; only PRs and queue changes need a person
- The guards live in the core, not the prompt, and are covered by 180 tests
- Next: Slack as the progress log (a state line, silent until something changes) and monitoring for the pricing Lambdas
Code is private — happy to walk through it.