Skip to content
Sophie Yin
← All work

Pricing on-call agent

2026 · Shipped · Built at Alt · Claude Code + MCP, Python, AWS

Case study

Problem

Pricing alerts land in one Slack channel, and each needs a manual dig across Airflow, Datadog, CloudWatch, SQS and GitHub. Most turn out to be the same few things — a transient failure, a pipeline waiting on another, a deploy that didn't finish, a task out of memory — and the on-call re-derives the diagnosis every time.

What it does

An agent in Claude Code or Codex, run hourly (or by hand) on the on-call's laptop, sweeps every unresolved pricing alert. For each, one call gathers the core evidence, with log searches as follow-ups; the agent matches it against a decision table written from real alert history and acts: it restarts the task, reruns failed CI jobs, opens a draft PR, proposes dead-letter queue changes, or escalates with the evidence. Restarts and reruns never wait for anyone; a person is needed only to merge a PR or approve a queue change. The guards — scope, caps, dry runs, approvals, where it may post — are enforced by a plain-Python core behind a thin MCP server, so the agent can't work around a refusal.

System

How an alert gets handled — and what runs underneath. Pick a problem below, or click Investigate to open it.

LLM step

One sweep · ⊕ opens a step

Pricing on-call sweepAn alert starts a sweep. The agent scans unresolved alerts, investigates each across every system involved, and decides: restart the task, rerun failed CI jobs, open a draft PR, propose dead-letter queue changes, or escalate. It replies in the alert's own thread. A person is needed only to merge a PR or approve a queue change, and the next sweep picks that up.NEXT SWEEP PICKS UP THE MERGE OR ✓AlertAIRFLOW · ROOTLY · CIScanALERT-DRIVENInvestigateEVIDENCEDECIDERestart taskAUTOMATICRerun failed CIAUTOMATICDraft PRON-CALL MERGESDLQ proposalNEEDS A ✓EscalateWITH EVIDENCEReply in threadTHE ALERT'S OWNOn-callREVIEWS · ✓Pricing on-call sweepAn alert starts a sweep. The agent scans unresolved alerts, investigates each across every system involved, and decides: restart the task, rerun failed CI jobs, open a draft PR, propose dead-letter queue changes, or escalate. It replies in the alert's own thread. A person is needed only to merge a PR or approve a queue change, and the next sweep picks that up.AlertAIRFLOW · ROOTLY · CIScanALERT-DRIVENInvestigateEVIDENCEDECIDERestart taskAUTOMATICRerun failed CIAUTOMATICDraft PRON-CALL MERGESDLQ proposalNEEDS A ✓EscalateWITH EVIDENCEReply in threadTHE ALERT'S OWNOn-callREVIEWS · ✓
Restart · 01/05

An alert fires and the hourly sweep picks it up. Here: a pricing task failed in Airflow.

Architecture

  1. 01An alert fires: a failed pricing task, a pricing alert, or a failed CI run
  2. 02Hourly sweep collects every unresolved alert
  3. 03Investigate: the core evidence, then follow-up log searches if needed
  4. 04Decide against a decision table built from alert history
  5. 05Act through guarded tools — restart, rerun, draft PR, queue proposal, or escalate
  6. 06Reply in the alert's own thread; the next sweep picks up a merge or an approval

Stack

  • AgentClaude Code / Codex running an on-call skill
  • ToolsMCP server over a plain-Python core (14 tools)
  • GuardsRulebook + checks in code: scope, caps, dry runs, approvals
  • ReadsAirflow, Datadog, CloudWatch, SQS + Postgres, GitHub, MLflow
  • AlertsRootly + Slack threads
  • RuntimeOn-call laptop over Twingate, credentials from the on-call's AWS login

Outcomes

  • Merged and live on the pricing on-call, with writes on
  • Restarts and CI reruns happen without waiting for anyone; only PRs and queue changes need a person
  • The guards live in the core, not the prompt, and are covered by 180 tests
  • Next: Slack as the progress log (a state line, silent until something changes) and monitoring for the pricing Lambdas

Code is private — happy to walk through it.