Automated incident escalation for DevOps on-call teams, stop paging the wrong person
the shed // AGENTIC AI BRIEFING

Alert fatigue burns out your best engineers and the wrong page costs you MTTA every single time. Here is how to build an agent that triages, gathers context, and pages the right human instead of the whole rotation.

See the pattern in action, tap through the tabs below:




escalation-agent.yml






A disk-usage warning on a staging box pages the whole rotation. Nobody acked it in five minutes, so PagerDuty escalated. Three engineers are now awake reading the same dashboard. It happens again next week, and the week after, until the on-call engineer starts muting the channel by default, which is the exact moment a real page gets missed.

The fix isn’t a quieter alert threshold. It’s a triage step that gathers context, checks the pattern against history, and decides who actually needs to wake up, if anyone does.

tier_1_autonomous:
  match:
    pattern: known_and_reversible
    blast_radius: bounded
  actions:
    - rollback_last_known_good
    - scale_within_bounds
  verify: deterministic_healthcheck
  on_verify_fail: escalate_tier_2

tier_2_approval_gated:
  match:
    diagnosis_confidence: ">= 0.7"
    blast_radius: customer_facing
  actions:
    - propose_fix
    - page: service_owner
  requires: human_approval

tier_3_human_led:
  match:
    pattern: novel_or_security
  actions:
    - gather_evidence
    - page: primary_and_secondary
  agent_role: investigate_only

[02:47:13] alert.received service=checkout-api severity=warn

02:47:13  alert: disk_usage 91% on checkout-api-worker-7
02:47:14  agent: correlating against last 90d of incidents...
02:47:15  agent: match found -> 4 prior occurrences, all self-resolved
02:47:15  agent: last resolution = log rotation cron delayed
02:47:16  agent: checking service ownership -> team: payments
02:47:16  agent: tier assigned = 1 (known pattern, reversible)
02:47:17  agent: action -> trigger log rotation, verify disk < 70%
02:47:34  agent: verify passed, disk_usage now 61%
02:47:34  agent: incident closed, no page sent
02:47:35  agent: summary posted to #payments-ops

Pick the noisiest rotation first. The service where on-call already mutes half the channel is your best pilot.

Start with Tier 2 only. Agent investigates and proposes, a human approves. No autonomous actions in week one.

Add exactly one Tier 1 action. A rollback or bounded restart. Fully reversible, nothing else.

Track wrong-page rate, not MTTA. A faster wrong page is still a wrong page.

Get a month of clean runs before widening scope to a second service or a second action.

The 3 AM page that shouldn't have gone to you

Every on-call engineer has the same story. A disk-usage alert fires on a staging box, PagerDuty escalates because nobody acked it in five minutes, and now three people are awake reading the same Grafana dashboard trying to figure out if this is real. It wasn't real. It almost never is.

The cost of that page isn't the five minutes it took to check. It's the twenty minutes it takes to fall back asleep, the context switch that follows the next engineer into tomorrow's sprint, and the slow burn of a rotation where everyone quietly dreads their week. Alert fatigue is not a training problem. It is a routing problem, and routing is exactly what agents are good at.

Automated incident escalation does not mean handing an agent the pager and hoping for the best. It means putting a triage step in front of the page: something that reads the alert, pulls the context a human would spend the first ten minutes gathering anyway, checks it against what's actually happened before, and only then decides who gets woken up, if anyone does.

What "automated" should actually mean here

The teams getting this right are not building one agent that owns the whole incident. They're building a triage layer with an explicit ceiling on what it's allowed to decide alone, then widening that ceiling as it earns trust. A framework that's held up well in practice splits the work into three tiers:

  • Tier 1, autonomous response. Reserved for incidents with a known pattern and a reversible fix: a bounded autoscale, a container rollback to the last known-good image. The agent acts, verifies deterministically, and escalates automatically the moment verification fails.
  • Tier 2, approval-gated action. The agent investigates, forms a diagnosis, and proposes the fix, but a human has to approve before anything executes. This is where most customer-facing incidents should live: fast investigation, a human still on the hook for the call.
  • Tier 3, human-led investigation. Novel failures, cascading outages, anything security-adjacent. The agent's job shrinks to evidence gathering and hypothesis testing. It does not get to act.

The tier a given alert lands in is decided by policy, not by how confident the model sounds. That distinction matters more than anything else in this post: the autonomy decision has to live outside the LLM, in deterministic rules keyed on blast radius, reversibility, and whether the pattern has been seen before. A model that's 95% sure is still wrong 1 time in 20, and you do not want that 1 time to be a database failover it triggered on its own.

How it's built

The plumbing is less exotic than it sounds, and most of it you already have:

  • Observability as the input. The agent needs read access to whatever you already page from: Datadog, Prometheus, your APM. It's correlating signals a human would open five tabs to check, not inventing new telemetry.
  • An MCP-connected triage agent. PagerDuty's SRE Agent is a live example of this pattern: it correlates alerts against observability data and incident history to build what they call a context flywheel, and it plugs into Claude, Cursor, and LangChain over Model Context Protocol rather than a bespoke integration per tool.
  • On-call schedule and ownership data. The single biggest MTTA win isn't remediation, it's routing. An agent that can read your on-call schedule and service ownership map can page the engineer who actually owns the failing service instead of whoever's on primary rotation this week.
  • A policy engine sitting in front of execution. This is the piece teams skip and then regret. Tier assignment, approval gates, and rollback triggers belong in code you can read and test, not in a prompt.
  • A written escalation policy the agent enforces, not invents. See the config tab above. It's boring on purpose. Boring is what you want at 3 AM.

Increasingly this context lives closer to the code too. Pre-commit risk scoring that checks a diff against historical incident data, surfaced right inside the IDE, is showing up as a companion pattern: catch the deploy that looks like the last three outages before it ships, not after.

Start here

Don't automate escalation for everything at once. Pick the service with the noisiest rotation, the one where the on-call engineer already mutes half the channel. Start Tier 1 with exactly one remediation: a rollback or a bounded restart, something fully reversible. Get a month of clean escalations before you add a second automated action. The goal in week one isn't fewer pages, it's zero wrong pages.

FAQ

Does an incident escalation agent replace on-call engineers?
No, and treating it that way is how teams get burned. It replaces the ten minutes of manual triage before a human gets involved, and it can execute a narrow set of pre-approved, reversible fixes. Novel or high-blast-radius incidents still route straight to a person.

What's the first metric to watch after rolling this out?
Wrong-page rate, not MTTA. A system that pages faster but still wakes up the wrong person, or wakes anyone for a non-incident, has made the rotation worse even if the dashboard looks better.

Can this run without giving the agent write access to production?
Yes, and that's the recommended starting point. Run Tier 2 only for the first few weeks: the agent investigates and proposes, a human clicks approve. Add Tier 1 autonomy only after the proposals have been consistently correct.

Want this built out for your team's actual stack instead of a generic template? The Ruby Blocks for DevOps Engineers course walks through wiring agent automation into a real on-call workflow end to end.