Automated postmortem drafting workflow for SRE and DevOps incident reviews
the shed // AGENTIC AI BRIEFING

Most incidents get resolved in an hour and written up three weeks later, badly. Automated postmortem drafting closes that gap by turning the incident’s own telemetry into a reviewable draft before anyone goes back to bed.

See the pattern in action, tap through the tabs below:




postmortem-crew

[bottleneck] incident closed 02:14 UTC // doc still empty 19 days later

The incident is the easy part. Paging works. The bridge fills. Someone spots the bad config, traffic recovers, everyone goes back to bed.

What does not happen is the writeup. The timeline lives in a Slack thread nobody wants to reread. The action items never get filed. In November the same failure happens again, and in February an auditor asks for evidence of a review process you technically have and practically do not.

The window where this is cheap to fix is about 30 minutes wide, right after resolution, while every artifact is still addressable.

[artifact] crew.yaml // four roles, one supervisor, zero publish rights

collector:
  role: Incident artifact collector
  goal: Fetch every artifact inside the blast window
  window: [start - 30m, end + 15m]
  tools: [pagerduty_incident, metrics_query, git_log, deploy_events, chat_export]
  constraints:
    - redact: [email, customer_name, api_key]
    - max_messages: 400

timeline:
  role: Timeline normalizer
  goal: One UTC timeline, deduped, gaps flagged
  rules:
    - every_entry_needs: artifact_id
    - mark: [detection, first_mitigation, resolution]
    - emit_gap_if_silence_exceeds: 12m

analysis:
  role: Contributing factor analyst
  goal: Propose factors, never declare root cause
  rules:
    - each_factor_requires: 1+ artifact_id
    - if_uncited: write_open_question_instead
    - forbidden: [naming_individuals, blame_language]

drafter:
  role: Template renderer
  goal: Render org postmortem template, status=draft
  output: markdown
  publish: false   # human owns this button

[session] INC-4471 // sev2 // checkout latency // 71 min

02:14:06  hook      incident.resolved received (INC-4471, sev2)
02:14:09  collector window set 01:33:00 -> 02:29:00 UTC
02:14:31  collector 6 alert events, 2 deploys, 118 chat messages
02:14:33  collector redacted 9 spans (customer_name), 1 span (email)
02:15:02  timeline  118 messages -> 23 timeline entries after dedupe
02:15:04  timeline  detection 01:37:12  first_mitigation 02:02:48
02:15:05  timeline  GAP 01:41:10 -> 01:56:40 (15m silence) flagged
02:15:44  analysis  factor 1: connection pool cap unchanged since 2024
                    cite: git:8f2a1c9, metrics:pool_saturation
02:15:46  analysis  factor 2: canary ran 4 min, p99 regression at 6 min
                    cite: deploy:rel-2291, dash:checkout-p99
02:15:49  analysis  OPEN QUESTION: what drove the 15m response gap?
                    reason: no artifact covers this interval
02:16:20  drafter   rendered postmortem-INC-4471.md (1,140 words)
02:16:21  drafter   4 draft action items, owners proposed, NOT assigned
02:16:22  drafter   status=draft, publish=false, queued for review

review queue: 1 item   human edit time logged next morning: 22 min

[rollout] three weeks, read-only first

Week 1. Collector only. On resolve, it posts a normalized timeline back into the incident channel. No analysis, no draft. You are checking whether the artifacts are even reachable.

Week 2. Add the analysis agent. Humans still write the postmortem, then compare their factors against the proposed ones. This is where you tune the citation rules.

Week 3. Turn on the drafter for sev1 and sev2 only. Draft lands in the review queue. A human owns the publish button, permanently.

Measure one number: minutes of human editing per postmortem. If it is not under 30 by week four, your evidence contract is too loose.

The bottleneck: the writeup nobody has time for

Your team is good at incidents. Paging works, the bridge fills up, someone finds the bad config, traffic comes back. Then the postmortem sits half-written for two weeks, and by the time anyone opens it again the timeline exists only in a Slack thread nobody wants to reread.

The cost is not the missing document. The cost is the action items that never get filed, the identical outage two months later, and the auditor who asks for evidence of a review process you technically have and practically do not.

Automated postmortem drafting attacks this at the one point where it is cheap: the moment the incident closes, while every artifact is still addressable. Alerts, deploy events, dashboard queries, chat transcript, the commits inside the blast radius. A human still owns the analysis. The machine owns the assembly.

What “done” looks like

The outcome you are actually buying, stated plainly:

  • A draft postmortem in the review queue within 30 minutes of resolution
  • A timeline built from timestamps, not from memory
  • Contributing factors proposed with citations, each linked to the artifact it came from
  • Draft action items with a proposed owner and priority, unassigned until a human confirms
  • An editor who spends 20 minutes correcting instead of 3 hours reconstructing

Note what is not on that list. The agent does not declare root cause, does not assign blame, and does not publish. It builds the scaffolding that makes human review tractable. Google’s SRE Workbook chapter on postmortem culture is still the best statement of why that boundary matters.

How the workflow is built

Implementation detail starts here, and it is deliberately boring.

The trigger

A webhook from your incident manager firing on state change to resolved. PagerDuty, Opsgenie and incident.io all emit the same useful payload: incident ID, severity, start and end timestamps, responder list. That payload defines the blast window, and the blast window defines everything downstream.

The four agents

Run them as a sequential crew, not a free-for-all. Each one has a narrow job and a narrow toolset:

  • Collector. Pulls artifacts inside a fixed window, typically incident start minus 30 minutes to end plus 15. Metrics queries, git log for the affected services, CI deploy events, the incident channel export. Redaction happens here, not later.
  • Timeline agent. Normalizes everything to one UTC sequence, dedupes the alert storm, marks detection and first mitigation, and flags silence gaps longer than about 12 minutes. Those gaps are usually the most interesting part of the review.
  • Analysis agent. Proposes contributing factors. Never root cause, never a person’s name.
  • Drafting agent. Renders your org’s existing template and stops. Status stays draft.

For orchestration, CrewAI and LangGraph both handle this cleanly. Pick LangGraph if you want an explicit state machine you can replay, CrewAI if you want role definitions your SREs will actually read. n8n works too if your team would rather see the graph than the code. The choice matters far less than the next section.

The evidence contract

This is the rule that makes the whole thing usable: every factual sentence in the draft carries a citation ID that resolves to a fetched artifact. If the analysis agent cannot cite, it writes an open question instead of a claim.

Skip this and you get a fluent, confident, subtly wrong timeline, which is worse than no timeline at all. One hallucinated timestamp in a sev1 review and your team will never trust the tool again. The citation requirement is not a nice-to-have, it is the entire trust model.

Where this goes wrong

  • Chat export scope creep. Incident channels contain customer names, ticket contents, occasionally credentials someone pasted in a panic. Redact at collection time, with an allowlist, not a regex you hope covers it.
  • Severity blindness. Auto-drafting every sev4 buries the reviews that matter. Gate on sev1 and sev2 to start.
  • Template drift. If the drafter invents its own headings, your postmortems stop being comparable. Pin the template and diff against it in CI.
  • The publish button. There is exactly one correct answer here and it is a human.
  • Cost. A 400-message incident channel plus dashboards is a lot of context. Cap message count and summarize per-source before the analysis step.

The honest limitation

This workflow cannot tell you why your team waited 15 minutes before escalating. It can only tell you that the gap exists and that nothing in the artifact set explains it. That is still worth a great deal, because the questions worth asking in a review are almost always the ones nobody thought to write down.

Treat the draft as a well-organized pile of evidence with some hypotheses attached. The judgment stays yours.

FAQ

Will an AI-drafted postmortem hold up in a SOC 2 or ISO audit?

The artifact an auditor wants is evidence that a documented review happened, with participants, findings and tracked remediation. A drafted-then-human-reviewed postmortem satisfies that as long as the review and approval are recorded. Log who edited and who approved, and keep the draft version alongside the final.

Does automating the writeup damage blameless culture?

It helps, if you constrain the analysis agent correctly. Agents have no social incentive to protect a teammate or shift responsibility, so a factor list built from artifacts tends to be less political than one built from memory. Forbid naming individuals in the analysis prompt and the drafts come out structurally blameless.

Which model should run the analysis step?

Use your strongest reasoning model for the analysis agent and a cheap fast model for collection and normalization. The collector is doing API calls and string work, which does not need frontier reasoning. Splitting the two typically cuts cost by more than half with no quality loss.

Start with the collector

Pick your last sev2, run a collector over it by hand, and see how much of the timeline you can rebuild from artifacts alone. That one exercise tells you whether this workflow is worth building for your team.

If you want the surrounding skills, our DevOps and cybersecurity track covers incident response and the agent tooling around it, and the full catalog lives on our courses page if you would rather have this built with your team than by yourself.