Most incidents get resolved in an hour and written up three weeks later, badly. Automated postmortem drafting closes that gap by turning the incident’s own telemetry into a reviewable draft before anyone goes back to bed.
See the pattern in action, tap through the tabs below:
[bottleneck] incident closed 02:14 UTC // doc still empty 19 days later
The incident is the easy part. Paging works. The bridge fills. Someone spots the bad config, traffic recovers, everyone goes back to bed.
What does not happen is the writeup. The timeline lives in a Slack thread nobody wants to reread. The action items never get filed. In November the same failure happens again, and in February an auditor asks for evidence of a review process you technically have and practically do not.
The window where this is cheap to fix is about 30 minutes wide, right after resolution, while every artifact is still addressable.
[artifact] crew.yaml // four roles, one supervisor, zero publish rights
collector:
role: Incident artifact collector
goal: Fetch every artifact inside the blast window
window: [start - 30m, end + 15m]
tools: [pagerduty_incident, metrics_query, git_log, deploy_events, chat_export]
constraints:
- redact: [email, customer_name, api_key]
- max_messages: 400
timeline:
role: Timeline normalizer
goal: One UTC timeline, deduped, gaps flagged
rules:
- every_entry_needs: artifact_id
- mark: [detection, first_mitigation, resolution]
- emit_gap_if_silence_exceeds: 12m
analysis:
role: Contributing factor analyst
goal: Propose factors, never declare root cause
rules:
- each_factor_requires: 1+ artifact_id
- if_uncited: write_open_question_instead
- forbidden: [naming_individuals, blame_language]
drafter:
role: Template renderer
goal: Render org postmortem template, status=draft
output: markdown
publish: false # human owns this button
[session] INC-4471 // sev2 // checkout latency // 71 min
02:14:06 hook incident.resolved received (INC-4471, sev2)
02:14:09 collector window set 01:33:00 -> 02:29:00 UTC
02:14:31 collector 6 alert events, 2 deploys, 118 chat messages
02:14:33 collector redacted 9 spans (customer_name), 1 span (email)
02:15:02 timeline 118 messages -> 23 timeline entries after dedupe
02:15:04 timeline detection 01:37:12 first_mitigation 02:02:48
02:15:05 timeline GAP 01:41:10 -> 01:56:40 (15m silence) flagged
02:15:44 analysis factor 1: connection pool cap unchanged since 2024
cite: git:8f2a1c9, metrics:pool_saturation
02:15:46 analysis factor 2: canary ran 4 min, p99 regression at 6 min
cite: deploy:rel-2291, dash:checkout-p99
02:15:49 analysis OPEN QUESTION: what drove the 15m response gap?
reason: no artifact covers this interval
02:16:20 drafter rendered postmortem-INC-4471.md (1,140 words)
02:16:21 drafter 4 draft action items, owners proposed, NOT assigned
02:16:22 drafter status=draft, publish=false, queued for review
review queue: 1 item human edit time logged next morning: 22 min
[rollout] three weeks, read-only first
Week 1. Collector only. On resolve, it posts a normalized timeline back into the incident channel. No analysis, no draft. You are checking whether the artifacts are even reachable.
Week 2. Add the analysis agent. Humans still write the postmortem, then compare their factors against the proposed ones. This is where you tune the citation rules.
Week 3. Turn on the drafter for sev1 and sev2 only. Draft lands in the review queue. A human owns the publish button, permanently.
Measure one number: minutes of human editing per postmortem. If it is not under 30 by week four, your evidence contract is too loose.
The bottleneck: the writeup nobody has time for
Your team is good at incidents. Paging works, the bridge fills up, someone finds the bad config, traffic comes back. Then the postmortem sits half-written for two weeks, and by the time anyone opens it again the timeline exists only in a Slack thread nobody wants to reread.
The cost is not the missing document. The cost is the action items that never get filed, the identical outage two months later, and the auditor who asks for evidence of a review process you technically have and practically do not.
Automated postmortem drafting attacks this at the one point where it is cheap: the moment the incident closes, while every artifact is still addressable. Alerts, deploy events, dashboard queries, chat transcript, the commits inside the blast radius. A human still owns the analysis. The machine owns the assembly.
What “done” looks like
The outcome you are actually buying, stated plainly:
- A draft postmortem in the review queue within 30 minutes of resolution
- A timeline built from timestamps, not from memory
- Contributing factors proposed with citations, each linked to the artifact it came from
- Draft action items with a proposed owner and priority, unassigned until a human confirms
- An editor who spends 20 minutes correcting instead of 3 hours reconstructing
Note what is not on that list. The agent does not declare root cause, does not assign blame, and does not publish. It builds the scaffolding that makes human review tractable. Google’s SRE Workbook chapter on postmortem culture is still the best statement of why that boundary matters.
How the workflow is built
Implementation detail starts here, and it is deliberately boring.
The trigger
A webhook from your incident manager firing on state change to resolved. PagerDuty, Opsgenie and incident.io all emit the same useful payload: incident ID, severity, start and end timestamps, responder list. That payload defines the blast window, and the blast window defines everything downstream.
The four agents
Run them as a sequential crew, not a free-for-all. Each one has a narrow job and a narrow toolset:
- Collector. Pulls artifacts inside a fixed window, typically incident start minus 30 minutes to end plus 15. Metrics queries, git log for the affected services, CI deploy events, the incident channel export. Redaction happens here, not later.
- Timeline agent. Normalizes everything to one UTC sequence, dedupes the alert storm, marks detection and first mitigation, and flags silence gaps longer than about 12 minutes. Those gaps are usually the most interesting part of the review.
- Analysis agent. Proposes contributing factors. Never root cause, never a person’s name.
- Drafting agent. Renders your org’s existing template and stops. Status stays draft.
For orchestration, CrewAI and LangGraph both handle this cleanly. Pick LangGraph if you want an explicit state machine you can replay, CrewAI if you want role definitions your SREs will actually read. n8n works too if your team would rather see the graph than the code. The choice matters far less than the next section.
The evidence contract
This is the rule that makes the whole thing usable: every factual sentence in the draft carries a citation ID that resolves to a fetched artifact. If the analysis agent cannot cite, it writes an open question instead of a claim.
Skip this and you get a fluent, confident, subtly wrong timeline, which is worse than no timeline at all. One hallucinated timestamp in a sev1 review and your team will never trust the tool again. The citation requirement is not a nice-to-have, it is the entire trust model.
Where this goes wrong
- Chat export scope creep. Incident channels contain customer names, ticket contents, occasionally credentials someone pasted in a panic. Redact at collection time, with an allowlist, not a regex you hope covers it.
- Severity blindness. Auto-drafting every sev4 buries the reviews that matter. Gate on sev1 and sev2 to start.
- Template drift. If the drafter invents its own headings, your postmortems stop being comparable. Pin the template and diff against it in CI.
- The publish button. There is exactly one correct answer here and it is a human.
- Cost. A 400-message incident channel plus dashboards is a lot of context. Cap message count and summarize per-source before the analysis step.
The honest limitation
This workflow cannot tell you why your team waited 15 minutes before escalating. It can only tell you that the gap exists and that nothing in the artifact set explains it. That is still worth a great deal, because the questions worth asking in a review are almost always the ones nobody thought to write down.
Treat the draft as a well-organized pile of evidence with some hypotheses attached. The judgment stays yours.
FAQ
Will an AI-drafted postmortem hold up in a SOC 2 or ISO audit?
The artifact an auditor wants is evidence that a documented review happened, with participants, findings and tracked remediation. A drafted-then-human-reviewed postmortem satisfies that as long as the review and approval are recorded. Log who edited and who approved, and keep the draft version alongside the final.
Does automating the writeup damage blameless culture?
It helps, if you constrain the analysis agent correctly. Agents have no social incentive to protect a teammate or shift responsibility, so a factor list built from artifacts tends to be less political than one built from memory. Forbid naming individuals in the analysis prompt and the drafts come out structurally blameless.
Which model should run the analysis step?
Use your strongest reasoning model for the analysis agent and a cheap fast model for collection and normalization. The collector is doing API calls and string work, which does not need frontier reasoning. Splitting the two typically cuts cost by more than half with no quality loss.
Start with the collector
Pick your last sev2, run a collector over it by hand, and see how much of the timeline you can rebuild from artifacts alone. That one exercise tells you whether this workflow is worth building for your team.
If you want the surrounding skills, our DevOps and cybersecurity track covers incident response and the agent tooling around it, and the full catalog lives on our courses page if you would rather have this built with your team than by yourself.


