Automated long-running task handoffs: AI agent ledger and shift-change handoff for multi-day DevOps tasks
the shed // AGENTIC AI BRIEFING

Multi-day infrastructure work dies at the shift change. Here is how teams give an AI agent durable memory so a 40-database migration survives context resets, restarts, and the 2 AM handoff.

Every platform team has a job that cannot finish in one sitting. A Postgres major-version upgrade across 40 databases. A Kubernetes 1.30 to 1.33 rollout across 25 clusters. A TLS certificate rotation touching 300 services. These tasks run for days, and the thing that breaks them is rarely the technical step. It is the handoff: the engineer who started it goes off shift, the AI agent that was driving it hits its context limit, and whoever picks it up starts by asking “wait, which ones are done?”

Automated long-running task handoffs fix that specific failure. The agent does not remember by re-reading a giant transcript. It remembers by writing state to a ledger that any future agent session, or any human, can resume from in under a minute.

See the pattern in action, tap through the tabs below:




pg16-migration/handoff

Day 2, 02:10. The agent session that upgraded 17 of 40 databases has been restarted after a context reset. The new session has no idea db-18 was mid-replication when the old one died. The on-call engineer who was watching it is asleep. The Slack thread is 400 messages long.

What usually happens: someone re-runs step one on a database that was already cut over, or skips one that never finished. Both are worse than the original outage they were trying to avoid.

What should happen: the new session reads a 60-line ledger, sees db-18 is in state replicating with a checkpoint LSN recorded, verifies that LSN, and continues. No transcript archaeology.

# ledger.yaml  (the agent's only source of truth)
task: pg14-to-pg16-fleet-upgrade
owner_shift: platform-oncall
started: 2026-09-09T14:00Z
invariants:
  - never run cutover on a db whose state != verified
  - never touch more than 2 dbs concurrently
  - pause if error_rate_5m > 0.5%
targets:
  db-17: {state: done, verified_at: "2026-09-10T01:52Z"}
  db-18: {state: replicating, checkpoint: "0/9A3F2C10",
          last_action: "pg_upgrade --link ok, logical replication started",
          next_action: "verify lag < 5s then cutover"}
  db-19: {state: queued}
  db-20: {state: blocked, reason: "pending app-team ack (ticket PLAT-4412)"}
handoff_note: >
  Started db-18 at 01:58Z. Replication lag was 2.1s at 02:07Z.
  Do NOT retry pg_upgrade on db-18, the --link step is irreversible.
  Nothing else is in flight.

$ agent resume --ledger ledger.yaml
[02:11:04] loaded ledger: 17 done, 1 replicating, 20 queued, 1 blocked, 1 failed
[02:11:04] handoff_note read. Rule: no pg_upgrade retry on db-18.
[02:11:05] db-18: verifying checkpoint 0/9A3F2C10 ... match
[02:11:06] db-18: replication lag 1.4s (under 5s) -> eligible for cutover
[02:11:06] invariant check: concurrency 1/2 ok, error_rate_5m 0.02% ok
[02:11:07] db-18: cutover requires approval (irreversible). Paging platform-oncall.
[02:11:07] ledger updated: db-18.state = awaiting_approval
[02:11:08] db-19: starting pg_upgrade --check (dry run, reversible)
[02:14:31] db-19: check passed. ledger updated: db-19.state = checked
[02:14:31] session budget 78% used. Writing handoff_note and exiting cleanly.
[02:14:32] handoff_note: "db-18 waiting on human approval for cutover.
           db-19 checked, ready for --link. Nothing in flight."

Start here (one afternoon):

1. Pick one multi-day task you already dread. Write its ledger schema by hand first: states, invariants, a handoff_note field.

2. Make every step idempotent or explicitly marked irreversible. Irreversible steps require a human approval gate.

3. Have the agent write to the ledger after every step and read it before every session. The ledger is the memory; the transcript is disposable.

4. Add a session budget. When it hits 75-80%, the agent writes its handoff note and exits on purpose instead of dying mid-step.

5. Run it in dry-run mode for one full cycle before you let it touch production.

The bottleneck: work that outlives the worker

An AI agent driving a long task has three ways to lose its mind. Its context window fills and older steps get truncated. The process restarts because of a deploy, a rate limit, or a laptop lid. Or the human supervising it changes, and the new person has a different mental model of where things are.

Human runbooks solved this decades ago with a checklist and a shift log. The mistake teams make with agents is assuming the model's conversation history is that shift log. It is not. It is a lossy, unstructured, expensive-to-reread pile of tokens, and once it gets summarized or truncated, the one line that said "db-18's link step is irreversible" is exactly the kind of detail that vanishes.

The workflow: ledger-first execution

The pattern has four parts, and none of them require a specific framework.

1. Externalized state, not conversational memory

Every unit of work gets an entry in a ledger the agent owns: a YAML or JSON file in the task's repo, or a row in a small Postgres table. Each entry has a state machine (queued, checked, in_progress, replicating, awaiting_approval, done, failed, blocked), a last_action, a next_action, and any checkpoint value needed to verify that the world still matches the ledger (a replication LSN, a Helm release revision, a certificate serial).

The rule is simple: the agent reads the ledger at the start of every session and writes it after every step. If the ledger and reality disagree, the agent stops and asks.

2. Idempotent steps and explicit irreversibility

Idempotency is what makes resumption safe. pg_upgrade --check can run ten times. pg_upgrade --link cannot. Tag each step in the ledger schema as reversible or irreversible. Irreversible steps get an approval gate, and the ledger records who approved and when, which doubles as your audit trail.

3. Deliberate session endings

Most long-task failures happen because the session died mid-step. So make endings intentional. Give the agent a budget (tokens, wall-clock, or step count). When it crosses roughly 75 percent, it finishes the current reversible step, writes a short handoff note in plain English, and exits. The next session starts by reading that note, then the ledger. This is the same discipline a good on-call engineer uses when they hand off at shift change: state, what's in flight, what not to touch.

4. Invariants the agent cannot argue with

Put your guardrails in the ledger, not the prompt. "Never more than two targets in flight," "pause if error rate exceeds 0.5 percent," "never cut over an unverified target." The agent checks invariants before each action, and a checker outside the model (a tiny script in CI or a pre-action hook) enforces them too. Prompts drift; a hook does not.

How it's built

You can run this with any agent runtime that supports tools and hooks. A few concrete stacks that work:

  • Claude Code or the Claude Agent SDK with a pre-tool hook that reads the ledger and blocks any action against a target whose state does not permit it. The agent's scratchpad file is the ledger. When the session ends, a stop hook forces the handoff note to be written.
  • LangGraph with a Postgres checkpointer. Each ledger target is a node in the graph; the checkpointer persists graph state between runs, so a restarted process resumes at the exact node it left. Human approval is an interrupt on the irreversible edge.
  • Temporal for the durable-execution layer. If you already run Temporal, model the whole migration as a workflow, each target as an activity, and let the agent decide the next activity. Temporal handles retries, timeouts, and restarts; the agent handles judgment.

Whatever you pick, keep the model out of the durability path. The framework remembers; the model reasons. Teams that get this backwards end up with an agent that is very confident and completely wrong about what it did yesterday.

What this looks like on a real week

A platform team running the Postgres fleet upgrade above went from three engineers babysitting a two-week rollout to one engineer approving cutovers from Slack. Not because the agent was smarter, but because the resume cost dropped to nearly zero. Any session, any person, any time of day could pick up the ledger and know exactly where things stood. The approval gate on irreversible steps meant nobody lost sleep over the agent being autonomous, because it was only autonomous on the reversible parts.

If you want a deeper grounding in the CI/CD, Kubernetes, and observability fundamentals this pattern sits on, the DevOps Boot Camp covers them end to end. For a broader look at how agents fit into SRE and security workflows, see our full course list.

FAQ

How is this different from agent memory features in frameworks?

Built-in memory (summaries, vector recall) is optimized for conversation, not operations. It is fuzzy by design. A ledger is exact, human-readable, and diffable in git, which is what you want when the question is "did db-18 finish or not."

What if the ledger and the real system disagree?

That is the most valuable moment in the whole workflow. The agent stops, records the discrepancy, and pages a human. A silent drift between "what we think we did" and "what happened" is how migrations turn into incidents.

Can this work with a single-agent setup, or does it need multiple agents?

It works with one. The handoff is between sessions and between humans, not necessarily between agents. Add a second reviewer agent later if you want an independent check on the handoff note before a human sees it.