tha-shed // field notes

Autopsy Mode: Watching an AI Agent Debug Production

Plus a fact-check nobody asked for: does Ruby’s JIT actually close the gap on Python in 2026? Step through a simulated incident, then run the numbers yourself.

An incident just paged. Watch the agent work it.

AWS’s DevOps Agent went GA this year on Bedrock AgentCore, and Google’s Agent Development Kit now wires straight into CI/CD, observability, and ticketing. Both point at the same shift: agents that read telemetry, form a hypothesis, and propose a fix — before a human opens a dashboard. Here’s a stylized walk-through of that loop. Click through it.

agent-session — checkout-api — sev2 idle
0 / 8
✓ Root cause identified

Connection pool exhaustion from an un-batched N+1 query introduced in the last deploy. The agent’s suggested fix: roll back the migration, batch the query, and — its actual recommendation — add a pool-saturation alert at 80% instead of 95% next time. It didn’t fix it unsupervised. It cut the median investigation time from ~40 minutes to under 3.

Pattern based on GA capabilities described for AWS DevOps Agent and Google’s Agent Development Kit — dialogue is illustrative, not a live transcript.

YJIT closed the gap. By how much depends what you’re running.

Pick a workload, then flip YJIT on and off. The numbers below are drawn from Shopify’s published benchmark work and Ruby 3.4 production reports — illustrative of the shape of the gap, not a live benchmark run.

Rails request (I/O-bound)
Tight CPU loop
JSON parse + serialize
Ruby JIT: YJIT enabled
# enable YJIT — Ruby 3.3+ RUBY_YJIT_ENABLE=1 ruby app.rb # or in a Rails initializer RubyVM::YJIT.enable

Figures reflect Shopify’s YJIT production benchmarks and 2026 CPython/CRuby comparison writeups. Your mileage depends entirely on your workload — this is a decision aid, not a guarantee.

01

If you’re on Rails and not running YJIT yet, it’s a config flag, not a rewrite — 15–25% median response-time drop is the realistic range in production.

02

Python still wins tight CPU-bound loops by roughly 1.2–1.5x. For I/O-bound web apps, that gap is noise next to your database and network latency.

03

AI incident agents aren’t replacing on-call — they’re replacing the first 30 minutes of grep-and-guess before a human forms a hypothesis.

04

The pattern to watch: agents wired into CI/CD, observability, and ticketing as one loop, not three separate tools you stitch together yourself.