Field Notes — Sept 24, 2026

Agent Bay & the Decision Surface

Two things landed this week that belong in the same post: GitHub shipped a desktop app for running several coding agents at once, and Rails World 2026 closed with an official Rails benchmark answering a question everyone’s been arguing about — does “convention over configuration” actually make a codebase easier for an AI agent to work in? Click through both below.

GitHub’s new Copilot app (technical preview, agent-native desktop) runs multiple agents on the same repo in parallel by giving each one its own isolated git worktree, then hands finished work to Agent Merge, which drives the PR through CI, review, and merge on its own. Spin up three agents below and watch the worktrees and merge queue move.

idle — 3 worktrees ready
💡
Why this matters: before isolated worktrees, two agents editing the same repo would stomp on each other’s checkout. GitHub reports commits on github.com nearly doubled year-over-year to ~1.4B/month and Actions usage passed 2B minutes/week — scale that pushed them toward per-agent worktrees and an automated merge step rather than one shared working copy. Source: GitHub Blog, Copilot app announcement.
you-can-do-this-today.sh — plain git, no app required
# one worktree per agent/feature, off the same repo & .git
git worktree add ../wt-agent-1 -b feature/avatar-upload
git worktree add ../wt-agent-2 -b feature/rate-limit-api
git worktree add ../wt-agent-3 -b fix/flaky-webhook-test

# each agent (or you) works its own lane, zero collision:
cd ../wt-agent-1 && bundle exec rspec

# clean up once its PR merges
git worktree remove ../wt-agent-1

DHH opened Rails World 2026 in Austin arguing 37signals has gone “pencils down” on hand-written code because Rails’ conventions are unusually agent-friendly. Rails’ own team put a number on it: 25 models, 60 runs each, September 2026, on real Rails feature tickets. This is that table — sort it, and decide for yourself.

“Convention over configuration set the path for 20+ years of great training data for AI to use today. Not only does this mean agents do great with Rails, but also that squishy humans can quickly and confidently review the output.”
— DHH source ↗
“Convention is compression: it lets a finite context window hold the shape of your entire application.” Bespoke stacks turn every choice — folders, naming, wiring — into a guess an agent has to get right on its own. — Andrew Kwak source ↗
sort: accuracy $/run time accuracy per $ 25 models · 60 runs each · capped 90min / 400 steps / $60
Model Accuracy Median time Mean tokens Mean cost Acc. per $
Takeaway The most accurate model (GPT-6 Astramax, 53.3%) costs ~130× more per run than the cheapest one tested. For everyday Rails feature tickets, look at the “accuracy per $” column above rather than the accuracy column alone — several mid-effort models land within a few points of the top score at a fraction of the cost. Re-run this eval on your own ticket backlog before standardizing on one model tier.