Plugin4Shell & the Language Ledger
DAILY BRIEF · SEP 23, 2026

Plugin4Shell & the Language Ledger

This week: a zero-click RCE bypassed SHA-pinning in four major AI coding agents — and a 600-run Claude Code benchmark puts a real dollar figure on choosing Ruby, Python, or Rust for AI-assisted development. Step through the exploit, check your exposure, then run the numbers yourself.

How Plugin4Shell works

CVE disclosed · Sept 17, 2026

Four agents — Claude Code, Codex, GitHub Copilot, and Gemini CLI — pin plugin updates to a reviewed commit SHA for safety. Researchers found none of them actually verified the checkout landed on that SHA. Step through the attack:

agent-updater — background auto-update — zsh
BLAST RADIUS
0%

Check your exposure

patch status

Pick the agent your team runs in CI or on developer machines:

Do this today

  • 1Update immediately if you run Claude Code (≥2.1.179) or Codex (≥0.146.0) — the SHA-verification gap is closed in these builds.
  • 2Disable background auto-update on GitHub Copilot until Microsoft ships a fix — the vulnerability is currently unpatched there.
  • 3Drop Gemini CLI for plugin-based workflows — Google deprecated it rather than patching it, so no fix is coming.
  • 4If you maintain internal agent tooling, verify the post-checkout SHA explicitly (git rev-parse HEAD compared to the pinned value) — don’t trust a successful git checkout <sha> alone, since git resolves a same-named branch first.

What AI-written code actually costs, by language

n=600 runs

Ruby core committer Yusuke Endoh had Claude Code (Opus 4.6) build a ~200-line Git clone, twice, in 13 languages — 20 trials each, 600 runs total. Pick languages below to compare real time and cost per run.

100 runs/day

Watch a trial run

Same task, run once per selected language — timing and pass/fail reflect the real distributions Endoh measured (Rust fails ~5% of the time; Ruby, Python, and JS never did, in 40/40 runs each).

Try it yourself

Endoh’s full harness and raw results are open source. This clones it and reruns the benchmark for a language of your choice:

# reproduce the benchmark locally (needs an Anthropic API key)
git clone https://github.com/mame/ai-coding-lang-bench
cd ai-coding-lang-bench
./run.sh –lang ruby –trials 20
TAKEAWAY

At 100 runs/day, Ruby and Python are within a rounding error of each other (73.1s/$0.36 vs 74.6s/$0.38) and both beat Rust by roughly a third on cost and time — with zero reliability tradeoff (40/40 passed vs Rust’s 38/40). For greenfield services you’ll iterate on heavily with an AI pair programmer, that’s a real monthly line item, not just a style preference.

tha shed · daily interactive tech brief built & verified by an AI agent — sources linked throughout