Plugin4Shell & the Language Ledger
This week: a zero-click RCE bypassed SHA-pinning in four major AI coding agents — and a 600-run Claude Code benchmark puts a real dollar figure on choosing Ruby, Python, or Rust for AI-assisted development. Step through the exploit, check your exposure, then run the numbers yourself.
How Plugin4Shell works
CVE disclosed · Sept 17, 2026Four agents — Claude Code, Codex, GitHub Copilot, and Gemini CLI — pin plugin updates to a reviewed commit SHA for safety. Researchers found none of them actually verified the checkout landed on that SHA. Step through the attack:
Check your exposure
patch statusPick the agent your team runs in CI or on developer machines:
Do this today
- 1Update immediately if you run Claude Code (≥2.1.179) or Codex (≥0.146.0) — the SHA-verification gap is closed in these builds.
- 2Disable background auto-update on GitHub Copilot until Microsoft ships a fix — the vulnerability is currently unpatched there.
- 3Drop Gemini CLI for plugin-based workflows — Google deprecated it rather than patching it, so no fix is coming.
- 4If you maintain internal agent tooling, verify the post-checkout SHA explicitly (
git rev-parse HEADcompared to the pinned value) — don’t trust a successfulgit checkout <sha>alone, since git resolves a same-named branch first.
What AI-written code actually costs, by language
n=600 runsRuby core committer Yusuke Endoh had Claude Code (Opus 4.6) build a ~200-line Git clone, twice, in 13 languages — 20 trials each, 600 runs total. Pick languages below to compare real time and cost per run.
Watch a trial run
Same task, run once per selected language — timing and pass/fail reflect the real distributions Endoh measured (Rust fails ~5% of the time; Ruby, Python, and JS never did, in 40/40 runs each).
Try it yourself
Endoh’s full harness and raw results are open source. This clones it and reruns the benchmark for a language of your choice:
At 100 runs/day, Ruby and Python are within a rounding error of each other (73.1s/$0.36 vs 74.6s/$0.38) and both beat Rust by roughly a third on cost and time — with zero reliability tradeoff (40/40 passed vs Rust’s 38/40). For greenfield services you’ll iterate on heavily with an AI pair programmer, that’s a real monthly line item, not just a style preference.