Two major AI coding agents just got fresh model upgrades within 24 hours of each other, and the latest benchmarks show a split decision: Claude Code wins on overall score, Codex wins on cost and speed by a wide margin.
What the benchmarks show
Artificial Analysis’ Coding Agent Index v1.5 combines DeepSWE v1.1, Terminal-Bench 4.0, and SWE-Atlas-QnA into a composite score. Running each agent at its current representative configuration:
| Metric | Claude Code + Sonnet 5.5 | Codex + GPT-6.1 Sol |
|---|---|---|
| Coding Agent Index v1.5 | 68 | 63 |
| DeepSWE v1.1 | 72% | 73% |
| Terminal-Bench 4.0 | 66% | 55% |
| SWE-Atlas-QnA | 67% | 61% |
| Cost per task | $14.19 | $1.04 |
| Average time per task | 1.5 hours | 15.5 min. |
| Tokens per task | 27.7 million | 3.2 million |
Claude Code leads three of four benchmark components. Codex leads DeepSWE by one point, finishes tasks in 15.5 minutes versus 1.5 hours, and costs $1.04 per task versus $14.19.
The effort-equalized comparison
At matched xhigh effort settings, the composite score gap disappears: both agents score 63. At that level, Claude Code drops to $3.33 per task and 27 minutes. Codex stays at $1.04 and 15.5 minutes.
The context for both updates: Anthropic released Sonnet 5.5 on Sept. 28, reporting more than 30% throughput improvement over Sonnet 5 and up to 30% lower task costs for most workloads. OpenAI followed on Sept. 29 with GPT-6.1 Sol, positioned for complex coding and professional work at lower cost than GPT-6 Astra.
Security is not a settled question for either
According to a VentureBeat investigation, a Codex vulnerability could expose a GitHub OAuth token through a crafted branch name. Claude Code had permission bypasses that Anthropic later patched. Neither platform has a clean record here.
Anthropic reports that Claude Code users approve 93% of permission prompts. In internal testing, its auto-mode classifier cut false positives on benign actions to 0.4%, but produced a 17% false-negative rate across 52 real “overeager” actions.
OpenAI documents sandboxing, approvals, and network controls for Codex. Anthropic offers managed settings for tool permissions, file access, and MCP servers on Team and Enterprise plans.
The operator takeaway
If you’re running high-volume automation against real repositories, the $13.15 cost difference per task adds up fast. If you need the highest benchmark score for complex terminal or QnA tasks, Claude Code holds that edge for now. Test both against your actual codebase before committing either to broader access, and verify credential scope and command-approval policies before you do.
