AppWizzy launched an MCP server that connects AI coding agents directly to its project infrastructure. Agents can create, deploy, and monitor cloud projects through a secure OAuth connection.
AppWizzy launched an MCP server that connects AI coding agents directly to its project infrastructure. Agents can create, deploy, and monitor cloud projects through a secure OAuth connection.
Supabase released supabase/evals under Apache-2.0 to benchmark Claude Code, Codex, and OpenCode on real engineering tasks inside containerised sandboxes.
A controlled study on 116 Python tasks found that pairing the weaker Codex as reviewer over Claude cut accuracy and more than doubled cost. Hierarchy matters.
A Microsoft Azure and AI MVP lays out how tiered inference routing splits coding requests across on-device, on-prem, and cloud models to cut cloud token usage without dropping quality.
Context isolation and task coordination are the two structural limits hitting autonomous AI coding agents in multi-project environments. Here is what the failure looks like.
AWS released Kiro Crew, an open-source orchestration platform that runs multi-agent engineering workflows across repos and sessions without an AWS account required.
OpenAI is testing an ad format that skips the destination website entirely. Clicking an ad opens a business-specific ChatGPT agent that answers questions, recommends products, and captures leads.
Researchers built AgenTag, a framework that identifies which AI coding agent authored a pull request with 0.96 weighted F1, using commit messages and PR descriptions rather than code diffs.
OpenAI's July 28 field report on eight real genomics deployments shows where AI coding agents break down — and what Astra must prove by September 2026.
Claude Code, OpenAI Codex, and Google Vertex AI all offer ZDR, but the fine print varies significantly. Here is what healthcare engineering teams need to know.
65% of developers use AI coding tools weekly, and the resulting quality gap is spawning a specialist role: the engineer who cleans up what AI gets subtly wrong.
A tech lead with 8 years in games breaks down which development tasks justify AI assistance and which shift hidden costs downstream. Key data from GDC 2026 and METR included.