Supabase open sources an eval framework for AI coding agents

3D rendered ai text on dark digital background

Supabase shipped a new open source benchmark called supabase/evals under the Apache-2.0 licence. The goal is straightforward: give developers a reproducible way to measure how well AI coding agents actually perform on real engineering work, not toy prompts.

What the Framework Tests

Scenarios are derived from real Supabase support tickets and GitHub issues, which makes the benchmark harder to game than synthetic test suites. Each task runs inside a containerised sandbox, so the environment is isolated and consistent across runs.

The framework currently targets Claude Code, OpenAI Codex, and OpenCode. Results feed into leaderboards, giving teams a comparable score across agents rather than relying on vendor benchmarks alone.

Why This Matters for Operators

If you’re choosing between AI coding agents for your stack, vendor-supplied benchmarks are marketing. An open, Apache-2.0 framework built on real support tickets is closer to what your agents will actually face. Supabase running this on their own engineering problems adds credibility the synthetic leaderboards lack.

The repo is public at github.com/supabase/evals.

Stay on top of AI & Automation with BizStack Newsletter
BizStack  —  Entrepreneur’s Business Stack
Logo