Supabase has released an open source benchmark framework called Evals, which evaluates the performance of AI coding agents such as Claude Code, Codex, and OpenCode on real-world tasks. The framework runs these agents against Supabase tasks, including schema building, debugging, and policy fixing, and scores their performance using deterministic checks and a language model-based evaluation system. The benchmark is designed to run in containerized environments and is released under the Apache-2.0 license. This development could help improve the efficiency and effectiveness of AI-powered coding tools in real-world applications. It matters because it may lead to more reliable and accurate AI-powered coding solutions.