Supabase released Evals, an Apache-2.0 licensed benchmark framework for evaluating AI coding agents on authentic engineering tasks. The system executes evaluations in containerized Docker environments using CLI and Model Context Protocol interfaces, combining deterministic checks with LLM-based scoring. Testing revealed top models like Opus 5 achieved 100% pass rates, while loading Supabase skills improved smaller model performance from 78% to 100% on build-stage tasks.