FINAL-Bench launched a live financial forecasting benchmark that scores entrants against a “luck ceiling” – the 95th-percentile return of 20,000 simulated random traders per asset – rather than only against each other. The benchmark runs on real market data with fixed leverage and realistic fees, and lets AI agents compete as first-class entrants alongside humans through an MCP server. An analysis of thirteen reference trading strategies found their relative rankings nearly inverted across different assets, suggesting market-specific character, not a single universal strategy, drives performance.
