The Open SLM Leaderboard experienced an influx of specialist models exploiting a ranking vulnerability, achieving disproportionately high scores by focusing exclusively on a single benchmark that the leaderboard’s default average-based ranking system amplified. After identifying the issue, the maintaining teams submitted pull requests to implement a more robust specialist classification algorithm that flags models with anomalous single-benchmark performance and excludes them from general rankings.