The Navid AI team investigated suspected vote manipulation on its Arabic TTS Arena leaderboard after one model won 91 of 92 battles in a single day. After reviewing audio samples and prompts, the team found the votes were legitimate: the model simply outperformed others specifically on Gulf Arabic dialect content. Using a dialect classifier, they determined that a single unified ranking had been masking large performance differences across five Arabic dialects (MSA, Gulf, Egyptian, Levantine, and Maghrebi), and responded by introducing dialect-specific leaderboards and more diverse dialectal test sentences rather than removing any votes.
