FINAL-Bench released LEADBOARD, a drug-property-prediction benchmark spanning 21 boards and more than 18,000 held-out test compounds. The team found that switching from a random to a time-based data split alone reduced AUROC scores by 0.21 points, showing how much methodology choices affect benchmark comparability. The project publishes baseline results, cross-laboratory noise floors, and explicit data-split classifications before accepting submissions, aiming for more transparent benchmarking practices.