A new benchmark called OSReward evaluates how reliable vision-language models are when used as automated judges of computer-using agent trajectories across web, mobile, Ubuntu, and Windows environments. The authors find that even state-of-the-art VLM judges share a systematic leniency bias that mislabels failed agent runs as successful, and that the few models reliable enough to trust are too costly to run at scale. To close the gap, they release OS-Shepherd, open reward models trained on 100,000 human-annotated trajectories that match commercial judge accuracy at a fraction of the inference cost.