A new benchmark investigates whether current video generation models can function as authentic world models by accurately reproducing the distribution of possible physical outcomes, not just a single plausible one. The researchers introduce PAWBench, a benchmark of fifty scenarios, along with an evaluation protocol called PAWEval that analyzes repeated video generations to assess whether models recover correct behavioral distributions. Testing eleven state-of-the-art systems, the study finds that no model consistently matches reference probabilities while also recovering the full range of valid behaviors, revealing a significant gap between current capabilities and genuine probabilistic world modeling.