Researchers introduced SparsePR, a training-free block-sparse attention method for video transformers and world models that combines Response-Coupled Partitioning with Probe-Fitted Residual Reconstruction. The work shows that attention concentration alone doesn’t specify a valid sparse operator, since queries sharing a block route can have poorly overlapping supports, and addresses this by evaluating a small set of query rows exactly and fitting a call-specific affine correction to the sparse output. The approach aims to accelerate video-generation and world-model inference without any retraining.