RECAP-Forcing is a training-free inference method for long video generation that organizes memory by content novelty rather than temporal recency, retaining key-value cache tied to newly appearing subjects and scenes via an attention sink and an optical-flow-based novelty bank. This appearance-indexed memory improves visual quality and semantic consistency across video generation baselines without adding learnable parameters.
