Parallel Tube Decoding (PTD) is a generative approach to spatio-temporal video grounding that decomposes localization into a temporal block followed by simultaneously decoded spatial blocks, paired with Decoupled Block Attention that enables parallel spatial generation while retaining context access. The authors report a 79-times reduction in Tube Completion Latency alongside improved grounding accuracy on multiple video-understanding tasks.
