Researchers present a system for accelerating reinforcement learning training by improving speculative decoding through online co-training of draft models. The work addresses two technical challenges: enabling branch attention in context-parallel implementations and transporting intermediate features across pipeline-parallel stages. The approach achieves substantial rollout and end-to-end speedups while maintaining alignment with baseline policies, and the authors demonstrate it scales to 122-billion-parameter models with long contexts.