Researchers introduced TAVR, a talking-avatar generation method that uses short video clips rather than a single reference image to preserve a subject’s identity across different scenes. Built on a Wan2.1 diffusion backbone, the method combines token selection, reference self-attention, and audio cross-attention to aggregate identity cues from multiple reference frames, and is trained in three stages: same-scene pretraining, cross-scene fine-tuning, and reinforcement learning. On a new 158-pair cross-scene benchmark the researchers introduced, TAVR scored 16.42 on overall quality versus 14.13 for the next-best baseline; the paper was accepted at SIGGRAPH Asia 2026.