KATok is a transformer-based VAE with an adaptive token selector for more efficient video compression, evaluating each token’s informativeness and removing redundant ones while maintaining spatial consistency through position-prediction strategies. The authors report superior compression ratios versus existing methods by reducing spatio-temporal redundancy in video data.
