BLARM is a feed-forward method for generating animated 3D meshes from monocular video without explicit rigging or skeletal annotations, representing motion through a compact set of learned, time-varying rigid motion components blended by predicted vertex weights and conditioned on video features via spatial-temporal attention. Training combines trajectory reconstruction, entropy regularization, and motion-aware contrastive learning to produce temporally coherent, interpretable animations.
