This paper introduces MineAmongUs, a 3D multimodal sandbox environment based on Among Us in which AI agents attempt deception through both verbal and non-verbal means, alongside ARIA, a configurable vision-language-model agent framework supporting ablations. Using an annotation scheme grounded in deception taxonomies, the authors find non-verbal channels are the more decisive contributor to winning across both harness ablations and cross-VLM evaluation.
