Lucida converts real indoor scenes into editable 3D simulation environments through a three-stage pipeline: parsing video into scene graphs with per-instance evidence, generating complete assets for each instance, and placing them with GizmoAct, a vision-language-model policy that casts placement as multi-turn GUI interaction. The method reports a 69% mAP gain over baseline approaches on scene-level object detection.
