Code-as-World represents physical environments as executable code, capturing object states, physical parameters, and dynamics in a compact, quantitatively grounded format. An agentic discovery loop inspired by abductive reasoning proposes, executes, renders, verifies, and refines world hypotheses from multimodal observations such as language or video. The authors train Code-as-World-VL, a vision-language model that reaches state-of-the-art results on quantitative physical-reasoning benchmarks.
