New research investigates whether multimodal large language models can translate local visual perception into effective spatial navigation within complex urban environments. The researchers built UrbanGround, described as the first sandbox that makes this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. Their evaluation found that while current multimodal agents demonstrate competence in visual recognition and short-range reasoning, they struggle with orientation and pedestrian-aware movement, ultimately failing to sustain goal-directed behavior across extended exploration as errors accumulate without correction.