LightNav-0 is a compact embodied-navigation model that leverages the spatial reasoning of pretrained vision-language models through a unified interface combining dual-channel pointing and a residual vector-quantized action tokenizer. It unifies instruction following, object navigation, and visual tracking without task-specific prediction heads, training on over 2,000 scenes and 4,000 hours of navigation data. The model reaches state-of-the-art results across 10 simulation benchmarks and generalizes zero-shot to diverse real-world robot embodiments.
