Researchers present Gander, an end-to-end omni interaction agent that integrates multimodal perception, real-time conversation, and complex task execution within a single unified framework. The system uses a ‘Brain-Cerebellum’ collaborative design, where a lightweight cerebellum handles streaming audio-visual input and responsive dialogue while a reasoning-focused brain manages planning and tool use asynchronously. This architecture enables continuous, bidirectional human-AI interaction across modalities rather than turn-based exchanges, keeping conversational latency low alongside more sophisticated reasoning.
