PonderPounce is a robot-control system using a pretrained multimodal LLM as episodic memory: “Ponder,” a reasoning-focused MLLM, accumulates observations and reasoning in its native context, while “Pounce,” a lightweight vision-language-action model, receives current observations alongside Ponder’s cognition tokens. On the RoboMME benchmark the approach reaches 60.83% accuracy with a 9B model versus 44.51% for baseline methods.
