This paper introduces DroneCATS-Agent, an architecture letting multimodal LLMs control drones directly through natural-language prompts, and DroneCATS, a benchmark spanning approaching visible targets, tracking moving objects, searching beyond initial view, and commanding multi-drone fleets. Testing frontier and open models down to 2B parameters shows spatial perception is robust, but smaller models fail by prematurely declaring task completion or not terminating correctly. The authors conclude that action-protocol execution and knowing when to stop, not navigation itself, separate deployable edge models from frontier systems.
