The vLLM Semantic Router team introduced a “looper” runtime that runs bounded micro-agent collaboration patterns inside the model-serving layer while preserving a single OpenAI-compatible API. It supports five looper algorithms – Confidence, Ratings, ReMoM, Fusion, and Workflows – selected dynamically based on task characteristics rather than a single fixed approach. Evaluations on LiveCodeBench, GPQA-Diamond, and Humanity’s Last Exam showed router-orchestrated model collaboration could match or exceed individual frontier models while preserving infrastructure-level control over budgets, timeouts, and failure handling.