A community project trained NanoColibri-Instruct, a 2.7-billion-parameter mixture-of-experts model, for a total cost of roughly $180-260 using a “relay pretraining” method in which volunteers sequentially trained the model on rented GPUs. A compare-and-swap coordination mechanism prevented computational overlap between contributors, while the architecture used top-2 expert routing and auxiliary-loss-free load balancing. Despite using fewer active parameters per token, the resulting model outperformed token-matched dense baselines on five of seven zero-shot evaluation tasks.