Researchers evaluated 13 open-source language models on Swiss legal tasks across German, French, and Italian using three specialized benchmarks. The study found the top five models finish within 2.1 points of each other on a composite score, with performance varying significantly by task type rather than by overall model ranking. The team used LLM-based judges rather than lexical metrics and found that a 31-billion-parameter Gemma 4 model delivers competitive results on a single GPU, while larger models require multi-GPU servers.