“Acquire, Repair, Preserve” is a three-stage post-training recipe for small language models playing interactive dialogue games: supervised fine-tuning for broad game participation, targeted preference learning to fix mechanical failures like repeated guesses, and techniques to preserve general capabilities. Tested on the LM Playschool Challenge with a 2B model, the recipe substantially improves game performance while largely preserving static benchmark scores.
