Developer DedeProGames reverse-engineered a social-deduction game’s engine and 59×34 map from six log files to build a benchmark for evaluating language models’ reasoning and deception abilities. The project ran 90 games across six models in the 27-31 billion parameter range, separating each model’s private reasoning from its public output. The work documented specific failure modes, including “role amnesia” and “stage direction leakage,” where models playing impostor roles broke character or leaked internal reasoning into public chat.