Researchers introduced NCP-Bench, a benchmark of 100 interactive narrative environments derived from film summaries, designed to evaluate whether language-model agents can maintain long-horizon consistency. Each environment includes structured trajectories, commitments, and facts that are automatically verified as a player agent and narrator agent interact over many turns. Current state-of-the-art models struggled significantly with narrative integrity, with the best-performing model achieving only a 42% survival rate after 20 turns and fact-conflict rates between 40% and 68%, highlighting an understudied weakness in extended agentic interaction.