This paper introduces FACE-Eval, a 5,100-sample benchmark testing how chain-of-thought faithfulness varies based on whether preference cues originate from the user message or tool returns, and whether they’re presented as direct summaries or raw artifacts. Evaluating 15 open-weight models across eight families, the authors find all models show lower commitment to cues delivered through tools or implicit formats, with detection worsening as unverbalized adoption increases; proposed source-attribution prompt interventions only partially close the gap.
