ElephantBench is a benchmark of 1,094 questions testing whether LLMs can recall multiple divergent accounts of long-tail facts, built via a graph-based pipeline that mines disagreements from low-exposure corpora to create multi-account question-answer records. The authors find even the strongest models recover both accounts on only 52.4% of questions, revealing systematic incompleteness in parametric memory of factually disputed topics.