This article introduces a metacognition benchmark and leaderboard designed to measure whether large language models can recognize their own errors before making them. The framework evaluates models along two axes: vulnerability to logical traps and the performance gain achievable from dedicated error-detection adapters. The authors released a 400-problem benchmark, ranked 24 models, and built 11 frozen-base adapters, finding that even top-performing models struggle with self-awareness on free-form writing tasks.