VIDRAFT’s AX-Ray diagnostic framework evaluates AI models for deployment safety beyond standard capability benchmarks, using 117 structured diagnostic items across three axes: MODEL-SCAN, AX-SCAN, and AGENT-SCAN. Applying it to two public models, Zamba2-1.2B and Nemotron-H-8B-Base-8K, the researchers found causal-leakage defects where future or suffix information changes prefix hidden states, logits, or scoring behavior in ways standard benchmarks miss. The framework maps these technical failures to governance requirements across multiple jurisdictions, treating causal leakage as a deployment-blocking defect regardless of aggregate capability scores.
