Researchers present EXPL-FR, a method using a lightweight adapter to align vision-language model encoders with frozen face recognition embedding spaces, enabling semantic attribute interpretation without requiring architectural access to the underlying face-recognition model. Evaluated across four face-recognition backbones and two vision-language encoders, the method identifies 100 maximally detectable attributes forming a readable semantic signature that surpasses full-vocabulary performance, and its fully prompt-driven audit ranks face-recognition models by ethnicity-specific errors without requiring human labels.