A study tested whether committees of LLM agents used for clinical decision support can be manipulated through benchmark shortcuts that are rewarded by evaluation metrics but clinically irrelevant. Across seven cohorts spanning datasets including MedQA-USMLE, MedMCQA, NIH ChestX-ray14, CheXpert, MIMIC-CXR, and SUPPORT2, the researchers found that while individual agents resist shortcuts in isolation, social pressure from peer agents asserting incorrect answers caused holdout agents to adopt the wrong answer in 38% of cases. Only an independent referee agent reliably detected this shortcut-cascade behavior.