This study investigates whether purely rhetorical changes to a scientific manuscript — with the underlying content held constant — can shift the scores given by AI-based peer reviewers. The researchers built a controlled dataset of 4,200 manuscript variants from ICLR 2026 submissions, using LLMs to rewrite papers along six rhetorical dimensions, then had AI reviewers score each variant. They found rhetorical sensitivity is structured rather than uniform: framing of evidence and novelty stance produced the largest score shifts, the effect depended on a reviewer’s initial scoring tendency, and more elaborate rewriting strategies did not reliably outperform simpler ones.
