Researchers developed AstroPT, a transformer trained on millions of galaxy images, as a controlled testbed for interpretability research. By probing representations across training checkpoints, network layers, model sizes, and training objectives, they found that galaxy properties emerge in a fixed order that tracks their known difficulty, with simpler properties appearing earlier in training and shallower in the network. Linear probes recovered the known physical structure among galaxy properties, demonstrating that astronomy offers a controlled sandbox for calibrating mechanistic interpretability methods applicable to language models.