A study examines how many times high-quality domain-specific data should be repeated during LLM pretraining as models and token budgets scale, since such data is otherwise diluted by the much larger volume of general web text. The researchers found that the optimal repetition count increases only mildly with model size at a fixed tokens-per-parameter ratio, and that domains achieving lower validation loss can tolerate more repetition, while the amount of available unique domain data only weakly predicts the optimal repeat count. They conclude that repetition counts tuned on smaller proxy models can reasonably estimate the right setting for larger models.