ByteDance Seed paper studies how high-quality domain data should repeat when scaling LLMs

ByteDance

Research official + media 1 src. ~1 min

A ByteDance Seed team posted Scaling Domain Data Repetition in LLM Pretraining (arXiv 2608.14071, HF paper-page submission Aug 17, arXiv listing Aug 14). At fixed tokens-per-parameter the optimal repetition count mildly increases with model size; optimal repetition is strongly negatively correlated with validation loss across domains; counts tuned on smaller proxy models at the same TPP transfer to larger models.

Why it matters

Concrete, transferable rule for tuning how much to repeat scarce high-quality domain data as LLMs grow — directly actionable for training domain-specialised models on small corpora, since the recipe scales without retuning the search.

Importance: 3/5

open-weights / GA marker

Sources