Replaying pre-training data improves fine-tuning

#1 Replaying pre-training data improves fine-tuning [PDF⁴] [Copy] [Kimi⁸] [REL]

To obtain a language model for a target domain (e.g. math), the current paradigm is to pre-train on a vast amount of generic web text and then fine-tune on the relatively limited amount of target data. Typically, generic data is only mixed in during fine-tuning to prevent catastrophic forgetting of the generic domain. We surprisingly find that replaying the generic data during fine-tuning can actually improve performance on the (less related) target task. Concretely, in a controlled pre-training environment with 4M target tokens, 4B total tokens, and 150M parameter models, generic replay increases target data efficiency by up to $1.87\times$ for fine-tuning and $2.06\times$ for mid-training. We further analyze data schedules that introduce target data during pre-training and find that replay helps more when there is less target data present in pre-training. We demonstrate the success of replay in practice for fine-tuning 8B parameter models, improving agentic web navigation success by $4.5\%$ and Basque question-answering accuracy by $2\%$.

Subjects: Computation and Language , Machine Learning

Publish: 2026-03-05 09:00:49 UTC

2603.04964

#1 Replaying pre-training data improves fine-tuning [PDF4] [Copy] [Kimi8] [REL]

#1 Replaying pre-training data improves fine-tuning [PDF⁴] [Copy] [Kimi⁸] [REL]