Continue the same self-supervised pretraining on a large domain or language corpus to teach a base model knowledge it never saw.
ConceptWhat it is
Continued pre-training, also called domain-adaptive pretraining or continual pretraining, resumes a base model's original self-supervised objective, next-token prediction over raw text, but on a large corpus from a target domain or language rather than a labeled task dataset. It exists because supervised fine-tuning and prompting reshape how a model uses what it already knows, yet cannot efficiently teach it facts, vocabulary, or a language it barely saw in the original run; continued pre-training genuinely adds that knowledge by updating all weights over billions of new tokens.
Because it reuses the same unlabeled objective as the original pretraining run, it needs no human annotation, only a large and clean domain corpus. It is nonetheless a heavyweight adaptation: every parameter moves, the output is a new domain-adapted base checkpoint per model, and you almost always follow it with instruction tuning or alignment before the model is usable as an assistant.
How it worksThe mechanics
You start from a released base checkpoint, assemble a large domain corpus, often tens to hundreds of billions of tokens, and tokenize it with the model's existing tokenizer; you then resume the causal-language-modeling objective of predicting the next token, running gradient updates over every parameter at a lower learning rate than the original run and usually mixing in a slice of general pretraining data as replay to limit forgetting; after one or a few passes, with held-out perplexity and a general benchmark tracked to catch regressions, the result is saved as a new domain-adapted base model that is handed on to instruction tuning and alignment.
At a glanceSee it
The decision the pipeline assumes away — when a missing capability actually warrants continued pre-training versus retrieval, fine-tuning, or prompting.
Inside the single training-resume box — tokenizer extension, learning-rate re-warming, and replay mixing, each guarding against a specific failure.
When to use itWhere it fits
- Injecting a domain the base model barely saw, such as dense legal, biomedical, or financial text with its own vocabulary and conventions.
- Teaching a language that was underrepresented in the original pretraining mix, so the model reads and writes it fluently.
- Adapting to a large proprietary corpus, like internal codebases or technical manuals, where the knowledge itself, not just the task format, is missing.
- When you have billions of clean domain tokens and want a reusable base to fine-tune many downstream tasks from.
When NOT to use itLimits & anti-patterns
- You only need to change output format, tone, or task behavior, where instruction tuning or LoRA does the job far more cheaply.
- Your domain corpus is small, megabytes rather than billions of tokens, so continued pretraining overfits and catastrophically forgets general skills.
- The knowledge changes frequently or must stay current, where retrieval keeps facts fresh without any retrain.
- You lack the GPU budget and MLOps to run a multi-billion-token job and evaluate it for regressions.
Trade-offsAdvantages & costs
Advantages
- Genuinely adds knowledge, vocabulary, and even a new language into the weights, which lighter methods cannot do.
- Reuses the same unlabeled self-supervised objective, so it needs no human annotation, only raw domain text.
- Produces a reusable domain base checkpoint that many downstream fine-tunes and tasks can build on.
- Can lift quality across a whole domain at once, rather than one narrow task at a time.
Trade-offs & costs
- High compute cost, since billions of tokens over all parameters means large GPU clusters and long runs.
- Risks catastrophic forgetting of general ability unless you replay general data and tune the learning rate carefully.
- Any knowledge update means retraining, so you cannot patch facts incrementally the way retrieval can.
- Yields a full-size checkpoint per model to store and serve, and still needs alignment before it is a usable assistant.
ExampleIn the real world
A team building an assistant for radiology reports finds that a general open-weight base model misreads clinical abbreviations and imaging terminology. They gather a de-identified corpus of tens of billions of tokens from radiology literature, guidelines, and internal reports, then run continued pre-training on the base checkpoint with a slice of general web text mixed in to preserve fluency. They track perplexity on held-out clinical text alongside a general benchmark to catch forgetting, stop after two passes once domain perplexity plateaus, and hand the resulting domain-adapted base off to instruction tuning so it can answer questions and summarize reports.
ToolsHow to implement it
- NVIDIA NeMo / Megatron-LMlarge-scale distributed pretraining frameworks built for continued-pretraining-sized runs.
- PyTorch FSDP / DeepSpeedsharded training that fits multi-billion-parameter models across many GPUs.
- Hugging Face Transformers and Datasetscausal-LM training loop and streaming-corpus tooling to resume the pretraining objective.
- Weights and Biasestrack training loss, held-out perplexity, and general-benchmark regression to detect forgetting.
Cost & effortWhat it takes
High cost and the heaviest of the adaptation methods: expect a multi-GPU cluster running for days over billions of tokens, plus the data engineering to collect, clean, and deduplicate the corpus. There is no annotation cost because the objective is self-supervised, but the compute, the storage per checkpoint, and the evaluation effort to guard against forgetting make it worth doing only when you genuinely need new knowledge in the weights and can amortize the base across many downstream models.