Home › Fine-Tuning & Alignment › Full fine-tuning
🎯 · Models

Full fine-tuning

Update every weight in the model on your own data, at full compute cost.

In one line

The most thorough and most expensive way to specialize a model.

ConceptWhat it is

Full fine-tuning continues training on all of a pretrained model's parameters using a task-specific dataset, rather than freezing most of the network. It exists for cases where lighter adaptation methods like LoRA or prompting do not move the model far enough, such as deep domain shifts in language, style, or reasoning pattern.

Because every weight is updated, it can reshape the model's behavior more thoroughly than adapter-based methods, at the cost of needing full training infrastructure and a much larger compute and data budget.

How it worksThe mechanics

The pretrained model's weights are loaded, gradients are computed and backpropagated through the entire network on batches of task-specific data, an optimizer updates every parameter over multiple epochs, and the result is saved as a new full-size checkpoint distinct from the base model.

At a glanceSee it

Full fine-tuning diagram
Full fine-tuning diagram 1

The decision path that earns full fine-tuning — only a deep domain shift with enough clean data and GPU budget justifies moving every weight.

Full fine-tuning diagram 2

Why touching every weight erases general ability, and the mix-in-data plus eval-gate loop that holds it back.

When to use itWhere it fits

  • Deep domain shifts, like adapting a general model to a specialized scientific or legal vocabulary at scale.
  • Large proprietary datasets big enough to justify full retraining economics.
  • Cases where LoRA and other lightweight methods have been tried and underperform on the target task.
  • Organizations with the infrastructure and budget to own and serve a fully custom checkpoint.

When NOT to use itLimits & anti-patterns

  • Small datasets, since full fine-tuning on limited data easily overfits and can catastrophically forget general capabilities.
  • Teams without serious GPU infrastructure, since full fine-tuning of large models requires distributed training setups.
  • Style or narrow-task adaptation, where PEFT methods achieve comparable results far more cheaply.

Trade-offsAdvantages & costs

Advantages
  • Maximum flexibility to reshape model behavior, since every parameter can change.
  • Often the strongest ceiling on task performance when data is abundant.
  • No adapter architecture constraints limiting what can be learned.
  • Produces a single self-contained checkpoint with no runtime dependency on a base model.
Trade-offs & costs
  • Very high compute cost, often requiring multi-GPU clusters for anything beyond small models.
  • Risk of catastrophic forgetting of general capabilities if not carefully regularized.
  • Large storage footprint per fine-tuned variant, one full model copy each.
  • Long iteration cycles make experimentation slow and expensive.

ExampleIn the real world

BloombergGPT was trained by continuing pretraining and fine-tuning on a mix of financial documents and general text, reshaping the full model rather than bolting on lightweight adapters.

ToolsHow to implement it

  • DeepSpeeddistributed training library that makes full fine-tuning of large models feasible.
  • PyTorch FSDPfully sharded data parallel training to fit large models across GPUs.
  • Hugging Face Trainerstandard training loop abstraction for supervised fine-tuning runs.
  • Weights and Biasesexperiment tracking to monitor loss and catastrophic forgetting across epochs.

Cost & effortWhat it takes

High cost, high latency to train, and significant data volume needed; substantial engineering effort for distributed training infrastructure, justified only when lighter methods fall short.

A living map of modern AI — kept current every morning