🎯 · Models

BitFit

BitFit freezes every weight matrix and fine-tunes only the bias terms for ultra-light task adaptation.

In one line

BitFit adapts a pretrained transformer by training only its bias vectors while leaving every weight matrix frozen.

ConceptWhat it is

BitFit (Bias-terms Fine-tuning) is a parameter-efficient fine-tuning (PEFT) method that freezes the entire pretrained model and updates only the bias terms — the additive b vectors inside attention projections, the feed-forward layers, and the layer-norm modules. Every weight matrix stays exactly as it was after pretraining; the only things that move during training are these small bias vectors, which typically amount to well under a tenth of a percent of the model's parameters.

It exists to answer a narrow but common question: how little can you change and still specialize a model to a task? Because the biases are a tiny, named subset, BitFit gives you a minimal-footprint adaptation you can store and swap per task while keeping one shared frozen backbone. It works by nudging behavior — reshaping how existing representations are combined — rather than by injecting new knowledge, which is why it shines on smaller datasets and has a comparatively low capacity ceiling when the task demands large shifts.

How it worksThe mechanics

Load the pretrained model, then walk its named parameters and set requires_grad = False on everything except parameters whose name marks them as a bias. Attach a standard task head and run ordinary supervised fine-tuning: forward pass, task loss, backprop. Gradients still flow through the whole network, but the optimizer only holds and updates the unfrozen bias vectors, so the frozen weights never change. After training you save just the (small) set of updated biases as the task delta; at serving time you reload the shared frozen backbone and apply the saved biases for whichever task you want, swapping bias sets to switch tasks.

At a glanceSee it

BitFit diagram
BitFit diagram 1

Where the trainable biases actually live — every weight matrix in attention, the feed-forward layers, and layer-norm stays frozen while only its additive bias vector moves.

BitFit diagram 2

The practical payoff — one frozen backbone serves many tasks by swapping in a few-megabyte bias checkpoint per task instead of storing a full model copy.

When to use itWhere it fits

  • When you need the smallest possible per-task footprint and want many tasks to share one frozen backbone, storing only bias deltas per task.
  • When the labeled dataset is small to medium and the task is close to what the base model already knows, so light behavioral nudging is enough.
  • When compute and memory are tight — the optimizer state is tiny, so training fits comfortably on modest hardware.
  • As a fast, cheap baseline to establish before reaching for heavier PEFT methods like LoRA or full fine-tuning.

When NOT to use itLimits & anti-patterns

  • When the task requires learning genuinely new knowledge or a large distribution shift from pretraining — biases alone cannot supply that capacity.
  • When you have abundant task data, where full fine-tuning or higher-capacity PEFT will pull meaningfully ahead in accuracy.
  • When you need the last few points of quality on a hard benchmark; BitFit trades peak performance for extreme thrift.
  • On architectures with few or no bias parameters, where there is little to tune in the first place.

Trade-offsAdvantages & costs

Advantages
  • Extremely parameter-efficient: only bias vectors change, so per-task storage is tiny and dozens of tasks can ride one frozen backbone.
  • Very low training cost and small optimizer footprint, since gradients update only a small named subset of parameters.
  • Simple and transparent to implement — a requires_grad mask over bias parameters, no new modules or architecture changes.
  • Preserves the pretrained weights exactly, reducing the risk of catastrophic forgetting and keeping the base reusable.
Trade-offs & costs
  • Low capacity ceiling: adapting only biases cannot represent large changes, so it underperforms on hard or far-from-pretraining tasks.
  • Accuracy gap versus full fine-tuning tends to widen as dataset size and task difficulty grow.
  • Effectiveness depends on the architecture actually having meaningful bias terms; models with few biases leave little to adapt.
  • Still requires a full backward pass through the frozen network, so it saves optimizer memory but not the full forward-plus-backward compute.

ExampleIn the real world

A team maintaining a shared encoder wants sentiment and topic classifiers for an internal support-ticket tool, with only a few thousand labeled tickets each. Instead of fine-tuning a full copy per task, they freeze the encoder and train only its bias terms for each task, producing two bias sets that are a fraction of the model's size. At serving time they host one frozen backbone and load whichever bias set the incoming request needs, switching between sentiment and topic by swapping biases. On these small, in-domain datasets accuracy lands close to full fine-tuning, storage per task is negligible, and adding a third task later is just another short bias-only training run — with the understanding that if a future task needs the model to reason over genuinely new domain content, they will move that one to LoRA or full fine-tuning.

ToolsHow to implement it

  • Hugging Face Transformers with PyTorch — the standard path: iterate named_parameters() and set requires_grad = False for all non-bias parameters, then train with the usual Trainer loop.
  • The authors' open-source BitFit reference implementation, useful for reproducing the original masked-language-model results and setup.
  • The broader PEFT ecosystem — Hugging Face PEFT and the AdapterHub adapters library — for standing BitFit up against LoRA, adapters, and prompt tuning as baselines.
  • Standard PyTorch optimizers (AdamW) which, given only bias parameters, keep optimizer state and memory very small.

Cost & effortWhat it takes

BitFit sits at the very low end of the cost curve. Engineering effort is minimal — a few lines to freeze weights and unfreeze biases — and there is little to tune beyond the usual learning rate and epochs. Training compute is a full forward-and-backward pass but with a tiny optimizer footprint, so it runs on modest GPUs and finishes quickly on small data. Storage and operational cost per task are near-negligible because you keep only bias deltas over a single shared frozen backbone. The real cost is a quality ceiling: you are trading peak accuracy and capacity for footprint, so budget for the chance that harder tasks will need to graduate to LoRA or full fine-tuning.

A living map of modern AI — kept current every morning