Home › LLMs & Foundation Models › Base / pre-trained
🤖 · Models

Base / pre-trained

The raw, next-token model before any instruction tuning shapes its behavior.

In one line

A base model just predicts the next token from raw text and has not been taught to follow instructions.

ConceptWhat it is

A base model (or pre-trained model) is the direct output of large-scale next-token prediction training on a broad corpus of text and code, before any instruction tuning or reinforcement learning shapes it toward following commands or chatting. It exists as the foundational asset: an enormous amount of world knowledge and language pattern compressed into weights, ready to be specialized further.

Prompted directly, a base model tends to continue text in the style it was given rather than answer a question helpfully, which is why almost no end user interacts with a raw base model directly.

How it worksThe mechanics

Training feeds trillions of tokens of text and code through the model with a single objective: predict the next token given everything before it, adjusting weights via gradient descent to minimize prediction error at massive scale. The result is a model that has learned grammar, facts, reasoning patterns, and style purely from exposure, but has no explicit notion of being an assistant or following a user's intent.

At a glanceSee it

Base / pre-trained diagram
Base / pre-trained diagram 1

The fork every base model faces — steer it in-context with a few-shot prompt, or spend compute post-training it into an instruction-following assistant.

Base / pre-trained diagram 2

How few-shot prompting actually works — the base model just continues the pattern in the prompt, so you tune the examples rather than the weights.

When to use itWhere it fits

  • As the starting point for further instruction tuning or domain-specific fine-tuning.
  • Research into raw model capabilities, biases, or emergent behavior before alignment.
  • Few-shot completion tasks where showing examples in the prompt works better than instructions.
  • Building a custom fine-tune where full control over the training recipe matters.

When NOT to use itLimits & anti-patterns

  • Direct end-user chat or assistant products, where a base model gives inconsistent, unhelpful, or unsafe completions.
  • Tasks needing reliable instruction following, since base models were never trained to comply with commands.
  • Safety-sensitive deployments, since base models lack the alignment tuning that reduces harmful outputs.

Trade-offsAdvantages & costs

Advantages
  • Maximum flexibility as a foundation for any downstream fine-tune.
  • Captures broad world knowledge and language patterns from massive pretraining data.
  • No behavioral constraints imposed by alignment tuning, useful for research.
  • Often cheaper to license or self-host than instruction-tuned frontier variants.
Trade-offs & costs
  • Unreliable at following direct instructions without careful prompting.
  • Can produce toxic, biased, or unsafe completions without alignment safeguards.
  • Requires additional fine-tuning investment to be product-ready.
  • Prompting technique matters far more than with a chat-tuned model.

ExampleIn the real world

Meta releases base checkpoints of Llama alongside instruction-tuned versions, letting research labs and companies build their own custom fine-tunes on top of the raw pretrained model.

ToolsHow to implement it

  • Hugging Face Hubhosts base checkpoints for Llama, Mistral, and other open-weight families.
  • Axolotl / torchtuneframeworks for fine-tuning base models into instruction-tuned variants.
  • EleutherAI's lm-eval-harnessbenchmarking raw base model capability.
  • Together AIhosts and fine-tunes open base models at scale.

Cost & effortWhat it takes

Pretraining from scratch costs millions of dollars; using an existing open base checkpoint for further fine-tuning is comparatively cheap; needs no additional data to use as-is, but needs curated data to specialize.

A living map of modern AI — kept current every morning