Home › Models
🎯 Models

Fine-Tuning & Alignment

Actually changing the model's weights for your task, style, or values.

At a glanceFine-Tuning & Alignment

Fine-Tuning & Alignment diagram

Rule of thumb: RAG for what the model should KNOW, fine-tuning for how it should BEHAVE.

CompareThe fine-tuning landscape, side by side

Every method worth knowing, not just the headline few — filter by family (PEFT, full-weight, alignment, distillation) or search. The chips further down open full deep-dive pages for the core methods.

20 of these 21 rows have a page so far — what it does, what it costs you, and how to tell when it is the thing biting you.

LandscapeTypes & approaches

Click a highlighted type to open its own page — concept, use case, and diagram.

01 — Start here

First, try not to

Fine-tuning is the most expensive way to change a model's behaviour and the easiest thing to reach for, because it feels like real engineering. Three weeks of it will often be beaten by an afternoon of prompt work. Before you build a training set, walk this table.

The failure pattern is always the same. Someone sees a bad output, concludes "the model doesn't know our domain", and starts collecting examples. But the model usually does know — it was never told precisely what you wanted, never shown an example, never given the document, or never allowed to think before answering. Fine-tuning fixes exactly one class of problem: behaviour the model can already almost produce, made consistent. It is very bad at everything else, and the research says so.

The sharpest published result here is on knowledge. Ovadia et al. compared fine-tuning against retrieval for injecting facts into an LLM and found RAG consistently outperformed fine-tuning — "both for existing knowledge encountered during training and entirely new knowledge", and that LLMs "struggle to learn new factual information through unsupervised fine-tuning" at all. If your complaint is that the model gets facts wrong, training is close to the worst tool available.

The symptom you actually haveCheaper fix, in orderWhy it beats fine-tuning
"It gets facts about our business wrong" RAG. Retrieve the document, put it in context, cite it. Directly measured: RAG beat fine-tuning for knowledge injection on both familiar and new facts. Facts also change — weights baked on Tuesday are stale by Friday, and a retrieval index is a database write.
"It won't return valid JSON / breaks my schema" Constrained decoding / Structured Outputs. This is solved deterministically, not statistically. OpenAI's Structured Outputs guarantees "the model will always generate responses that adhere to your supplied JSON Schema, so you don't need to worry about the model omitting a required key, or hallucinating an invalid enum value." A grammar cannot emit an invalid key. A fine-tune can, just less often. Never train for syntax you can enforce.
"The tone/shape is close but drifts" Few-shot: put 3–8 real gold outputs in the prompt. Costs one afternoon and zero infrastructure. It is also the honest pre-test for fine-tuning: if the model cannot copy the pattern when the pattern is sitting in its context window, gradient descent on 500 examples is unlikely to teach it. OpenAI's own guidance: "If 50 examples have no impact, rethink your task or prompt before adding training data."
"It ignores half my instructions" Decompose the prompt. One instruction per line, put the hard constraint last, split into two calls. A 40-line prompt with a buried constraint is a prompt bug, not a weights bug. Splitting one over-loaded call into extract → format is usually a same-day fix.
"It's wrong on hard reasoning" A bigger/reasoning model, or let it think first. Reasoning is the capability most reliably bought with model choice and inference-time compute. Fine-tuning a small model rarely creates reasoning it did not have; it mostly creates confident wrong answers in your house style.
"It's inconsistent run to run" Drop temperature. Pin the model version. Free. Check this before anything else — a shocking amount of "model quality" work is temperature 1.0 and a floating model alias.
"It can't do multi-step work / can't act" Tools and an agent loop. You cannot fine-tune the ability to call your API into existence. Give it the tool.
"It's too expensive" Prompt caching, shorter context, cheap model with routing, batch API. Try these first — they are config changes. Distilling into a tuned small model is the real answer at volume (see §02), but only after the free levers are pulled.
"It's unsafe / says the wrong thing" Guardrails: input+output classifiers, refusal rules in the system prompt. An external classifier can be audited, versioned and hot-fixed in minutes. A refusal trained into weights cannot be explained to a regulator or patched before Monday.

The rule. Fine-tuning changes how the model behaves. RAG changes what it knows. Tools change what it can do. Guardrails change what you let through. If your problem is knowledge, capability or safety, you are in the wrong section of this site — and the wrong section is cheaper.

02 — The honest list

When it is the only wayWhen fine-tuning is the only thing that works

Four cases. They are real, they are worth real money, and nothing on the table above touches them. If you are not in one of these four, close the tab.

Format and style at scale✓

Not "valid JSON" — a grammar does that. This is the unspecifiable part: your house tone, your report structure, the ten thousand small conventions your senior analyst applies without being able to write them down. You cannot enumerate it in a prompt because nobody can articulate the rule. You can only demonstrate it. That is precisely what SFT is: learning a mapping you can exhibit but not describe. When the style guide is longer than the task, train it.

Domain vocabulary genuinely absent✓

Rare, and usually misdiagnosed. This is not "the model hasn't read our wiki" (RAG). It is a token distribution the base model never saw enough of to represent well: proprietary instrument codes, internal shorthand, a low-resource language, a genomics or telemetry notation. The tell is that the model mistokenises or paraphrases the terms rather than merely getting them wrong. Note the qualifier from the same paper that says RAG wins for facts: fine-tuning on new knowledge can work if you expose the model "to numerous variations of the same fact during training" — one paraphrase is not enough. That is a heavy data programme, not a weekend.

Latency and cost at volume✓

The strongest ROI case, and the most under-used. You have a narrow, high-volume task running on a frontier model with a 2,000-token prompt. Fine-tune a small model on the frontier model's outputs (distillation) and delete the prompt — the instructions are in the weights now. The published proof of the shape: InstructGPT's 1.3B model was preferred by human labelers over 175B GPT-3, "despite having 100x fewer parameters." Task-specialised small beats general large, on that task. At a million calls a month this is the whole business case.

Behaviour prompting cannot reach✓

Preferences that are comparative, not statable: which of two correct answers is better. Hedging vs committing. How to refuse this class of request in your voice. You can rank these but not specify them, so you train on the ranking — DPO or RLHF. This is also the only route when the behaviour must survive an adversarial user, since a system prompt is a request and weights are not.

Note what is not on this list: knowledge, reasoning, tool use, factuality, safety-as-compliance, and "it feels a bit off". Those have cheaper owners.

03 — The gate

PrerequisitesPrerequisites. If you fail one, stop.

This section is designed to stop you. A fine-tune without these is not a risky project — it is an unmeasurable one, which is worse, because it cannot be shown to have failed and therefore never gets cancelled.

GatePass looks likeIf you fail
1. An eval set that PREDATES the tuning A frozen, held-out set of real inputs with a scoring function you wrote before you saw any tuned output. Written down. In version control. Stop. Build the eval first. This is not negotiable and it is the gate people skip. An eval written after you see results is a description of your results. OpenAI's guide leads with exactly this: "Good evals first! Only invest in fine-tuning after setting up evals."
2. A baseline number you are trying to beat "Best prompt + few-shot + RAG scores 71% on the eval set." One number. Dated. Stop. Without it, any post-tuning number is unfalsifiable — you cannot tell a win from a wash. And measuring the baseline frequently ends the project happily: the prompt gets to 85% and the training budget goes back.
3. Enough clean examples OpenAI: floor of 10, but "We recommend starting with 50 well-crafted demonstrations and evaluating the results", with improvements typically appearing around 50–100. For open-weight LoRA on a real style/format task, plan for hundreds to a few thousand. Clean beats numerous: consistency in the labels matters more than count. Not ready. Go label. But run the 50-example probe first — OpenAI's own stop condition is blunt: "If 50 examples have no impact, rethink your task or prompt before adding training data." Fifty examples that move nothing is a signal about your task, not about your dataset size.
4. Your examples agree with each other Two labellers given the same input produce substantially the same output. You measured this. Stop. If your humans disagree, you do not have a task — you have an unresolved product decision. The model will faithfully learn the average of your disagreement, which is a style nobody wanted. Resolve the spec, then label.
5. Someone owns the retrain A named person, a cadence, a budget line for the next 12 months. Reconsider. A fine-tune is not a deliverable, it is a dependency. Base models get deprecated, your data drifts, your product changes. An orphaned adapter is a slowly-rotting liability that nobody can retrain because the person who built the pipeline left.
6. You can serve it A concrete answer to: who hosts the weights, what does an idle GPU cost, how do you roll back? Reconsider. A managed API has no idle cost. A self-hosted 7B adapter does, whether or not anyone calls it. This line often destroys the ROI case on its own.

Gates 1 and 2 are the same discipline stated twice, deliberately. The eval set is the deliverable. The tuned model is a by-product. Teams that build the eval first frequently never fine-tune — the eval shows them which of the §01 fixes to apply, they apply it, and they are done in a week. That is the good outcome, not a failure to ship.

04 — End to end

The steps, in order

Data prep → split → baseline → method → train → eval → deploy → monitor. The order matters more than any individual step: baseline before training, split before looking.

1 · Data prep. Decide the exact shape first, because it determines the method. TRL's SFTTrainer accepts prompt-completion ({"prompt": …, "completion": …}) or conversational ({"messages": [{"role": …, "content": …}]}) formats, standard or chat-templated. Preference methods need a different shape entirely: {"prompt": …, "chosen": …, "rejected": …}. Two decisions people get wrong here: compute loss on the right tokens — for prompt-completion data TRL computes loss on the completion only by default (completion_only_loss), and for chat data assistant_only_loss=True restricts it to assistant turns; training on the user's words teaches the model to write your users' questions. And match the chat template exactly — a mismatched EOS token is the single most common cause of a model that will not stop generating.

2 · Split — before you look at anything. Train / validation / test. Split by entity, not by row: if one customer, document or ticket thread appears in both train and test, your test score is measuring memorisation and will be beautiful and worthless. The test set is opened once, at the end. If you tune hyperparameters against your test set you no longer have one.

3 · Baseline. Run gate 2. Score the best prompt-only system on the eval set and write the number in the run log. Everything downstream is measured against this single number.

4 · Method choice. §05 is the decision. Default for almost everyone: LoRA, or QLoRA if the model does not fit your GPU. Full fine-tuning is a deliberate, justified exception, not a starting point.

5 · Train. Start with the smallest thing that could work: one epoch, a known-good config (§06), the smallest viable base model. Watch the validation loss, not the training loss. Training loss going down while validation loss goes up is memorisation — stop and take the earlier checkpoint. Save checkpoints often enough that "stop early" is an option you actually have.

6 · Eval against the baseline. Two questions, both mandatory. Did it beat the baseline on the target task? And — the one everyone skips — what did it break? Run a held-back general capability check. Biderman et al. found full fine-tuning degrades performance outside the target domain measurably more than LoRA does; LoRA "better maintains the base model's performance on tasks outside the target domain." Something is usually worse. Find out what, and decide whether you accept it, before a user does.

7 · Deploy. Ship behind a flag, on a fraction of traffic, with the prompt-only path still live and one config change away. For LoRA you have a genuine choice: serve the adapter separately (swap adapters per tenant on one base model — the modularity that makes LoRA a product decision, not just a training trick) or call merge_and_unload() to fold the weights in and pay zero inference latency. The LoRA paper's claim is precisely that: "no additional inference latency," because the update is merged back into the frozen weights.

8 · Monitor drift. Three clocks are ticking. Your input distribution drifts (new customers, new products). Your base model gets deprecated — and your adapter is welded to a specific base checkpoint, so its deprecation notice is your retraining deadline. Your eval ages, because it was drawn from last year's traffic. Re-run the eval on a schedule, not when someone complains. Keep the prompt-only baseline runnable forever: the day it catches up with your tuned model — and model releases do this regularly — you should find out from a dashboard and delete the pipeline with a clear conscience.

05 — The matrix

Your situation → the method

Rows are what you actually said in the meeting. Columns are the options. Read your row; the verdict includes why. Note how many rows resolve to the last column.

Your situationLoRAQLoRAFull FTInstruction-tune (SFT)DPORLHFVerdict
"It won't hold my output format" ✅ Best fit — if the format is a style you can only demonstrate ✅ Same, when GPU-bound ⚠️ Overkill ✅ This is SFT — LoRA is how you run it ❌ Wrong tool ❌ Absurd Usually DON'T. If the format is machine-checkable — JSON, a schema, an enum — use constrained decoding and stop. Structured Outputs guarantees schema adherence; training only makes violation rarer. Fine-tune only for the unspecifiable house style that survives the schema. → LoRA + SFT.
"Domain terms it has never seen" ⚠️ Works if the terms exist but are weak ⚠️ Same ⚠️ Only for genuinely new token distributions, and expensive ⚠️ Only via many paraphrases ❌ ❌ DON'T — first. 90% of the time this is knowledge, and RAG "consistently outperforms" fine-tuning for it, on both old and new facts. Only if the model mis-tokenises the vocabulary is this real; then continued pretraining on a lot of in-domain text, plus "numerous variations of the same fact." Budget months.
"Facts are wrong / out of date" ❌❌❌❌❌❌ DON'T. Ever. Weights are a snapshot; facts change. This is RAG's job by measurement, not by opinion — and unsupervised fine-tuning barely learns new facts at all.
"Too slow/expensive at my volume" ✅ Tune a small model on the big one's outputs ✅ Cheapest way to run it ✅ Justified here — you want max quality from a small model you will serve forever ✅ Distillation is SFT on teacher outputs ⚠️ Only to polish ❌ The best case for fine-tuning. Precedent for the shape: 1.3B InstructGPT was preferred to 175B GPT-3. Pull the free levers first (caching, shorter prompts, routing, batch). If they fail and volume is high, distil. Check the maths: idle GPU cost vs per-token API cost — this is where the case usually dies.
"Needs to refuse differently" ⚠️ Carries the adapter ⚠️ Same ❌ Disproportionate ⚠️ Only if you can write the ideal refusal ✅ Right tool — you can rank refusals even when you can't specify them ⚠️ Only at frontier-lab scale Guardrail first, DPO second. A classifier is auditable, versionable and hot-fixable; weights are none of those. Train the refusal only when it must survive an adversarial user, or when "better" is comparative. Then: DPO on LoRA.
"Two answers are both right, one is better" ⚠️ The vehicle ⚠️ The vehicle ⚠️ ❌ SFT can't express "better" ✅ Exactly what DPO is for ⚠️ Better ceiling, far more machinery DPO. Preference data (prompt / chosen / rejected) encodes comparative quality that SFT structurally cannot. DPO is "stable, performant, and computationally lightweight, eliminating the need for sampling from the LM during fine-tuning" — no reward model, no PPO loop. SFT first, then DPO on top.
"Model ignores instructions / rambles" ⚠️⚠️❌ ✅ Only if you are starting from a base (non-instruct) model ❌❌ DON'T. Use an instruct-tuned checkpoint — someone already did this properly with more data than you have. Then fix the prompt. Instruction-tuning is for turning a base model into a chat model, and if you are asking this question you should not be starting from a base model.
"Model won't fit on my GPU" ❌ Not the constraint LoRA solves ✅ Precisely this ❌ Impossible n/an/an/a QLoRA — 4-bit NF4 base + LoRA adapters. QLoRA finetunes a 65B model on a single 48GB GPU while preserving "full 16-bit finetuning task performance." Costs ~30% wall-clock in one published comparison. Memory problem, memory answer.
"Need maximum quality, cost no object" ⚠️ Underperforms at low rank ❌ ✅ Genuinely better on target ✅✅ Stack after ⚠️ Full FT — with eyes open. Biderman et al.: "in the standard low-rank settings, LoRA substantially underperforms full finetuning" on code and math, because "full finetuning learns perturbations with a rank that is 10-100X greater than typical LoRA configurations." The price is forgetting: full FT degrades off-domain more. You are trading generality for peak. Also try LoRA at high rank first (§06).
"Need many variants — per client / per tenant" ✅ The killer feature ✅ ❌ One full copy of the weights each. Don't. ✅⚠️❌ LoRA, decisively. "You can have multiple lightweight and portable LoRA models for various downstream tasks built on top of them" — many adapters, one base model in memory, swapped per request. This is an architecture decision as much as a training one.
"We have no eval set" ❌❌❌❌❌❌ DON'T. Go to §03. Whatever you train, you will not be able to tell whether it worked. "Good evals first! Only invest in fine-tuning after setting up evals."
"We have <50 examples" ❌❌❌❌❌❌ DON'T — put them in the prompt. At that count they are worth more as few-shot examples than as gradients. OpenAI recommends starting at 50 demonstrations and treats no-impact-at-50 as a signal to rethink the task, not to label more.
"We want our own ChatGPT-quality alignment" ⚠️⚠️⚠️✅ Step one ✅ Do this instead ⚠️ Almost certainly not DPO, not RLHF. RLHF works — InstructGPT is the proof — but it needs a reward model, a PPO loop, sampling during training and a labeling operation. DPO gets the same objective "with only a simple classification loss" and "matches or improves response quality... while being substantially simpler." Choose RLHF/online RL only with a dedicated team and a reason.

How to read the ✅/⚠️/❌. ✅ = the right tool for this row. ⚠️ = works, but it is not what is standing between you and the outcome. ❌ = will not fix your problem no matter how well you execute it. The most important cells on this page are the ❌ ones.

06 — Knobs

ParametersParameters, and what each one actually trades off

Defaults below are the documented library defaults from PEFT and TRL, or values from a specific cited experiment. Read the honesty note at the end of this section before you copy any of it into a config file.

What each knob buys you.

KnobWhat it doesThe trade-offDocumented value
LoRA rank r Rank of the two low-rank update matrices. Sets capacity of the adaptation. Low r = fewer trainable params, less memory, less capacity, more regularisation. High r = closer to full FT, more memory, more overfitting risk on small data. PEFT: "Lower rank results in smaller update matrices with fewer trainable parameters." PEFT example uses r=16. Raschka swept 8→2048 and found r=256 best for that setup.
lora_alpha Scaling factor. Effective scale is lora_alpha/r (or lora_alpha/√r with rsLoRA). Because it is a ratio, raising r without raising alpha silently shrinks your update. Too high destabilises; too low means the adapter barely moves the model. Rule of thumb alpha = 2 × r — Raschka: "The common recommendation of choosing alpha as two-times the rank... indeed yielded the best results," and deviating either way was worse. PEFT's own example uses alpha = r.
use_rslora Switches scaling to lora_alpha/√r. Makes high ranks usable — under the default alpha/r, raising r fights itself. PEFT: "stabilizes the adapters and unlocks the increased performance potential from higher ranks." Turn on if r > 64.
lora_dropout Dropout on the adapter path. Regularisation vs learning speed. Matters most on small datasets where overfitting is the live risk. 0.1 in PEFT's documented example.
target_modules Which matrices get an adapter. Attention-only (q_proj,v_proj) is cheapest; all-linear costs more memory and does better. PEFT: LoRA "is typically only applied to the attention blocks in Transformer models" for "simplicity and further parameter efficiency" — i.e. the default is a cost choice, not a quality one. target_modules="all-linear" is QLoRA-style and per PEFT gives "performance equal to a fully finetuned model." Raschka independently found enabling LoRA on all layers beat Q/V-only.
bias Whether bias terms train. Anything but "none" means disabling the adapter no longer returns you the base model exactly — you lose clean A/B. "none" in every PEFT example.
Learning rate (SFT) Step size. The single most destructive knob. Too high on full FT = catastrophic forgetting. Adapters take a much higher LR than full FT because only new parameters are learning. TRL SFTConfig default 2e-5 (vs 5e-5 in base TrainingArguments). TRL tip for adapters: "≈1e-4".
Learning rate (DPO) Step size for preference training. Preference training is far more fragile than SFT — an order of magnitude lower. TRL DPOConfig default 1e-6. TRL tip for adapters: "≈1e-5".
DPO beta How far the policy may drift from the reference model. Higher β = less deviation, safer, less effect. Lower β = stronger preference signal, more drift, more reward hacking and degeneration. TRL default 0.1. Documented as "Higher β means less deviation from the reference model."
Epochs Passes over your data. More epochs = more fit and more memorisation. This is where small datasets die. Start at 1. Raschka: two passes over the 50k Alpaca set was worse than one across all benchmarks. Trust the validation curve, not the count.
Batch size / grad accumulation Examples per optimizer step. Larger = smoother gradients, more VRAM. Gradient accumulation buys the effective batch without the memory — at wall-clock cost. Mostly a hardware knob, not a quality knob. Set it to whatever fits, then use accumulation to reach the effective batch you want.
max_length / packing Sequence length; packing groups short examples into fixed-length blocks. Truncation silently destroys training data — if your examples exceed max_length, you are training on their first half. Packing improves throughput. TRL default max_length=1024, packing=False. Check your token length distribution before accepting 1024.

Starting configs by scenario. These are documented defaults and cited values assembled into sensible starting points — they are a place to begin a sweep, not answers.

ScenarioMethodStarting configBasis
House style / format, a few hundred examples LoRA + SFT r=16, alpha=32, dropout=0.1, target=["q_proj","v_proj"], bias="none", lr=1e-4, epochs=1, loss on completion only PEFT example values; TRL's ≈1e-4 adapter tip; alpha=2r rule. Low rank is deliberate — small data, and low rank regularises.
Harder task, thousands of examples, quality matters LoRA + SFT r=64–256, alpha=2r, use_rslora=True, target="all-linear", lr=1e-4, epochs=1–2 Raschka's sweep (r=256/alpha=512 best; all-layers > Q,V-only); PEFT on rsLoRA unlocking high rank and on all-linear giving "performance equal to a fully finetuned model."
Model doesn't fit the GPU QLoRA 4-bit NF4 + double quantization + paged optimizers, then LoRA as above. Consider init_lora_weights="loftq". QLoRA paper (65B on one 48GB GPU at "full 16-bit finetuning task performance"); PEFT on LoftQ minimising quantization error. Expect ~30% slower (21.33GB→14.18GB, 6,685s→10,059s in Raschka's run).
Preference / "which answer is better" DPO on a LoRA SFT first. Then beta=0.1, lr=1e-5 (adapter) or 1e-6 (full), loss_type="sigmoid", max_length per your data TRL DPOConfig defaults + TRL's adapter LR tip. Sigmoid is the original DPO loss; TRL ships ~15 variants — do not start there.
Distil a big model into a small one SFT (LoRA or full) Generate teacher outputs on real traffic → SFT the small model → eval vs teacher, not vs the old small model The InstructGPT 1.3B-beats-175B result is the precedent for the shape. Full FT is defensible here: one model, served forever, quality is the product.
Low rank isn't enough but you can't afford full FT LoRA variants use_dora=True, or init_lora_weights="pissa" PEFT: DoRA "can improve the performance of LoRA especially at low ranks" but "introduces a bigger overhead than pure LoRA"; PiSSA "converges more rapidly than LoRA and ultimately attain superior performance."

Honesty note — read this before you trust the table above. You will see pages that claim a specific rank and learning rate for every model × task, with a before/after number attached. Those numbers are almost always invented, or copied from one experiment and presented as a law. Here is what is actually true:

  • Published before/after numbers do not exist for most model × scenario pairs.Nobody has run the grid. What exists is: library defaults (PEFT, TRL), a handful of paper-scale results on specific benchmarks, and a small number of careful public sweeps on one model and one dataset.
  • The r=256 / alpha=512 result is one experiment.Raschka's sweep is a single model on a single instruction dataset. It is real, it is well documented, and it is not a universal setting. It is evidence that r=16 may be leaving quality on the table — not proof that 256 is your answer.
  • "99.3% of ChatGPT" needs its asterisk.QLoRA's Guanaco result — 99.3% of ChatGPT's level after 24 hours of finetuning on one GPU — is judged by GPT-4 and human raters on a specific benchmark, not on your task. Cite it for the memory claim, which is a hard engineering result. Treat the quality percentage as a headline, not a forecast.
  • LoRA vs full FT depends on the domain and the data budget.The most careful public comparison (Biderman et al., programming and math; ~100K instruction pairs and 20B tokens of continued pretraining) found LoRA "substantially underperforms full finetuning" on target-domain tasks in standard low-rank settings — while forgetting less, and regularising better than weight decay or dropout. Both halves are the finding. Anyone quoting only one half is selling something.
  • The only number that matters is yours.Your baseline, on your eval set, from §03. Every value in this section is a starting point for a sweep you still have to run.
References

Sources

Every external number on this page comes from one of these. Where a claim has no citation, it is an engineering judgement and is labelled as one.

  1. Hugging Face PEFT — LoRA conceptual guide
    https://huggingface.co/docs/peft/main/en/conceptual_guides/lora
  2. Hugging Face PEFT — LoRA developer guide (rank, alpha, rsLoRA, DoRA, LoftQ, init)
    https://huggingface.co/docs/peft/main/en/developer_guides/lora
  3. Hugging Face PEFT — library overview
    https://huggingface.co/docs/peft/main/en/index
  4. Hugging Face TRL — SFTTrainer / SFTConfig
    https://huggingface.co/docs/trl/main/en/sft_trainer
  5. Hugging Face TRL — DPOTrainer / DPOConfig
    https://huggingface.co/docs/trl/main/en/dpo_trainer
  6. Hu et al. 2021 — LoRA: Low-Rank Adaptation of Large Language Models (arXiv:2106.09685)
    https://arxiv.org/abs/2106.09685
  7. Dettmers et al. 2023 — QLoRA: Efficient Finetuning of Quantized LLMs (arXiv:2305.14314)
    https://arxiv.org/abs/2305.14314
  8. Biderman et al. 2024 — LoRA Learns Less and Forgets Less (TMLR; arXiv:2405.09673)
    https://arxiv.org/abs/2405.09673
  9. Ovadia et al. 2023 — Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs (arXiv:2312.05934)
    https://arxiv.org/abs/2312.05934
  10. Ouyang et al. 2022 — Training language models to follow instructions with human feedback / InstructGPT (arXiv:2203.02155)
    https://arxiv.org/abs/2203.02155
  11. Rafailov et al. 2023 — Direct Preference Optimization (arXiv:2305.18290)
    https://arxiv.org/abs/2305.18290
  12. OpenAI — Supervised fine-tuning guide
    https://developers.openai.com/api/docs/guides/supervised-fine-tuning
  13. OpenAI — Structured Outputs guide
    https://developers.openai.com/api/docs/guides/structured-outputs
  14. Raschka 2023 — Practical Tips for Finetuning LLMs Using LoRA (Lightning AI)
    https://lightning.ai/pages/community/lora-insights/

GlossaryKey terms

CheckCheck your understanding

When is fine-tuning worth it?

When you need a consistent style/format/skill prompting can't reliably produce, or to shrink a large model into a cheaper specialised one. Not for facts - use RAG.

What is LoRA?

It freezes the base model and trains small adapter matrices - cutting cost and enabling many task-specific variants on one base model.

RAG, fine-tuning, or both?

Both is common: fine-tune for behaviour/format, RAG for current knowledge. Prompt-engineer first.

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page A domain-specific reasoning model built on someone else's open weights is a concrete example of the build-on-open-weights path.

    TechCrunch AI · 14 Sep 2026 · source

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning