Train from scratch only at frontier-lab scale, run open weights when data gravity or sustained volume demands it, and rent an API for everything else.
ConceptWhat it is
From scratch vs open-weight vs API is the model-sourcing decision: pre-train your own model from raw data, download published weights and run them on your own GPUs, or rent a frontier model through a provider's hosted API. The three routes are one gradient of ownership — at one end you choose the parameter count, the tokenizer and the context window yourself; at the other, every one of those is defined, versioned and continuously improved by the provider.
It exists because the dimensions that matter — control, security, compute, cost, scalability, governance — all move together along that gradient, so the real question is never which model is best but which layer of the stack you are prepared to own.
How it worksThe mechanics
From scratch means owning every layer. You pick the architecture and parameter count, train your own tokenizer, set the context window, pre-train on your own corpus across thousands of GPU-months, then build the full inference, scaling and governance stack around the result. Nothing is inherited and nothing is shared — data never leaves your boundary because no one else's boundary is involved.
Open-weight inherits the expensive part. The architecture, tokenizer and context window are fixed by the release (context is sometimes extensible), but the weights are downloadable and modifiable: you fine-tune on domain data, deploy into your own VPC on your own GPUs, and own scaling and governance for your deployment while the lab owns the pre-training.
API inherits everything. The provider defines the model, runs inference on elastic infrastructure, improves the weights continuously, and bills per token; your compute requirement collapses to near zero, and security and governance become a shared responsibility bounded by the provider's terms and SLAs.
At a glanceSee it
When to use itWhere it fits
- From scratchwhen you operate at frontier-lab or national scale, hold a large differentiated corpus no existing model was trained on, and intend to own every layer from tokenizer to serving.
- Open-weightwhen data gravity rules — the data cannot leave your boundary, so the model must come to it — or when sustained volume pushes unit economics past the per-token crossover.
- APIwhen time-to-value dominates: frontier quality on day one, usage-based cost that scales to zero when idle, and no MLOps capacity required.
- Mixed estates are normal — an API for the long tail of features, and an open-weight deployment for the one high-volume, data-sensitive workload.
When NOT to use itLimits & anti-patterns
- From scratch below frontier scale — it almost never pays. Choosing your own parameters and tokenizer sounds like control, but a fine-tuned open model delivers most of the capability for a small fraction of the compute, and the frontier moves faster than a bespoke model can be retrained.
- Open-weight without the volume or the MLOps capacity to keep GPUs busy — an idle reserved GPU erases the unit-economics argument that justified it.
- API where residency or governance rules forbid data crossing to a third party, or at sustained volume where the per-token bill quietly exceeds what your own infrastructure would cost.
- Treating the choice as permanent — crossovers move with every price cut and model release, so the decision should be re-run, not archived.
Trade-offsAdvantages & costs
Advantages
- From scratch gives maximum control — parameters, tokenizer, context, training data and governance are all yours, and data never leaves your boundary.
- Open-weight gives high control at medium cost — weights you can modify, private-VPC deployment, and per-token economics that beat APIs at sustained volume.
- API gives the best model quality per unit of engineering effort — elastic scale, continuous improvement, and pure usage-based OPEX with no upfront spend.
- The three compose — each workload can route to the cheapest tier that satisfies its constraints.
Trade-offs & costs
- From scratch carries very high upfront compute plus a full inference, scaling and governance stack you must build before the first useful token.
- Open-weight shifts serving, security and capacity planning onto your team, and open releases can lag the closed frontier on hard tasks.
- API limits control — provider-defined parameters, tokenizer and context, shared-responsibility security, and governance bounded by the provider's SLAs.
- Each route's economics degrade outside its zone: idle GPUs, runaway token bills, or a bespoke model the frontier laps within a year.
ExampleIn the real world
BloombergGPT was pre-trained from scratch on decades of proprietary financial data — and within months, general frontier models matched it on financial benchmarks it was built for. Meanwhile banks with the same privacy constraints fine-tune Llama-family models inside their own VPCs, and most startups ship on hosted APIs from day one: the same decision, answered three ways by what each owner could afford to own.
ToolsHow to implement it
- Megatron-LM / NVIDIA NeMothe distributed-training stacks a from-scratch pre-train actually requires.
- Hugging Face Hub + vLLMdownload open weights and serve them at high throughput on your own GPUs.
- OpenAI / Anthropic / Google APIshosted frontier models billed per token with zero infrastructure.
- AWS Bedrock / Google Vertex AIthe middle ground: open and closed models under managed billing inside your cloud boundary.
Cost & effortWhat it takes
From scratch is very high upfront — typically tens of millions in compute plus a training team before anything ships; open-weight is a medium, mostly fixed infrastructure cost that pays off only at sustained utilization; API is pure usage-based OPEX from the first call — cheapest at low or spiky volume, most expensive at sustained scale.
What changedWhat changed here
Updated this page A domain-specific reasoning model built on someone else's open weights is a concrete example of the build-on-open-weights path.
Three kinds of claim, strongest first. Signal runs every morning.