Tip: print this page (or save as PDF) for a portable one-sheet. Every chip opens its own page; “Show all” reveals the rest of a topic.
What is different about this stack10
Ten places where building on models departs from ordinary software engineering. Every row has pages below it.
| Concern | Traditional software | A production AI system |
|---|---|---|
| Correctness | Deterministic. The same input returns the same output, every time. | Probabilistic. The same input returns different output, so correctness is a rate over a sample rather than a property of a run. |
| Testing | Unit tests assert exact values. A pass is a boolean. | Eval sets score a held-out sample. A pass is a threshold on a rate, and it moves when the model, the prompt or the corpus moves. |
| How it fails | It throws. You get a stack trace pointing at the line. | It answers confidently and wrongly, with no error anywhere. The dangerous failure is the one that looks exactly like success. |
| Cost shape | Fixed per deploy, scaling with infrastructure and traffic. | Per call and per token, so a longer answer costs more than a shorter one and a retry costs the whole call again. |
| Latency | Milliseconds, and roughly constant for the same operation. | Seconds, varying with how much the model chooses to say. Time to first token and total time are different products. |
| Security boundary | Around the process and the network. | Around every value that enters the context. Retrieved documents and tool results are attacker-controlled input in the same token stream as your instructions. |
| What changes under you | Only the code you deployed. | The model version, the provider's defaults, the corpus and the prompt — four moving parts, three of which you do not own. |
| Improvement | Reproduce the bug, fix it, add a regression test. | Read a sample of real failures, label what went wrong, turn the recurring ones into an eval set, change one thing, re-measure. |
| Data | A schema you designed and control. | A corpus with a licence, a stated purpose and an erasure obligation, plus derived copies in an index, a cache and possibly some weights. |
| Observability | Logs and metrics: what was called, how long it took. | Traces including what was retrieved and which prompt was assembled. Without those an answer cannot be explained after the fact. |
Foundations3
the data-science bedrock every model stands on
🛤️ The Road to LLMs
How seven decades of AI led to models that understand and generate language.
Left to right: each era removed a limitation of the one before it, until models learned language itself.
📊 Data Science & ML Fundamentals
The discipline of turning data into predictions - the bedrock under all AI.
Most enterprise 'AI' still lives in this loop; an LLM is one option inside 'Train', not a replacement for it.
🧠 Neural Networks & the Transformer
The architecture family - and the one design - that made LLMs possible.
Read bottom-up: self-attention lets every token weigh every other token, in parallel.
Models3
the engine you call — and how to change what it does
🤖 LLMs & Foundation Models
Large models pre-trained on vast data, adaptable to almost any task.
The training pipeline explains why a raw base model and a helpful chat model behave so differently.
🧭 Embeddings & Vector Search
Turning text and images into meaning-carrying numbers you can search.
Similar meanings land close together, so 'search by meaning' becomes 'find the nearest vectors'.
Show all 33
🎯 Fine-Tuning & Alignment
Actually changing the model's weights for your task, style, or values.
You are changing the weights, not the prompt. Reach for it to move behaviour, style and format — facts belong in retrieval, which is cheaper and updatable.
Show all 20
Ground2
making a general model useful for your job
✍️ Prompt Engineering
Steering the model with how you ask - no training required.
Prompting is the cheapest, fastest lever; always try it before RAG or fine-tuning.
Show all 22
📚 RAG - Retrieval-Augmented Generation
Give the model your data at answer-time instead of retraining it.
The model never learns your data - you fetch the right passages at question-time and put them in the prompt.
Show all 17
Build5
wiring models into real, multi-step products
🔗 Orchestration Frameworks
The glue that chains models, data, and tools into a real workflow.
Frameworks wire these steps together, add memory, and handle branching and retries.
🕹️ Agents & Tool Use
Models that plan, call tools, and take multi-step actions toward a goal.
An agent decides which tools to call and loops on results. Autonomy adds power - and reliability risk you must contain.
Show all 27
Agent Skills
What a skill actually is, which surface it runs on, and what it is constantly confused with — every claim dated, because provider surfaces move monthly.
A skill is a folder of instructions the model loads when it needs them. Which surface it runs on decides what it can actually do, and those surfaces move monthly.
Show all 35
🗄️ Data for AI
The pipelines, ownership and quality work that decide whether any of the rest can work.
Forward-Deployed Engineering
Engineering that starts at the customer's workflow and ends at a measured outcome — coding is only part of the job.
Four nested loops: a deployment cycles until validation clears, delivery repeats it across real workflows, and field learning feeds the roadmap — so the next deployment starts further ahead.
Operate5
making it trustworthy, measurable, affordable and production-ready
✅ Evals & Testing
How you know the system is actually good - and staying good.
No evals = shipping on vibes. A regression set lets you change a prompt or model with confidence.
🛡️ Guardrails & Responsible AI
Keeping outputs safe, accurate, private, and compliant.
Guardrails wrap the model on both sides. Design for the failure modes, not just the happy path.
⚙️ Deployment, Inference & LLMOps
Running models reliably and affordably at scale - MLOps for LLMs.
LLMOps keeps the system alive and affordable after launch, when real traffic and cost curves hit.
Show all 26
What Happens After You Hit Enter
The forty-two stages between your API call and the first character on screen — one fixed sequence, not a menu, and any one of them can be why the answer is wrong or slow.
Forty-two stages sit between your call and the first character on screen, in one fixed sequence. Any single one of them can be why an answer is wrong or slow.
Show all 42
🔒 Security
Prompt injection, data exfiltration, and the tool permissions an agent should never hold.
Reference1
consult, don't read — look it up and run the numbers
🗺️ Frontier Models
Every model competing at the frontier, each with its own page: context window, price, what it is for.