text-embedding-3 lets you shorten the vector at request time, trading a little accuracy for a large cut in storage and search cost.
ConceptWhat it is
text-embedding-3 is OpenAI's embedding family, in a small and a large variant. Embeddings are the least glamorous part of a retrieval system and among the most consequential: the model decides what counts as similar, and every retrieval quality number inherits that decision.
Its distinguishing feature is that the vector can be truncated. The models are trained so that the first N dimensions remain useful on their own, so a caller can ask for a shorter vector and keep most of the quality — a knob that is unusual and directly changes infrastructure cost.
How it worksThe mechanics
Text is sent to the endpoint and a float vector comes back, at the model's native dimensionality unless a shorter one is requested. Because the shortening is trained in rather than applied afterwards, a truncated vector behaves like a real embedding rather than a damaged one, and the quality loss is gradual rather than sudden.
The operational rules are the ones that catch teams out. The same model and the same dimension count must be used for the corpus and for the query, since vectors from different models are not comparable at all. Changing either means re-embedding the whole corpus, which makes the choice a commitment rather than a setting.
At a glanceSee it
The dimension knob is trained in, so a shortened vector degrades gradually. Corpus and query must use the same model and the same dimension, or the numbers are not comparable.
When to use itWhere it fits
- As a strong, low-effort default when there is no reason to self-host.
- When index size or search latency is the binding constraint and the dimension knob can relieve it.
- For multilingual corpora, where the family performs respectably without a separate model per language.
- When you want a known quantity to measure a self-hosted alternative against.
When NOT to use itLimits & anti-patterns
- Where data cannot leave your network, which rules out any hosted embedding endpoint.
- At very large corpus scale, where per-token embedding cost may favour a self-hosted open model.
- For a narrow technical domain, where a fine-tuned or domain-specific model can beat a general one.
- When you are unwilling to re-embed later, since the choice is effectively locked in by the index.
Trade-offsAdvantages & costs
Advantages
- Strong general-purpose quality with essentially no setup.
- The dimension knob trades quality for cost on a curve you can measure rather than guess.
- Reliable, well-documented and supported by every vector store and framework.
- Cheap enough that embedding cost is rarely the dominant line in a retrieval budget.
Trade-offs & costs
- Data leaves your network, which is a hard blocker in some settings.
- Model changes are the provider's decision, and a silently updated model invalidates an index.
- A general model can be beaten by a domain-specific one on narrow corpora.
- Re-embedding to switch is a full corpus rebuild, so the lock-in is real even though the API is simple.
ExampleIn the real world
A retrieval system over a few million chunks finds vector search latency is the bottleneck and the index no longer fits comfortably in memory. Rather than sharding, the team re-embeds at a reduced dimension. The index shrinks substantially, search speeds up, and the measured recall drop is small enough to accept — a decision that was only available because the truncation is trained in.
ToolsHow to implement it
- The OpenAI embeddings endpointbatched requests, since per-call overhead dominates for short texts.
- A dimension sweep on your own eval setthe only way to know where the quality curve bends for your corpus.
- The model id recorded with the indexso a mismatch between corpus and query vectors is detectable rather than mysterious.
- An open alternative benchmarked alongsideto know what the hosted convenience is actually costing.
Cost & effortWhat it takes
Priced per token and low enough that embedding a substantial corpus is usually a modest one-off. The recurring cost is query embedding, one short call per question. The cost that dominates is storage and search over the vectors, which is exactly what the dimension knob addresses.