Home › Embeddings & Vector Search › stella_en_1.5B_v5
🧭 · Models

stella_en_1.5B_v5

stella_en_1.5B_v5: a compact open English text embedding model with selectable output dimensions for retrieval

In one line

A free, self-hostable 1.5-billion-parameter English text embedding model that produces strong retrieval vectors, with a selectable output dimension from a compact 1024 up to 8192.

ConceptWhat it is

stella_en_1.5B_v5 is an open-weights text embedding model from the community: it maps a piece of English text into a single dense vector so that semantically similar texts land close together, which is the core primitive behind semantic search, RAG retrieval, clustering, and deduplication. At roughly 1.5 billion parameters it is small enough to run on a single modest GPU, yet it scores near the top of open English models on the MTEB retrieval leaderboard.

It exists to give teams near-hosted-API quality without the hosted API: the weights are free to download and self-host, so data never leaves your environment and there is no per-call fee. Its distinguishing feature is a range of selectable output dimensions: trained with a Matryoshka Representation Learning style objective, it exposes dimension-specific projection heads so one model serves many storage budgets, from a compact 1024-dim vector up to a maximal 8192-dim one. That lets you dial storage and latency against recall by picking a dimension rather than switching to a different model.

How it worksThe mechanics

Input text is tokenized (with a recommended context of about 512 tokens) and, for a query, a short task instruction prompt is prepended while passages are embedded plain — an asymmetric encoding scheme. The tokens pass through the transformer encoder, and the token hidden states are mean-pooled into a single base vector that a dimension-specific projection head maps to your chosen output size, so the result can be a compact 1024-dim vector or a maximal 8192-dim one depending on the head you select. Vectors are L2-normalized and compared by cosine similarity, so retrieval is an approximate-nearest-neighbor lookup in a vector store.

At a glanceSee it

stella_en_1.5B_v5 diagram

When to use itWhere it fits

  • English-only or English-dominant retrieval and RAG where you want top-tier open-model quality without a hosted API.
  • You must self-host for data privacy, cost control, or air-gapped and offline use.
  • You want to tune the storage, latency, and recall trade-off by choosing an output dimension instead of switching models.
  • Semantic search, near-duplicate detection, or clustering over English text chunks around the model's recommended 512-token input length.

When NOT to use itLimits & anti-patterns

  • Multilingual or cross-lingual corpora, where an explicitly multilingual embedding model will retrieve far better.
  • You have no GPU appetite and want zero ops, in which case a hosted embedding API is simpler.
  • You need image, audio, or other multimodal embeddings, since this model is text-only.
  • Very tight memory or latency budgets where a much smaller embedding model is good enough.

Trade-offsAdvantages & costs

Advantages
  • Strong English retrieval quality, competitive with larger and hosted models on MTEB, at only about 1.5B parameters.
  • Selectable output dimensions let one model span 1024 to 8192 dims, so you trade storage and speed for recall without switching models.
  • Free open weights and full self-hosting mean no per-call fees and your data stays in your environment.
  • Instruction-tuned with distinct query and passage prompts, so it is optimized for asymmetric retrieval out of the box.
Trade-offs & costs
  • English-centric: quality drops on other languages and on cross-lingual matching.
  • Short input window: the recommended context is about 512 tokens, so long documents must be chunked before embedding.
  • Self-hosting overhead is yours to own, including GPU capacity, serving, batching, and uptime; and high output dimensions such as 8192 multiply index size and search cost unless you pick a smaller one.
  • Peak quality depends on using the correct query-versus-passage instruction prompts, which is easy to get wrong.

ExampleIn the real world

A platform team wants semantic search across roughly half a million English documentation chunks drawn from internal wikis and runbooks. They download the open weights and serve the model on a single mid-range GPU, embedding each chunk at 1024 dimensions to keep the index compact, and store the vectors in pgvector alongside their existing Postgres. At query time they prepend the retrieval instruction prompt to the user's question, embed it, run an approximate-nearest-neighbor cosine search for the top 50 candidates, then rerank. Because the model exposes several output dimensions, when a few high-value collections need better recall the team re-embeds just those at 8192 dimensions without swapping models, spending extra storage only where it measurably pays off.

ToolsHow to implement it

  • Sentence Transformers and the Hugging Face transformers library to load and run the model.
  • FAISS, Qdrant, or pgvector as the vector store for approximate-nearest-neighbor search.
  • Text Embeddings Inference (TEI) or vLLM for efficient batched serving of the embeddings.
  • The MTEB benchmark to compare it head-to-head against alternative embedding models on your task types.

Cost & effortWhat it takes

There is no license or per-call fee: the weights are free, so the real cost is the GPU you run inference on plus the storage and search cost of the vectors. At about 1.5B parameters the model fits on a single modest GPU, though large batches raise memory needs and high output dimensions such as 8192 multiply index size, so choosing a smaller output dimension like 1024 is the main lever for keeping storage and query latency cheap. Effort concentrates in standing up and maintaining the serving stack, batching for throughput, chunking inputs to the roughly 512-token window, and applying the correct query and passage prompts; once that pipeline exists, marginal cost per embedding is low and fully under your control.

A living map of modern AI — kept current every morning