Home › Embeddings & Vector Search › Redis (Redis Stack)
🧭 · Models

Redis (Redis Stack)

Redis Stack adds in-memory vector search to the store many teams already run for caching.

In one line

Redis Stack bolts HNSW and FLAT vector search onto an in-memory database, trading RAM cost for millisecond-scale, filterable retrieval.

ConceptWhat it is

Redis is best known as an in-memory key-value store used for caching, sessions, and queues. Redis Stack and its underlying RediSearch module extend it with a secondary index engine that includes a vector field type, turning the same server into a vector database. Embeddings are stored as fields on ordinary Redis hashes or JSON documents, and a search index makes them queryable by nearest-neighbor similarity.

It exists so teams already running Redis can add semantic retrieval without standing up a separate system. Because the index lives in RAM, queries typically return in single-digit milliseconds, and vector similarity can be combined in one call with metadata filters and full-text search (hybrid search), keeping the retrieval layer on infrastructure the team already operates and monitors.

How it worksThe mechanics

You define an index with FT.CREATE, declaring a VECTOR field plus its algorithm — FLAT for an exact brute-force scan or HNSW for an approximate graph search — along with the distance metric (cosine, L2, or inner product) and dimensionality. Your pipeline embeds each chunk and writes it, with its metadata, to a Redis hash or JSON key; the index updates as keys are written. At query time you embed the user's text and issue FT.SEARCH or FT.AGGREGATE with a KNN clause, optionally pre-filtering on tags, numeric ranges, or text so only matching documents are scored. Redis returns the top-k nearest keys with their fields, which the app hands to the LLM as grounding context.

At a glanceSee it

Redis (Redis Stack) diagram

When to use itWhere it fits

  • You already run Redis for caching or queues and want semantic search without adding a new datastore to operate.
  • Latency is critical — real-time recommendations, agent memory, or chat where retrieval must stay in the low-millisecond range.
  • Your working set of vectors fits comfortably in RAM, on a single node or across a cluster's memory budget.
  • You need hybrid queries that mix vector similarity with tag, numeric, or full-text filters in a single round trip.

When NOT to use itLimits & anti-patterns

  • You hold tens or hundreds of millions of vectors where keeping everything in memory is prohibitively expensive.
  • Cost-per-vector matters more than latency, and a disk-backed store would be dramatically cheaper.
  • You need heavy analytical filtering, large-scale disk offload, or advanced index types that purpose-built engines handle better.
  • Your team wants a fully managed, batteries-included vector service and has no Redis operational experience.

Trade-offsAdvantages & costs

Advantages
  • Single-digit-millisecond query latency because the index is entirely in memory.
  • Consolidates cache, session store, and vector search on one familiar system, shrinking operational surface.
  • Supports both HNSW and FLAT, hybrid search, and pre-filtering, with mature client libraries in most languages.
  • Horizontal scale and high availability come through Redis Cluster and replication that ops teams already know.
Trade-offs & costs
  • RAM-bound: every vector and its index structure consumes memory, making large corpora costly to hold.
  • Recall and speed depend on HNSW tuning such as M and EF, and rebuilding or resizing indexes can be operationally involved.
  • Redis licensing shifted to source-available terms in 2024, prompting the Valkey fork — check which build and license you actually deploy.
  • Less specialized than dedicated vector databases for very large scale, disk tiering, or advanced index families.

ExampleIn the real world

A support team building a documentation assistant already uses Redis to cache API responses. They add a Redis Stack index over roughly 200,000 help-article chunks, each stored as a JSON document with an embedding, a product tag, and a language field. When a user asks a question in the chat widget, the app embeds the query and runs one FT.SEARCH that pre-filters on the user's product and language, then returns the top eight nearest chunks in a few milliseconds. Those chunks are passed to the LLM to draft a grounded answer, and because retrieval shares the Redis the app already scales, no new database had to be provisioned.

ToolsHow to implement it

  • redis-pyand node-redis — official clients that expose the search and vector query commands.
  • RedisVLa Python library purpose-built for defining schemas and running vector queries on Redis.
  • LangChainand LlamaIndex — both ship Redis vector store integrations for RAG pipelines.
  • Valkeythe community OSS fork of Redis, relevant when source-available licensing is a concern.

Cost & effortWhat it takes

The dominant cost is RAM. Because vectors and their HNSW graphs live in memory, spend scales with corpus size times dimensionality, and a large index can demand expensive high-memory instances or a multi-node cluster. Managed offerings such as Redis Cloud or a cloud provider's Redis service trade some of that cost for less operational effort. For a team already running Redis the incremental engineering effort is modest — define an index and write embeddings — but capacity planning for memory and HNSW tuning for recall are the ongoing work.

A living map of modern AI — kept current every morning