Redis Stack bolts HNSW and FLAT vector search onto an in-memory database, trading RAM cost for millisecond-scale, filterable retrieval.
ConceptWhat it is
Redis is best known as an in-memory key-value store used for caching, sessions, and queues. Redis Stack and its underlying RediSearch module extend it with a secondary index engine that includes a vector field type, turning the same server into a vector database. Embeddings are stored as fields on ordinary Redis hashes or JSON documents, and a search index makes them queryable by nearest-neighbor similarity.
It exists so teams already running Redis can add semantic retrieval without standing up a separate system. Because the index lives in RAM, queries typically return in single-digit milliseconds, and vector similarity can be combined in one call with metadata filters and full-text search (hybrid search), keeping the retrieval layer on infrastructure the team already operates and monitors.
How it worksThe mechanics
You define an index with FT.CREATE, declaring a VECTOR field plus its algorithm — FLAT for an exact brute-force scan or HNSW for an approximate graph search — along with the distance metric (cosine, L2, or inner product) and dimensionality. Your pipeline embeds each chunk and writes it, with its metadata, to a Redis hash or JSON key; the index updates as keys are written. At query time you embed the user's text and issue FT.SEARCH or FT.AGGREGATE with a KNN clause, optionally pre-filtering on tags, numeric ranges, or text so only matching documents are scored. Redis returns the top-k nearest keys with their fields, which the app hands to the LLM as grounding context.
At a glanceSee it
When to use itWhere it fits
- You already run Redis for caching or queues and want semantic search without adding a new datastore to operate.
- Latency is critical — real-time recommendations, agent memory, or chat where retrieval must stay in the low-millisecond range.
- Your working set of vectors fits comfortably in RAM, on a single node or across a cluster's memory budget.
- You need hybrid queries that mix vector similarity with tag, numeric, or full-text filters in a single round trip.
When NOT to use itLimits & anti-patterns
- You hold tens or hundreds of millions of vectors where keeping everything in memory is prohibitively expensive.
- Cost-per-vector matters more than latency, and a disk-backed store would be dramatically cheaper.
- You need heavy analytical filtering, large-scale disk offload, or advanced index types that purpose-built engines handle better.
- Your team wants a fully managed, batteries-included vector service and has no Redis operational experience.
Trade-offsAdvantages & costs
Advantages
- Single-digit-millisecond query latency because the index is entirely in memory.
- Consolidates cache, session store, and vector search on one familiar system, shrinking operational surface.
- Supports both HNSW and FLAT, hybrid search, and pre-filtering, with mature client libraries in most languages.
- Horizontal scale and high availability come through Redis Cluster and replication that ops teams already know.
Trade-offs & costs
- RAM-bound: every vector and its index structure consumes memory, making large corpora costly to hold.
- Recall and speed depend on HNSW tuning such as M and EF, and rebuilding or resizing indexes can be operationally involved.
- Redis licensing shifted to source-available terms in 2024, prompting the Valkey fork — check which build and license you actually deploy.
- Less specialized than dedicated vector databases for very large scale, disk tiering, or advanced index families.
ExampleIn the real world
A support team building a documentation assistant already uses Redis to cache API responses. They add a Redis Stack index over roughly 200,000 help-article chunks, each stored as a JSON document with an embedding, a product tag, and a language field. When a user asks a question in the chat widget, the app embeds the query and runs one FT.SEARCH that pre-filters on the user's product and language, then returns the top eight nearest chunks in a few milliseconds. Those chunks are passed to the LLM to draft a grounded answer, and because retrieval shares the Redis the app already scales, no new database had to be provisioned.
ToolsHow to implement it
- redis-pyand node-redis — official clients that expose the search and vector query commands.
- RedisVLa Python library purpose-built for defining schemas and running vector queries on Redis.
- LangChainand LlamaIndex — both ship Redis vector store integrations for RAG pipelines.
- Valkeythe community OSS fork of Redis, relevant when source-available licensing is a concern.
Cost & effortWhat it takes
The dominant cost is RAM. Because vectors and their HNSW graphs live in memory, spend scales with corpus size times dimensionality, and a large index can demand expensive high-memory instances or a multi-node cluster. Managed offerings such as Redis Cloud or a cloud provider's Redis service trade some of that cost for less operational effort. For a team already running Redis the incremental engineering effort is modest — define an index and write embeddings — but capacity planning for memory and HNSW tuning for recall are the ongoing work.