Home › Reference
Reference

Open-source resources

The repositories people actually build LLM products on, and the situation each one is for.

How to use thisA directory, not a library

Everything on this page belongs to somebody else. We link to their repository and say what it is for; we do not reproduce their code, their documentation or their diagrams. Their version is current and maintained, and ours would be neither.

What this page adds

A link list is easy and mostly useless — the hard part is knowing which of forty repositories to open. So every row carries the situation you would reach for it in, and projects that solve the same problem sit next to each other so the choice is visible. That judgement is the only thing here that is ours.

Two things you will not find. There are no star counts, because they move weekly and a number written into prose is a future lie. And there is no ranking, because these solve different problems and a league table would send you to the wrong one.

Agent frameworksWhen one model call is not enough

They differ less in what they can do than in what they make explicit. Pick by which part of the problem you want the framework to hold.

ProjectWhat it isReach for it when
LangGraph
LangChain
An agent runtime built around an explicit state graph rather than a chain.The run has branches, retries and human approval steps you need to see and resume.
CrewAI
CrewAI Inc
Multi-agent orchestration where each agent has a role and they collaborate on a task.The work genuinely divides into roles and you want that structure to be the code.
smolagents
Hugging Face
A deliberately small agent loop — the whole thing is readable in an afternoon.You want one agent, minimal abstraction, and to understand every line in the loop.
Semantic Kernel
Microsoft
An agent and plugin framework with first-class .NET and enterprise integration.You are in a .NET shop and the surrounding platform matters more than the loop.
Mastra
Mastra
A TypeScript-first agent framework from the team behind Gatsby.Your stack is TypeScript end to end and you do not want a Python service in the middle.
VoltAgent
VoltAgent
TypeScript agent framework with a large worked-example set and its own observability.You learn fastest from runnable examples and want tracing built in rather than bolted on.
Dify
LangGenius
A low-code platform: visual agent building with RAG, tool calling and deployment.Non-engineers need to change the flow without a pull request.

Ordered from most explicit about control flow to most abstracted away from it.

RetrievalGetting your own data in front of the model

The frameworks and the stores are separate decisions. A framework that is right for messy scanned documents is not the one that is right for a clean corpus you control.

ProjectWhat it isReach for it when
Haystack
deepset
Modular, production-oriented pipelines with evaluation tooling included.You are moving a demo to production and want to swap components without a rewrite.
LlamaIndex
LlamaIndex
Data-centric framework for connecting private data to a model.The hard part is the connectors and the indexing strategy, not the generation.
RAGFlow
InfiniFlow
A RAG engine built for messy real documents, with agentic retrieval and citation grounding.The corpus is scanned PDFs and inconsistent enterprise documents, not clean markdown.
txtai
NeuML
An embeddings database that also carries pipelines and orchestration.You want vector store, text processing and workflow in one dependency.
Chroma
Chroma
An embedded vector database that runs in-process with no server to operate.You are prototyping, or the corpus is small enough that a service is overhead.
Qdrant
Qdrant
A vector database in Rust with filtering, quantisation and hybrid search.You have outgrown embedded storage and need filtered search at real scale.

The last two are vector stores rather than frameworks; they sit underneath the others rather than competing with them.

EvaluationKnowing whether a change helped

This is the category teams skip and then wish they had not, because without it every model or prompt change is decided on impression.

ProjectWhat it isReach for it when
Ragas
Exploding Gradients
Metrics for retrieval-augmented systems — faithfulness, relevance, answer quality.You need to know whether a RAG change helped, and by how much.
promptfoo
promptfoo
Test and compare prompts, models and agent behaviour, designed to sit in CI.You want a model or prompt change to fail a build the way a code change would.
DeepEval
Confident AI
An evaluation framework with many metrics, LLM-as-judge, and guardrail scanners.You want a broad metric set and jailbreak or PII checks in the same harness.
Inspect
UK AI Safety Institute
An evaluation framework built for rigorous, reproducible model assessment.The evaluation itself has to stand up to scrutiny, not just guide your iteration.

OperationsTracing, routing and the bill

ProjectWhat it isReach for it when
Langfuse
Langfuse
Tracing, metrics, prompt management and evals for LLM applications.You cannot answer 'why did this call cost that much' or 'what did it actually send'.
Opik
Comet
Open-source tracing and evaluation for LLM and agent applications.You want tracing and scoring in one tool rather than stitching two together.
OpenLLMetry
Traceloop
OpenTelemetry instrumentation for LLM calls.You already run OTel and want model calls in the same traces as everything else.
LiteLLM
BerriAI
One interface across many providers, with routing, fallbacks, budgets and logging.You want the routing seam from the LLMs page without writing it yourself.

LiteLLM is the routing seam described on the LLMs page, already built.

ServingRunning open weights yourself

Only relevant once you have decided to self-host. The choice is mostly about the hardware you are targeting and the shape of your traffic.

ProjectWhat it isReach for it when
vLLM
vLLM project
A high-throughput inference server built around paged attention.You are self-hosting and throughput per GPU is the number that decides the bill.
Ollama
Ollama
Run open-weight models locally with one command.You want a model on your laptop for development, or data cannot leave the machine.
llama.cpp
ggml
Efficient CPU and consumer-GPU inference, the engine under much of the local ecosystem.The target is commodity hardware, an edge device, or no GPU at all.
SGLang
SGLang
A serving runtime focused on structured generation and fast multi-call programs.Your workload is many short structured calls rather than a few long ones.

Tool accessThe Model Context Protocol

ProjectWhat it isReach for it when
Model Context Protocol servers
Anthropic and contributors
Reference MCP servers — the standard way to give a model tools and data.You want tool access that is not bespoke to one framework.
MCP specification
Model Context Protocol
The protocol itself, with the schema and the reasoning behind it.You are writing your own server and need the contract, not a wrapper.

Going widerThe lists other people maintain

When this page does not have what you need, these do. Each is maintained by someone who watches its corner more closely than we do.

ProjectWhat it isReach for it when
Awesome LLM Apps
Shubham Saboo
A large collection of runnable app and agent examples across RAG, MCP and voice.You learn by running something end to end before reading the theory.
Awesome LLMOps
TensorChord
A curated index of the operations side — serving, monitoring, gateways, registries.You know the model works and now have to run it.
Awesome LLM Inference Engine
sihyeong
An index of inference engines and optimisation techniques, alongside a survey.You are choosing between serving runtimes and want the landscape, not a blog post.
Awesome Harness Engineering
ai-boost
The scaffolding around an agent — evals, memory, MCP, permissions, observability.The model is fine and the harness is what keeps breaking.
Awesome AI Agents 2026
ARUNAGIRINATHAN-K
A broad agent index with comparison guides.You want breadth first and will narrow later.

ReferencesSources

Every project links to its own repository above. These are the indexes and reviews consulted while assembling this page.

  1. Awesome LLMOps — curated LLMOps tooling (TensorChord)
  2. Awesome Open Source LLMOps
  3. Awesome LLM Inference Engine — engines and optimisation techniques
  4. Awesome Harness Engineering — agent harnesses, evals, memory, MCP, observability
  5. Awesome AI Agents 2026 — agent frameworks and comparisons
  6. Awesome LLM Apps — runnable app and agent examples (Shubham Saboo)
  7. VoltAgent — TypeScript agent framework and examples
  8. Best open-source agent frameworks, 2026 review (Firecrawl)
  9. Best open-source RAG frameworks, 2026 review (Firecrawl)
  10. Open-source RAG frameworks compared (Olostep)
  11. Popular agentic open-source tools, 2026 edition (You.com)

What changedWhat changed here

Nothing in the daily brief has touched this page since 2026-09-25. The sweep runs every morning and checks every page on this site; when it finds something for this one, it lands here.

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning