LlamaIndex Ships legal-kb to Replace Single-Shot RAG With Agentic Filesystem Retrieval
LlamaIndex ships legal-kb, an open-source reference app showing how to give AI agents filesystem-style tools to autonomously navigate large document knowledge bases
- LlamaIndex released legal-kb, an open-source reference app demonstrating agentic retrieval over large document knowledge bases.
- Introduces the "Retrieval Harness" pattern: agents get four filesystem-style tools (retrieve, findFiles, readFile, grepFile) instead of a single embedding lookup.
- Built on LlamaCloud Index v2 (LlamaParse Platform), with automatic document parsing, indexing, and background sync per project.
- Supports document versioning: re-uploading a file creates v1, v2, v3 side by side, queryable by version metadata filter.
- Bring-your-own-keys model: supports OpenAI and Anthropic per turn; LlamaCloud offers 10k free credits to start.
- Live demo at legal-kb.dev; self-hosting requires PostgreSQL, WorkOS, and a LlamaCloud account.
Most RAG (retrieval-augmented generation) systems work the same way: a user asks a question, the system runs one embedding search, pulls the top-k chunks, and hands them to the model. That works fine for simple Q&A. It breaks down the moment your knowledge base is large, evolving, and the questions require multi-step reasoning across many documents. LlamaIndex just shipped a blueprint for what comes next.
The Retrieval Harness
LlamaIndex has published legal-kb, a public reference application on GitHub described as a knowledge base for legal documents, powered by LlamaIndex Index v2 (the LlamaParse Platform). The real payload here is not the legal domain specifically. The project demonstrates a pattern the team calls a Retrieval Harness for agentic retrieval. The approach differs from single-shot retrieval: instead of one embedding search per query, an agent is given filesystem-style tools and can then crawl a large, evolving knowledge base to solve a task.
The harness provides a persistent data pipeline that can connect to a data source, index and update a large knowledge base, and expose a broad set of tools akin to filesystem operations. You can plug this into any of your agents to let them autonomously crawl an arbitrary knowledge base to solve a task with any complexity.
Four Tools, One Agent
The tools mirror operations engineers already know: semantic and keyword search, regex grep, file search, and read. Concretely, the agent gets:
- retrieve: hybrid semantic + keyword search over the index, with optional reranking and metadata filters
- findFiles: locate documents by name or glob pattern
- readFile: pull the full content of a specific document
- grepFile: run a regex pattern across the corpus
The order matters. The system prompt enforces a sequence: the agent must call findFiles first to establish the document inventory, then narrow with retrieve, and confirm exact wording with readFile or grepFile before citing. This structured approach is what produces grounded, citation-backed answers rather than hallucinated summaries.
How It Is Built
legal-kb is a working TanStack Start web app, not a library. You sign in, create a project, upload files, and chat with an agent. Each project is mirrored as a managed LlamaCloud Index v2, and uploaded files are parsed and indexed automatically in the background.
The full stack under the hood:
- Parsing and indexing: LlamaCloud Index v2 (LlamaParse), one index per project
- Agent runtime: the
ToolLoopAgentfrom Vercel AI SDK 6. You pick OpenAI or Anthropic per turn and bring your own keys. Reasoning is streamed: Claude models use extended thinking; OpenAI reasoning models use medium reasoning effort. - Auth: WorkOS AuthKit
- Storage: PostgreSQL via Prisma, with AES-256-GCM encryption for user API keys at rest
- Frontend: React 19, TanStack Start, Tailwind v4, Vite, Bun
The upload pipeline is straightforward. Bytes are pushed to the project's LlamaCloud source directory, a File and ProjectFile row are written to PostgreSQL via Prisma, and an index sync is triggered but not awaited; the UI polls status until ready.
Version Control for Your Knowledge Base
One underrated feature is document versioning. Versioning is scoped to the (project, filename) pair. Re-uploading nda.pdf to the same project produces v1, v2, v3 side by side. The retrieval layer filters on the version metadata field, giving version control over the knowledge base itself. For legal and compliance use cases where documents are amended frequently, this is significant: you can query against a specific version of a contract without losing history.
Where It Shines, Where It Does Not
The agentic retrieval pattern is genuinely better for complex, multi-hop questions across a focused document set. LlamaIndex's own research found that filesystem agents are more accurate (2 points higher on correctness) but RAG is faster (3.81 seconds quicker), and at 100-1000 documents RAG wins on scale. The honest framing: filesystem agents excel with smaller, focused document sets where accuracy trumps speed.
There are also real infrastructure requirements. Running legal-kb yourself means standing up a PostgreSQL database, a WorkOS account, a LlamaCloud account, and supplying your own OpenAI or Anthropic keys. Advanced parsing features require a paid LlamaCloud subscription. The live demo at legal-kb.dev is the easiest way to evaluate it without the setup overhead.
The Broader Shift
This release is part of a deliberate repositioning by LlamaIndex. LlamaCloud is transitioning to LlamaParse across docs and the website. Over the past year, the enterprise platform has evolved into a truly agentic document processing platform, with LlamaParse at its core. Renaming reflects that evolution and sharpens focus on powerful document parsing and agentic workflows for document-centric automation.
The deeper argument being made is architectural. Traditional RAG assumes retrieval is a single, stateless lookup. The Retrieval Harness treats the knowledge base more like a filesystem that an agent can explore iteratively. As one community member put it: hybrid search across the whole index with read, grep, and find is basically giving an agent the same tools a developer would use on a local filesystem. That framing is useful: if you have ever used grep and find to navigate an unfamiliar codebase, you already understand the intuition.
Getting Started
To self-host, clone the repo and follow these steps:
- Run
bun installand copy.env.exampleto.env.local - Set your
DATABASE_URL,ENCRYPTION_KEY, and WorkOS credentials - Run
bun run db:pushto apply the Prisma schema - Run
bun run devto start athttp://localhost:5173 - Add your LlamaCloud, OpenAI, and/or Anthropic keys via the profile page
The production build outputs a self-contained Node server via Nitro, compatible with Render, Fly.io, Vercel, Netlify, Cloudflare, and AWS Lambda. LlamaCloud offers 10k free credits to get started with the indexing backend.
legal-kb is a reference implementation, not a finished product. The license is still TBD and the repo has just 18 commits. But as a concrete, runnable demonstration of how agentic retrieval differs from classic RAG, it is one of the clearest examples the community has seen. If you are building document-heavy applications in legal, fintech, or compliance, the Retrieval Harness pattern is worth understanding now.