LlamaIndex's LiteParse v2.1 Converts PDFs to Markdown 45x Faster Without AI
LiteParse v2.1 adds LLM-free markdown output that beats every open-source competitor across three standard benchmarks at 3ms per page
- LiteParse v2.1 adds LLM-free, model-free markdown output to the open-source PDF parser from LlamaIndex.
- It tops all three standard benchmarks (ParseBench, opendataloader-bench, olmOCR-bench) against every model-free competitor.
- Speed is 3.16 ms/page -- roughly 45x faster than pymupdf4llm and 58x faster than markitdown.
- Runs in Python, Node, Rust, and natively in the browser via WASM under an Apache-2.0 license.
- Weak on scanned docs, math-heavy PDFs, and charts -- those still require a vision model like LlamaParse.
- Available now: GitHub | Docs | Browser demo
LiteParse v2.1 just landed with the one feature its users kept asking for: markdown output. The twist is how it delivers it -- no LLM, no cloud call, no OCR model. Just a fast, heuristic pipeline that parses PDFs directly into structured markdown, and does it faster than every open-source alternative tested.
The gap it fills
A few weeks ago, LiteParse 2.0 launched as the fastest tool for converting PDFs to text. But two questions kept coming up: where are the benchmarks, and does it output markdown? Both are answered in v2.1. By building this markdown pipeline, the team was also able to measure and improve overall extraction quality -- so the markdown mode doubled as a forcing function for better parsing across the board.
The broader context matters here. Most PDF-to-markdown tools fall into two camps: heavy LLM-based parsers (like LlamaParse itself, Docling, or cloud OCR services) that are accurate but slow and expensive, and lightweight rule-based tools that are fast but historically couldn't produce clean markdown. LiteParse v2.1 positions itself as the fastest open-source, model-free, PDF-to-markdown pipeline -- a gap that was genuinely underserved.
How the pipeline works
PDFs carry a ton of data: font family, font size, text location, and more. All of these are treated as input signals to classify text into specific markdown elements like paragraphs, tables, lists, and headers. Think of it as a hand-crafted feature extractor: instead of training a neural net to recognize document structure, the team encodes that logic as deterministic rules.
LiteParse uses a custom PDFium fork to capture as much signal as possible, and then combines that with signals from its existing grid-projection algorithm -- a technique that maps text blocks onto a spatial grid to infer layout. The result is a purely heuristic approach with no model weights to load and no inference latency.
Benchmark numbers
LiteParse achieved top overall scores on all three benchmarks when measured against model-free approaches: opendataloader-bench at 0.875, olmOCR-bench at 0.391, and ParseBench at 0.3279. Here's how it stacks up against the main competitors across the three benchmarks:
| Tool | ParseBench (Overall) | opendataloader-bench (Overall) | olmOCR-bench (Overall) | Speed (ms/page) |
|---|---|---|---|---|
| LiteParse | 0.328 | 0.871 | 39.2% | 3.16 |
| pymupdf4llm | 0.310 | 0.732 | 32.9% | 141.5 |
| opendataloader | 0.294 | 0.831 | 32.7% | 66.3 |
| pdf-inspector | 0.266 | 0.792 | 30.5% | 3.83 |
| markitdown | 0.186 | 0.589 | 28.7% | 182.5 |
The speed gap is striking. pymupdf4llm clocks in at 141.5 ms/page and markitdown at 182.5 ms/page, while LiteParse processes a page in 3.16 ms -- roughly 45x faster than pymupdf4llm. That difference compounds quickly on large document sets.
Where it struggles
The team is upfront about the ceiling. This approach prioritizes speed, but also has to accept an upper bound on accuracy -- it isn't going to do better than LlamaParse. Concretely:
- For complex documents -- dense tables, multi-column layouts, charts, handwritten text, or scanned PDFs -- you'll get significantly better results with LlamaParse.
- Math-heavy documents score 0% on olmOCR-bench's arxiv_math and old_scans_math categories -- expected, since those require actual OCR or a vision model.
- Chart understanding is essentially zero for all model-free tools; scoring charts requires extracting structured data from images, which needs a vision model.
- The three benchmarks don't always agree on what "good" markdown looks like -- tuning for one benchmark often regresses another. LiteParse deliberately targets balanced performance rather than gaming any single leaderboard.
Licensing and where it runs
This is where LiteParse has a real structural advantage over its closest competitor. LiteParse is permissively licensed under Apache-2.0 and runs as a single engine across four ecosystems, including natively in the browser via WASM. The Python-only tools can't go where a browser or a Node service needs them, and pymupdf4llm inherits PyMuPDF's AGPL-3.0 copyleft, which is a non-starter for many commercial codebases without a paid license.
That runtime portability is a genuine differentiator. You can run the exact same parsing engine in a Python backend, a Node.js service, a Rust binary, or directly in the browser -- no server round-trip needed for the WASM build.
Getting started
Install it in whichever runtime you're working in:
# Python
pip install liteparse
# Node
npm i @llamaindex/liteparse
# Rust
cargo install liteparse
# Browser (WASM)
npm i @llamaindex/liteparse-wasmParsing a PDF to markdown takes two lines in Python:
from liteparse import LiteParse
lp = LiteParse(output_format="markdown")
result = lp.parse("doc.pdf")
print(result.text)Or from the CLI:
lit parse doc.pdf --format markdownThere's also a live browser demo running entirely via WASM -- no backend involved. The source code and full documentation are both public.
The practical sweet spot
LiteParse v2.1 is the right tool when you need fast, local, cost-free PDF-to-markdown conversion and your documents are reasonably well-structured. The ideal use cases:
- RAG pipelines -- structured markdown with headings, tables, lists, images, and links is great for feeding LLMs and RAG pipelines.
- Agent workflows -- parse documents locally without a network call, keeping latency and cost at zero.
- Browser-based tools -- the WASM build lets you parse PDFs client-side, which opens up privacy-preserving document apps that never send files to a server.
- Commercial products -- the Apache-2.0 license means no legal friction, unlike the AGPL-3.0 on pymupdf4llm.
If your documents are scanned, math-heavy, or packed with complex visual layouts, the LLM-based LlamaParse is still the right call. But for the large middle ground of standard business and technical PDFs, v2.1 makes a strong case for going model-free.