LlamaIndex Rewrites LiteParse v2.0 in Rust, Making it 100x Faster
LlamaIndex rewrote LiteParse in Rust, delivering up to 100x faster document parsing that now runs natively in Python, Node, Rust, and the browser via WASM
- Full Rust rewrite: LiteParse v2.0 is rebuilt from scratch in Rust, delivering up to 100x faster parsing on small documents.
- Runs everywhere: Native packages now available for Python (
pip install liteparse), Node/TS, Rust, and browser/edge via WASM. - Browser support: A new
@llamaindex/liteparse-wasmpackage enables fully local, in-browser PDF parsing with zero server dependencies. - Benchmark leader: Parses a 457-page, 100MB PDF in 0.777s, outperforming pymupdf, pypdf, markitdown, and pdftotext.
- Agent-native: Installable as a skill directly in Claude Code, Codex, and OpenCode for agentic document workflows.
- Open source & free: Apache 2.0 license, no API keys or cloud calls required. GitHub repo | Blog post
LiteParse v2.0 is a complete ground-up rewrite of LlamaIndex's open-source document parser, and the headline change is the language it's written in: Rust. The result is a parser that's up to 100x faster on small documents, ships as a native package across four ecosystems, and can now run entirely inside a browser tab with no server required.
The Problem With v1
LiteParse v1 was a Node/TypeScript project. That was fine if you lived in the JavaScript world, but it created a hard dependency on a Node install for everyone else. The team added a Python package that just wrapped the CLI, but this wasn't going to remove the hard dependency on a Node install. They explored compiling the TypeScript code into a binary, but the complex system dependencies made this impossible.
The old version had a practical ceiling: it came from the Node/TypeScript world, which was fine if your app lived there, but less fine if you wanted the same parser inside Python workflows, Rust services, desktop apps, browser contexts, edge runtimes, or agent tools. The only real fix was a full rewrite.
One Rust Core, Four Ecosystems
LiteParse is now available as a native Rust, Node, Python, and WASM package. The team rewrote the entire project in Rust and adapted it to run anywhere: Rust, Python, Node, in the browser, and even on edge runtimes. The install commands are as simple as they get:
# Python
pip install liteparse
# Node / TypeScript
npm i @llamaindex/liteparse
# Rust
cargo install liteparse
# Browser / Edge (WASM)
npm i @llamaindex/liteparse-wasmThe project is a Rust workspace with the core library and language-specific binding crates, including liteparse-napi for Node.js (via napi-rs), liteparse-python (via PyO3), and liteparse-wasm (via wasm-bindgen). Changes to the Rust core propagate automatically to every language binding, which means no more version drift between the Python and Node packages.

How Fast Is It?
The performance story depends on document size, but the numbers are compelling across the board. Previously the runtime was dominated by spinning up a Node process, so small documents see a 5-100x speedup. For larger documents, the team observed around a 3x speedup.
Comparing to other PDF parsing utilities, LiteParse is blazing fast on most documents, with just 0.777s for a 457-page, 100MB document. Against common Python alternatives like pymupdf, pypdf, markitdown, and pdftotext, it consistently leads on throughput. The Rust implementation uses a custom fork and build of PDFium, and is compiled against a build of tesseract-rs as the default OCR implementation.

The WASM Story Is the Real Surprise
Browser support was a community-driven effort before v2. Simon Willison showed that it was possible to run LiteParse in the browser, but you needed to stub out a lot of functions using Vite. The team took that work and made it the official path to browser support in their docs. With the Rust rewrite, they went further.
They created a WASM target for LiteParse and packaged it as @llamaindex/liteparse-wasm, which runs in both the browser and edge runtimes. You can open a demo webpage and try out the WASM build yourself, with all parsing running locally in your browser.
There is one trade-off to know about: to accomplish this, the team still had to stub out system dependencies, meaning the WASM version only parses file bytes directly, and OCR is instead provided through a callback that is passed in (e.g. calling tesseract-js) rather than being a built-in default. For most use cases this is a non-issue, but it's worth knowing if you're building a fully offline browser app that needs OCR.
What It Handles (and What It Doesn't)
LiteParse is an open-source document parsing library that parses text with spatial layout information and bounding boxes. Written in Rust, it runs entirely on your machine with no cloud dependencies, no LLMs, and no API keys. It is designed specifically for use cases that require fast, accurate text parsing: real-time applications, coding agents, and local workflows.
Supported input formats include:
- PDFs -- native text extraction with spatial layout reconstruction
- Office documents -- DOCX, XLSX, PPTX via LibreOffice conversion
- Images -- JPG, PNG, TIFF, WebP, SVG via ImageMagick + OCR
- 50+ document types in total
The OCR layer is pluggable. Built-in Tesseract works out of the box, but you can swap in EasyOCR or PaddleOCR via a simple HTTP server interface, or implement your own by following the OCR API spec.
Where LiteParse hits its limits is equally clear. For dense tables, handwriting, scanned messes, charts, multi-column layouts, and high-stakes extraction, you can still reach for a stronger parser like LlamaParse. LiteParse is explicitly the fast, free, local-first tier -- not a replacement for cloud-based document AI.
Agent-Native From the Start
LiteParse is the less glamorous layer underneath: a fast local parser that can turn PDFs and office documents into usable text, page structure, screenshots, and spatial information without sending the file to an LLM. That last point matters more than it sounds. In agentic pipelines, every document that gets parsed is a potential latency bottleneck and a privacy exposure. A parser that runs in-process, with no API call, changes the economics of how often an agent can afford to read a file.
The tool also ships with first-class agent skill support. You can install it directly into Claude Code, Codex, or OpenCode with a single command:
npx skills add run-llama/llamaparse-agent-skills --skill liteparseFree and Open Source
LiteParse is open-source and runs entirely on your machine with no cloud dependencies, no LLMs, and no API keys. It is licensed under Apache 2.0. The GitHub repo is public, and packages are live on PyPI, npm, and crates.io right now. For teams that need to handle complex layouts, scanned documents, or production-grade accuracy, LlamaIndex's cloud-based LlamaParse remains the recommended upgrade path.
The Rust rewrite is the kind of infrastructure decision that pays dividends quietly. If your parser is fast but trapped in one ecosystem, adoption gets awkward. If your parser is portable, embeddable, and exposed through the languages and runtimes developers already use, it can disappear into products. That is exactly what v2.0 is going for.