
23 Sep 2026
LiteParse is the open-source document parser from the LlamaIndex team. It turns PDFs, Office files and images into plain text, JSON with bounding boxes, or Markdown, entirely on your own machine. No model, no GPU, no cloud call.
Since version 2, released on 25th May 2026, the core is written in Rust on top of Google's PDFium. One engine now serves Python, Node.js, Rust and the browser, with the same lit command line everywhere.
| Maker | LlamaIndex (run-llama) |
| Licence | Apache 2.0 |
| Language | Rust core (about 84% of the code), with bindings for Python, Node.js and WebAssembly |
| Latest | 2.14.7, 22nd September 2026 |
| Traction | 12,402 stars, 855 forks |
| Inputs | PDF natively. Word, Excel, PowerPoint and images through conversion (LibreOffice for Office files) |
| Outputs | Text, JSON with bounding boxes, Markdown, PNG page screenshots |
What changed since April
My first note, further down, described version 1: TypeScript on PDF.js, with Tesseract.js for OCR. Version 2 replaced all of that. Most of the April details no longer apply.
- Rust and PDFium instead of TypeScript and PDF.js. Text extraction runs at about 2 to 5 ms per page, according to the README.
- Markdown output: headings, tables, lists, images and links rebuilt from the page layout. Version 1 only had text and JSON.
- Python, Rust and browser packages, next to the npm one.
- A complexity check,
lit is-complex, that tells you before a full parse whether a document needs OCR. - A worker pool in Python and Node.js, for real parallel parsing with a hard timeout per document.
- An agent skill, so a coding agent can call it directly.
The April note lists a Homebrew install. The current README doesn't mention Homebrew any more: install it with pip, npm or cargo instead.
For business people
The problem: before an LLM can use a document, something has to turn the PDF into clean text. Cloud parsers charge per page and send your documents to someone else's server. LiteParse does the job locally, for free, in well under a second for a typical report.
Where it fits: as the first pass on every document. The complexity check sorts the easy files from the hard ones. Easy files go straight through LiteParse. Scans, dense tables and charts go to a heavier tool.
The limits: it follows rules, not a model, so it can't read a chart and struggles with complex tables and handwriting. LlamaIndex says so openly and points to its paid cloud parser, LlamaParse, for the hard cases. LiteParse is also the free entry point to that paid product.
How it compares: on 3 public benchmarks, LiteParse beats the other free tools that run without a model. These are LlamaIndex's own numbers, run on their own machine, and the first benchmark, ParseBench, is their own too.
| Benchmark | LiteParse | + PaddleOCR | Best other model-free tool | markitdown |
|---|---|---|---|---|
| ParseBench, 2,049 documents | 0.364 | 0.389 | 0.283 (pdf-inspector) | 0.185 |
| opendataloader-bench, 200 documents | 0.886 | 0.901 | 0.842 (opendataloader) | 0.589 |
| olmOCR-bench, 1,403 pages | 39.6% | 42.2% | 33.7% (pdf-inspector) | 28.7% |
Microsoft's markitdown, the tool most people reach for first, comes last on all 3. The absolute scores stay low for every tool: charts score close to zero across the board.
For technical people
How it works
- Conversion. Office files go through LibreOffice, and images through Rust image crates, to become PDFs.
- Extraction. PDFium reads the native text layer with positions.
- Selective OCR. Only pages or regions without usable text go to OCR.
- Merge and grid projection. Native text and OCR results are merged, then projected onto a grid to rebuild the reading order and layout.
- Rendering. The result comes out as text, JSON or Markdown.
OCR is pluggable. Tesseract is bundled and needs no setup. Any OCR engine can plug in through a small HTTP API. The best benchmark results use PaddleOCR models, served by the included RapidOCR server on ONNX Runtime.
Install
pip install liteparse # Python, also installs the lit CLI
npm i -g @llamaindex/liteparse # Node.js
cargo install liteparse # Rust CLI
Command line
lit parse report.pdf --format markdown -o report.md
lit parse report.pdf --target-pages "1-5,10" --no-ocr
lit is-complex report.pdf --quiet && lit parse report.pdf --no-ocr
lit batch-parse ./inbox ./parsed
lit screenshot report.pdf --dpi 300 -o ./pages
is-complex exits non-zero when any page needs OCR, so it works as a shell test. Each page gets a verdict and a reason: scanned, no text, sparse text, embedded images, garbled text.
Python
from liteparse import LiteParse
parser = LiteParse(output_format="markdown", ocr_enabled=False)
result = parser.parse("report.pdf")
print(result.total_pages)
print(result.text)
The JSON output can carry much more on request: layout blocks with coordinates, table cells with their own boxes, form fields, annotations, vector lines and the tagged-PDF structure. All of it is off by default.
⚠️ WARNING: PDFium processes one parse at a time per process. For parallel parsing, use the worker pool mode, not threads.
My test
23 Sep 2026 On my Mac Studio (M1 Ultra), with version 2.14.7 in a fresh Python virtual environment. The package is about 26 MB. The test file was the Transformer paper from arXiv: 15 pages, 2.2 MB.
| Step | Result |
|---|---|
lit is-complex --quiet |
0.13 seconds, exit 0: no page needs OCR |
lit parse --no-ocr --format markdown |
0.52 seconds for all 15 pages, 6,583 words, 10 tables |
| Headings, abstract, body text | Clean, in the right reading order |
| A complex table (2-row header) | Scrambled: the header merged into the first row |
| Exponents | Lost: "1.0 · 10²⁰" came out as "1.0 · 1020" |
| Line-break hyphens | 1 lost: "English-to-German" became "Englishto-German" |
The speed claim holds. Plain prose comes out ready to use. Tables with merged headers and scientific notation need a check by eye.
Value
- A free, private first pass. For documents I wouldn't send to a cloud parser, LiteParse is fast and good enough on prose.
- Routing built in.
is-complexanswers the question every document pipeline asks first: cheap path or expensive path? - Better than markitdown on these benchmarks, and it keeps positions, which markitdown doesn't.
- A candidate engine for my own document-to-Markdown skill (see The NicAI skills catalogue ), with a rule to review any table before it goes into the knowledge base.
The April 2026 note (version 1)
23 Apr 2026
What it does
LiteParse is an open-source document parser from the LlamaIndex team. It performs spatial text parsing with bounding boxes entirely locally, no cloud dependencies. Built on PDF.js for fast native PDF parsing with optional Tesseract.js OCR for scanned documents.
Language
TypeScript (71.9%) with Python components (26.5%) for the OCR server.
Install
# npm
npm i -g @llamaindex/liteparse
# Homebrew (macOS/Linux)
brew install llamaindex-liteparse
Also available to build from source.
Key features
- Fast text parsing using PDF.js
- Flexible OCR with built-in Tesseract.js support
- Multiple output formats - JSON and text
- Precise bounding boxes for text positioning
- Screenshot generation for LLM agents
- Multi-platform - Linux, macOS, Windows
- No cloud required - runs fully standalone
Supported formats
Beyond native PDFs, LiteParse handles automatic conversion for:
- Office documents (Word, PowerPoint, spreadsheets)
- Images (JPG, PNG, GIF, TIFF, WebP)
Value
A solid alternative to cloud-based document parsers like LlamaParse (also from the LlamaIndex team, but cloud-hosted). The local-first approach is good for privacy-sensitive workflows and air-gapped environments. The bounding box output is useful for layout-aware RAG pipelines where you need to know where text sits on the page, not just what it says.
Further Reading
