LiteParse

LlamaIndex's open-source local document parser, rewritten in Rust: PDFs and Office files to text, JSON or Markdown in milliseconds, with no model and no cloud.

23 Sep 2026

LiteParse is the open-source document parser from the LlamaIndex team. It turns PDFs, Office files and images into plain text, JSON with bounding boxes, or Markdown, entirely on your own machine. No model, no GPU, no cloud call.

Since version 2, released on 25th May 2026, the core is written in Rust on top of Google's PDFium. One engine now serves Python, Node.js, Rust and the browser, with the same lit command line everywhere.

Maker LlamaIndex (run-llama)
Licence Apache 2.0
Language Rust core (about 84% of the code), with bindings for Python, Node.js and WebAssembly
Latest 2.14.7, 22nd September 2026
Traction 12,402 stars, 855 forks
Inputs PDF natively. Word, Excel, PowerPoint and images through conversion (LibreOffice for Office files)
Outputs Text, JSON with bounding boxes, Markdown, PNG page screenshots

What changed since April

My first note, further down, described version 1: TypeScript on PDF.js, with Tesseract.js for OCR. Version 2 replaced all of that. Most of the April details no longer apply.

  • Rust and PDFium instead of TypeScript and PDF.js. Text extraction runs at about 2 to 5 ms per page, according to the README.
  • Markdown output: headings, tables, lists, images and links rebuilt from the page layout. Version 1 only had text and JSON.
  • Python, Rust and browser packages, next to the npm one.
  • A complexity check, lit is-complex, that tells you before a full parse whether a document needs OCR.
  • A worker pool in Python and Node.js, for real parallel parsing with a hard timeout per document.
  • An agent skill, so a coding agent can call it directly.

The April note lists a Homebrew install. The current README doesn't mention Homebrew any more: install it with pip, npm or cargo instead.

For business people

The problem: before an LLM can use a document, something has to turn the PDF into clean text. Cloud parsers charge per page and send your documents to someone else's server. LiteParse does the job locally, for free, in well under a second for a typical report.

Where it fits: as the first pass on every document. The complexity check sorts the easy files from the hard ones. Easy files go straight through LiteParse. Scans, dense tables and charts go to a heavier tool.

The limits: it follows rules, not a model, so it can't read a chart and struggles with complex tables and handwriting. LlamaIndex says so openly and points to its paid cloud parser, LlamaParse, for the hard cases. LiteParse is also the free entry point to that paid product.

How it compares: on 3 public benchmarks, LiteParse beats the other free tools that run without a model. These are LlamaIndex's own numbers, run on their own machine, and the first benchmark, ParseBench, is their own too.

Benchmark LiteParse + PaddleOCR Best other model-free tool markitdown
ParseBench, 2,049 documents 0.364 0.389 0.283 (pdf-inspector) 0.185
opendataloader-bench, 200 documents 0.886 0.901 0.842 (opendataloader) 0.589
olmOCR-bench, 1,403 pages 39.6% 42.2% 33.7% (pdf-inspector) 28.7%

Microsoft's markitdown, the tool most people reach for first, comes last on all 3. The absolute scores stay low for every tool: charts score close to zero across the board.

For technical people

How it works

  1. Conversion. Office files go through LibreOffice, and images through Rust image crates, to become PDFs.
  2. Extraction. PDFium reads the native text layer with positions.
  3. Selective OCR. Only pages or regions without usable text go to OCR.
  4. Merge and grid projection. Native text and OCR results are merged, then projected onto a grid to rebuild the reading order and layout.
  5. Rendering. The result comes out as text, JSON or Markdown.

OCR is pluggable. Tesseract is bundled and needs no setup. Any OCR engine can plug in through a small HTTP API. The best benchmark results use PaddleOCR models, served by the included RapidOCR server on ONNX Runtime.

Install

pip install liteparse               # Python, also installs the lit CLI
npm i -g @llamaindex/liteparse      # Node.js
cargo install liteparse             # Rust CLI

Command line

lit parse report.pdf --format markdown -o report.md
lit parse report.pdf --target-pages "1-5,10" --no-ocr
lit is-complex report.pdf --quiet && lit parse report.pdf --no-ocr
lit batch-parse ./inbox ./parsed
lit screenshot report.pdf --dpi 300 -o ./pages

is-complex exits non-zero when any page needs OCR, so it works as a shell test. Each page gets a verdict and a reason: scanned, no text, sparse text, embedded images, garbled text.

Python

from liteparse import LiteParse

parser = LiteParse(output_format="markdown", ocr_enabled=False)
result = parser.parse("report.pdf")
print(result.total_pages)
print(result.text)

The JSON output can carry much more on request: layout blocks with coordinates, table cells with their own boxes, form fields, annotations, vector lines and the tagged-PDF structure. All of it is off by default.

⚠️ WARNING: PDFium processes one parse at a time per process. For parallel parsing, use the worker pool mode, not threads.

My test

23 Sep 2026 On my Mac Studio (M1 Ultra), with version 2.14.7 in a fresh Python virtual environment. The package is about 26 MB. The test file was the Transformer paper from arXiv: 15 pages, 2.2 MB.

Step Result
lit is-complex --quiet 0.13 seconds, exit 0: no page needs OCR
lit parse --no-ocr --format markdown 0.52 seconds for all 15 pages, 6,583 words, 10 tables
Headings, abstract, body text Clean, in the right reading order
A complex table (2-row header) Scrambled: the header merged into the first row
Exponents Lost: "1.0 · 10²⁰" came out as "1.0 · 1020"
Line-break hyphens 1 lost: "English-to-German" became "Englishto-German"

The speed claim holds. Plain prose comes out ready to use. Tables with merged headers and scientific notation need a check by eye.

Value

  • A free, private first pass. For documents I wouldn't send to a cloud parser, LiteParse is fast and good enough on prose.
  • Routing built in. is-complex answers the question every document pipeline asks first: cheap path or expensive path?
  • Better than markitdown on these benchmarks, and it keeps positions, which markitdown doesn't.
  • A candidate engine for my own document-to-Markdown skill (see The NicAI skills catalogue ), with a rule to review any table before it goes into the knowledge base.

The April 2026 note (version 1)

23 Apr 2026

What it does

LiteParse is an open-source document parser from the LlamaIndex team. It performs spatial text parsing with bounding boxes entirely locally, no cloud dependencies. Built on PDF.js for fast native PDF parsing with optional Tesseract.js OCR for scanned documents.

Language

TypeScript (71.9%) with Python components (26.5%) for the OCR server.

Install

# npm
npm i -g @llamaindex/liteparse

# Homebrew (macOS/Linux)
brew install llamaindex-liteparse

Also available to build from source.

Key features

  • Fast text parsing using PDF.js
  • Flexible OCR with built-in Tesseract.js support
  • Multiple output formats - JSON and text
  • Precise bounding boxes for text positioning
  • Screenshot generation for LLM agents
  • Multi-platform - Linux, macOS, Windows
  • No cloud required - runs fully standalone

Supported formats

Beyond native PDFs, LiteParse handles automatic conversion for:

  • Office documents (Word, PowerPoint, spreadsheets)
  • Images (JPG, PNG, GIF, TIFF, WebP)

Value

A solid alternative to cloud-based document parsers like LlamaParse (also from the LlamaIndex team, but cloud-hosted). The local-first approach is good for privacy-sensitive workflows and air-gapped environments. The bounding box output is useful for layout-aware RAG pipelines where you need to know where text sits on the page, not just what it says.

Further Reading

NicAI
Written by NicAI, Nic's AI assistant, for his personal knowledge base. Researched and drafted by the model, not hand-written by Nic. Verify anything you plan to act on.