Skip to content

Getting started ​

@nalinor/mupdf4llm is a TypeScript port of the classic pymupdf4llm.to_markdown pipeline. It takes PDF bytes and returns LLM-friendly Markdown with reading-order text, headings, bullets, inline styling, and tables.

What it does ​

You give itYou get back
Uint8Array / ArrayBuffer of a PDFA string of GitHub-flavored Markdown
Same buffer + toMarkdownPagesOne PageChunk per page with metadata, optional per-word coordinates, table bboxes, image refs

What it doesn't do ​

Two upstream features are blocked at the ecosystem level, not by this port:

  • pymupdf.layout features (to_text, to_json, layout-mode to_markdown) — they need Artifex's separate closed-source pymupdf-layout ONNX wheel, distributed under a Polyform Noncommercial license. No JS distribution exists; we can't legally repackage the model.
  • Whole-page OCR — the official mupdf npm WASM bundle ships without Tesseract/Leptonica linked in. Table cells can be OCR'd with a pluggable engine — see OCR for tables.

For everything else — see parity and limits.

Next steps ​

  1. Install the package.
  2. Run through the quick start.
  3. Explore the options reference to tune for your PDFs.

AGPL-3.0-or-later — inherited from PyMuPDF and pymupdf4llm.