Getting started
@nalinor/mupdf4llm is a TypeScript port of the classic pymupdf4llm.to_markdown pipeline. It takes PDF bytes and returns LLM-friendly Markdown with reading-order text, headings, bullets, inline styling, and tables.
What it does
| You give it | You get back |
|---|---|
Uint8Array / ArrayBuffer of a PDF | A string of GitHub-flavored Markdown |
Same buffer + toMarkdownPages | One PageChunk per page with metadata, optional per-word coordinates, table bboxes, image refs |
What it doesn't do
Two upstream features are blocked at the ecosystem level, not by this port:
pymupdf.layoutfeatures (to_text,to_json, layout-modeto_markdown) — they need Artifex's separate closed-sourcepymupdf-layoutONNX wheel, distributed under a Polyform Noncommercial license. No JS distribution exists; we can't legally repackage the model.- Whole-page OCR — the official
mupdfnpm WASM bundle ships without Tesseract/Leptonica linked in. Table cells can be OCR'd with a pluggable engine — see OCR for tables.
For everything else — see parity and limits.
Next steps
- Install the package.
- Run through the quick start.
- Explore the options reference to tune for your PDFs.