Skip to content

Installation ​

Requirements ​

  • Node ≥ 20 or Bun ≥ 1.0
  • ESM only: use import. From CommonJS, await import("@nalinor/mupdf4llm") (the mupdf dependency uses top-level await, which require() cannot load).
  • The only runtime dependency is mupdf (Artifex's official WASM bindings). No native build, no system libs.
  • Optional, only for OCR of table cells with the default RapidOCR engine: npm i ppu-paddle-ocr onnxruntime-node (ONNX Runtime is a native package of a few hundred MB).

Install ​

sh
npm i @nalinor/mupdf4llm
sh
bun add @nalinor/mupdf4llm
sh
pnpm add @nalinor/mupdf4llm
sh
yarn add @nalinor/mupdf4llm

Optional: LlamaIndex adapter ​

If you want the PDFMarkdownReader integration, install llamaindex alongside — it's declared as an optional peer dependency:

sh
npm i llamaindex

Without llamaindex, the adapter still loads — it just returns plain { text, metadata } objects with the same shape as a LlamaIndex Document (the metadata field is always metadata, never extra_info).

Optional: parity tests ​

To run the parity suite locally you also need Python with pymupdf4llm:

sh
pip install pymupdf4llm

The tests spawn python3 -c "..." to compare the TS output against the upstream. CI sets this up automatically — only needed if you contribute to the library itself.

Verify ​

ts
import { toMarkdown } from "@nalinor/mupdf4llm";
import { readFileSync } from "node:fs";

const md = await toMarkdown(readFileSync("any.pdf"));
console.log(md.slice(0, 80));

If you see Markdown coming out, you're set. Continue to the quick start.

AGPL-3.0-or-later — inherited from PyMuPDF and pymupdf4llm.