Drop-in for pymupdf4llm
Same algorithms as the Python reference (`pymupdf4llm.helpers.pymupdf_rag.to_markdown`). Validated by parity tests that diff every fixture against the live Python output.
TypeScript port of pymupdf4llm on top of the official mupdf WASM package. Tracks the Python upstream closely — exact on synthetic fixtures, with documented divergences on real-world PDFs.
The Python pymupdf4llm is excellent, but RAG pipelines increasingly live in Node / Bun / browser-adjacent stacks. Spinning up a Python sidecar just to call to_markdown is friction. This package gives you the same output, the same options, and the same fixtures — straight from the JavaScript runtime you already have.
npm i @nalinor/mupdf4llmbun add @nalinor/mupdf4llmpnpm add @nalinor/mupdf4llmimport { toMarkdown } from "@nalinor/mupdf4llm";
import { readFileSync } from "node:fs";
const md = await toMarkdown(readFileSync("paper.pdf"));
console.log(md);Need per-page chunks for RAG? Switch to toMarkdownPages:
import { toMarkdownPages } from "@nalinor/mupdf4llm";
const chunks = await toMarkdownPages(readFileSync("paper.pdf"), {
extractWords: true,
});
for (const c of chunks) console.log(c.metadata.page, c.text.length);See the quick start for more.