Skip to content

Options reference ​

Every option on MarkdownOptions. Defaults match pymupdf4llm.helpers.pymupdf_rag.to_markdown.

Page selection ​

KeyTypeDefaultNotes
pagesnumber[]all pages0-based indices.
marginsnumber | [t, b] | [l, t, r, b]0Crop the page before extraction.
filenamestring""Surfaced in PageChunk.metadata.file_path and used in image filenames.

Output mode ​

KeyTypeDefaultNotes
pageSeparatorsbooleanfalseInsert --- end of page=N --- markers in joined output.

Use toMarkdownPages instead of toMarkdown when you want per-page chunks — the choice of return shape isn't an option, it's a separate entry point.

Text and styling ​

KeyTypeDefaultNotes
ignoreCodebooleanfalseSuppress fenced code blocks for monospaced spans.
forceTextbooleantrueEmit text even on top of images.
hdrInfoIdentifyHeaders | TocHeaders | function | falseauto IdentifyHeadersSee Headers.

Tables ​

KeyTypeDefaultNotes
tableStrategy"lines_strict" | "lines" | "text" | "explicit" | "pixels" | null"lines_strict"See Tables. null disables detection.
explicitTableGrids{ hLines: number[]; vLines: number[] }[][]For "explicit" only.
textSource"pdf" | "ocr" | "auto""auto""auto". See OCR.
ocrOcrEngineRapidOCRNeeds ppu-paddle-ocr + onnxruntime-node unless you pass one.
ocrDpinumber300Render resolution for "pixels" and OCR.
detectOrientationbooleantrueWith "pixels", turn a page without a text layer scanned sideways. See Page rotation.

Images ​

KeyTypeDefaultNotes
writeImagesbooleanfalseSave each detected image region to imagePath.
embedImagesbooleanfalseInline images as data: URIs in the Markdown.
imagePathstring""Output directory for writeImages. Created if missing.
imageFormat"png" | "jpg" | "jpeg""png"Encoding format.
dpinumber150Rasterization DPI.
imageSizeLimitnumber0.05Skip images smaller than this fraction of a page edge.

Words / chunks ​

KeyTypeDefaultNotes
extractWordsbooleanfalsePopulate PageChunk.words with per-word bboxes.

Page rotation ​

KeyTypeDefaultNotes
removeRotationbooleantrueDerotate the page (visual-preserving) before processing, like PyMuPDF's remove_rotation().

Filtering ​

KeyTypeDefaultNotes
fontsizeLimitnumberundefinedSkip spans whose font size is below this (in pt). Mirrors upstream FONTSIZE_LIMIT.
ignoreAlphabooleanfalseAccept invisible text. No-op today — mupdf.js doesn't expose per-char alpha; field reserved for upstream parity.

Element selection ​

KeyTypeDefaultNotes
elementsMarkdownElement[]all (omitted)Whitelist of markdown constructs to emit. Anything not listed becomes plain text.

When elements is omitted, every construct is emitted (the original behavior). When provided, only the listed constructs are produced; everything else falls back to plain text. Valid values:

ValueControls
"bold"**…** around bold spans
"italic"_…_ around italic spans
"inlineCode"`…` around monospaced spans (outside code blocks)
"codeBlock"``` fenced blocks for monospaced lines
"header"#-prefixed headings
"bulletList"- bullet-list conversion
"link"[text](url) hyperlinks
"table"Markdown tables
"image"![](…) image references
"lineBreak"<br> between wrapped lines inside table cells

elements combines with the legacy toggles (ignoreCode, hdrInfo: false, tableStrategy: null, writeImages/embedImages): a construct is emitted only if both the whitelist allows it and no legacy option disabled it.

ts
// Plain text, headers and tables — but no bold/italic and no <br> in cells.
await toMarkdown(buf, {
  elements: ["header", "table", "link", "bulletList", "italic"],
});

// Strip all markup — just text.
await toMarkdown(buf, { elements: [] });

Misc ​

KeyTypeDefaultNotes
showProgressbooleanfalseRender ProgressBar to stderr while iterating pages.

For the canonical types, see src/helpers/types.ts.

AGPL-3.0-or-later — inherited from PyMuPDF and pymupdf4llm.