Skip to content

index ​

Classes ​

ClassDescription
IdentifyHeaders-
OcrSetupErrorOCR cannot run at all (package missing, bad model, a service that rejects the credentials). It aborts the conversion instead of marking every cell as failed; custom engines throw it for such fatal problems.
Point-
ProgressBarMinimal text progress bar — mirrors pymupdf4llm.helpers.progress.ProgressBar.
Rect-
TocHeadersAssign header levels from the document's table of contents.

Interfaces ​

InterfaceDescription
CellText-
FormField-
ImageInfo-
MarkdownOptions-
OcrEnginePluggable OCR backend. The library calls recognize once per table cell and uses the returned text as the cell content. An exception or empty text marks the cell as "failed".
OcrImageAn 8-bit grayscale image handed to an OcrEngine.
OcrResultWhat an OcrEngine read in an image.
PageChunk-
RapidOcrModelModel file locations (paths, URLs or buffers), as accepted by ppu-paddle-ocr.
RapidOcrOptions-
WordPer-word record produced by extractWords, matching the structure of page.get_text("words") in PyMuPDF: (x0, y0, x1, y1, text, block, line, word).

Type Aliases ​

Type AliasDescription
CellSourceWhere a table cell's text came from: the PDF text layer, OCR, or OCR that failed.
MarkdownElement-
TextSourceWhere table cell text is taken from. See MarkdownOptions.textSource.

Functions ​

FunctionDescription
clusterStripesGroup rectangles into horizontal stripes where consecutive rectangles vertically overlap. Ports pymupdf4llm.helpers.utils.cluster_stripes — the building block used by reading-order detection.
computeReadingOrderSort rectangles into a single-page reading order: top-to-bottom by stripe, then left-to-right within each stripe. Mirrors pymupdf4llm.helpers.utils.compute_reading_order.
createRapidOcrDefault OCR engine: the RapidOCR stack (PaddleOCR PP-OCR models on ONNX Runtime) through the optional peer dependencies ppu-paddle-ocr and onnxruntime-node. Models are downloaded and cached on first use.
extractWordsExtract every word from a page, grouped on whitespace boundaries. Mirrors page.get_text("words") in PyMuPDF: the word index resets per line, and block advances on every text block in document order.
getKeyValuesExtract every form field from a PDF as a flat list, one entry per widget.
getPageRotationRead the /Rotate entry on a PDF page. Returns 0 if missing.
removeRotationStrip rotation from a page while preserving its visual appearance, then return the previous rotation value.
setPageRotationSet (or clear) the /Rotate entry on a PDF page. Pass 0 to remove rotation.
toMarkdownOpen a PDF from bytes and convert to markdown.
toMarkdownPagesOpen a PDF from bytes and convert to page-chunk JSON.

AGPL-3.0-or-later — inherited from PyMuPDF and pymupdf4llm.