index
Classes
| Class | Description |
|---|---|
| IdentifyHeaders | - |
| OcrSetupError | OCR cannot run at all (package missing, bad model, a service that rejects the credentials). It aborts the conversion instead of marking every cell as failed; custom engines throw it for such fatal problems. |
| Point | - |
| ProgressBar | Minimal text progress bar — mirrors pymupdf4llm.helpers.progress.ProgressBar. |
| Rect | - |
| TocHeaders | Assign header levels from the document's table of contents. |
Interfaces
| Interface | Description |
|---|---|
| CellText | - |
| FormField | - |
| ImageInfo | - |
| MarkdownOptions | - |
| OcrEngine | Pluggable OCR backend. The library calls recognize once per table cell and uses the returned text as the cell content. An exception or empty text marks the cell as "failed". |
| OcrImage | An 8-bit grayscale image handed to an OcrEngine. |
| OcrResult | What an OcrEngine read in an image. |
| PageChunk | - |
| RapidOcrModel | Model file locations (paths, URLs or buffers), as accepted by ppu-paddle-ocr. |
| RapidOcrOptions | - |
| Word | Per-word record produced by extractWords, matching the structure of page.get_text("words") in PyMuPDF: (x0, y0, x1, y1, text, block, line, word). |
Type Aliases
| Type Alias | Description |
|---|---|
| CellSource | Where a table cell's text came from: the PDF text layer, OCR, or OCR that failed. |
| MarkdownElement | - |
| TextSource | Where table cell text is taken from. See MarkdownOptions.textSource. |
Functions
| Function | Description |
|---|---|
| clusterStripes | Group rectangles into horizontal stripes where consecutive rectangles vertically overlap. Ports pymupdf4llm.helpers.utils.cluster_stripes — the building block used by reading-order detection. |
| computeReadingOrder | Sort rectangles into a single-page reading order: top-to-bottom by stripe, then left-to-right within each stripe. Mirrors pymupdf4llm.helpers.utils.compute_reading_order. |
| createRapidOcr | Default OCR engine: the RapidOCR stack (PaddleOCR PP-OCR models on ONNX Runtime) through the optional peer dependencies ppu-paddle-ocr and onnxruntime-node. Models are downloaded and cached on first use. |
| extractWords | Extract every word from a page, grouped on whitespace boundaries. Mirrors page.get_text("words") in PyMuPDF: the word index resets per line, and block advances on every text block in document order. |
| getKeyValues | Extract every form field from a PDF as a flat list, one entry per widget. |
| getPageRotation | Read the /Rotate entry on a PDF page. Returns 0 if missing. |
| removeRotation | Strip rotation from a page while preserving its visual appearance, then return the previous rotation value. |
| setPageRotation | Set (or clear) the /Rotate entry on a PDF page. Pass 0 to remove rotation. |
| toMarkdown | Open a PDF from bytes and convert to markdown. |
| toMarkdownPages | Open a PDF from bytes and convert to page-chunk JSON. |