Skip to content

Function: extractWords() ​

ts
function extractWords(page): Word[];

Defined in: src/helpers/text/extractWords.ts:36

Extract every word from a page, grouped on whitespace boundaries. Mirrors page.get_text("words") in PyMuPDF: the word index resets per line, and block advances on every text block in document order.

Parameters ​

page ​

Page

Returns ​

Word[]

AGPL-3.0-or-later — inherited from PyMuPDF and pymupdf4llm.