pdftext
Section titled “pdftext”Deterministic PDF text extraction (pdfplumber) for figure-extraction jobs — the counterpart to markitdown <file>, deliberately not a replacement for it.
Project
Section titled “Project”- Location:
~/utils/pdftext/(WSL) - Package:
uv-managed; installed as a CLI viauv tool install .→~/.local/bin/pdftext
Quick Start
Section titled “Quick Start”pdftext <file.pdf> # text to stdoutpdftext <file.pdf> -o out.txtpdftext <file.pdf> --page 3 # single page, 1-indexedpdftext <file.pdf> --layout # preserve horizontal spacing (when columns matter)As a library, import pdfplumber directly — see ~/utils/acct-qbo/pdf_text.py.
When to use this vs MarkItDown
Section titled “When to use this vs MarkItDown”Per AGENTS.md “Document extraction — pick by job”:
| Job | Tool |
|---|---|
| Document → markdown for an LLM to read (summarize, RAG, “what does this say”) | markitdown |
| Pulling specific figures deterministically, where a number lands in the books or drives a decision | pdftext |
Why: MarkItDown’s markdown shape varies between files of the same document type. Harmless when a model reads it; dangerous when a regex extracts money. Measured across 26 real BMO statements (2026-08-09), that variance produced five distinct parsing bugs — including one statement silently parsed as zero transactions, and an amount emitted on the line before its row, which deleted a real charge and invented a credit. Both looked like success.
pdfplumber returns the page’s own text lines in reading order, so a table row is a line. Verified: all 26 statements clean, 8.5s for the whole corpus.
Gotcha it does NOT solve
Section titled “Gotcha it does NOT solve”Choosing the right extractor reduces the failure rate; it does not make a parse trustworthy. Gate figure extraction on the document’s own arithmetic — cross-check parsed rows against the printed subtotal/summary and fail loudly on mismatch. That gate is what actually catches errors, and it caught two the extractor swap alone would have missed.
Consumers
Section titled “Consumers”~/utils/acct-qbo/—/hst-assistancestatement parsers (viapdf_text.py)/process-expenses— bill/invoice field extraction
Related
Section titled “Related”- External tool comparison (MarkItDown, pdfplumber, Docling, pdf-inspector): Utilities
- Do not use firecrawl/pdf-inspector for financial extraction — it returned confident, well-formed output containing no figures at all on a real statement, with no error.