Skip to content

Deterministic PDF text extraction (pdfplumber) for figure-extraction jobs — the counterpart to markitdown <file>, deliberately not a replacement for it.

  • Location: ~/utils/pdftext/ (WSL)
  • Package: uv-managed; installed as a CLI via uv tool install . → ~/.local/bin/pdftext
Terminal window
pdftext <file.pdf> # text to stdout
pdftext <file.pdf> -o out.txt
pdftext <file.pdf> --page 3 # single page, 1-indexed
pdftext <file.pdf> --layout # preserve horizontal spacing (when columns matter)

As a library, import pdfplumber directly — see ~/utils/acct-qbo/pdf_text.py.

Per AGENTS.md “Document extraction — pick by job”:

JobTool
Document → markdown for an LLM to read (summarize, RAG, “what does this say”)markitdown
Pulling specific figures deterministically, where a number lands in the books or drives a decisionpdftext

Why: MarkItDown’s markdown shape varies between files of the same document type. Harmless when a model reads it; dangerous when a regex extracts money. Measured across 26 real BMO statements (2026-08-09), that variance produced five distinct parsing bugs — including one statement silently parsed as zero transactions, and an amount emitted on the line before its row, which deleted a real charge and invented a credit. Both looked like success.

pdfplumber returns the page’s own text lines in reading order, so a table row is a line. Verified: all 26 statements clean, 8.5s for the whole corpus.

Choosing the right extractor reduces the failure rate; it does not make a parse trustworthy. Gate figure extraction on the document’s own arithmetic — cross-check parsed rows against the printed subtotal/summary and fail loudly on mismatch. That gate is what actually catches errors, and it caught two the extractor swap alone would have missed.

  • ~/utils/acct-qbo/ — /hst-assistance statement parsers (via pdf_text.py)
  • /process-expenses — bill/invoice field extraction
  • External tool comparison (MarkItDown, pdfplumber, Docling, pdf-inspector): Utilities
  • Do not use firecrawl/pdf-inspector for financial extraction — it returned confident, well-formed output containing no figures at all on a real statement, with no error.