Skip to content

PDF compression wrapper. Two modes:

  • Default (lossy images, lossless text) — Ghostscript /screen preset. Text stays vector/100%-exact; embedded raster images get downsampled. This is the recommended default per Talbot’s call (2026-08-17): text accuracy is the only hard requirement, degraded images are fine.
  • --lossless — qpdf recompress only (no image touching). Use when the images themselves must also stay byte-exact (e.g. a scanned signature page where pixel fidelity matters).
  • Location: ~/utils/pdfcompress/ (WSL) — pdfcompress (bash) + verify_pdf_text.py (PEP 723 self-contained uv run --script, no separate venv needed).
  • Binaries: qpdf in ~/.local/bin/ (extracted from the Ubuntu .deb via dpkg-deb -x, no sudo needed); gs (Ghostscript) installed system-wide via sudo apt install ghostscript (Talbot ran this directly, 2026-08-17).
Terminal window
pdfcompress in.pdf out.pdf # lossy images, lossless text (default, biggest win)
pdfcompress --lossless in.pdf out.pdf # qpdf only, nothing touches images
pdfcompress --batch src_dir/ dst_dir/ # recursive: every *.pdf under src_dir,
# mirrored into dst_dir, same subfolders
pdfcompress --lossless --batch src_dir/ dst_dir/ # batch + lossless

Batch mode prints a before/after line per file and a total (bytes + ratio) at the end. It does not modify src_dir — always writes to dst_dir, creating subfolders as needed.

Every compression is verified automatically (verify_pdf_text.py, in the same folder) — see “Verification method” below. A file that fails is deleted from dst_dir and listed at the end; pdfcompress --batch exits non-zero if anything failed. --no-verify skips the check (not recommended).

Findings (2026-08-17, PDF-Compression task)

Section titled “Findings (2026-08-17, PDF-Compression task)”

qpdf-only (lossless) ceiling: ~7–15% — not enough of a win; the size gaps Talbot noticed originate in embedded image content, and lossless recompression can’t touch pixel data.

Ghostscript /screen (lossy images, lossless text) — the real win. Tested on all 12 of Talbot’s 2025 BMO Investorline statements — full comparison, /screen vs /ebook:

FileOrig/screenratio/ebookratio
Jan 2025727,797 B182,205 B4.0x262,211 B2.8x
Feb 2025723,641 B177,828 B4.1x257,843 B2.8x
Mar 2025723,718 B177,992 B4.1x258,007 B2.8x
Apr 2025691,860 B183,790 B3.8x263,785 B2.6x
May 2025691,128 B183,164 B3.8x263,179 B2.6x
Jun 2025691,260 B183,221 B3.8x263,227 B2.6x
Jul 2025691,779 B184,010 B3.8x264,013 B2.6x
Aug 2025697,562 B191,910 B3.6x271,895 B2.6x
Sep 2025700,110 B190,749 B3.7x270,737 B2.6x
Oct 2025702,511 B194,671 B3.6x274,668 B2.6x
Nov 20254,478,301 B126,578 B35.4x197,966 B22.6x
Dec 20252,948,191 B117,088 B25.2x188,503 B15.6x
TOTAL14.47MB2.09MB6.9x (14.5%)3.04MB4.8x (21.0%)

Not a minimal gap — /ebook runs ~45% larger than /screen across the set (3.04MB vs 2.09MB). Recommendation: /screen, since Talbot’s stated bar is text-exact + images-can-degrade, and /screen’s degradation was already visually confirmed acceptable (page 1 of August statement — banner + all figures legible). /printer (300dpi) exists as a third, milder option (~8x on the worst-case file) if image fidelity ever needs to be higher than /screen gives.

  • Text verified 100% lossless on every one of the 12 files: page.get_text() string-equality, original vs. compressed, all pages, via PyMuPDF — despite Ghostscript emitting harmless PDFDocEncoding XX is undefined warnings during conversion (font-encoding notices, not data loss — confirmed by the text match).
  • Root cause of the size gaps, confirmed: each BMO statement page embeds a 2550×300px header banner image — print-resolution artwork, wildly oversized for on-screen/archival use. /screen’s downsampling is exactly the right fix; it’s not losing statement data, it’s discarding resolution the banner graphic never needed.
  • Image quality at /screen: visually inspected page 1 of the August statement post-compression — banner gradient and all numbers/text fully legible (see task log for the rendered PNG).

Verification method (built into the tool, 2026-08-17)

Section titled “Verification method (built into the tool, 2026-08-17)”

Original method (raw get_text() string-equality) had false positives — retired. Retesting across a broader document set (legacy BMO/Wealthsimple statements, an SSRN academic paper, Globe & Mail articles, a book excerpt) found two benign cases where Ghostscript’s re-encoding changes the extracted text without losing anything:

  • Font-glyph-mapping artifacts — a legacy BMO Investorline statement’s original PDF had a font whose glyph-to-Unicode mapping made PyMuPDF misread “Dollars” as “Dolars” (missing letter); Ghostscript’s re-encoding incidentally fixed the extraction. Visually identical either way.
  • Table re-flow whitespace — an SSRN paper’s table headers gained one extra blank line under re-encoding; no words or numbers changed.

Both would fail a strict string-equality check despite zero real data loss — and did, until this was caught in File-Naming-Vault-Wide-Rollout round 9, where a workaround for a permission-classifier block accidentally overwrote 22 flagged-but-actually-fine files without the mismatch skip-list. Files were confirmed fine content-wise, but the near-miss is why the check moved into the tool itself instead of staying an ad-hoc external script.

Current method (verify_pdf_text.py): page count must match, and every numeric token (\d[\d,]*\.?\d*) in the original must appear the same number of times in the compressed output — this is the actual invariant that matters for a financial reference artifact (per AI-Testing-Standards.md’s “gate on the document’s own facts” principle). Prose/whitespace is allowed to differ; a changed or missing number fails and the compressed file is deleted, never silently kept.

Tested clean (numbers exact, whitespace-only or zero diff) across: BMO Chequing/CreditCard/Investorline statements (2015–2026), Wealthsimple statements, an SSRN academic paper (15MB→1.6MB), a Globe & Mail article, and a book-excerpt PDF — 140 files, 0 failures, ~42% average size reduction on this mixed set.

(For the earlier lossless-only qpdf pass, pixel-hash (SHA-256 of rendered samples, 150–200 DPI) was also checked and matched exactly — see git/task history for that data.)