pdfcompress
Section titled “pdfcompress”PDF compression wrapper. Two modes:
- Default (lossy images, lossless text) — Ghostscript
/screenpreset. Text stays vector/100%-exact; embedded raster images get downsampled. This is the recommended default per Talbot’s call (2026-08-17): text accuracy is the only hard requirement, degraded images are fine. --lossless— qpdf recompress only (no image touching). Use when the images themselves must also stay byte-exact (e.g. a scanned signature page where pixel fidelity matters).
Project
Section titled “Project”- Location:
~/utils/pdfcompress/(WSL) —pdfcompress(bash) +verify_pdf_text.py(PEP 723 self-containeduv run --script, no separate venv needed). - Binaries:
qpdfin~/.local/bin/(extracted from the Ubuntu.debviadpkg-deb -x, no sudo needed);gs(Ghostscript) installed system-wide viasudo apt install ghostscript(Talbot ran this directly, 2026-08-17).
Quick Start
Section titled “Quick Start”pdfcompress in.pdf out.pdf # lossy images, lossless text (default, biggest win)pdfcompress --lossless in.pdf out.pdf # qpdf only, nothing touches images
pdfcompress --batch src_dir/ dst_dir/ # recursive: every *.pdf under src_dir, # mirrored into dst_dir, same subfolderspdfcompress --lossless --batch src_dir/ dst_dir/ # batch + losslessBatch mode prints a before/after line per file and a total (bytes + ratio) at the end. It does not modify src_dir — always writes to dst_dir, creating subfolders as needed.
Every compression is verified automatically (verify_pdf_text.py, in the same folder) — see “Verification method” below. A file that fails is deleted from dst_dir and listed at the end; pdfcompress --batch exits non-zero if anything failed. --no-verify skips the check (not recommended).
Findings (2026-08-17, PDF-Compression task)
Section titled “Findings (2026-08-17, PDF-Compression task)”qpdf-only (lossless) ceiling: ~7–15% — not enough of a win; the size gaps Talbot noticed originate in embedded image content, and lossless recompression can’t touch pixel data.
Ghostscript /screen (lossy images, lossless text) — the real win. Tested on all 12 of Talbot’s 2025 BMO Investorline statements — full comparison, /screen vs /ebook:
| File | Orig | /screen | ratio | /ebook | ratio |
|---|---|---|---|---|---|
| Jan 2025 | 727,797 B | 182,205 B | 4.0x | 262,211 B | 2.8x |
| Feb 2025 | 723,641 B | 177,828 B | 4.1x | 257,843 B | 2.8x |
| Mar 2025 | 723,718 B | 177,992 B | 4.1x | 258,007 B | 2.8x |
| Apr 2025 | 691,860 B | 183,790 B | 3.8x | 263,785 B | 2.6x |
| May 2025 | 691,128 B | 183,164 B | 3.8x | 263,179 B | 2.6x |
| Jun 2025 | 691,260 B | 183,221 B | 3.8x | 263,227 B | 2.6x |
| Jul 2025 | 691,779 B | 184,010 B | 3.8x | 264,013 B | 2.6x |
| Aug 2025 | 697,562 B | 191,910 B | 3.6x | 271,895 B | 2.6x |
| Sep 2025 | 700,110 B | 190,749 B | 3.7x | 270,737 B | 2.6x |
| Oct 2025 | 702,511 B | 194,671 B | 3.6x | 274,668 B | 2.6x |
| Nov 2025 | 4,478,301 B | 126,578 B | 35.4x | 197,966 B | 22.6x |
| Dec 2025 | 2,948,191 B | 117,088 B | 25.2x | 188,503 B | 15.6x |
| TOTAL | 14.47MB | 2.09MB | 6.9x (14.5%) | 3.04MB | 4.8x (21.0%) |
Not a minimal gap — /ebook runs ~45% larger than /screen across the set (3.04MB vs 2.09MB). Recommendation: /screen, since Talbot’s stated bar is text-exact + images-can-degrade, and /screen’s degradation was already visually confirmed acceptable (page 1 of August statement — banner + all figures legible). /printer (300dpi) exists as a third, milder option (~8x on the worst-case file) if image fidelity ever needs to be higher than /screen gives.
- Text verified 100% lossless on every one of the 12 files:
page.get_text()string-equality, original vs. compressed, all pages, via PyMuPDF — despite Ghostscript emitting harmlessPDFDocEncoding XX is undefinedwarnings during conversion (font-encoding notices, not data loss — confirmed by the text match). - Root cause of the size gaps, confirmed: each BMO statement page embeds a 2550×300px header banner image — print-resolution artwork, wildly oversized for on-screen/archival use.
/screen’s downsampling is exactly the right fix; it’s not losing statement data, it’s discarding resolution the banner graphic never needed. - Image quality at
/screen: visually inspected page 1 of the August statement post-compression — banner gradient and all numbers/text fully legible (see task log for the rendered PNG).
Verification method (built into the tool, 2026-08-17)
Section titled “Verification method (built into the tool, 2026-08-17)”Original method (raw get_text() string-equality) had false positives — retired. Retesting across a broader document set (legacy BMO/Wealthsimple statements, an SSRN academic paper, Globe & Mail articles, a book excerpt) found two benign cases where Ghostscript’s re-encoding changes the extracted text without losing anything:
- Font-glyph-mapping artifacts — a legacy BMO Investorline statement’s original PDF had a font whose glyph-to-Unicode mapping made PyMuPDF misread “Dollars” as “Dolars” (missing letter); Ghostscript’s re-encoding incidentally fixed the extraction. Visually identical either way.
- Table re-flow whitespace — an SSRN paper’s table headers gained one extra blank line under re-encoding; no words or numbers changed.
Both would fail a strict string-equality check despite zero real data loss — and did, until this was caught in File-Naming-Vault-Wide-Rollout round 9, where a workaround for a permission-classifier block accidentally overwrote 22 flagged-but-actually-fine files without the mismatch skip-list. Files were confirmed fine content-wise, but the near-miss is why the check moved into the tool itself instead of staying an ad-hoc external script.
Current method (verify_pdf_text.py): page count must match, and every numeric token (\d[\d,]*\.?\d*) in the original must appear the same number of times in the compressed output — this is the actual invariant that matters for a financial reference artifact (per AI-Testing-Standards.md’s “gate on the document’s own facts” principle). Prose/whitespace is allowed to differ; a changed or missing number fails and the compressed file is deleted, never silently kept.
Tested clean (numbers exact, whitespace-only or zero diff) across: BMO Chequing/CreditCard/Investorline statements (2015–2026), Wealthsimple statements, an SSRN academic paper (15MB→1.6MB), a Globe & Mail article, and a book-excerpt PDF — 140 files, 0 failures, ~42% average size reduction on this mixed set.
(For the earlier lossless-only qpdf pass, pixel-hash (SHA-256 of rendered samples, 150–200 DPI) was also checked and matched exactly — see git/task history for that data.)