Skip to content

  • shows graphically folder/file sizes, to help manage system structure and backups
  • Superseded for large-volume scans (2026-08-06) by WizTree below — kept installed for its graphical treemap view, but WizTree is now diskcheck’s attribution engine and the recommended tool for a fresh full-drive look.
  • https://diskanalyzer.com/
  • Reads the NTFS MFT directly instead of walking the tree — orders of magnitude faster than WinDirStat/du on a large volume. Measured 2026-08-06: full D: scan (2TB drive, 1.17TB used, 146,531 files) via CLI export in 93s, non-elevated — vs. WSL du -shx over the same drives not finishing in 10+ minutes (disk-space-monitoring task).
  • Installed: C:\Program Files\WizTree\WizTree64.exe
  • CLI export usage: WizTree64.exe D:\ /export=out.csv — writes a CSV (File Name,Size,Allocated,Modified,Attributes,Files,Folders) with a Generated by WizTree ... banner line before the header. Non-elevated: cannot see wbAdmin-ACL’d folders (System_Images, System Volume Information) — read those with an elevated run instead.
  • Used by diskcheck’s attribution module (D:\FSS\Software\Utils\PythonUtils\diskcheck\src\diskcheck\attribution.py) to report which folders grew.

→ Changing tabs or workspace defaults? Core/IT/Utils/Custom/herdr-workspace.md → Configuration (first section) — not below.

Terminal multiplexer (tmux-like) with built-in AI-agent state awareness (blocked/working/idle/done sidebar). Runs agents in persistent terminals that survive sleep/restart; reconnect from any device. Apache 2.0. Replaces Talbot’s ~/.tmux/work.sh tmux setup as the daily driver, auto-starting on every new Ubuntu terminal via ~/.bashrc (exec herdr).

  • Site: https://herdr.dev/ · Docs: https://herdr.dev/docs/
  • Installed (0.8.0, ~/.local/bin/herdr) — 2026-08-04
  • Install: curl -fsSL https://herdr.dev/install.sh | sh (WSL/Linux)
  • Claude integration: herdr integration install claude
  • Config: ~/.config/herdr/config.toml (keybindings/theme/UI — static, edit freely, herdr server reload-config picks up changes)
  • Session state (workspaces/tabs/pane cwd): ~/.config/herdr/session.json — live server-persisted. Never hand-edit while the server is running; it is overwritten from memory on every change and read back only when the server starts fresh. Change it via in-app hotkeys or the CLI API (herdr tab|workspace <subcommand>), not by editing the file or by re-running the bootstrap script against an already-existing workspace (see Core/IT/Utils/Custom/herdr-workspace.md for the gotcha this caused).
  • Bootstrap script (workspace layout): ~/utils/herdr-workspace.sh — see Core/IT/Utils/Custom/herdr-workspace.md
  • Kill/restart server (test config changes without rebooting Windows): herdr server stop — stops the server via the API socket; next herdr invocation (new terminal, or run manually) starts a fresh server and re-reads config.toml. Live sessions in session.json are restored except any drift you were trying to clear — for state changes, use the CLI/hotkeys above instead of a restart.
  • Plugin — claude-session-title: mirrors Claude Code’s pane title (from /rename or auto-summary) into Herdr’s pane metadata, so the same name shows in both the terminal multiplexer and the Remote Control session list.
  • Related external repos: herdr-file-viewer (file viewer plugin), herdr-spreader (spreader plugin)
  • Fast ext4 disk-usage analyzer — no 9p-mount penalty when run against native WSL paths (unlike du/ncdu/dust crossing into /mnt/c or /mnt/d, which are equally slow regardless of tool — the bottleneck is the mount, not the tool).
  • https://github.com/dundee/gdu
  • Install: sudo apt install gdu (in Ubuntu 24.04 repos, not yet installed as of 2026-08-06 — see disk-space-monitoring task Next Steps)
  • Usage: gdu / (or any native WSL path) for an interactive TUI browser

CLI wrapper (Ghostscript /screen default, --lossless via qpdf, --batch for whole folders) — full findings, benchmarks, and usage: pdfcompress.

Evaluated 2026-08-09 against all 26 real FY26 BMO statements (13 credit-card + 13 chequing). Recommendation is split because the two jobs have opposite priorities: feeding an LLM rewards breadth and tolerates fuzziness; extracting figures for the books rewards determinism and demands loud failure.

JobUseWhy
General use — any document → markdown for AI/LLM reading, RAG, summarizingMarkItDownBroadest format coverage (PDF, DOCX, PPTX, XLSX, images w/ OCR, audio, HTML), one CLI, fast, already installed. Output shape varies between files, which is fine when an LLM reads it
Quality-first — extracting figures deterministically (accounting, anything reconciled)pdfplumberReturns the page’s own text lines in reading order, so a statement row is a line. No markdown layer to parse, therefore no layout variance to defend against
Scanned / image-only PDFs, complex multi-column layout, OCRDoclingThe only evaluated option that handles these. Reach for it when the others return nothing
”Is this PDF scanned or text-based?” triagepdf-inspectorFast, accurate classification. Do not use it to extract figures — see below
  • https://github.com/microsoft/markitdown · uv tool install 'markitdown[all]' (WSL, ~/.local/bin/markitdown), 2026-07-28
  • Usage: markitdown <file> → markdown to stdout; markitdown <file> -o out.md
  • Known limitation: its markdown shape varies per file for the same document type. Five of the accounting project’s eighteen gotchas came from that alone — multi-card table truncation, debit/credit column position, date-format cell variance, a statement rendered with no pipe delimiters at all, and an amount emitted on the line before its row (which silently deleted a real charge and invented a credit). All are cosmetic for LLM reading; all are dangerous for bookkeeping.
  • https://github.com/jsvine/pdfplumber · MIT · pure Python, no ML models · uv add pdfplumber
  • page.extract_text() gives clean lines; page.extract_tables() and word-level coordinates are there if structure is ever needed
  • Verified: all 26 statements produce clean one-row-per-line text, zero run-together amounts, 8.5s for the whole corpus
  • In use by ~/utils/acct-qbo/pdf_text.py (/hst-assistance, /process-expenses successors)
  • https://github.com/docling-project/docling · MIT · maintained by IBM Research Zurich
  • Genuinely capable: PDF/DOCX/PPTX/XLSX/HTML/EPUB/audio/video, VLM document understanding, strong OCR, LangChain/LlamaIndex integrations
  • Costs measured here: pip install docling pulls 5.2 GB (torch et al.), and it took 344 seconds for a single 3-page statement — ~2.5 hours for the corpus pdfplumber does in 8.5s. It also invoked OCR on a text-based PDF unprompted
  • It got all 5 figures right. It is simply the wrong shape of tool for clean text PDFs — keep it for scanned documents and genuinely hard layouts

pdf-inspector — do not use for financial extraction

Section titled “pdf-inspector — do not use for financial extraction”
  • https://github.com/firecrawl/pdf-inspector · MIT · Rust, with Node/Python/WASM bindings · npm install @firecrawl/pdf-inspector
  • Very fast, deterministic, no ML, and good at classifying scanned vs text-based
  • Disqualified for accounting on evidence: across 4 test statements it recovered the real dollar figures 3/3, 3/3, 2/3 — and 0/3 on BMO_2026-06-25.pdf, returning well-formed markdown containing no financial figures at all, through every API it exposes (extractText returned 2 characters). No error, no warning. Inconsistent silent data loss that looks like success is the single worst failure mode for bookkeeping — a loud failure is recoverable, a confident empty result is not.

The transferable lesson: for money, prefer the tool whose output you can gate on the document’s own arithmetic, and always build that gate. The accounting parsers cross-check every parse against the statement’s own printed subtotals; that gate is what caught both the orphaned-amount corruption and, once added to the bank side, a two-transactions-merged-into-one error that had been recorded as a confirmed fact for months.