Skip to content

2026-08-05 — my_backup false verification alert + D: disk emergency

Section titled “2026-08-05 — my_backup false verification alert + D: disk emergency”

Daily CRITICAL “Backup Verification” alert, contradicting a successful backup log. Root cause + 10 rounds of follow-on work, escalating into a same-day disk-space emergency and a near-miss data-loss mistake.

Root cause of the false alert (3 independent defects)

Section titled “Root cause of the false alert (3 independent defects)”
  1. check_backups re-added to WSL cron as bare uv run (2026-08-03) — ran under Linux Python, which can’t see D:/UNC paths. 11 of 12 reported errors were phantom.
  2. B2 offsite mirror checked with os.path.exists() on an rclone remote — always false. B2 had never actually been verified since the mirror was added.
  3. Duplicate config blocks (mirrors: vs health_checks “Mirror Freshness”) checked the same files with different severity rules — one [SKIP], one CRITICAL, same run.

Fix: platform guard (check_environment() refuses to run off Windows, exit 2, distinct “MISCONFIGURED” alert); check_rclone_mirror_age() replacing the broken B2 check; duplicate health-check group removed; verifier rescheduled via my_backup-check-task.bat through cmd.exe at 11:00. Full detail: repo CHANGELOG.md (commits 98e83d1, 21711ac).

Second failure found: monthly system image silently empty since March

Section titled “Second failure found: monthly system image silently empty since March”

create_system_image() trusted wbAdmin’s exit 0 without measuring the artifact — two of the last three “successful” monthly images were 0 bytes. Retention also sorted image folders alphabetically (AU < JL < JN), which would have deleted the newest real image on the next prune. Fixed: artifact size verification, chronological sort, create-then-prune ordering, free-space pre-flight, retention floor of 1. Commits 1b40783, 0943915, cc3fa92, 0bcbff3.

Traced to Windows’ own Backup and Restore (D:\DESKTOP-TA, 245.7 GB) running redundantly alongside my_backup, plus VSS shadow storage capped at 559 GB (30% of the drive) actually consuming ~480 GB. Windows Backup disabled, VSS cap dropped to 50 GB, DESKTOP-TA deleted after confirming a newer verified system image existed. End state: D: 29 GB → 607 GB free.

Near-miss: deleted two real system images on a false “empty” reading

Section titled “Near-miss: deleted two real system images on a false “empty” reading”

2026JL01/2026AU01 read as 0 bytes from WSL du, unelevated PowerShell, and WinDirStat — all three tools shared the same blind spot (wbAdmin’s ACL denies traversal to unprivileged readers; Path.rglob() swallows PermissionError and returns empty). Deleted based on that false corroboration. The tell was visible and missed: reclaimed space (480 GB) vastly exceeded VSS’s own reported max (~16 GB) — a reconciliation check that would have caught it before the delete. Guarded in code (os.scandir + explicit PermissionError handling; unreadable = occupied, never empty). Commit a149725. Full writeup in repo LESSONS.md — worth reading before any future “is this folder empty” decision on a Windows drive from WSL.

  • Notepad++ configured to default new files to Unix LF (fixes the CRLF-from-editor half of the WSL/Windows line-ending problem); TextPad confirmed to have no equivalent setting.
  • retention_count set to 2, held at 2 per Talbot (recovery depth valued over disk space; upgrade drive first if needed).

pytest 18/18. Live dry runs clean (0 errors, 0 warnings). 11:00 verification email arrived clean on 2026-08-06 — no CRITICAL alert. D: settled at 579+ GB free with two verified system images (2026AU05 186 GB, 2026JN24 216.9 GB).

  • wsl-windows-boundary-guards — generalize the wrong-side-interpreter guard, CRLF detect/repair util, PreToolUse hook for .venv poisoning, extend to rename_receipts/diskcheck/folder_structure.
  • disk-space-monitoring (Core/IT/Projects/System-Maintenance/Tasks/) — weekly monitor still can’t see a 12 GB/day slide; carries the ACL-blind-spot warning forward.
  • my_backup-silent-failure-gaps — scheduler↔job_monitor reconciliation (highest leverage — this is what hid the 4-month verification outage), 7-Zip case-variant pre-flight, NAS-robocopy-fatal-despite-optional bug, snapshot-size-drop guard untested in production, DST edge in staleness check.

Lesson candidates (repo LESSONS.md already carries full detail — checked via qmd, no dup elsewhere in vault)

Section titled “Lesson candidates (repo LESSONS.md already carries full detail — checked via qmd, no dup elsewhere in vault)”
  • “Empty” can mean “I lack permission to look” — never delete on an unreadable-equals-empty reading from tools sharing one access mechanism. Candidate for AGENTS.md’s existing “verify mass updates against raw storage” rule if this pattern recurs outside my_backup.