my_backup: 3 CRITICAL alerts (7-Zip, job-monitor, virus scan)
Section titled “my_backup: 3 CRITICAL alerts (7-Zip, job-monitor, virus scan)”Trigger: 3 CRITICAL alert emails over 2 days — [CRITICAL] Virus Scan Alert (MpCmdRun.exe not found), 7-Zip: WSL creation failed, job-monitor my_backup_daily/my_backup_full_maintenance exit code 1.
Root causes and fixes
Section titled “Root causes and fixes”-
7-Zip WSL archive fatal failure — two things compounded:
- Case-variant
.claude/projects/dirs (-mnt-d-FSS-KB-Corevs-mnt-d-FSS-KB-core) → 7-Zip’s “Duplicate filename on disk” abort. Merged the 8 session files into the correctly-cased dir, removed the stale one. - Real underlying pattern: 7-Zip archiving
~/over\\wsl$treats a dangling symlink as a fatal error (rc≥2, “Creation failed”) but a valid-but-external symlink (venv shims,.codegraph/current, etc.) as only a tolerated warning (rc=1). Two dangling symlinks were the actual recurring trigger:~/.claude/debug/latest(stale since March) and~/.kb-source -> /mnt/d/FSS/KB/Business(dead since the KB dept-first restructure, confirmed unreferenced before removing). Both removed. - Shipped a permanent fix in
my_backup/src/my_backup/tasks.py: pre-flight scan for dangling symlinks (viawsl.exe ... find ~ -xtype l) before the 7-Zip step, auto-excluding any found so future occurrences degrade to a warning instead of a CRITICAL alert. Filtered the warning against existing excludes so it only flags genuinely new finds (first pass surfaced 44 pre-existing/already-excluded noise items; filtered version correctly showed just the 1 new one in testing). Tested against a synthetic broken symlink end-to-end before shipping. Also fixed 7-Zip’s error logging, which truncated stderr to 100–200 chars and hid which file actually caused the issue — now logs full stderr. - Commit:
my_backup344e0f7.
- Case-variant
-
job-monitor exit-1 alerts — same root cause as #1 (the 7-Zip step failing mid-pipeline); no separate bug.
-
Virus scan “MpCmdRun.exe not found” — the crontab had two
virus_scanentries: a correct weekly one (native WSL) and a broken duplicate monthly one invoked viacmd.exe(Windows-native).virus_scan.py’s Defender-path lookup uses WSL-only/mnt/c/...paths that don’t resolve under native Windows Python. Explains every non-Sunday day-1 failure in the history (Mar 1, May 1, Jul 1, Aug 1). Removed the broken duplicate cron entry — the weekly WSL-native scan already covers it.
Bonus finding: dead independent verifier
Section titled “Bonus finding: dead independent verifier”While applying Utility-Reliability-Standards, found check_backups.py — the independent-verifier script the standard itself calls for — already existed and worked correctly, but had no crontab entry at all; logs/check_backups.log hadn’t updated since March. Re-scheduled (WSL cron, daily 11:00 — 3hrs after the 08:00 main backup) and registered in job_monitor/config.yaml. Commit: job_monitor fac4062.
Side effect (disclosed, not a bug)
Section titled “Side effect (disclosed, not a bug)”Repeated live verification runs during this session burned through Backblaze B2’s daily transaction cap, triggering one additional (expected, self-resolving) Mirror: Backblaze B2 (Active) Failed alert. Not a pipeline defect.
Follow-up spun off
Section titled “Follow-up spun off”my_backup-reliability-standards-full — Talbot confirmed the remaining Utility-Reliability-Standards items (dry-run mode, test suite, escalation thresholds, annual fire drill, etc.) are warranted as their own task, not bundled into this bug-fix.
Lessons captured
Section titled “Lessons captured”5 lessons added to my_backup/LESSONS.md: dangling-vs-external symlink behavior in 7-Zip-over-\\wsl$; stderr truncation destroying diagnosability; exclude-filtering a noisy detection feature before shipping; “exists in code” ≠ “actually scheduled” for verifiers; duplicate cron entries invoking the same script from different platforms can silently break platform-specific path logic.