Skip to content

my_backup: 3 CRITICAL alerts (7-Zip, job-monitor, virus scan)

Section titled “my_backup: 3 CRITICAL alerts (7-Zip, job-monitor, virus scan)”

Trigger: 3 CRITICAL alert emails over 2 days — [CRITICAL] Virus Scan Alert (MpCmdRun.exe not found), 7-Zip: WSL creation failed, job-monitor my_backup_daily/my_backup_full_maintenance exit code 1.

  1. 7-Zip WSL archive fatal failure — two things compounded:

    • Case-variant .claude/projects/ dirs (-mnt-d-FSS-KB-Core vs -mnt-d-FSS-KB-core) → 7-Zip’s “Duplicate filename on disk” abort. Merged the 8 session files into the correctly-cased dir, removed the stale one.
    • Real underlying pattern: 7-Zip archiving ~/ over \\wsl$ treats a dangling symlink as a fatal error (rc≥2, “Creation failed”) but a valid-but-external symlink (venv shims, .codegraph/current, etc.) as only a tolerated warning (rc=1). Two dangling symlinks were the actual recurring trigger: ~/.claude/debug/latest (stale since March) and ~/.kb-source -> /mnt/d/FSS/KB/Business (dead since the KB dept-first restructure, confirmed unreferenced before removing). Both removed.
    • Shipped a permanent fix in my_backup/src/my_backup/tasks.py: pre-flight scan for dangling symlinks (via wsl.exe ... find ~ -xtype l) before the 7-Zip step, auto-excluding any found so future occurrences degrade to a warning instead of a CRITICAL alert. Filtered the warning against existing excludes so it only flags genuinely new finds (first pass surfaced 44 pre-existing/already-excluded noise items; filtered version correctly showed just the 1 new one in testing). Tested against a synthetic broken symlink end-to-end before shipping. Also fixed 7-Zip’s error logging, which truncated stderr to 100–200 chars and hid which file actually caused the issue — now logs full stderr.
    • Commit: my_backup 344e0f7.
  2. job-monitor exit-1 alerts — same root cause as #1 (the 7-Zip step failing mid-pipeline); no separate bug.

  3. Virus scan “MpCmdRun.exe not found” — the crontab had two virus_scan entries: a correct weekly one (native WSL) and a broken duplicate monthly one invoked via cmd.exe (Windows-native). virus_scan.py’s Defender-path lookup uses WSL-only /mnt/c/... paths that don’t resolve under native Windows Python. Explains every non-Sunday day-1 failure in the history (Mar 1, May 1, Jul 1, Aug 1). Removed the broken duplicate cron entry — the weekly WSL-native scan already covers it.

While applying Utility-Reliability-Standards, found check_backups.py — the independent-verifier script the standard itself calls for — already existed and worked correctly, but had no crontab entry at all; logs/check_backups.log hadn’t updated since March. Re-scheduled (WSL cron, daily 11:00 — 3hrs after the 08:00 main backup) and registered in job_monitor/config.yaml. Commit: job_monitor fac4062.

Repeated live verification runs during this session burned through Backblaze B2’s daily transaction cap, triggering one additional (expected, self-resolving) Mirror: Backblaze B2 (Active) Failed alert. Not a pipeline defect.

my_backup-reliability-standards-full — Talbot confirmed the remaining Utility-Reliability-Standards items (dry-run mode, test suite, escalation thresholds, annual fire drill, etc.) are warranted as their own task, not bundled into this bug-fix.

5 lessons added to my_backup/LESSONS.md: dangling-vs-external symlink behavior in 7-Zip-over-\\wsl$; stderr truncation destroying diagnosability; exclude-filtering a noisy detection feature before shipping; “exists in code” ≠ “actually scheduled” for verifiers; duplicate cron entries invoking the same script from different platforms can silently break platform-specific path logic.