Skip to content

Two job_monitor failures from a screenshot (Full Maintenance exit 1, DiskSpaceCheck never-ran) — both root-caused, fixed, and verified live same session.

  • Full Maintenance (exit 1, last run 8/1) — no log evidence survived (backup.wrapper.log gone, backup.log had zero 8/1 entries, likely cleared during the 8/5 reliability-fix pass). Most probable cause: the 7-zip fatal-symlink bug (fixed 344e0f7, 2026-08-03) or a system-image bug fixed 2026-08-05 — both are inside the pipeline stages --maintenance full runs. Re-ran the real monthly job (cmd.exe /c my_backup-maintenance-task.bat) — completed clean, exit 0, all 8 stages OK.
  • DiskSpaceCheck (never ran) — not a bug, a scheduling gap. The disk-space-monitoring task (2026-08-06) rewrote diskcheck to write heartbeat.json, but the Windows Scheduled Task hadn’t fired since (last run 8/2, before the rewrite). Triggered it manually (schtasks /run /tn DiskSpaceCheck) — clean, exit 0, heartbeat written; ran disk-verify to relay it into job_monitor.
  • job_monitor heartbeat for my_backup_full_maintenance updated by hand to reflect the successful re-run (next scheduled run isn’t until 2026-09-01).
  • Talbot kept DiskSpaceCheck’s Task Scheduler trigger at weekly (Sunday 10AM), declining the daily-08:00 change disk-space-monitoring had recommended. Flagged once: the weekly cadence already missed a real 3-day, 12GB/day slide (64GB→29GB, 08-02→08-05) that a Sunday-only check couldn’t catch mid-week — up to 6 days of blind spot if that recurs. Decision stands; not re-litigated.

job_monitor status.json: all 11 jobs ok, overall healthy (was: degraded, 2 failures).

  • A “cleanup”/fix pass that touches log files can silently destroy the evidence for the bug it’s fixing — backup.wrapper.log and backup.log’s 8/1 entries were both gone by the time this task investigated, leaving only inference from commit timeline instead of a confirmed root cause. → my_backup/LESSONS.md
  • A monitored job’s heartbeat can lag its own monitoring-code rewrite until the next natural schedule fires — DiskSpaceCheck looked “never ran” for 2+ days after the 8/6 diskcheck heartbeat rewrite simply because its weekly Windows task hadn’t fired again yet. General rule: after wiring new heartbeat/monitoring code into an existing scheduled job, trigger it once manually rather than waiting for the natural schedule. → candidate for Utility-Reliability-Standards.md