my_backup-failed-issues
Section titled “my_backup-failed-issues”Two job_monitor failures from a screenshot (Full Maintenance exit 1, DiskSpaceCheck never-ran) — both root-caused, fixed, and verified live same session.
What was found
Section titled “What was found”- Full Maintenance (exit 1, last run 8/1) — no log evidence survived (
backup.wrapper.loggone,backup.loghad zero 8/1 entries, likely cleared during the 8/5 reliability-fix pass). Most probable cause: the 7-zip fatal-symlink bug (fixed344e0f7, 2026-08-03) or a system-image bug fixed 2026-08-05 — both are inside the pipeline stages--maintenance fullruns. Re-ran the real monthly job (cmd.exe /c my_backup-maintenance-task.bat) — completed clean, exit 0, all 8 stages OK. - DiskSpaceCheck (never ran) — not a bug, a scheduling gap. The
disk-space-monitoringtask (2026-08-06) rewrotediskcheckto writeheartbeat.json, but the Windows Scheduled Task hadn’t fired since (last run 8/2, before the rewrite). Triggered it manually (schtasks /run /tn DiskSpaceCheck) — clean, exit 0, heartbeat written; randisk-verifyto relay it intojob_monitor. job_monitorheartbeat formy_backup_full_maintenanceupdated by hand to reflect the successful re-run (next scheduled run isn’t until 2026-09-01).
Decision
Section titled “Decision”- Talbot kept
DiskSpaceCheck’s Task Scheduler trigger at weekly (Sunday 10AM), declining the daily-08:00 changedisk-space-monitoringhad recommended. Flagged once: the weekly cadence already missed a real 3-day, 12GB/day slide (64GB→29GB, 08-02→08-05) that a Sunday-only check couldn’t catch mid-week — up to 6 days of blind spot if that recurs. Decision stands; not re-litigated.
Result
Section titled “Result”job_monitor status.json: all 11 jobs ok, overall healthy (was: degraded, 2 failures).
Lessons (candidates — not yet promoted)
Section titled “Lessons (candidates — not yet promoted)”- A “cleanup”/fix pass that touches log files can silently destroy the evidence for the bug it’s fixing —
backup.wrapper.logandbackup.log’s 8/1 entries were both gone by the time this task investigated, leaving only inference from commit timeline instead of a confirmed root cause. →my_backup/LESSONS.md - A monitored job’s heartbeat can lag its own monitoring-code rewrite until the next natural schedule fires —
DiskSpaceChecklooked “never ran” for 2+ days after the 8/6diskcheckheartbeat rewrite simply because its weekly Windows task hadn’t fired again yet. General rule: after wiring new heartbeat/monitoring code into an existing scheduled job, trigger it once manually rather than waiting for the natural schedule. → candidate forUtility-Reliability-Standards.md