job_monitor
Section titled “job_monitor”Custom Python cron/task watchdog. Built over healthchecks.io (SaaS) — WSL always runs, no external dependency, richer diagnostics (exit code + staleness, not just “missed ping”).
Project
Section titled “Project”- Location:
~/utils/system/job_monitor/(WSL), standaloneuvpackage, own git repo - Installed as CLI:
uv tool install .→job-monitorin PATH (/home/ta/.local/bin/job-monitor) - Config:
config.yaml— jobs withmax_age_hoursthresholds
Quick Start
Section titled “Quick Start”job-monitor # full check + alert on failure, run daily via cron (6 AM)uv run job-monitor --no-alert # dry run, prints results, no email — use for verificationEach monitored cron job calls heartbeat.sh <name> $? on completion. main.py checks staleness + exit codes against config.yaml, fires CRITICAL alert via notify_manager, writes status.json.
Schedule reconciliation (added 2026-08-12)
Section titled “Schedule reconciliation (added 2026-08-12)”Heartbeat checking alone only sees jobs already declared in config.yaml — a job never added is invisible (hid the 4-month check_backups outage). schedules.py runs inside the daily pass and diffs the real triggers (WSL crontab -l + Windows schtasks /query) against config:
| Check | Catches |
|---|---|
| scheduled but UNMONITORED | real trigger, no config entry |
| monitored but NOT SCHEDULED | watched job whose trigger was deleted/disabled |
| CADENCE MISMATCH | max_age_hours shorter than real interval (false alarm), or >3x looser (dead job stays quiet for days) |
Join key: the heartbeat.sh <name> call, read from the cron line or Windows wrapper .bat. Jobs whose heartbeat is written elsewhere declare cron_match: / scheduled_as: instead. Deliberately-unmonitored triggers → reconcile.ignore_triggers. Cron intervals measured by simulating the spec over 400 days, not pattern-matched.
Config fields: monitored_since: (grace period, no “never ran” alert before first scheduled fire), scheduled_as:, cron_match:, unscheduled_ok:.
Gotcha: the monitor has no monitor of its own
Section titled “Gotcha: the monitor has no monitor of its own”job_monitor itself ran unmonitored from inception until 2026-08-12 — a dead monitor sends no alerts, indistinguishable from “all healthy.” Its own config.yaml entry is only a weak self-check (catches “ran, exited non-zero”). Real detector is disk_verify (disk_verify) — different trigger, own alert path, verifies status.json freshness (written at END of run, so a fresh timestamp proves completion).
Gotcha: an ignore_triggers entry is invisible to reconciliation, forever
Section titled “Gotcha: an ignore_triggers entry is invisible to reconciliation, forever”reconcile.ignore_triggers was built for jobs whose silent death costs nothing (desktop reminders). \TaTasks\WSL Keepalive was put there too, but for a different reason (it has no heartbeat to check) — and reconciliation treats every ignored trigger identically: it never looks at it again. Result (found 2026-09-17, my_backup-job-monitors-stale): the keepalive task itself stopped existing in Task Scheduler for 5+ weeks — the exact thing that holds the WSL2 VM up for pre-login cron jobs — and nothing noticed, because it was ignored, not monitored. Fix: reconcile.required_triggers (schedules.py, reconcile()) — a label in this list must exist as a live, enabled trigger in schtasks/cron regardless of heartbeat; missing → REQUIRED TRIGGER MISSING issue. A daemon/keepalive with no completion to report goes in required_triggers, not just ignore_triggers (it still needs ignore_triggers too, to suppress the generic “scheduled but unmonitored” complaint once it exists again).
Current jobs monitored (see config.yaml for authoritative list)
Section titled “Current jobs monitored (see config.yaml for authoritative list)”my_backup_daily, virus_scan, create_system_image, my_backup_full_maintenance, send_status_report, asset_history_update, mbr_health_check, mbr_daily_run, job_monitor (self), MonthlyDriveCleanup, send_manual_reminders.
Ignore-listed non-critical: \TaTasks\WSL Keepalive, \TaTasks\Focus Reminder-Mission, \TaTasks\Notify util, \TaTasks\test.
Related
Section titled “Related”- Dashboard follow-on (not yet built): Cron Health widget on MBR ops dashboard —
mbr-ops-dashboard-cron-health.md - KB task: my_backup-silent-failure-gaps
- Build history / decision rationale: commit
00d5ed2(initial),3da44d9(KB task files),fcdd475(reconciliation)