Skip to content

Custom Python cron/task watchdog. Built over healthchecks.io (SaaS) — WSL always runs, no external dependency, richer diagnostics (exit code + staleness, not just “missed ping”).

  • Location: ~/utils/system/job_monitor/ (WSL), standalone uv package, own git repo
  • Installed as CLI: uv tool install . → job-monitor in PATH (/home/ta/.local/bin/job-monitor)
  • Config: config.yaml — jobs with max_age_hours thresholds
Terminal window
job-monitor # full check + alert on failure, run daily via cron (6 AM)
uv run job-monitor --no-alert # dry run, prints results, no email — use for verification

Each monitored cron job calls heartbeat.sh <name> $? on completion. main.py checks staleness + exit codes against config.yaml, fires CRITICAL alert via notify_manager, writes status.json.

Schedule reconciliation (added 2026-08-12)

Section titled “Schedule reconciliation (added 2026-08-12)”

Heartbeat checking alone only sees jobs already declared in config.yaml — a job never added is invisible (hid the 4-month check_backups outage). schedules.py runs inside the daily pass and diffs the real triggers (WSL crontab -l + Windows schtasks /query) against config:

CheckCatches
scheduled but UNMONITOREDreal trigger, no config entry
monitored but NOT SCHEDULEDwatched job whose trigger was deleted/disabled
CADENCE MISMATCHmax_age_hours shorter than real interval (false alarm), or >3x looser (dead job stays quiet for days)

Join key: the heartbeat.sh <name> call, read from the cron line or Windows wrapper .bat. Jobs whose heartbeat is written elsewhere declare cron_match: / scheduled_as: instead. Deliberately-unmonitored triggers → reconcile.ignore_triggers. Cron intervals measured by simulating the spec over 400 days, not pattern-matched.

Config fields: monitored_since: (grace period, no “never ran” alert before first scheduled fire), scheduled_as:, cron_match:, unscheduled_ok:.

Gotcha: the monitor has no monitor of its own

Section titled “Gotcha: the monitor has no monitor of its own”

job_monitor itself ran unmonitored from inception until 2026-08-12 — a dead monitor sends no alerts, indistinguishable from “all healthy.” Its own config.yaml entry is only a weak self-check (catches “ran, exited non-zero”). Real detector is disk_verify (disk_verify) — different trigger, own alert path, verifies status.json freshness (written at END of run, so a fresh timestamp proves completion).

Gotcha: an ignore_triggers entry is invisible to reconciliation, forever

Section titled “Gotcha: an ignore_triggers entry is invisible to reconciliation, forever”

reconcile.ignore_triggers was built for jobs whose silent death costs nothing (desktop reminders). \TaTasks\WSL Keepalive was put there too, but for a different reason (it has no heartbeat to check) — and reconciliation treats every ignored trigger identically: it never looks at it again. Result (found 2026-09-17, my_backup-job-monitors-stale): the keepalive task itself stopped existing in Task Scheduler for 5+ weeks — the exact thing that holds the WSL2 VM up for pre-login cron jobs — and nothing noticed, because it was ignored, not monitored. Fix: reconcile.required_triggers (schedules.py, reconcile()) — a label in this list must exist as a live, enabled trigger in schtasks/cron regardless of heartbeat; missing → REQUIRED TRIGGER MISSING issue. A daemon/keepalive with no completion to report goes in required_triggers, not just ignore_triggers (it still needs ignore_triggers too, to suppress the generic “scheduled but unmonitored” complaint once it exists again).

Current jobs monitored (see config.yaml for authoritative list)

Section titled “Current jobs monitored (see config.yaml for authoritative list)”

my_backup_daily, virus_scan, create_system_image, my_backup_full_maintenance, send_status_report, asset_history_update, mbr_health_check, mbr_daily_run, job_monitor (self), MonthlyDriveCleanup, send_manual_reminders.

Ignore-listed non-critical: \TaTasks\WSL Keepalive, \TaTasks\Focus Reminder-Mission, \TaTasks\Notify util, \TaTasks\test.

  • Dashboard follow-on (not yet built): Cron Health widget on MBR ops dashboard — mbr-ops-dashboard-cron-health.md
  • KB task: my_backup-silent-failure-gaps
  • Build history / decision rationale: commit 00d5ed2 (initial), 3da44d9 (KB task files), fcdd475 (reconciliation)