Skip to content

Built Option B (custom Python utility) over Option A (healthchecks.io SaaS). Rationale: WSL always runs, external dependency concern, notify_manager already in stack, richer diagnostics (exit code + staleness vs. just “missed ping”).

/home/ta/utils/system/job_monitor/ — standalone uv Python package, git repo initialized.

FilePurpose
heartbeat.shCalled by each cron job: ; heartbeat.sh <name> $?
src/job_monitor/main.pyChecks staleness + exit codes, fires CRITICAL alert, writes status.json
config.yaml7 jobs with max_age_hours thresholds
README.mdSetup + “adding a new job” guide

Installed as CLI tool: job-monitor in PATH.

JobMax Age
virus_scan192h (8 days)
create_system_image840h (35 days)
my_backup_full_maintenance840h (35 days)
send_status_report192h (8 days)
asset_history_update192h (8 days)
mbr_health_check26h
mbr_daily_run26h

All 7 jobs got ; heartbeat.sh <name> $? appended. New daily monitor entry:

0 6 * * * /home/ta/.local/bin/job-monitor >> /home/ta/utils/system/job_monitor/logs/job_monitor.log 2>&1

Every run writes status.json (overall: ok/degraded, per-job status/exit_code/issue). Follow-on task drafted to add Cron Health widget to MBR ops dashboard: mbr-ops-dashboard-cron-health.md.

First 6 AM run alerts all 7 jobs “never ran” — expected, not a failure. Heartbeats populate as jobs run on their normal schedules.

  • job_monitor repo: 00d5ed2 — initial utility
  • KB Business: 3da44d9 — task files

Schedule reconciliation (added 2026-08-12)

Section titled “Schedule reconciliation (added 2026-08-12)”

src/job_monitor/schedules.py — runs inside the daily 14:00 pass. Heartbeat checking only sees jobs already declared in config.yaml; a job that was never added is invisible (this is what hid the four-month check_backups outage). Reconciliation enumerates the real triggers — WSL crontab -l plus Windows schtasks /query — and diffs them against the config:

CheckCatches
scheduled but UNMONITOREDa real trigger with no config entry
monitored but NOT SCHEDULEDa watched job whose trigger was deleted or disabled
CADENCE MISMATCHmax_age_hours shorter than the real interval (guaranteed false alarm), or >3x looser (a dead job stays quiet for days)

Join key is the heartbeat.sh <name> call, read from the cron line or from inside the Windows wrapper .bat. Jobs whose heartbeat is written elsewhere declare cron_match: / scheduled_as: instead. Deliberately-unmonitored triggers go in reconcile.ignore_triggers. Cron intervals are measured by simulating the spec over 400 days, not pattern-matched.

Config fields added: monitored_since: (grace period so a newly-added job does not alert “never ran” before its first scheduled fire), scheduled_as:, cron_match:, unscheduled_ok:.

CLI: uv run job-monitor --no-alert runs the full check and prints results without sending email — use this for verification instead of the bare command.

job_monitor ran unmonitored from inception until 2026-08-12. A dead monitor sends no alerts, and that silence is indistinguishable from “all healthy”. Its own config entry is only a weak self-check (catches “ran and exited non-zero”). The real detector is disk_verify, which runs on a different trigger with its own alert path and verifies status.json freshness — written at the END of a run, so a fresh timestamp proves completion.

job_monitor (self), MonthlyDriveCleanup (heartbeat emitted by cleanup-c-drive.ps1), send_manual_reminders. Ignore-listed as non-critical: \TaTasks\Focus Reminder-Mission, \TaTasks\Notify util, \TaTasks\test.

KB task: my_backup-silent-failure-gaps · commit fcdd475