Skip to content

WSL-side companion to Windows diskcheck. Owns everything that shouldn’t live on the Windows side of the boundary: WSL ext4 monitoring (native — no 9p mount penalty), independent verification that the Windows DiskSpaceCheck task actually ran, and all job_monitor heartbeat writes for disk monitoring (kept WSL-native on purpose, so it never depends on \\wsl$ being reachable from a Windows process).

Built as part of disk-space-monitoring (2026-08-06) — see that task for the architecture decision (Windows main check + WSL independent verifier, mirroring my_backup’s 2am/8am defense-in-depth pattern documented in Utility-Reliability-Standards).

  • Location: ~/utils/system/disk_verify/
  • Package: uv-managed, WSL-native (safe to uv run directly — not on the shared-drive exposure list in wsl-windows-boundary-guards)
Terminal window
cd ~/utils/system/disk_verify
uv sync
uv run disk-verify
  1. WSL ext4 free space (/) — absolute GB + percentage floor, config.yaml.
  2. Windows diskcheck heartbeat — reads /mnt/d/FSS/Software/Utils/PythonUtils/diskcheck/heartbeat.json (boundary-agnostic, always reachable via /mnt/d regardless of \\wsl$ state). If stale beyond max_staleness_hours (default 30h), fires a CRITICAL — this is the exact “monitor silently stopped running” failure class that motivated the whole task. Read failures (missing/unreadable) are logged as warnings, never escalated to “the Windows job is broken” — mirrors check_backups.py’s “cannot verify != verified bad” pattern.

Writes two heartbeats every run, both from WSL-native code:

  • DiskSpaceCheck — timestamp copied from the Windows heartbeat’s last_run, so job_monitor’s own staleness math reflects the Windows job’s actual freshness
  • disk_verify — its own run

Registered in ~/utils/system/job_monitor/config.yaml, max_age_hours: 30 each.

Daily WSL cron, 09:00 — 1h after the Windows DiskSpaceCheck task (target 08:00, pending Talbot’s Task Scheduler change — see disk-space-monitoring Next Steps):

0 9 * * * cd /home/ta/utils/system/disk_verify && uv run disk-verify >> logs/cron.log 2>&1

Update 2026-08-12 — daily Windows free-space + job_monitor liveness

Section titled “Update 2026-08-12 — daily Windows free-space + job_monitor liveness”

DiskSpaceCheck is weekly, not daily. The Task Scheduler entry runs Sundays; max_staleness_hours was 30. That guaranteed a false “task may have stopped running” alert 5-6 days out of every 7 — it fired a CRITICAL on 2026-08-12 and was also the cause of that day’s second alert (disk_verify exit 1). Limit raised to 192h. The pending “change the Windows task to daily 08:00” item from disk-space-monitoring is resolved differently and closed: the task stays weekly (Talbot’s call).

Daily coverage now comes from windows_drives: — disk_verify reads the Windows volumes directly through drvfs (/mnt/c, /mnt/d report real volume figures), so the number that matters is checked daily with no second Task Scheduler job. Floors: C: 20GB/5%, D: 150GB/5% (D: needs headroom for a ~72GB monthly system image). Measured 2026-08-12: C: 35.7GB free (15%), D: 560.9GB free (30%) — the “D: at 99%” figure in older notes is stale, superseded by the 480GB reclaimed on 2026-08-05.

Independent job_monitor liveness check (job_monitor.status_path): a monitor cannot detect its own death — if WSL cron stops, every alert stops with it. disk_verify runs on a different trigger with its own notify_manager path, so it is the outside party. Alerts if status.json is older than 30h.

KB task: my_backup-silent-failure-gaps · commit 71620e5