disk-space-monitoring
Section titled “disk-space-monitoring”Background
Section titled “Background”Talbot, 2026-08-05: “Low disk space itself is a persistent issue and frustration. D: is a 2 TB drive, which should provide lots of room for my current needs, including backups. Let’s review existing issues related to disk space (C: and D:) and prep a dedicated task for the IT department to address this issue and upgrade the current monitoring system as needed.”
Surfaced from my_backup-verification-alert, where D: at 99% turned out to be masking two system-image retention bugs that would have failed the 2026-09-01 run.
This is the third occurrence of the same problem: 2026-02-06 (both drives exhausted → I/O starvation, 30-second file opens), 2026-07-21 (C: down to 7.9 GB, and diskcheck had been silently failing for 5 months), now 2026-08-05 (D: at 29 GB). See System-Maintenance LESSONS.md.
Current state — RESOLVED 2026-08-05, monitoring work remains
Section titled “Current state — RESOLVED 2026-08-05, monitoring work remains”The acute space emergency that triggered this task is fixed. D: went from 29 GB free (99% full) to ~580 GB over the course of my_backup-verification-alert. What was done, and what the 743 GiB “unaccounted” gap turned out to be:
| Action | Reclaimed |
|---|---|
VSS shadow-copy cap on D: lowered 559 GB → 50 GB | ~480 GB — this was the bulk of the mystery gap; vssadmin reported only 65.6 GB “used” |
D:\DESKTOP-TA deleted (Windows Backup and Restore target; task disabled) | 245.7 GB |
Two System_Images folders removed | ⚠️ these were NOT empty — see below |
Do not repeat the folder deletion mistake. 2026JL01 and 2026AU01 were deleted on a false reading: WSL du/find, non-elevated PowerShell and WinDirStat all reported them as 0 bytes, because wbAdmin writes those folders with an ACL that denies traversal to unprivileged callers. They held real ~190 GB images. Full write-up in my_backup/LESSONS.md → “Empty” can mean “I lack permission to look”. Any disk tooling this task builds must read D: from an elevated context, or it will systematically under-report.
Recovery points now: 2026AU05 system image (186.0 GB, verified 2026-08-05) and 2026JN24 (216.9 GB).
Still true and still the point of this task: the slide from 64.1 GB (08-02) to 29 GB (08-05) — roughly 12 GB/day — ran for three days under a weekly monitor that never fired. Space was reclaimed by hand; nothing would catch the next one.
| Volume | Size | Free (post-cleanup) |
|---|---|---|
C: | 238 GB | 44 GB (82% used) |
D: | 1.9 TB | ~580 GB |
WSL ext4 (/) | 1007 GB | 911 GB (5% used) — unmonitored |
1. Upgrade the monitor — diskcheck (the whole point of this task)
Section titled “1. Upgrade the monitor — diskcheck (the whole point of this task)”Six concrete gaps, all found 2026-08-05:
- Cadence is too coarse. Weekly (Sunday 09:00) cannot catch a 12 GB/day slide —
D:will hit zero between checks. Move to daily, and wire it intojob_monitorso the monitor is itself monitored (it silently failed for 5 months in 2026; the existing heartbeat gap is exactly whatUtility-Reliability-Standardsexists to prevent). - Rate-of-change detection. The useful alert is “D: is losing 12 GB/day — 2 days to full”, not “D: is below a line”.
diskcheck.logalready holds the history needed to compute this. - Thresholds are hardcoded in
src/diskcheck/__init__.py(C: 20,D: 50) and are absolute GB only. Move toconfig.yamland add a percentage floor — 50 GB is 25% of C: but 2.5% of D:. (Already on System-MaintenanceUPGRADES.mdas “diskcheck as a proper util”.) - The monitor was never broken — it was correct and too slow. It logged
D: 64.1GBon 08-02, above its 50 GB threshold and therefore silent, and by 08-05 the drive was at 29 GB. A weekly check cannot see a 3-day slide. Post-cleanupD:is back to ~580 GB, so this task is about cadence and trend, not the threshold value. - Severity. A drive days-from-full alerts at
WARNING. Escalate toCRITICALbelow a hard floor or when the trend predicts exhaustion inside N days. - WSL’s ext4 volume is unmonitored — 1007 GB with 911 GB free today, but it is a separate VHDX that can fill independently, and a full WSL disk breaks everything on the Linux side.
- No attribution in the alert. “D: low” doesn’t say what grew. Include the top N folders by delta since the previous run so the alert is actionable on its own.
- Apply the util-quality items from
UPGRADES.mdwhile in here:config.yaml, structured logging (setup_logging), no hardcoded log path. - ⚠️
diskcheckalso has the shared-drive venv exposure — its live Scheduled Task runsuv.exe run diskcheckfromD:\...with no isolatedUV_PROJECT_ENVIRONMENT. That fix is assigned to wsl-windows-boundary-guards; coordinate so the two tasks don’t both rewrite the scheduled task.
2. Pick a disk-usage tool and register it
Section titled “2. Pick a disk-usage tool and register it”Talbot uses WinDirStat manually and asked whether dust (https://github.com/bootandy/dust) or something better should replace it.
Measured 2026-08-05, before recommending anything:
du -shxover the Windows volumes ran for more than 10 minutes and was still working through the second volume when it was stopped — it had completedD:top-level,D:\FSSandD:\Taby then. Anything walking/mnt/dfrom WSL crosses the 9p filesystem, which makes per-filestatcalls slow. This applies todust,ncdu,gduandduequally — the bottleneck is the mount, not the tool.dustis not in the Ubuntu 24.04 repos (du-dust: no candidate) andcargois not installed, so it needs a GitHub release binary.- In-repo alternatives that are available:
ncdu1.19,gdu5.25.0,duf0.8.1. - Neither WSL nor non-elevated PowerShell can read
D:\DESKTOP-TAorD:\System Volume Information— the ~700 GB hole above. Any WSL-side tool will under-report D: by ~40%.
Recommendation to validate, not to assume — split by side of the boundary:
- Windows drives (C:, D:) → evaluate WizTree rather than
dust. It reads the NTFS MFT directly instead of walking the tree, which is typically orders of magnitude faster than WinDirStat on a 2 TB volume, and run elevated it can see the folders that defeated every measurement above. This is the tool that actually answers Talbot’s question. - WSL ext4 →
gduordust(fast on native ext4, where the 9p penalty doesn’t apply). - Verify the WizTree speed claim on the real
D:before recommending it — do not repeat a vendor claim as fact. - Register the chosen tool(s) in
Core/IT/Utils/External/Utilities.md(Talbot asked for this explicitly). Follow the existing format — name, URL, one-line purpose, install command, usage. Add under## Windowsnext to the existing WinDirStat entry, and under## WSLfor the Linux-side pick. Update or annotate the WinDirStat entry if it is superseded.
Success Criteria
Section titled “Success Criteria”- The ~700 GB is accounted for, and
D:has a documented, deliberate allocation rather than an unexplained 99%. - A monitor that would have caught a 12 GB/day slide before the drive filled — daily, trend-aware,
job_monitor-heartbeated, covering C:, D: and WSL ext4. - Alerts name the folders that grew, not just the drive.
- A disk-usage tool chosen on measured evidence and registered in
Utilities.md. - Talbot has a clear decision on whether Windows Backup and Restore stays enabled.
Context
Section titled “Context”- Origin + all measurements: my_backup-verification-alert, 2026-08-05.
- Project SSOT:
Core/IT/Projects/System-Maintenance/— readSTATUS.md(issues 1, 2, 5, 7) andLESSONS.md(the 2026-02-06 exhaustion and the 2026-07-21 silent-monitor root cause) before starting; much of the history is already written up there. - Deferred items this task should close:
UPGRADES.md→ “diskcheck as a proper util”. - Coordinate with: wsl-windows-boundary-guards (owns the
DiskSpaceCheckwrapper/venv fix), my_backup-silent-failure-gaps. Note thecreate_system_imagefree-space pre-flight listed there was already shipped 2026-08-05 (my_backup0bcbff3) and proved itself on the live run; the 7-Zip archive step still has none. - Repo:
D:\FSS\Software\Utils\PythonUtils\diskcheck\— thresholds currently insrc/diskcheck/__init__.py, log atdiskcheck.log, task name isDiskSpaceCheck(notdiskcheck— that mismatch once made it invisible toschtaskslookups).
Claude Response — 2026-08-06
Section titled “Claude Response — 2026-08-06”Architecture (confirmed with Talbot before building): split responsibility, mirroring my_backup’s defense-in-depth pattern — Windows diskcheck stays the main C:/D: check; a new WSL-side disk_verify companion owns WSL-ext4 monitoring, job_monitor heartbeat writes (kept WSL-native so it never depends on \\wsl$ being reachable), and independent verification that the Windows job actually ran. Folder attribution stays non-elevated (WizTree CLI) so the daily Task Scheduler job stays at RunLevel: Limited — no reopening the PATH-resolution failure in LESSONS.md.
Summary:
-
diskcheckrewritten —config.yaml(abs GB + % floor per drive),trend.py(trailing-window slope, days-to-exhaustion, WARNING/CRITICAL),attribution.py(WizTree CLI diff, top-5 grown folders),heartbeat.py(boundary-agnosticheartbeat.jsonon D:),logging_setup.pyperUtility-Reliability-Standards. 13 unit tests pass (isolated venv outside the shared drive — never ran bareuv runinside the repo from WSL, perwsl-windows-boundary-guards’s landmine). Committedf00d730. - New
disk_verifyWSL util (~/utils/system/disk_verify/) — ext4 threshold check, reads the Windows heartbeat via/mnt/dfor independent verification (treats “cannot verify” as a warning, never as “the Windows job is broken” — same pattern ascheck_backups.py), writes bothjob_monitorheartbeats. 8 tests pass, ran end-to-end for real (uv run disk-verify) — correctly logged the Windows heartbeat as “missing” (expected, new code hasn’t run on Windows yet) without false-escalating. Own git repo, committed31a81a6. Registered inCore/IT/Utils/Custom/disk_verify.md. -
job_monitorwired —DiskSpaceCheck+disk_verifyadded toconfig.yaml,max_age_hours: 30. Ranjob-monitor— correctly flaggedDiskSpaceCheckas “never ran” (accurate — awaiting the Windows-side run) and picked updisk_verify’s fresh heartbeat. - WSL cron added for
disk_verify— daily 09:00 (1h after the Windows task’s target 08:00). - Tool evaluation, on measured evidence, not vendor claims: WizTree CLI export of the real
D:(2TB, 1.17TB used, 146,531 files) took 93s non-elevated —du -shxover the same drives didn’t finish in 10+ minutes. WizTree registered inUtilities.md(superseding WinDirStat for scans, WinDirStat kept for its GUI treemap);gduregistered for WSL-ext4 (not yet installed — needssudo, see Next Steps). - Attribution parser fixed against real WizTree output — the actual export has a
Generated by WizTree ...banner line before the CSV header that my first draft didn’t handle; caught by running the real scan, not assumed from docs. Verified the fixed parser against the real captured CSV (correctly foundD:\bak\814GB,D:\Ta\247GB, etc., and correctly showsSystem Volume Informationas 0GB — the known non-elevated blind spot, not silently wrong). -
~/700 GB already accounted forand the Windows Backup/Restore call — not re-touched, per the task file’s own “already resolved” note.
What could NOT be tested from this WSL session (flagging per Testing Requirements, not claiming untested behavior works):
- Real end-to-end run of the Windows-side
diskcheck—shutil.disk_usage("C:\\")etc. don’t work from Linux, and running it via WSL would risk the shared-drive.venvpoisoning. All Windows-side logic was unit-tested against injected fixtures instead. First real Windows run is unverified until Talbot (or a future elevated session) triggers it. - The WizTree CLI shellout inside
diskcheckitself (only the parser was exercised against real WizTree output, not throughdiskcheck’s ownattribute_growth()end-to-end on Windows).
Next Steps for Talbot:
- Trigger the real
DiskSpaceChecktask once (Start-ScheduledTask -TaskName "DiskSpaceCheck") and checklogs/diskcheck.log+heartbeat.jsonappear inD:\FSS\Software\Utils\PythonUtils\diskcheck\— closes the “unverified on Windows” gap above. If it errors, the log will say why (structured now, not silent). * - Change the Task Scheduler trigger to daily 08:00 (leaves
Execute/RunLevel: Limiteduntouched —wsl-windows-boundary-guardsowns rewritingExecuteinto an isolated wrapper, don’t let two tasks fight over the same task object):Terminal window $Trigger = New-ScheduledTaskTrigger -Daily -At 8:00AMSet-ScheduledTask -TaskName "DiskSpaceCheck" -Trigger $TriggerGet-ScheduledTaskInfo -TaskName "DiskSpaceCheck" - Install
gdu(needs sudo, couldn’t run non-interactively):sudo apt install gdu* - Confirm Windows Backup and Restore stays disabled — background section says it was disabled during the 2026-08-05 cleanup; this task didn’t re-touch it. Just need your yes/no for the record. *
- Optional: periodic elevated WizTree run if you want visibility into
System_Images/System Volume Informationgrowth occasionally — the daily automated check deliberately can’t see those (non-elevated, to keep the Task Scheduler job atRunLevel: Limited). Not automated; a manual “run WizTree as admin every month or so” habit would close the gap if you want it closed. *