Skip to content

Talbot, 2026-08-05: “Low disk space itself is a persistent issue and frustration. D: is a 2 TB drive, which should provide lots of room for my current needs, including backups. Let’s review existing issues related to disk space (C: and D:) and prep a dedicated task for the IT department to address this issue and upgrade the current monitoring system as needed.”

Surfaced from my_backup-verification-alert, where D: at 99% turned out to be masking two system-image retention bugs that would have failed the 2026-09-01 run.

This is the third occurrence of the same problem: 2026-02-06 (both drives exhausted → I/O starvation, 30-second file opens), 2026-07-21 (C: down to 7.9 GB, and diskcheck had been silently failing for 5 months), now 2026-08-05 (D: at 29 GB). See System-Maintenance LESSONS.md.

Current state — RESOLVED 2026-08-05, monitoring work remains

Section titled “Current state — RESOLVED 2026-08-05, monitoring work remains”

The acute space emergency that triggered this task is fixed. D: went from 29 GB free (99% full) to ~580 GB over the course of my_backup-verification-alert. What was done, and what the 743 GiB “unaccounted” gap turned out to be:

ActionReclaimed
VSS shadow-copy cap on D: lowered 559 GB → 50 GB~480 GB — this was the bulk of the mystery gap; vssadmin reported only 65.6 GB “used”
D:\DESKTOP-TA deleted (Windows Backup and Restore target; task disabled)245.7 GB
Two System_Images folders removed⚠️ these were NOT empty — see below

Do not repeat the folder deletion mistake. 2026JL01 and 2026AU01 were deleted on a false reading: WSL du/find, non-elevated PowerShell and WinDirStat all reported them as 0 bytes, because wbAdmin writes those folders with an ACL that denies traversal to unprivileged callers. They held real ~190 GB images. Full write-up in my_backup/LESSONS.md → “Empty” can mean “I lack permission to look”. Any disk tooling this task builds must read D: from an elevated context, or it will systematically under-report.

Recovery points now: 2026AU05 system image (186.0 GB, verified 2026-08-05) and 2026JN24 (216.9 GB).

Still true and still the point of this task: the slide from 64.1 GB (08-02) to 29 GB (08-05) — roughly 12 GB/day — ran for three days under a weekly monitor that never fired. Space was reclaimed by hand; nothing would catch the next one.

VolumeSizeFree (post-cleanup)
C:238 GB44 GB (82% used)
D:1.9 TB~580 GB
WSL ext4 (/)1007 GB911 GB (5% used) — unmonitored

1. Upgrade the monitor — diskcheck (the whole point of this task)

Section titled “1. Upgrade the monitor — diskcheck (the whole point of this task)”

Six concrete gaps, all found 2026-08-05:

  • Cadence is too coarse. Weekly (Sunday 09:00) cannot catch a 12 GB/day slide — D: will hit zero between checks. Move to daily, and wire it into job_monitor so the monitor is itself monitored (it silently failed for 5 months in 2026; the existing heartbeat gap is exactly what Utility-Reliability-Standards exists to prevent).
  • Rate-of-change detection. The useful alert is “D: is losing 12 GB/day — 2 days to full”, not “D: is below a line”. diskcheck.log already holds the history needed to compute this.
  • Thresholds are hardcoded in src/diskcheck/__init__.py (C: 20, D: 50) and are absolute GB only. Move to config.yaml and add a percentage floor — 50 GB is 25% of C: but 2.5% of D:. (Already on System-Maintenance UPGRADES.md as “diskcheck as a proper util”.)
  • The monitor was never broken — it was correct and too slow. It logged D: 64.1GB on 08-02, above its 50 GB threshold and therefore silent, and by 08-05 the drive was at 29 GB. A weekly check cannot see a 3-day slide. Post-cleanup D: is back to ~580 GB, so this task is about cadence and trend, not the threshold value.
  • Severity. A drive days-from-full alerts at WARNING. Escalate to CRITICAL below a hard floor or when the trend predicts exhaustion inside N days.
  • WSL’s ext4 volume is unmonitored — 1007 GB with 911 GB free today, but it is a separate VHDX that can fill independently, and a full WSL disk breaks everything on the Linux side.
  • No attribution in the alert. “D: low” doesn’t say what grew. Include the top N folders by delta since the previous run so the alert is actionable on its own.
  • Apply the util-quality items from UPGRADES.md while in here: config.yaml, structured logging (setup_logging), no hardcoded log path.
  • ⚠️ diskcheck also has the shared-drive venv exposure — its live Scheduled Task runs uv.exe run diskcheck from D:\... with no isolated UV_PROJECT_ENVIRONMENT. That fix is assigned to wsl-windows-boundary-guards; coordinate so the two tasks don’t both rewrite the scheduled task.

Talbot uses WinDirStat manually and asked whether dust (https://github.com/bootandy/dust) or something better should replace it.

Measured 2026-08-05, before recommending anything:

  • du -shx over the Windows volumes ran for more than 10 minutes and was still working through the second volume when it was stopped — it had completed D: top-level, D:\FSS and D:\Ta by then. Anything walking /mnt/d from WSL crosses the 9p filesystem, which makes per-file stat calls slow. This applies to dust, ncdu, gdu and du equally — the bottleneck is the mount, not the tool.
  • dust is not in the Ubuntu 24.04 repos (du-dust: no candidate) and cargo is not installed, so it needs a GitHub release binary.
  • In-repo alternatives that are available: ncdu 1.19, gdu 5.25.0, duf 0.8.1.
  • Neither WSL nor non-elevated PowerShell can read D:\DESKTOP-TA or D:\System Volume Information — the ~700 GB hole above. Any WSL-side tool will under-report D: by ~40%.

Recommendation to validate, not to assume — split by side of the boundary:

  • Windows drives (C:, D:) → evaluate WizTree rather than dust. It reads the NTFS MFT directly instead of walking the tree, which is typically orders of magnitude faster than WinDirStat on a 2 TB volume, and run elevated it can see the folders that defeated every measurement above. This is the tool that actually answers Talbot’s question.
  • WSL ext4 → gdu or dust (fast on native ext4, where the 9p penalty doesn’t apply).
  • Verify the WizTree speed claim on the real D: before recommending it — do not repeat a vendor claim as fact.
  • Register the chosen tool(s) in Core/IT/Utils/External/Utilities.md (Talbot asked for this explicitly). Follow the existing format — name, URL, one-line purpose, install command, usage. Add under ## Windows next to the existing WinDirStat entry, and under ## WSL for the Linux-side pick. Update or annotate the WinDirStat entry if it is superseded.
  • The ~700 GB is accounted for, and D: has a documented, deliberate allocation rather than an unexplained 99%.
  • A monitor that would have caught a 12 GB/day slide before the drive filled — daily, trend-aware, job_monitor-heartbeated, covering C:, D: and WSL ext4.
  • Alerts name the folders that grew, not just the drive.
  • A disk-usage tool chosen on measured evidence and registered in Utilities.md.
  • Talbot has a clear decision on whether Windows Backup and Restore stays enabled.
  • Origin + all measurements: my_backup-verification-alert, 2026-08-05.
  • Project SSOT: Core/IT/Projects/System-Maintenance/ — read STATUS.md (issues 1, 2, 5, 7) and LESSONS.md (the 2026-02-06 exhaustion and the 2026-07-21 silent-monitor root cause) before starting; much of the history is already written up there.
  • Deferred items this task should close: UPGRADES.md → “diskcheck as a proper util”.
  • Coordinate with: wsl-windows-boundary-guards (owns the DiskSpaceCheck wrapper/venv fix), my_backup-silent-failure-gaps. Note the create_system_image free-space pre-flight listed there was already shipped 2026-08-05 (my_backup 0bcbff3) and proved itself on the live run; the 7-Zip archive step still has none.
  • Repo: D:\FSS\Software\Utils\PythonUtils\diskcheck\ — thresholds currently in src/diskcheck/__init__.py, log at diskcheck.log, task name is DiskSpaceCheck (not diskcheck — that mismatch once made it invisible to schtasks lookups).

Architecture (confirmed with Talbot before building): split responsibility, mirroring my_backup’s defense-in-depth pattern — Windows diskcheck stays the main C:/D: check; a new WSL-side disk_verify companion owns WSL-ext4 monitoring, job_monitor heartbeat writes (kept WSL-native so it never depends on \\wsl$ being reachable), and independent verification that the Windows job actually ran. Folder attribution stays non-elevated (WizTree CLI) so the daily Task Scheduler job stays at RunLevel: Limited — no reopening the PATH-resolution failure in LESSONS.md.

Summary:

  • diskcheck rewritten — config.yaml (abs GB + % floor per drive), trend.py (trailing-window slope, days-to-exhaustion, WARNING/CRITICAL), attribution.py (WizTree CLI diff, top-5 grown folders), heartbeat.py (boundary-agnostic heartbeat.json on D:), logging_setup.py per Utility-Reliability-Standards. 13 unit tests pass (isolated venv outside the shared drive — never ran bare uv run inside the repo from WSL, per wsl-windows-boundary-guards’s landmine). Committed f00d730.
  • New disk_verify WSL util (~/utils/system/disk_verify/) — ext4 threshold check, reads the Windows heartbeat via /mnt/d for independent verification (treats “cannot verify” as a warning, never as “the Windows job is broken” — same pattern as check_backups.py), writes both job_monitor heartbeats. 8 tests pass, ran end-to-end for real (uv run disk-verify) — correctly logged the Windows heartbeat as “missing” (expected, new code hasn’t run on Windows yet) without false-escalating. Own git repo, committed 31a81a6. Registered in Core/IT/Utils/Custom/disk_verify.md.
  • job_monitor wired — DiskSpaceCheck + disk_verify added to config.yaml, max_age_hours: 30. Ran job-monitor — correctly flagged DiskSpaceCheck as “never ran” (accurate — awaiting the Windows-side run) and picked up disk_verify’s fresh heartbeat.
  • WSL cron added for disk_verify — daily 09:00 (1h after the Windows task’s target 08:00).
  • Tool evaluation, on measured evidence, not vendor claims: WizTree CLI export of the real D: (2TB, 1.17TB used, 146,531 files) took 93s non-elevated — du -shx over the same drives didn’t finish in 10+ minutes. WizTree registered in Utilities.md (superseding WinDirStat for scans, WinDirStat kept for its GUI treemap); gdu registered for WSL-ext4 (not yet installed — needs sudo, see Next Steps).
  • Attribution parser fixed against real WizTree output — the actual export has a Generated by WizTree ... banner line before the CSV header that my first draft didn’t handle; caught by running the real scan, not assumed from docs. Verified the fixed parser against the real captured CSV (correctly found D:\bak\ 814GB, D:\Ta\ 247GB, etc., and correctly shows System Volume Information as 0GB — the known non-elevated blind spot, not silently wrong).
  • ~/700 GB already accounted for and the Windows Backup/Restore call — not re-touched, per the task file’s own “already resolved” note.

What could NOT be tested from this WSL session (flagging per Testing Requirements, not claiming untested behavior works):

  • Real end-to-end run of the Windows-side diskcheck — shutil.disk_usage("C:\\") etc. don’t work from Linux, and running it via WSL would risk the shared-drive .venv poisoning. All Windows-side logic was unit-tested against injected fixtures instead. First real Windows run is unverified until Talbot (or a future elevated session) triggers it.
  • The WizTree CLI shellout inside diskcheck itself (only the parser was exercised against real WizTree output, not through diskcheck’s own attribute_growth() end-to-end on Windows).

Next Steps for Talbot:

  • Trigger the real DiskSpaceCheck task once (Start-ScheduledTask -TaskName "DiskSpaceCheck") and check logs/diskcheck.log + heartbeat.json appear in D:\FSS\Software\Utils\PythonUtils\diskcheck\ — closes the “unverified on Windows” gap above. If it errors, the log will say why (structured now, not silent). *
  • Change the Task Scheduler trigger to daily 08:00 (leaves Execute/RunLevel: Limited untouched — wsl-windows-boundary-guards owns rewriting Execute into an isolated wrapper, don’t let two tasks fight over the same task object):
    Terminal window
    $Trigger = New-ScheduledTaskTrigger -Daily -At 8:00AM
    Set-ScheduledTask -TaskName "DiskSpaceCheck" -Trigger $Trigger
    Get-ScheduledTaskInfo -TaskName "DiskSpaceCheck"
  • Install gdu (needs sudo, couldn’t run non-interactively): sudo apt install gdu *
  • Confirm Windows Backup and Restore stays disabled — background section says it was disabled during the 2026-08-05 cleanup; this task didn’t re-touch it. Just need your yes/no for the record. *
  • Optional: periodic elevated WizTree run if you want visibility into System_Images/System Volume Information growth occasionally — the daily automated check deliberately can’t see those (non-elevated, to keep the Task Scheduler job at RunLevel: Limited). Not automated; a manual “run WizTree as admin every month or so” habit would close the gap if you want it closed. *