my_backup: Daily Backup silently failing 8 days
Section titled “my_backup: Daily Backup silently failing 8 days”Date: 2026-07-26
What happened
Section titled “What happened”my_backup STATUS check flagged “no backup runs detected in last week.” The Windows Scheduled Task \TaTasks\Daily Backup had been failing (exit code 2) every day from 07-19 through 07-26.
Root cause
Section titled “Root cause”.venv’s own pyvenv.cfg showed a Linux Python home — only a WSL uv can write that. Cross-referenced against git log + Core/IT/Logs/2026-07-18_my_backup-notification-email-investigation.md:20: on 2026-07-18 at 16:45, a manual uv run send_status_report (run from WSL, verifying an unrelated fix) rebuilt the project’s default .venv — which lives on the shared D: drive both WSL and Windows touch — as a Linux venv. Windows uv.exe couldn’t sync it afterward (Access is denied removing the Linux lib64 symlink), so the Windows Scheduled Task failed silently every day after.
Not a code regression — the WSL cron path was never at risk (it already used UV_PROJECT_ENVIRONMENT=/home/ta/.venvs/my_backup, off D:). The Windows task called bare uv.exe run my_backup with no override, so it silently depended on the ambiguous shared default path. That architecture was the actual root cause; the manual command was just the trigger.
- Deleted the poisoned
.venv, rebuilt clean, verified a full backup run (exit 0). - Guardrail: gave the Windows Scheduled Task its own isolated
UV_PROJECT_ENVIRONMENT(C:\Users\Admin\.venvs\my_backup, offD:) via a new wrapper scriptD:\FSS\Software\Utils\Windows\my_backup-daily-task.bat, mirroring the isolation WSL already had. Neither production entrypoint touches the shared default.venvanymore — a future strayuv runin the repo dir just poisons an unused folder. - Verified the guardrail under the real execution context (
schtasks /run, Admin/InteractiveToken), not just a manual re-run: exit 0, isolated venv confirmed Windows-native, unchanged after the run. - Disabled Maximizer SQL dump (
config.yaml) — deprecated 2026-07-22, Twenty CRM is SSOT — and removed it from the Critical Files health check so it doesn’t alarm forever. - Registered
my_backup_dailyinjob_monitor(max_age_hours: 26) — this job was never monitored, which is why an 8-day outage took a week to surface instead of alarming same-day.
Gotchas discovered (promoted to GlobalDevRules.md § 3.9)
Section titled “Gotchas discovered (promoted to GlobalDevRules.md § 3.9)”- Shared-drive default
.venvis a silent WSL/Windows collision trap when neither entrypoint setsUV_PROJECT_ENVIRONMENT. schtasks /query /xmloutput piped through WSL declaresencoding="UTF-16"but writes plain single-byte bytes — re-importing via/create /xmlfails with “unable to switch the encoding” until properly re-encoded.
Commits
Section titled “Commits”my_backupe8a85e0(Maximizer disable), CHANGELOG updatedjob_monitorafe81db(heartbeat registration)- KB
f898e72
Follow-up (deferred, not blocking)
Section titled “Follow-up (deferred, not blocking)”Artifact-freshness checking (independent verifier pattern) for my_backup_daily — Talbot confirmed weekly status-report cadence is sufficient for now.