Skip to content

Process lessons for KB-OS ongoing project (KB operating system: SMTM, dept structure, routing, dashboards). Captures what bit us, what to do different. Portal: KB-OS.

  • Structural routing changes must update every folder-scoped view same pass (SMTM v5.2, 2026-06-25): dept-based task routing moved tasks into dept Tasks/ folders, but _DASHBOARD.md Bases queries stayed pinned to file.inFolder("Core/_WorkingOn/Tasks") — dept-routed tasks vanished from dashboard silently. When files move location, sweep every query/dashboard/hook/skill assuming old location, same change. Folder move not done until views reading that folder updated.

  • Dogfood new structure on real task immediately (2026-06-25): filing SMTM-v5.2 task itself into Core/Processes/Tasks/ (first task under new dept-routing rule) surfaced broken dashboard. Run new convention on real unit of work right away; gaps show in use, not review.

  • Vault-wide Bases queries: prefer path pattern over enumerating folders (2026-06-25): with 20+ /Tasks/ folders growing, file.path.contains("/Tasks/") auto-covers new dept folders zero maintenance; or of file.inFolder(...) brittle, needs edit per new dept. (Caveat: Bases renders only in Obsidian — validate query changes visually in-app.)

  • git mv doesn’t stage content edits made after move (2026-07-07): during Ongoing-Projects Core/<Dept>/Projects/ migration, several files git mv-ed then content fixed (broken internal links, stale paths) — edits sat unstaged until follow-up git status sweep caught them, almost shipped incomplete commit. After any git mv + content-edit sequence, re-run git status before commit — don’t assume rename staged everything.

  • Flat, project-named notes sidestep naming-collision debates free (2026-07-07): consolidating scattered Ongoing-Project portals into Core/<Dept>/Projects/, naming each note after own project/repo (ai-config.md, monorepo.md, my_backup.md) rather than uniform generic name (DASHBOARD.md for all) avoided shared-folder naming collision that blocked earlier uniform-naming attempt (KB-OS-portal-rename.md, 2026-07-06). When multiple items need portal notes in one folder, name-by-identity beats name-by-role when role name would collide.

  • Templater corruption — open item, root cause unconfirmed (2026-07-07): Templates/Folder Note.md and Templates/SMTM Task.md found with <% tp.file.title %> / <% tp.date.now(...) %> placeholders replaced by literal rendered values (“ai-config”, a date) — restored via git checkout. Best guess: Templater fired against template file itself rather than target note, maybe during heavy git mv churn while Obsidian had files open/cached. Not reproduced — if recurs, check whether templates_folder/enabled_templates_hotkeys in .obsidian/plugins/templater-obsidian/data.json point at file also open in active pane during folder restructure.

  • Obsidian Bases: nested and/or filter groups broken, not just unconfirmed (2026-07-07): Core/DASHBOARD.md “Current Focus” query nested or: (matching two path patterns) inside top-level and: list, to combine recency filter with two-path scope one block. Rendered as table (no visible error) but silently ignored file.mtime recency filter — every matching file showed regardless of age. Splitting into two separate flat and:-only blocks, exact same date filter unchanged, fixed immediately. Never nest boolean groups in Bases filters; use one flat block per distinct path scope instead. Full gotcha entry: KB-OS-Usage.md.

  • Isolate one variable at a time when system fails silently (2026-07-07): broken Current Focus query had two plausible bugs at once (missing status != "complete" filter, nested or: above). Fixing both same pass would leave date-arithmetic question genuinely unresolved either way. Fixing status filter and un-nesting boolean group together, then getting Talbot confirmation it worked, still one round short of proof — only isolated for certain once date filter itself left completely unchanged across broken and fixed versions, making nesting only variable that changed. When silently-failing system has multiple candidate causes, change minimum needed to test one hypothesis, get explicit confirmation before claiming which fix worked.

  • “Recently touched” view self-defeating during heavy-editing session (2026-07-07): companion “Recently Touched Projects” Bases block (mtime, last 14 days) showed almost every project-adjacent file in vault. Not query bug — session’s own restructuring work (repeated git mv, portal edits, JOB_DESCRIPTION updates) bumped mtime across most of /Projects/ same day. Removed block; zero signal. Recency-based “what’s active” views least trustworthy right after (or during) big batch-edit session — editor’s own housekeeping pollutes signal it tries to surface.

  • Rename-for-consistency can silently multiply duplicate-filename problem (2026-07-07): renaming _DASHBOARD.md → DASHBOARD.md vault-wide (Talbot’s explicit call, overriding initial recommendation against it) brought vault-wide count of files literally named DASHBOARD.md to 7 (Core, MBR, SDC, FSS, KB-OS, plus 2 renamed). Every reference touched during rename made path-qualified, but this now standing risk for any future bare [DASHBOARD](/fss/dashboard/) link (first-occurrence-wins resolution — see existing Focus-Pages precedent in Core/CLAUDE.md). Before rename increasing filename’s existing occurrence count, flag new collision surface explicitly — permanent characteristic of vault from that point on, not one-time migration cost.

  • Templater file regex beats per-folder templates for any vault that grows (2026-07-08): 6 explicit folder-template entries in data.json (one per Tasks/ folder) needed manual update every time new dept folder added. Switching enable_file_templates: true + regex .*/Tasks/.* covers all current and future Tasks/ folders one rule. Prefer file regex over folder templates whenever same template applies across class of folders that will grow.

  • One adaptive template + tp.system.suggester() beats N specialized templates (2026-07-08): temptation was create 3 separate template files (Standalone / Repo-backed / KB-native). One file with wizard simpler: one SSOT, no divergence risk, no “which template do I pick?” decision for user. When variant behavior maps cleanly to menu selection, use one file with wizard rather than multiple files.

  • tp.file.rename() in Templater enables auto-naming from description prompt (2026-07-08): calling await tp.file.rename(slug) during template execution renames file before content written. Combined with slug = desc.toLowerCase().replace(/\s+/g, "-"), generates canonical <project>-<slug> filenames automatically — no manual file rename needed after creation.

  • New Ongoing Project must register in 3 places; automate or one gets missed (2026-07-08): Ongoing-Projects.md registry, SMTM Smart template repoProjects array, KB portal note. First two originally manual — smart template’s project list went stale immediately on first new project. Fix: /setup-ongoing-project Step 4 now writes all three same pass. Any system with hardcoded list mirroring registry must have update automated in same tool that creates registry entry.

  • Curated dashboard rows go stale moment source task closes — cross-check against _active.md before publishing (2026-07-09): first draft of Biggest Rocks table listed “ai-config-rules-upgrade Phase 1” as active rock, sourced from 3-day-old task file describing it as running — but _active.md showed project completed, closed 2026-07-07. Curated (non-query) view trades silent-failure risk for staleness risk; refresh step must include verifying each row’s source still live. When hand-curating any “what’s active/biggest” view, verify every row against master tracker same pass.

  • Verify archive completeness at commit time, not just move time (2026-07-09, mbr-content-consolidation): mbr-cleanup.md confirmed present in archive/2026-07/ right after Phase-1 mv, but vanished by commit time (deleted outside session — Obsidian or manual). Caught only by checking committed tree (git ls-tree) against decision table, then recovered from git history, re-archived in amended commit. Archive sweep not verified until commit tree checked against file list; filesystem state at move time not durable evidence.

  • Check whether brainstorm file already birthed SSOT before extracting it line-by-line (2026-07-09): MBR-Upgrades-CC.md (1993L) looked like huge extraction job, but own final responses show it created Filter+Focus.md (all 53 rocks) and Executive-Summary.md — right verification was marker-grep of ~12 distinctive ideas against those SSOTs (10/12 hits; only “Lazy Mode” + one story detail uncaptured). For any mega-file, first ask “did this session’s output already get consolidated?” verify by sampling markers — extraction effort should target gap, not file.

  • Talbot’s inline edits to presented decision table are first-class CEO directives (2026-07-09): approval came back partly as checkbox answers, partly as edits inside Verdict column (“confirmed by Ta”, changed verdicts, “Move. See below (2)”, scope carve-outs like “but keep Strategy/Notes/”). Reading only ## Talbot Response section would miss several decisions. When decision table is deliverable, diff table itself next session — table is response surface, not static Claude output.

  • Extraction passes must specifically hunt for last unanswered Talbot Response (2026-07-09): myMBR-OS.md ended with Talbot Response (QuickBooks 2017→QBO migration, client-data privacy separation, what-data-lives-where) that never got Claude reply — real open infrastructure item buried 3 months in stale task file. Extracting only “ideas and decisions” would skip it — it was a question, not content. When closing stale task file, check whether final Talbot Response was ever answered; unanswered questions route to owning dept Inbox as open items.

  • Promoting project to “ongoing” activates portal/task-location conventions — can supersede earlier ad-hoc verdict (2026-07-09): rate-scanner’s approved verdict (“relocate task to Projects/rate-scanner/Tasks/”) executed, then invalidated hours later when Talbot’s next directive made rate-scanner registered ongoing project — 2026-07-07 conventions then required dissolving _WorkingOn/Projects/rate-scanner/ (ongoing ≠ finite folder), moving task to dept MBR/IT/Tasks/ (repo-backed rule). “Make X ongoing project” not just registry row; walk full convention chain (portal home, task location, live-state SSOT in repo), and when it overrides previously-approved placement, flag supersession explicitly rather than silently re-moving.

  • Obsidian Bases not: filter must nest inside and:, not alongside it (2026-07-10): not: as peer key to and: under filters: causes Bases semantic error (raw YAML rendered instead of table). When need both, not: goes inside and: list as nested item: - not:\n - condition. Confirmed via Grok snippet — Bases supports not:, just not as sibling to and:.

  • Editing archive files bumps file.mtime into time-window Bases queries (2026-07-10): any file edit — including sed for checkbox fixes — updates file.mtime. If Bases query filters file.mtime > now() - "14 days", bulk-editing old archive files pulls them into active view. Wrong-layer fix (editing files) made problem worse than original. Diagnose which query layer matching before touching files.

  • file.path.contains("/Tasks/") matches /Tasks/archive/ subfolders — always add exclusion (2026-07-10): archive files live at .../Tasks/archive/..., so standard task-listing filter also picks them up. Every task-listing Bases query needs explicit - not:\n - file.path.contains("/archive/") inside and: block. Added to Core/DASHBOARD.md and MBR/DASHBOARD.md.

  • “MCP-primary” marketing ≠ MCP-only — check CLI surface before treating as CLI>MCP violation (2026-07-13, KB-OS-memory-upgrade): CodeGraph positioned MCP-first, initial review flagged conflict with standing CLI-over-MCP preference. But reading CLI reference revealed codegraph explore returns identical output to codegraph_explore MCP tool — adopted CLI-only, no MCP server, no preference exception needed. Had caveat taken at face value tool might’ve been rejected or forced through MCP unnecessarily. Before accepting/rejecting dual-use agent tool on stated interface, verify actual CLI surface (--help / CLI docs); marketing framing routinely hides first-class CLI.

  • Subagent’s disk writes survive API-error termination even when closing text doesn’t (2026-07-13, graphify-fullvault-rerun): session hit Claude usage limit twice mid-run, killing several of 23 parallel extraction subagents with “terminated early” errors. But every subagent already called Write to save output JSON to disk before generating final text summary — API error only cut off closing narration, not actual work. Checking disk state (valid JSON, real node/edge counts) before re-dispatching found all 23 chunks already complete; re-running would’ve wasted full second pass. When Agent-tool subagent “fails,” check what it already wrote to disk before assuming work lost — write-then-summarize ordering means partial failures often not partial at all.

  • Spot-check for templated/auto-generated content before committing to expensive multi-subagent extraction (2026-07-13, graphify-fullvault-rerun): whole-vault graphify run’s file list included ~85 daily Rate-Scanner-Ops.md logs; diffing two samples showed near-byte-identical (same 7 sources, same structure, machine-generated). Excluding saved ~4 of 27 planned subagent chunks of deep-mode LLM extraction for content with near-zero unique-entity value — recall for that content already covered by qmd’s BM25 search. Before dispatching N subagents against large file list, diff couple files from any suspiciously-repetitive-looking cluster (daily logs, generated reports); excluding true templates nearly free, token savings scale with corpus size.

  • Task-file approvals don’t clear auto-mode classifier for self-modifying changes — ai-config’s post-commit hook deploys working tree, not commit (2026-07-14, KB-OS-resume-rate-limit): approval written inside task file’s ## Talbot Response not accepted as live confirmation for editing settings.wsl.json (hooks/settings = self-modifying; classifier wants explicit chat approval). Meanwhile committing only approved files triggered standing post-commit → deploy.sh hook, which copies whatever on disk — deliberately-unstaged settings change went live anyway as side effect of unrelated commit. For settings/hooks changes, get approval in chat, not just task file; never leave half-finished changes in ai-config working tree when committing anything — deploy hook publishes tree, not commit.

  • deploy.sh coverage per-platform, missing copy line fails silently (2026-07-14, KB-OS-resume-rate-limit): statusline-command.sh deployed correctly on WSL but Windows section of deploy.sh never had copy line for it — nothing errors when file simply not in list, Windows copy just goes stale. Caught only by diffing Windows target before closing round (fixed: b9fe886). When adding file to deploy.sh, add one copy line per platform section; when verifying deploy, diff file on every platform target, not just one touched.

  • Cited “pattern to mirror” isn’t guaranteed to itself be properly SSOT’d (2026-07-13, memory-phase3-sessionstart-hook): task said model new SessionStart hook on session-digest.sh — but that reference file itself hand-placed directly in ~/.claude/hooks/, untracked by git, no ai-config copy, no deploy.sh step at all. Building on it as-is would’ve silently propagated same gap into new hook. Fixed by migrating session-digest.sh into ai-config/claude/hooks/ (verified byte-identical via diff), adding missing deploy.sh step first. When told mirror existing pattern, verify that pattern’s own SSOT/deploy status before treating as safe foundation — citation as precedent doesn’t mean already correct.

  • Thin index doc (ROADMAP-style) restating content from own SSOT is structural drift trap — fix link-only, never restate (2026-07-15, mBR-Business-Plan-SSOT-Structure): MBR/Strategy/ROADMAP.md had grown own Mission/Vision wording, independently drifted from both Identity/Mission.md (actual SSOT) and Strategic Plan.md (correctly quoted SSOT) — third divergent copy nobody asked for. Fix wasn’t content correction; structural: ROADMAP now only links [Identity/Mission](/sdc/strategy/identity/mission/)/[Strategic Plan](/mbr/strategy/strategic-plan/), never restates them, so nowhere left for third copy to form. Any index/TOC doc sitting above content SSOT (ROADMAP above Strategic Plan, dashboard above policy doc, etc.) should link, never restate — restating invites silent redundant drift content fix alone won’t prevent recurring.

  • No desktop-control tool exists in Claude Code sessions; Obsidian CLI’s eval/screenshot is the actual render-verification path (2026-08-25, Obsidian-Guide-KB-OS-Upgrades): Talbot asked whether “computer use” could visually confirm a new Obsidian Bases file rendered correctly. Checked the tool list explicitly rather than assuming — no generic desktop/screen-control tool present; Playwright drives web browsers only, and Obsidian is an Electron desktop app, not a webpage, so Playwright can’t attach to it. The real mechanism is the official Obsidian CLI: eval runs JS inside the running app, a screenshot command returns a base64 PNG — same tool already flagged in KB-OS-Obsid-cli’s “not adopted (scoped)” verdict for a different question. WSL-to-Windows reachability of a locally-running Obsidian instance from that CLI is still unconfirmed — don’t claim visual verification happened until that’s actually tested.

  • A tool’s default scan-root assumes structural symmetry; a new homing decision can break it silently (2026-08-31, KB-OS-dept-structure): orgchart’s default --kb-core only walked Core/, correct while every wrapper SVP Staff file lived under a Core dept. Talbot’s later call to home SVP-Strategy/SVP-Risks/SVP-Offerings in SDC/ (not a new Core dept) broke that assumption invisibly — the tool would’ve kept running clean, just silently omitted the new agents. Caught only because the tree was re-checked against the actual file count after the homing decision, not because the tool errored. After any structural decision that changes where files live, re-verify a generator/scanner’s scope assumption before trusting its output — a clean run is not proof of completeness.

  • Move file overlap parallel session scope? Check git log dest path first — maybe already moved (2026-08-27, sdc-dept-scaffold-mbr-absorb): Investment-Debt-Providers.md move to SDC/Strategy/Research/ had already landed via a parallel sdc-levpro-sdmath-complete session (commit 2ce1cca) before this task got to it; git mv on top was a harmless no-op, but only confirmed safe after checking git log --oneline -- <dest> post-hoc. When two sessions run in parallel against the same repo, check git history at the destination path before theorizing about a stray duplicate — it may be a race, not corruption.

  • Designing a gate? Read every skill in the family first — the sibling’s gap decides the design (2026-08-31, kb-os-smtm-continuation-path-reliability): the task framed the question as Task-vs-Project construct — should multi-phase work move to SMTM Layer 2? The answer came from reading project-task-complete.md, which nobody had asked about: project closure had the identical missing-successor hole as task-complete.md. That single fact disproved the Project-retrofit option on evidence — a Layer-2 retrofit would not have prevented either documented incident, so the variable was never the construct, it was a missing close-time check. Asserting either way without reading the sibling would have been a coin flip dressed as analysis. Before designing any cross-cutting mechanism, read every skill in the family (task-* and project-*), not only the ones the task names — how the gap is distributed across siblings is often what picks the design. Full record: logs/2026-08-31_SMTM-Continuation-Gate.md.

  • Forced choice pre-filled from data → test the pre-filled value itself, not only the pointer it resolves (2026-09-23, KB-OS-ongoing-project-continuation): project-now verified all 12 ## Now pointers OK, yet 3 “next ROADMAP item” values wrong — stale KB-OS ROADMAP offered finished work, ai-config standing rule as checkbox, rate-scanner plain bullets offered nothing. Stale source = finished work offered as the one alternative.

  • qmd search dup-check: narrow exact-phrase query misses same idea, different words (2026-09-26, stray-file): searched “generic HTML artifact skill,” zero hit; missed Standalone-HTML-Artifact-Skill.md (2026-05-26, open, unstarted) — near-identical ask, different vocabulary (skill vs. tool, HTML artifact vs. artifact page). qmd is BM25 keyword-scored, not concept match — a dup-check on a project-idea/skill-scope question needs 2+ phrasings tried (synonym swap: skill/tool/utility, artifact/page/dashboard) plus a filename grep across Tasks folders, not one semantic query trusted as exhaustive.

  • Structural routing changes must update every folder-scoped view same pass → promoted to global CLAUDE.md 2026-06-26 (“WHEN FILES MOVE, UPDATE EVERY VIEW THAT READS THEM”)
  • Restructure audits must cover .obsidian/ plugin configs, not just skills/commands → promoted to ai-config/AGENTS.md 2026-07-07 (Skills Quality Standard section)
  • “MCP-primary” marketing ≠ MCP-only — verify CLI surface first → promoted to ai-config/AGENTS.md 2026-07-13 (CLI vs MCP section)
  • Task-file approvals ≠ live approval for self-modifying settings changes + post-commit deploy hook publishes working tree, not commit → promoted to ai-config/AGENTS.md 2026-07-14 (Hard Rules / SSOT section, after “De-duplicating into an SSOT”)
  • Evaluating a paid bundle vs build-it-yourself — check each component at its own source (GitHub/repo), not the sales page → promoted to ai-config/AGENTS.md 2026-08-25 (Hard Rules section, after “CHECK RELEASE ASSETS, NOT WRITE-UPS”)
  • A doc’s own version stamp can lag its own changelog — read the changelog tail before bumping → promoted to ai-config/AGENTS.md 2026-08-31 (Hard Rules section, after “A TOOL’S INSTALLER VERSION IS NOT ITS RUNNING VERSION”)
  • Agent-context upgrade? A/B it before adopting. Use a scratch copy, 2 cold agents, the same 5 questions, and a pass mark set beforehand (≥30% fewer reads/tokens). OKF indexes and descriptions on 87 notes: 5/5 in both arms, −2.6% tokens, +3 tool calls. Grep on a small corpus is already cheap, and indexes add navigation steps. (KB-OS-okf-evaluation, 2026-09-23 → log 2026-09-23_OKF-Evaluation)
  • A decision for Talbot goes inside the AskUserQuestion dialog, never in prose beside it → promoted to ai-config/AGENTS.md 2026-09-23 (Hard Rules section, before “A CLASSIFIER BLOCK OFTEN FLAGS REAL RISK”)