Lessons — KB-OS
Section titled “Lessons — KB-OS”Process lessons for KB-OS ongoing project (KB operating system: SMTM, dept structure, routing, dashboards). Captures what bit us, what to do different. Portal: KB-OS.
Lessons
Section titled “Lessons”-
Structural routing changes must update every folder-scoped view same pass (SMTM v5.2, 2026-06-25): dept-based task routing moved tasks into dept
Tasks/folders, but_DASHBOARD.mdBases queries stayed pinned tofile.inFolder("Core/_WorkingOn/Tasks")— dept-routed tasks vanished from dashboard silently. When files move location, sweep every query/dashboard/hook/skill assuming old location, same change. Folder move not done until views reading that folder updated. -
Dogfood new structure on real task immediately (2026-06-25): filing SMTM-v5.2 task itself into
Core/Processes/Tasks/(first task under new dept-routing rule) surfaced broken dashboard. Run new convention on real unit of work right away; gaps show in use, not review. -
Vault-wide Bases queries: prefer path pattern over enumerating folders (2026-06-25): with 20+
/Tasks/folders growing,file.path.contains("/Tasks/")auto-covers new dept folders zero maintenance;oroffile.inFolder(...)brittle, needs edit per new dept. (Caveat: Bases renders only in Obsidian — validate query changes visually in-app.) -
git mvdoesn’t stage content edits made after move (2026-07-07): during Ongoing-ProjectsCore/<Dept>/Projects/migration, several filesgit mv-ed then content fixed (broken internal links, stale paths) — edits sat unstaged until follow-upgit statussweep caught them, almost shipped incomplete commit. After anygit mv+ content-edit sequence, re-rungit statusbefore commit — don’t assume rename staged everything. -
Flat, project-named notes sidestep naming-collision debates free (2026-07-07): consolidating scattered Ongoing-Project portals into
Core/<Dept>/Projects/, naming each note after own project/repo (ai-config.md,monorepo.md,my_backup.md) rather than uniform generic name (DASHBOARD.mdfor all) avoided shared-folder naming collision that blocked earlier uniform-naming attempt (KB-OS-portal-rename.md, 2026-07-06). When multiple items need portal notes in one folder, name-by-identity beats name-by-role when role name would collide. -
Templater corruption — open item, root cause unconfirmed (2026-07-07):
Templates/Folder Note.mdandTemplates/SMTM Task.mdfound with<% tp.file.title %>/<% tp.date.now(...) %>placeholders replaced by literal rendered values (“ai-config”, a date) — restored viagit checkout. Best guess: Templater fired against template file itself rather than target note, maybe during heavygit mvchurn while Obsidian had files open/cached. Not reproduced — if recurs, check whethertemplates_folder/enabled_templates_hotkeysin.obsidian/plugins/templater-obsidian/data.jsonpoint at file also open in active pane during folder restructure. -
Obsidian Bases: nested
and/orfilter groups broken, not just unconfirmed (2026-07-07):Core/DASHBOARD.md“Current Focus” query nestedor:(matching two path patterns) inside top-leveland:list, to combine recency filter with two-path scope one block. Rendered as table (no visible error) but silently ignoredfile.mtimerecency filter — every matching file showed regardless of age. Splitting into two separate flatand:-only blocks, exact same date filter unchanged, fixed immediately. Never nest boolean groups in Bases filters; use one flat block per distinct path scope instead. Full gotcha entry:KB-OS-Usage.md. -
Isolate one variable at a time when system fails silently (2026-07-07): broken Current Focus query had two plausible bugs at once (missing
status != "complete"filter, nestedor:above). Fixing both same pass would leave date-arithmetic question genuinely unresolved either way. Fixing status filter and un-nesting boolean group together, then getting Talbot confirmation it worked, still one round short of proof — only isolated for certain once date filter itself left completely unchanged across broken and fixed versions, making nesting only variable that changed. When silently-failing system has multiple candidate causes, change minimum needed to test one hypothesis, get explicit confirmation before claiming which fix worked. -
“Recently touched” view self-defeating during heavy-editing session (2026-07-07): companion “Recently Touched Projects” Bases block (mtime, last 14 days) showed almost every project-adjacent file in vault. Not query bug — session’s own restructuring work (repeated
git mv, portal edits, JOB_DESCRIPTION updates) bumpedmtimeacross most of/Projects/same day. Removed block; zero signal. Recency-based “what’s active” views least trustworthy right after (or during) big batch-edit session — editor’s own housekeeping pollutes signal it tries to surface. -
Rename-for-consistency can silently multiply duplicate-filename problem (2026-07-07): renaming
_DASHBOARD.md→DASHBOARD.mdvault-wide (Talbot’s explicit call, overriding initial recommendation against it) brought vault-wide count of files literally namedDASHBOARD.mdto 7 (Core, MBR, SDC, FSS, KB-OS, plus 2 renamed). Every reference touched during rename made path-qualified, but this now standing risk for any future bare[DASHBOARD](/fss/dashboard/)link (first-occurrence-wins resolution — see existing Focus-Pages precedent inCore/CLAUDE.md). Before rename increasing filename’s existing occurrence count, flag new collision surface explicitly — permanent characteristic of vault from that point on, not one-time migration cost. -
Templater file regex beats per-folder templates for any vault that grows (2026-07-08): 6 explicit folder-template entries in
data.json(one perTasks/folder) needed manual update every time new dept folder added. Switchingenable_file_templates: true+ regex.*/Tasks/.*covers all current and futureTasks/folders one rule. Prefer file regex over folder templates whenever same template applies across class of folders that will grow. -
One adaptive template +
tp.system.suggester()beats N specialized templates (2026-07-08): temptation was create 3 separate template files (Standalone / Repo-backed / KB-native). One file with wizard simpler: one SSOT, no divergence risk, no “which template do I pick?” decision for user. When variant behavior maps cleanly to menu selection, use one file with wizard rather than multiple files. -
tp.file.rename()in Templater enables auto-naming from description prompt (2026-07-08): callingawait tp.file.rename(slug)during template execution renames file before content written. Combined withslug = desc.toLowerCase().replace(/\s+/g, "-"), generates canonical<project>-<slug>filenames automatically — no manual file rename needed after creation. -
New Ongoing Project must register in 3 places; automate or one gets missed (2026-07-08): Ongoing-Projects.md registry, SMTM Smart template
repoProjectsarray, KB portal note. First two originally manual — smart template’s project list went stale immediately on first new project. Fix:/setup-ongoing-projectStep 4 now writes all three same pass. Any system with hardcoded list mirroring registry must have update automated in same tool that creates registry entry. -
Curated dashboard rows go stale moment source task closes — cross-check against
_active.mdbefore publishing (2026-07-09): first draft of Biggest Rocks table listed “ai-config-rules-upgrade Phase 1” as active rock, sourced from 3-day-old task file describing it as running — but_active.mdshowed project completed, closed 2026-07-07. Curated (non-query) view trades silent-failure risk for staleness risk; refresh step must include verifying each row’s source still live. When hand-curating any “what’s active/biggest” view, verify every row against master tracker same pass. -
Verify archive completeness at commit time, not just move time (2026-07-09, mbr-content-consolidation):
mbr-cleanup.mdconfirmed present inarchive/2026-07/right after Phase-1mv, but vanished by commit time (deleted outside session — Obsidian or manual). Caught only by checking committed tree (git ls-tree) against decision table, then recovered from git history, re-archived in amended commit. Archive sweep not verified until commit tree checked against file list; filesystem state at move time not durable evidence. -
Check whether brainstorm file already birthed SSOT before extracting it line-by-line (2026-07-09):
MBR-Upgrades-CC.md(1993L) looked like huge extraction job, but own final responses show it createdFilter+Focus.md(all 53 rocks) andExecutive-Summary.md— right verification was marker-grep of ~12 distinctive ideas against those SSOTs (10/12 hits; only “Lazy Mode” + one story detail uncaptured). For any mega-file, first ask “did this session’s output already get consolidated?” verify by sampling markers — extraction effort should target gap, not file. -
Talbot’s inline edits to presented decision table are first-class CEO directives (2026-07-09): approval came back partly as checkbox answers, partly as edits inside Verdict column (“confirmed by Ta”, changed verdicts, “Move. See below (2)”, scope carve-outs like “but keep
Strategy/Notes/”). Reading only## Talbot Responsesection would miss several decisions. When decision table is deliverable, diff table itself next session — table is response surface, not static Claude output. -
Extraction passes must specifically hunt for last unanswered Talbot Response (2026-07-09):
myMBR-OS.mdended with Talbot Response (QuickBooks 2017→QBO migration, client-data privacy separation, what-data-lives-where) that never got Claude reply — real open infrastructure item buried 3 months in stale task file. Extracting only “ideas and decisions” would skip it — it was a question, not content. When closing stale task file, check whether final Talbot Response was ever answered; unanswered questions route to owning dept Inbox as open items. -
Promoting project to “ongoing” activates portal/task-location conventions — can supersede earlier ad-hoc verdict (2026-07-09): rate-scanner’s approved verdict (“relocate task to
Projects/rate-scanner/Tasks/”) executed, then invalidated hours later when Talbot’s next directive made rate-scanner registered ongoing project — 2026-07-07 conventions then required dissolving_WorkingOn/Projects/rate-scanner/(ongoing ≠ finite folder), moving task to deptMBR/IT/Tasks/(repo-backed rule). “Make X ongoing project” not just registry row; walk full convention chain (portal home, task location, live-state SSOT in repo), and when it overrides previously-approved placement, flag supersession explicitly rather than silently re-moving. -
Obsidian Bases
not:filter must nest insideand:, not alongside it (2026-07-10):not:as peer key toand:underfilters:causes Bases semantic error (raw YAML rendered instead of table). When need both,not:goes insideand:list as nested item:- not:\n - condition. Confirmed via Grok snippet — Bases supportsnot:, just not as sibling toand:. -
Editing archive files bumps
file.mtimeinto time-window Bases queries (2026-07-10): any file edit — includingsedfor checkbox fixes — updatesfile.mtime. If Bases query filtersfile.mtime > now() - "14 days", bulk-editing old archive files pulls them into active view. Wrong-layer fix (editing files) made problem worse than original. Diagnose which query layer matching before touching files. -
file.path.contains("/Tasks/")matches/Tasks/archive/subfolders — always add exclusion (2026-07-10): archive files live at.../Tasks/archive/..., so standard task-listing filter also picks them up. Every task-listing Bases query needs explicit- not:\n - file.path.contains("/archive/")insideand:block. Added toCore/DASHBOARD.mdandMBR/DASHBOARD.md. -
“MCP-primary” marketing ≠ MCP-only — check CLI surface before treating as CLI>MCP violation (2026-07-13, KB-OS-memory-upgrade): CodeGraph positioned MCP-first, initial review flagged conflict with standing CLI-over-MCP preference. But reading CLI reference revealed
codegraph explorereturns identical output tocodegraph_exploreMCP tool — adopted CLI-only, no MCP server, no preference exception needed. Had caveat taken at face value tool might’ve been rejected or forced through MCP unnecessarily. Before accepting/rejecting dual-use agent tool on stated interface, verify actual CLI surface (--help/ CLI docs); marketing framing routinely hides first-class CLI. -
Subagent’s disk writes survive API-error termination even when closing text doesn’t (2026-07-13, graphify-fullvault-rerun): session hit Claude usage limit twice mid-run, killing several of 23 parallel extraction subagents with “terminated early” errors. But every subagent already called
Writeto save output JSON to disk before generating final text summary — API error only cut off closing narration, not actual work. Checking disk state (valid JSON, real node/edge counts) before re-dispatching found all 23 chunks already complete; re-running would’ve wasted full second pass. When Agent-tool subagent “fails,” check what it already wrote to disk before assuming work lost — write-then-summarize ordering means partial failures often not partial at all. -
Spot-check for templated/auto-generated content before committing to expensive multi-subagent extraction (2026-07-13, graphify-fullvault-rerun): whole-vault graphify run’s file list included ~85 daily
Rate-Scanner-Ops.mdlogs; diffing two samples showed near-byte-identical (same 7 sources, same structure, machine-generated). Excluding saved ~4 of 27 planned subagent chunks of deep-mode LLM extraction for content with near-zero unique-entity value — recall for that content already covered by qmd’s BM25 search. Before dispatching N subagents against large file list, diff couple files from any suspiciously-repetitive-looking cluster (daily logs, generated reports); excluding true templates nearly free, token savings scale with corpus size. -
Task-file approvals don’t clear auto-mode classifier for self-modifying changes — ai-config’s
post-commithook deploys working tree, not commit (2026-07-14, KB-OS-resume-rate-limit): approval written inside task file’s## Talbot Responsenot accepted as live confirmation for editingsettings.wsl.json(hooks/settings = self-modifying; classifier wants explicit chat approval). Meanwhile committing only approved files triggered standingpost-commit→deploy.shhook, which copies whatever on disk — deliberately-unstaged settings change went live anyway as side effect of unrelated commit. For settings/hooks changes, get approval in chat, not just task file; never leave half-finished changes inai-configworking tree when committing anything — deploy hook publishes tree, not commit. -
deploy.shcoverage per-platform, missing copy line fails silently (2026-07-14, KB-OS-resume-rate-limit):statusline-command.shdeployed correctly on WSL but Windows section ofdeploy.shnever had copy line for it — nothing errors when file simply not in list, Windows copy just goes stale. Caught only by diffing Windows target before closing round (fixed:b9fe886). When adding file todeploy.sh, add one copy line per platform section; when verifying deploy,difffile on every platform target, not just one touched. -
Cited “pattern to mirror” isn’t guaranteed to itself be properly SSOT’d (2026-07-13, memory-phase3-sessionstart-hook): task said model new
SessionStarthook onsession-digest.sh— but that reference file itself hand-placed directly in~/.claude/hooks/, untracked by git, noai-configcopy, nodeploy.shstep at all. Building on it as-is would’ve silently propagated same gap into new hook. Fixed by migratingsession-digest.shintoai-config/claude/hooks/(verified byte-identical viadiff), adding missingdeploy.shstep first. When told mirror existing pattern, verify that pattern’s own SSOT/deploy status before treating as safe foundation — citation as precedent doesn’t mean already correct. -
Thin index doc (ROADMAP-style) restating content from own SSOT is structural drift trap — fix link-only, never restate (2026-07-15, mBR-Business-Plan-SSOT-Structure):
MBR/Strategy/ROADMAP.mdhad grown own Mission/Vision wording, independently drifted from bothIdentity/Mission.md(actual SSOT) andStrategic Plan.md(correctly quoted SSOT) — third divergent copy nobody asked for. Fix wasn’t content correction; structural: ROADMAP now only links[Identity/Mission](/sdc/strategy/identity/mission/)/[Strategic Plan](/mbr/strategy/strategic-plan/), never restates them, so nowhere left for third copy to form. Any index/TOC doc sitting above content SSOT (ROADMAP above Strategic Plan, dashboard above policy doc, etc.) should link, never restate — restating invites silent redundant drift content fix alone won’t prevent recurring. -
No desktop-control tool exists in Claude Code sessions; Obsidian CLI’s
eval/screenshot is the actual render-verification path (2026-08-25, Obsidian-Guide-KB-OS-Upgrades): Talbot asked whether “computer use” could visually confirm a new Obsidian Bases file rendered correctly. Checked the tool list explicitly rather than assuming — no generic desktop/screen-control tool present; Playwright drives web browsers only, and Obsidian is an Electron desktop app, not a webpage, so Playwright can’t attach to it. The real mechanism is the official Obsidian CLI:evalruns JS inside the running app, a screenshot command returns a base64 PNG — same tool already flagged inKB-OS-Obsid-cli’s “not adopted (scoped)” verdict for a different question. WSL-to-Windows reachability of a locally-running Obsidian instance from that CLI is still unconfirmed — don’t claim visual verification happened until that’s actually tested. -
A tool’s default scan-root assumes structural symmetry; a new homing decision can break it silently (2026-08-31, KB-OS-dept-structure):
orgchart’s default--kb-coreonly walkedCore/, correct while every wrapper SVP Staff file lived under a Core dept. Talbot’s later call to homeSVP-Strategy/SVP-Risks/SVP-OfferingsinSDC/(not a new Core dept) broke that assumption invisibly — the tool would’ve kept running clean, just silently omitted the new agents. Caught only because the tree was re-checked against the actual file count after the homing decision, not because the tool errored. After any structural decision that changes where files live, re-verify a generator/scanner’s scope assumption before trusting its output — a clean run is not proof of completeness. -
Move file overlap parallel session scope? Check git log dest path first — maybe already moved (2026-08-27, sdc-dept-scaffold-mbr-absorb):
Investment-Debt-Providers.mdmove toSDC/Strategy/Research/had already landed via a parallelsdc-levpro-sdmath-completesession (commit2ce1cca) before this task got to it;git mvon top was a harmless no-op, but only confirmed safe after checkinggit log --oneline -- <dest>post-hoc. When two sessions run in parallel against the same repo, check git history at the destination path before theorizing about a stray duplicate — it may be a race, not corruption. -
Designing a gate? Read every skill in the family first — the sibling’s gap decides the design (2026-08-31, kb-os-smtm-continuation-path-reliability): the task framed the question as Task-vs-Project construct — should multi-phase work move to SMTM Layer 2? The answer came from reading
project-task-complete.md, which nobody had asked about: project closure had the identical missing-successor hole astask-complete.md. That single fact disproved the Project-retrofit option on evidence — a Layer-2 retrofit would not have prevented either documented incident, so the variable was never the construct, it was a missing close-time check. Asserting either way without reading the sibling would have been a coin flip dressed as analysis. Before designing any cross-cutting mechanism, read every skill in the family (task-*andproject-*), not only the ones the task names — how the gap is distributed across siblings is often what picks the design. Full record:logs/2026-08-31_SMTM-Continuation-Gate.md. -
Forced choice pre-filled from data → test the pre-filled value itself, not only the pointer it resolves (2026-09-23, KB-OS-ongoing-project-continuation):
project-nowverified all 12## Nowpointers OK, yet 3 “next ROADMAP item” values wrong — stale KB-OS ROADMAP offered finished work, ai-config standing rule as checkbox, rate-scanner plain bullets offered nothing. Stale source = finished work offered as the one alternative. -
qmd searchdup-check: narrow exact-phrase query misses same idea, different words (2026-09-26, stray-file): searched “generic HTML artifact skill,” zero hit; missedStandalone-HTML-Artifact-Skill.md(2026-05-26, open, unstarted) — near-identical ask, different vocabulary (skill vs. tool, HTML artifact vs. artifact page). qmd is BM25 keyword-scored, not concept match — a dup-check on a project-idea/skill-scope question needs 2+ phrasings tried (synonym swap: skill/tool/utility, artifact/page/dashboard) plus a filename grep across Tasks folders, not one semantic query trusted as exhaustive.
Escalation Candidates
Section titled “Escalation Candidates”- Structural routing changes must update every folder-scoped view same pass → promoted to global CLAUDE.md 2026-06-26 (“WHEN FILES MOVE, UPDATE EVERY VIEW THAT READS THEM”)
- Restructure audits must cover
.obsidian/plugin configs, not just skills/commands → promoted toai-config/AGENTS.md2026-07-07 (Skills Quality Standard section) - “MCP-primary” marketing ≠ MCP-only — verify CLI surface first → promoted to
ai-config/AGENTS.md2026-07-13 (CLI vs MCP section) - Task-file approvals ≠ live approval for self-modifying settings changes +
post-commitdeploy hook publishes working tree, not commit → promoted toai-config/AGENTS.md2026-07-14 (Hard Rules / SSOT section, after “De-duplicating into an SSOT”) - Evaluating a paid bundle vs build-it-yourself — check each component at its own source (GitHub/repo), not the sales page → promoted to
ai-config/AGENTS.md2026-08-25 (Hard Rules section, after “CHECK RELEASE ASSETS, NOT WRITE-UPS”) - A doc’s own version stamp can lag its own changelog — read the changelog tail before bumping → promoted to
ai-config/AGENTS.md2026-08-31 (Hard Rules section, after “A TOOL’S INSTALLER VERSION IS NOT ITS RUNNING VERSION”) - Agent-context upgrade? A/B it before adopting. Use a scratch copy, 2 cold agents, the same 5 questions, and a pass mark set beforehand (≥30% fewer reads/tokens). OKF indexes and descriptions on 87 notes: 5/5 in both arms, −2.6% tokens, +3 tool calls. Grep on a small corpus is already cheap, and indexes add navigation steps. (KB-OS-okf-evaluation, 2026-09-23 → log 2026-09-23_OKF-Evaluation)
- A decision for Talbot goes inside the
AskUserQuestiondialog, never in prose beside it → promoted toai-config/AGENTS.md2026-09-23 (Hard Rules section, before “A CLASSIFIER BLOCK OFTEN FLAGS REAL RISK”)