OKF evaluation: don’t adopt (the A/B pilot showed no worthwhile gain)
Section titled “OKF evaluation: don’t adopt (the A/B pilot showed no worthwhile gain)”Task: KB-OS-okf-evaluation (2 rounds, Rock #6), spun out of KB-OS-cleanup.
Continuation: sdc-ip-drop-stale-after (SDC/IP/Projects/Strategies-Library/Tasks/), which removes stale_after from the Strategies Library and corrects the OKF claim in SCHEMA §9.
Decision
Section titled “Decision”OKF (Open Knowledge Format) is not adopted vault-wide. That covers per-folder index.md files, backfilled description:, type:, log.md, and the trust fields. Talbot’s bar was “adopt selectively, after the pilot proves it’s worth it”. The pilot didn’t pass.
The pilot: a cold-agent A/B test
Section titled “The pilot: a cold-agent A/B test”The test ran on a scratch copy of the pilot scope, Core/Processes/ minus Tasks/, Projects/, archive/, Templates/ and Logs/ (87 notes). Nothing in the vault was written.
- Arm A: the folder as it is.
- Arm B: the same folder with a drafted
description:on all 87 notes and 18 generated OKF §8index.mdfiles. The agent was told to navigate by the indexes.
Each arm got a fresh agent with Read/Glob/Grep only and no qmd. Both got the same 5 questions, worded on purpose so they didn’t match the notes’ own terms.
| Arm A (as is) | Arm B (OKF) | |
|---|---|---|
| Correct | 5/5 | 5/5 |
| Tool calls | 6 | 9 |
| Files opened | 3 | 6 |
| Tokens | 72,804 | 70,926 (−2.6%) |
The pass mark, set before the run: the same answers with ≥30% fewer reads/tokens. It failed. Grep over a few hundred KB is already cheap, and the indexes added navigation steps. Arm B had one real win: “which design targets have no guidance” was answered from Design/Targets/index.md alone. Limits: one run per arm and 87 notes. The full 1,238-note vault is untested, but grep scales there too, and in real sessions agents also have qmd.
Facts established (don’t re-derive)
Section titled “Facts established (don’t re-derive)”- OKF status: repo
GoogleCloudPlatform/open-knowledge-format, Apache-2.0, spec v0.2, no tagged releases. Since the 2026-08-21 commits, every timestamp must be an ISO datetime with an offset. §8index.mdfiles carry no frontmatter (only a bundle-rootokf_versionis allowed) and hold* [Title](link) - desclists. The reference generatorsrc/reference_agent/bundle/index.pyuses Gemini to write folder descriptions. - Vault frontmatter coverage (2026-09-23): 1,238 notes. 98% have
title:, 9% havedescription:(and 6 of the 14 inCore/Processesare placeholders), and 10% havetype:, with inconsistent values. - Waypoint isn’t installed in either vault. The 9 folder notes with Map of Contents blocks are frozen and stale.
- Why OKF
index.mdconflicts with KB-OS, in case this is ever revisited:- Any
index.mdinTasks/breaks SMTM. No-argument skills pick the newest.md, the Bases/Tasks/filters pick it up, and the Templater regex.*/Tasks/.*injects the task wizard. - Every indexed note gains a backlink, which adds noise to
Related-Notes.base. - Duplicate
indexbasenames. /audit-kb-osStep 2 checks only wikilinks.- qmd and graphify would need ignore entries.
SDC/IP/Strategies/**/index.mdbelong tostrategy-lintand are read-only.
- Any
stale_afterwas dropped on Talbot’s call: “extra admin with no payoff”.
What changed
Section titled “What changed”Folder Notes.mdandTemplates/Folder Note.mdwere corrected. Waypoint is documented as not installed, the Map of Contents blocks as frozen, and the create step now uses Notebook Navigator’s Create folder note.- Successor stub
sdc-ip-drop-stale-afterwas created.
Lessons (approved and promoted 2026-09-23)
Section titled “Lessons (approved and promoted 2026-09-23)”The lesson below went to KB-OS/LESSONS.md. A second lesson went to ai-config/AGENTS.md: a decision for Talbot goes inside the AskUserQuestion dialog, never in prose beside it. That one came from this close, where /task-complete 4b showed the lesson candidate as text beside the continuation dialog and it went unseen. The skill is fixed (ai-config c0fcad0).
Promoted lesson:
Agent-context upgrade? A/B it before adopting. Use a scratch copy, two cold agents, the same 5 questions, and a pass mark set beforehand (≥30% fewer reads/tokens). OKF indexes and descriptions on 87 notes: 5/5 in both arms, −2.6% tokens, +3 tool calls. Grep on a small corpus is already cheap, and indexes add navigation steps.