Skip to content

OKF evaluation: don’t adopt (the A/B pilot showed no worthwhile gain)

Section titled “OKF evaluation: don’t adopt (the A/B pilot showed no worthwhile gain)”

Task: KB-OS-okf-evaluation (2 rounds, Rock #6), spun out of KB-OS-cleanup. Continuation: sdc-ip-drop-stale-after (SDC/IP/Projects/Strategies-Library/Tasks/), which removes stale_after from the Strategies Library and corrects the OKF claim in SCHEMA §9.

OKF (Open Knowledge Format) is not adopted vault-wide. That covers per-folder index.md files, backfilled description:, type:, log.md, and the trust fields. Talbot’s bar was “adopt selectively, after the pilot proves it’s worth it”. The pilot didn’t pass.

The test ran on a scratch copy of the pilot scope, Core/Processes/ minus Tasks/, Projects/, archive/, Templates/ and Logs/ (87 notes). Nothing in the vault was written.

  • Arm A: the folder as it is.
  • Arm B: the same folder with a drafted description: on all 87 notes and 18 generated OKF §8 index.md files. The agent was told to navigate by the indexes.

Each arm got a fresh agent with Read/Glob/Grep only and no qmd. Both got the same 5 questions, worded on purpose so they didn’t match the notes’ own terms.

Arm A (as is)Arm B (OKF)
Correct5/55/5
Tool calls69
Files opened36
Tokens72,80470,926 (−2.6%)

The pass mark, set before the run: the same answers with ≥30% fewer reads/tokens. It failed. Grep over a few hundred KB is already cheap, and the indexes added navigation steps. Arm B had one real win: “which design targets have no guidance” was answered from Design/Targets/index.md alone. Limits: one run per arm and 87 notes. The full 1,238-note vault is untested, but grep scales there too, and in real sessions agents also have qmd.

  • OKF status: repo GoogleCloudPlatform/open-knowledge-format, Apache-2.0, spec v0.2, no tagged releases. Since the 2026-08-21 commits, every timestamp must be an ISO datetime with an offset. §8 index.md files carry no frontmatter (only a bundle-root okf_version is allowed) and hold * [Title](link) - desc lists. The reference generator src/reference_agent/bundle/index.py uses Gemini to write folder descriptions.
  • Vault frontmatter coverage (2026-09-23): 1,238 notes. 98% have title:, 9% have description: (and 6 of the 14 in Core/Processes are placeholders), and 10% have type:, with inconsistent values.
  • Waypoint isn’t installed in either vault. The 9 folder notes with Map of Contents blocks are frozen and stale.
  • Why OKF index.md conflicts with KB-OS, in case this is ever revisited:
    • Any index.md in Tasks/ breaks SMTM. No-argument skills pick the newest .md, the Bases /Tasks/ filters pick it up, and the Templater regex .*/Tasks/.* injects the task wizard.
    • Every indexed note gains a backlink, which adds noise to Related-Notes.base.
    • Duplicate index basenames.
    • /audit-kb-os Step 2 checks only wikilinks.
    • qmd and graphify would need ignore entries.
    • SDC/IP/Strategies/**/index.md belong to strategy-lint and are read-only.
  • stale_after was dropped on Talbot’s call: “extra admin with no payoff”.
  • Folder Notes.md and Templates/Folder Note.md were corrected. Waypoint is documented as not installed, the Map of Contents blocks as frozen, and the create step now uses Notebook Navigator’s Create folder note.
  • Successor stub sdc-ip-drop-stale-after was created.

Lessons (approved and promoted 2026-09-23)

Section titled “Lessons (approved and promoted 2026-09-23)”

The lesson below went to KB-OS/LESSONS.md. A second lesson went to ai-config/AGENTS.md: a decision for Talbot goes inside the AskUserQuestion dialog, never in prose beside it. That one came from this close, where /task-complete 4b showed the lesson candidate as text beside the continuation dialog and it went unseen. The skill is fixed (ai-config c0fcad0).

Promoted lesson:

Agent-context upgrade? A/B it before adopting. Use a scratch copy, two cold agents, the same 5 questions, and a pass mark set beforehand (≥30% fewer reads/tokens). OKF indexes and descriptions on 87 notes: 5/5 in both arms, −2.6% tokens, +3 tool calls. Grep on a small corpus is already cheap, and indexes add navigation steps.