Skip to content

Hyperresearch pilot: evaluation, install, and first full-tier run

Section titled “Hyperresearch pilot: evaluation, install, and first full-tier run”

Evaluated jordan-gibbs/hyperresearch (v0.11.1) as a candidate addition to the AI toolkit alongside the existing custom /research skill. Kept and installed — added to Core/AI/Skills/DASHBOARD.md Research section. Verdict: install standalone (doesn’t touch/replace /research), reserve for high-stakes questions, budget a full day per run, and plan to personally drive the last mile (steps 15–16 + run finish) — the pipeline does not self-finish.

  • Killed a false urgency premise twice: the task’s original “$110 API credit expiring” framing didn’t apply (hyperresearch runs on CC session usage, not the API-key pool); a real but separate CA$100 CC Pro promotional-credit expiry did apply and justified running the pilot when it did.
  • Ran /skill-security-review on the cloned source before install (external GitHub tool, no source-tier auto-pass per AGENTS.md). PASS — SSRF-safe HTTP layer, explicit prompt-injection fencing of fetched content, no injection language in its own skill markdown, reputable dependencies.
  • Installed via uv tool install hyperresearch (not raw pip) + hyperresearch install (wired 19 skill files, 15 subagent defs, a PreToolUse hook — left the existing /research skill untouched).
  • Piloted on a real $MART DEBT / SDC question: whether leveraged/borrow-to-invest strategies produce longer investor holding periods than unleveraged ones (tracing a specific pre-2015 David Chilton reference to a Dalbar/Morningstar/academic study).
  • The run classified itself full tier (not light), spawning dozens of subagents across 16 steps. Two real operational failures surfaced and were fixed inline:
    1. A round-3 background agent narrated “pilot launched” without actually executing anything — caught by checking ~/research/runs/ and ps aux directly, not trusting the agent’s own report.
    2. The genuine run hit a monthly Claude API spend-limit 429 mid-step-5, killing every in-flight subagent. Resumed from the manifest checkpoint (not a restart) rather than re-running from scratch.
  • The run then completed steps 1–14.5 (fetching, contradiction graph, loci analysis, depth investigation, 4 adversarial critics, patcher, cite-checker) and produced a complete 12,290-word draft on 2026-09-16 10:15 — but nobody ran the final two steps or the ship gate, so the draft sat unshipped for over a day. Talbot’s next message assumed the pilot had produced nothing at all; checking hyperresearch run status -j directly against the vault disproved that.
  • Finished the pipeline: ran polish (step 15) and readability audit (step 16), then the binding hyperresearch run finish ship gate — which failed twice (over length even with tolerance; 15 quote-integrity errors) before being fixed and passing clean. One flagged quote turned out to be the user’s own previously mangled published figure (“9.9 times out of 10”) — verified against the source note and restored verbatim rather than re-paraphrased.
  • Final report: research/notes/final_report_holding-period-leverage-behavioral-e78d21.md (/home/ta/research/notes/...), 11,987 words, 19.67 citations/1000 words, run finish → passed: true.

Findings that matter beyond this one pilot

Section titled “Findings that matter beyond this one pilot”
  • The pipeline does not self-finish. Completing the expensive analytical steps is not the same as shipping — the two cheap finishing steps and the binding gate need a human to notice and drive, especially after any interruption.
  • A passed content review (4 critics + patcher + cite-checker) is not the same as passing mechanical verification. The report cleared substantive review and still carried 15 unverifiable quotations at the ship gate.
  • Quote-integrity gate has a known false-positive shape: adjacent short scare-quoted single words (e.g. "leverage" next to "borrow") mis-pair under its sequential-quote regex, flagging ordinary prose between them as an unverifiable “quote.” Fix is italicizing word-mentions, not quoting them.
  • Cost is real but untracked by the tool itself — run report’s spend telemetry stayed at zero throughout (it only meters external API spend, not Claude Code session usage). No dollar figure is derivable from hyperresearch’s own instrumentation; the two finishing-step subagents alone consumed ~190k tokens for a single non-iterating pass each.
  • Full narrative, all pipeline-debugging detail, and the two LESSONS.md-worthy gotchas above are also recorded in ai-config/LESSONS.md (2026-09-17 entries) for future hyperresearch use.

Five ranked follow-up research questions for Market Down Wins / SDC (the no-margin-call feature as a testable product variable, CIRO’s uncollected leverage-flag KYC data, current-cycle drawdown behaviour, emotional risk-capacity screening instruments, and CIRO/FAIR Canada regulatory direction) are recorded in full in git history for this task file and in the shipped report’s own §C — deliberately not spun into a new task now; pick up when/if pursued.