Skip to content

Market timing — Red Team review (2026-09-28 → 09-29)

Section titled “Market timing — Red Team review (2026-09-28 → 09-29)”

Purpose: audit-level, independent attack on the market-timing research (L1–L10) before it leads a public report or the MDW offering. Talbot’s standard: “institution level confidence via institution level testing, including audit level proof.”

Project: ~/projects/cape-timing (README carries the full Red Team record and run order).

  1. Pass 1, Claude Opus 5.5. scripts/audit_pass1.py is an independent re-implementation that imports nothing from cape_timing.
    • It reproduced L1 (55% / 14%), L7 (512 / 373 / 196), the Dec 1999 figures and L3’s ρ 0.73.
    • It added a block bootstrap, CAPE-lag stress (1, 3, 6 and 12 months), era splits and leave-one-crash-out tests.
  2. Pass 2, OpenAI GPT-6 Sol via Codex CLI (ChatGPT Plus, headless, in a throwaway copy of the repo). redteam/pass2_codex/.
    • It reproduced every audited number.
    • Gemini CLI was tried first and failed: its free-tier API key allows 20 requests a day on Flash.
  3. Pass 3, OpenAI. L11 and L12, both sides. redteam/pass3_codex_l11_l12/.
  • L4: the ECY real-time percentile was computed on 1926+ data, so its first value was 1946 and 1929–45 was dropped. On the stated 1881 convention, the dearest band beat cash 16% of the time (EV −14%). “Today” had been reported three ways; it is the 19th percentile.
  • Japan: the yearly JST dividend yield was used from January of the same year, a look-ahead. The first <1% signal moves from 1985-01 to 1986-01, and the gain before the peak from +225% to +199%.
  • L1: the 1964–66 starts were top-5% CAPE, not “ordinary valuations”.
  • L8: the TSX 1999–2000 claim was a hand-typed literal; it is now computed (23 of 24 starts beat cash; Aug 2000 at CAPE 54 lost).
  • Overstatements corrected: “robust but modest” → the return edge is not proven (bootstrap interval includes zero); “ECY is best” → highest-ranked, not distinguishable from 1/CAPE.
  • Two certainty-equivalent methods are now shown (ratio-based, and each asset’s real wealth).
  • Provisional labels on the September 2026 data.
  • Latched exit (valuation starts a cut, only the trend ends it): cape_timing.vt.fired(latched=True) is the default in every script.
    • Halving, 100/0: worst decline −62% vs −84% buy-and-hold (−80% with the old exit), unchanged at every signal lag tested.
    • Every 10-year start: 316 ahead / 373 identical / 392 behind. By start era: 1926–59: 123 / 180 / 99; 1960–89: 1 / 193 / 166; 1990–2016: 192 / 0 / 127.
    • Return gain +0.58%/yr, bootstrap 90% interval −0.30% to +1.70%, so not established. Present it as protection.
  • MT-1 at risk level 1. Risks sign-off is pending. The MT-1 record is corrected and updated.
  • Framing: a smaller worst case, written in advance; the core message is to be refined in the rewrite task.
  • L11, “I have enough, is it wrong to wait?” (run_enough_wealth.py):
    • There are 6 real episodes. Waiting won 4 of them; it lost 1993–94 and the still-open 2013–26 episode, which has never re-entered.
    • Waiting from 2014: T-bills kept 0.88 of real wealth vs 3.51× for stocks. With 4% real withdrawals, 0.40× vs 2.48× (OpenAI).
    • Both reviewers’ judgment: a liability-matched real spending floor plus a permanent, smaller stock sleeve usually beats open-ended waiting.
  • L12, CAPE drift (run_cape_drift.py):
    • Median CAPE was about 14–15 before 1990 and 24.8 in 1990–2016.
    • Ranking still works (ρ 0.81), but calibration broke (out-of-sample R² −0.46).
    • A since-1881 threshold may never re-arm a “valid re-entry point”. The causes are unverified.
  • The Red Team brief is frozen at what the reviewers saw (it describes the old exit). Don’t rerun build_brief.py for the record.
  • Per-start tables treat each start as a fresh adopter. A continuously running latch gives different counts (for example, a 2009-02 start would already be halved).
  • Taxable Canada: the latched full-exit rule earned 7.45% vs 7.94% buy-and-hold after tax. The halving version after tax is untested (moved to the lab task).
  • Promoted lessons:
    • ai-config/AGENTS.md corollary: a “signed in” AI CLI ≠ usable for agentic work (commit cec7d84).
    • Core/AI/AI-Dev-Workflow.md: the headless external Red Team recipe.
  • Behavioural-factor-framework (continuation): Talbot’s reframe (the behavioural factor bf, four toll domains, Martin-based triggers, margin/leverage).
  • Market-timing-report-rewrite: prose, core message, artifact upgrades.
  • sdc-market-timing-lab: cost model with TIME, the household after-tax halving test, the spending-floor test.
  • ai-config-openai-codex-integration: Codex in ai-config and KB-OS.
  • Full history: git log -p -- SDC/IP/Tasks/Market-timing-Red-Team.md.