Market Timing Across Four Investor Domains: reward/risk signals and the behavioural factor (bf)
Section titled “Market Timing Across Four Investor Domains: reward/risk signals and the behavioural factor (bf)”Status: DRAFT (A1). Written by SVP-IP from Behavioural-factor-framework on 2026-09-29. The CEO confirms or renames bf. Anything touching leverage (Domain 4) routes through SDC/Risks before any external use.
Engine:
~/projects/cape-timing(\\wsl$\Ubuntu-24.04\home\ta\projects\cape-timing). Pre-registration:docs/prereg_behavioural_2026-09-29.md, commit6db45db, made before either study ran. Studies:scripts/run_bf_triggers.py→results/bf_triggers.json, andscripts/run_margin.py→results/margin.json. The figures here are from the post-Red-Team rerun: commits7ad7e9band08b58b5; the first runs were569c920ande5f29e0. Acronyms are defined in .
0. Round 2: the timing research (Talbot’s redirect, 2026-09-29)
Section titled “0. Round 2: the timing research (Talbot’s redirect, 2026-09-29)”“Re-examine the original premise and… move more in the desired direction of continuing market timing research on four (or more) domains… timing signals related to even the Martin ratio, perhaps re-entering equities when say the ECY hits an acceptable threshold.”
Study C is pre-registered in cape-timing/docs/prereg_reward_risk_timing_2026-09-29.md (commit 45a6a12, before any outcome was computed). Its results are in results/reward_risk_timing.json (commit 4b5d84f).
- The signal: ECY, the Excess CAPE Yield: 1/CAPE minus the real 10-year yield. It was verified to equal 1/CAPE − (GS10 − trailing 10-year inflation) exactly, so it was knowable at the time. It is used as a real-time percentile since 1881, lagged 1 month (3 months as a robustness check).
- The bands: POOR < 10th percentile; MARGINAL 10th–25th; ACCEPTABLE ≥ 25th. These come from round 10’s in-sample finding, which is a stated limitation. Today: the ~19th percentile (MARGINAL).
- Two co-primary scores, both reward/risk and not future value:
- the median 10-year Martin ratio, with Ulcer floored at 0.02;
- the certainty equivalent (CE, γ = 4): the sure multiple of cash a moderately risk-averse investor would trade the outcome for. A CE below 1 means “would rather hold cash”.
- The test subset is the decision state: starts with ECY below the 25th percentile (165 of 1,081 US starts, 1926–2016). That is today’s situation.
0.1 The fact that makes ECY re-entry viable
Section titled “0.1 The fact that makes ECY re-entry viable”ECY does not drift the way CAPE does. The CAPE-based waiting rule (L11) never re-entered after 2013, because CAPE’s since-1881 percentile stayed high. ECY, by contrast, was at or above its 25th percentile in 143 of 165 months since 2013: low real interest rates offset high CAPE.
The result: in Domains 1 and 3, no ECY-based wait ever failed to re-enter within 10 years (P(not deployed) = 0%). This is the key difference from the CAPE wait (27–33% never re-entered).
0.2 Results by domain (ECY < 25th percentile starts, lag 1)
Section titled “0.2 Results by domain (ECY < 25th percentile starts, lag 1)”Domain 1: new money at high valuations. ✅ The ECY re-entry idea passes its pre-registered test.
| Policy | CE (γ=4) | Median Martin | Median W | 10th pct W | Median months behind invest-now |
|---|---|---|---|---|---|
| N0 Invest now | 0.93 | 0.04 | 1.12 | 0.71 | 0 |
| N1 12-month DCA | 0.98 | 0.09 | 1.17 | 0.73 | 106 |
| N2 Wait in bills until ECY ≥ 25th | 1.05 | 0.07 | 1.15 | 0.79 | 15 |
| N3 Half now, half at ECY ≥ 25th | 1.01 | 0.07 | 1.16 | 0.77 | 15 |
| N4 Wait, with a 24-month backstop | 1.03 | 0.08 | 1.16 | 0.76 | 15 |
W is 10-year wealth relative to T-bills. Invest-now has a CE below 1 from these starts: a moderately risk-averse investor would prefer cash. All four alternatives pass all three pre-registered criteria at lag 1.
At lag 3, N2 and N4 fail the Martin criterion (they pass CE). N3 (half now, half at ECY ≥ 25th) and N1 pass at both lags: the robust choice.
The honest limits:
- Waiting protects the typical case, not the crash. ECY reaches its 25th percentile during big declines, not at their bottoms. A wait starting 1929-08 deployed in Dec 1929, after −32%, and the market then fell another −76%. A wait from 2000-03 deployed in Oct 2001, after −31%, then fell −20%. In 2007, ECY was already “acceptable” (low real rates), so the rule gave no warning of 2008 at all. Worst drawdown is unchanged (−84%).
- Modern era: since 2013, ECY has hovered at the boundary. Waits lasted 1–5 months, with outcomes of −11% to +11% relative to investing at once. Today’s wait (2026-06) had not yet re-entered by the end of the data. At today’s percentile the rule is close to a coin flip on timing; its value is the pre-commitment, not a forecast.
Domain 2: existing equity. ❌ No ECY rule beats MT-1.
| Policy | CE | Median Martin | Median W | 10th pct W | Worst drawdown (full period) | Now |
|---|---|---|---|---|---|---|
| E0 Buy and hold | 0.93 | 0.04 | 1.12 | 0.71 | −84% | 100% |
| E1 MT-1 (latched valuation + trend) | 1.15 | 0.21 | 1.43 | 0.80 | −62% | 100% |
| E2 ECY hysteresis (cut < 10th, restore ≥ 25th) | 1.07 | 0.15 | 1.23 | 0.79 | −75% | 100% |
| E3 MT-1 exit, re-enter on trend or ECY ≥ 25th | 1.09 | 0.24 | 1.37 | 0.79 | −78% | 100% |
| E4 ECY < 10th arms a trend exit, same re-entry | 1.04 | 0.12 | 1.20 | 0.78 | −77% | 100% |
| E5 Continuous tilt, 50–100% by ECY | 1.10 | 0.16 | 1.24 | 0.87 | −73% | 61% |
- Why ECY re-entry hurts here: the same mid-crash re-entry. E3 re-entered in 1930–31 on “acceptable value” and rode the rest of the fall down; its worst drawdown is −78% vs MT-1’s −62%. Valuation tells you the decade’s reward/risk, not the path. Only the trend tells you the path. Exit on valuation + trend, re-enter on trend (MT-1) remains best.
- E5 (continuous tilt) has the best 10th percentile (0.87) and matches MT-1’s all-starts CE (1.47). But it sits partly out of the market most of the time (54% of months), trailed buy-and-hold for 109 of 162 months since 2013, and is at 61% stocks today. It is a Domain 3 tool (a permanent reward/risk tilt), not a trading rule.
- The bf lens (simulated abandonment; mean 10-year wealth relative to bills, all starts): E3 is the best realized result at the 3%/3% hazard (2.19 vs MT-1’s 2.17), and ties MT-1 at 10%/10% (2.08). Its faster re-entry cuts the “years behind” that make followers quit. This is the one place where Talbot’s ECY re-entry idea earns something, and it earns it behaviourally, not mechanically.
- Lag 3: MT-1 is best on both CE (1.16) and Martin (0.21); E3 falls to 0.19.
Domain 3: the “enough” investor (real terms, a 4% real withdrawal; CAPE ≥ 90th armed starts, n = 299). A trade-off map, with no winner test.
| Policy | Median real end | 10th pct real end | P(real wealth falls below 70% of start) | Median real Ulcer | Never re-entered |
|---|---|---|---|---|---|
| H Hold 100% | 0.95 | 0.39 | 68% | 0.26 | — |
| L Permanent 40/60 | 0.77 | 0.54 | 45% | 0.19 | — |
| W-CAPE (L11: wait until CAPE < 75th) | 0.76 | 0.50 | 70% | 0.19 | 27% |
| W-ECY (wait until ECY ≥ 25th) | 0.87 | 0.45 | 75% | 0.24 | 0% |
| S-ECY sleeve: 40% stocks, 70% when ECY ≥ 25th | 0.91 | 0.55 | 63% | 0.20 | 0% |
- The ECY sleeve beats the CAPE wait on every column except Ulcer (0.20 vs 0.19). Against a permanent 40/60, it gives up floor-breach safety (63% vs 45%) for a higher median (0.91 vs 0.77) and a similar 10th percentile.
- An all-in ECY wait is the worst for the floor (75%): it re-enters 100% mid-decline.
- Reading: for an investor with “enough”, how much is in stocks matters more than when. ECY is best used to set the sleeve size, not as an in/out switch. The household spending-floor test (lab task) is the proper next step.
Domain 4: leverage (T-bill + 4%, maintenance 30%). ❌ ECY rules don’t beat MT-1 deleveraging.
| Policy | P(call), 33% LTV | 10th pct equity, 33% | P(call), 50% LTV | 10th pct equity, 50% |
|---|---|---|---|---|
| M0 Margin, no rule | 16.5% | 0.59 | 38.1% | 0.32 |
| M1 MT-1 delever, re-lever on trend | 7.4% | 0.97 | 31.3% | 0.51 |
| M2 MT-1 delever, re-lever on trend and ECY ≥ 25th | 7.5% | 0.99 | 31.3% | 0.51 |
| M3 ECY delever < 10th, re-lever ≥ 25th | 13.6% | 0.71 | 40.0% | 0.41 |
- Adding an ECY condition to re-levering changes almost nothing: ECY was usually already ≥ 25th when the trend recovered.
- Valuation-only deleveraging is much worse, for the same reason as above: it re-levers mid-crash.
- The primary control remains the loan-to-value ceiling (Study B).
0.3 What this round establishes
Section titled “0.3 What this round establishes”- Talbot’s ECY re-entry idea works for new money (Domain 1). The robust form is half now, half when ECY ≥ 25th (N3). The main benefit is avoiding the worst decade outcomes: CE 1.01 vs 0.93, 10th percentile 0.77 vs 0.71.
- For existing equity and leverage, valuation must not drive re-entry. ECY becomes “acceptable” partway down a crash. The trend is the re-entry signal; valuation is the arming signal. MT-1 stands.
- For the “enough” investor, use ECY to size a permanent sleeve, not to switch in and out.
- The “or better” candidate for the next pre-registration: re-enter on trend up AND ECY ≥ 25th, only after a valuation-armed exit (in Domain 4 this is M2, which is neutral). Also a drawdown-aware ECY: acceptable ECY and a completed decline of ≥ 20%. Neither was tested for Domains 1–2 this round; they are proposed, not claimed.
- Martin-based triggers (round 1: C3 and C4) failed partly because the signal first exists in 1966 and could not act in 1929. The ECY-percentile rules exist from 1901 and still did not beat MT-1. So the conclusion is not an artifact of signal start dates.
1. The idea in one line
Section titled “1. The idea in one line”Most analyses score a strategy as if a robot runs it. A proper expected value includes the probability that a human actually implements it, and what that human does instead. (Talbot, 2026-09-29: “A proper EV calculation SHOULD include a factor for (human) probability of implementation.”)
2. Definition
Section titled “2. Definition”bf (working name) is the share of a recommended strategy that is actually implemented. It is a number from 0 to 1.
It has two readings:
| Reading | Meaning | Example |
|---|---|---|
| Group | The probability that a member of a population implements it | 47% of a survey group reinvest their RRSP refund (figure unverified, see §4.1) |
| Individual | The share this person actually implemented | This client invested $250 of a $400 refund, so bf = 0.625 |
It also has two time shapes:
| Shape | Where it applies | Form |
|---|---|---|
| Constant | A decision made once: invest the refund, or redirect the freed cash flow | A single number |
| Path-dependent, bf(path) | A strategy that must be kept: holding through a crash, following a timing rule | A monthly abandonment hazard that depends on what the path has done so far (depth of the drawdown, how long you have been behind a benchmark). Realized bf is the share of months the plan was actually followed |
2.1 The general formula: bf needs a fallback
Section titled “2.1 The general formula: bf needs a fallback”bf alone is under-specified, because not implementing is never neutral. Something else happens instead:
E[V] = bf · V(implemented) + (1 − bf) · V(fallback)
The fallback depends on the domain:
- The RRSP refund: the fallback is “spent”, so V = 0 for retirement.
- Freed cash flow: the fallback is “absorbed into lifestyle”.
- A buy-and-hold investor: the fallback is to sell after the crash, sit in cash, and re-enter late.
- A timing-rule follower: the fallback is to quit the rule after years of lagging and chase buy-and-hold.
Study A (§5b) shows that the fallback often matters more than bf itself. A rule follower who quits to buy-and-hold loses little. A buy-and-hold investor who panics to cash loses a lot. So a strategy can have a lower realized bf and still higher realized wealth.
Two conditions for the formula to hold (Red Team pass 4):
- V(implemented) and V(fallback) must be the values for the people and paths that actually land in each state, not population averages.
- Who implements is not random: it depends on liquidity, risk capacity, returns already experienced, and advice.
- With unconditional averages, the formula misses a covariance term: Cov(implements, V_impl − V_fallback).
- So a group bf must not be applied to an individual without checking that the individual resembles the group.
- Compute the client’s objective once. If emotion or risk aversion is already inside a certainty equivalent, adding a separate “emotion toll” counts the same harm twice.
- “Spent” has zero retirement value, but it may have real current-consumption value.
Prior art, so the claim is framed correctly. Scoring what investors actually do is established:
- implementation shortfall (the CFA Institute trading-cost curriculum);
- dollar-weighted investor returns (Morningstar Mind the Gap);
- the behavioural life-cycle hypothesis (Shefrin & Thaler);
- commitment-savings experiments (Ashraf, Karlan & Yin);
- defaults (Madrian & Shea).
These are named by Red Team pass 4 and not yet verified here (§7).
What is plausibly distinctive is carrying bf and its fallback as a first-class input in every client-facing strategy comparison, so it is never left in a footnote. Whether that is novel is unverified.
For path-dependent strategies, the closed form above becomes a simulation: E[V] = the average over paths and simulated investors of V(the policy they actually followed).
2.2 A strategy’s behavioural value
Section titled “2.2 A strategy’s behavioural value”A strategy’s behavioural value is how much it raises E[V], either by raising bf on the paths that decide wealth, or by making the fallback less costly, net of its tolls. The tolls are:
- tax;
- carry (interest cost);
- time;
- the emotional strain itself.
2.3 An ethical constraint that the formula alone does not carry
Section titled “2.3 An ethical constraint that the formula alone does not carry”Raising bf is only good when the strategy suits the person. For a risk-magnifying strategy (anything that borrows to invest), the goal is the right bf: people who should not implement it, don’t.
That is Rule 1 (intervention strength is a function of risk). It is being worked out in sdc-behavioural-solutions-debt. That task decides how hard SDC may push; this framework measures what pushing is worth.
3. The four domains at a glance
Section titled “3. The four domains at a glance”| Domain | Common rule of thumb | What people actually do (sourced) | Tolls | Where bf enters | Strength of the bf evidence |
|---|---|---|---|---|---|
| 1. New investable money at high valuations (lump sums, windfalls) | Invest the lump sum now; or dollar-cost average | Well-intentioned people mostly don’t act: 86% of savers had made no change 4 months later (SMarT). Self-serve take-up is 4–10%, against 55–83% when the action arrives done | Regret, inertia (“waiting” becomes never), opportunity cost | bf of the re-entry decision: does parked money ever get invested? | Strong for inaction in general; untested for re-entry specifically |
| 2. Existing invested equity | Stay the course. Time in the market beats timing | The reported gap is −1.1 to −1.2 pts/yr (Morningstar, contested); the timing component alone is 0.10%/yr (Fulkerson et al. 2026). Allocation-fund investors capture ~97% of fund returns. Vanguard puts coaching at 100–200 bps, “intermittent… during periods of market duress or euphoria” | Capital-gains tax on switching, emotional strain, time | bf(path): a hazard driven by drawdown depth and time behind | Moderate: real but smaller than advertised |
| 3. A non-wealth objective (peace of mind, reward/risk, “enough”) | “Your age in bonds”, risk-tolerance questionnaires | Low lifetime experienced returns lower participation and the equity share (Malmendier & Nagel 2011). Default target-date investors: 85% were still fully invested 6 years later (EBRI/ICI). For an investor with enough, open-ended waiting had no reliable re-entry point (L11, Codex pass 3) | Regret, sleep, sequence risk, inflation on cash | bf and a different objective: the certainty equivalent, Martin ratio or spending-floor survival, not terminal wealth | Weak to moderate. No source measures abandonment by objective |
| 4. Leveraged, with or without margin call | At most ~20% of net worth, 10+ years, prefer a no-margin-call term loan (practitioner norms) | Investors using margin trade 2.2x as much (15.2% vs 6.9% monthly turnover). Accounts near a call sell +23.95% more of their holdings. In 68% of aggressive-leverage cases the client was in a loss (NB sweep, via FAIR Canada) | Carry, forced sale, emotion, flywheel | Under margin, bf is overridden by the broker: the exit is mechanical. With a no-margin-call loan, bf returns as a calendar-trigger risk | Strong on the mechanism, missing on retail call frequency |
Sources for every figure are in §7, §9 and the research note leverage_holding-period_behavioral_research. Its underlying source notes are in ~/hyperresearch/research/notes/. The figures used here were re-checked there on 2026-09-29: 0.10%/yr; 86%; 0.993; 23.95%, from the raw PDF.
4. Domain notes
Section titled “4. Domain notes”4.1 New investable money at high valuations: when to deploy
Section titled “4.1 New investable money at high valuations: when to deploy”- The question: a lump sum (a sale, an inheritance, a bonus, maturing savings) arrives while valuations are high. Should it be invested now, averaged in, or held until reward/risk is acceptable?
- The common rules of thumb: “invest the lump sum now, because time in the market beats timing”; “dollar-cost average to reduce regret”.
- The results: Study C, Domain 1 (§0.2).
- Where bf enters: a lump sum parked “until things look better” often never gets invested, because the re-entry decision is itself a decision nobody makes (Behavioural-Solutions §2–3). A pre-committed re-entry trigger (for example, ECY ≥ 25th) turns that open-ended wait into a rule.
- (Round 1 read this domain as RRSP refunds and freed cash flow. That was a misread; that material is now Appendix A.)
4.2 Existing invested equity: bf is real, but the measured cost is smaller than the headline
Section titled “4.2 Existing invested equity: bf is real, but the measured cost is smaller than the headline”- Morningstar reports a large gap: −1.1 pts/yr for the decade to 2023, −1.2 for the decade to 2024. Fulkerson, Jordan, Riley & Yan (FAJ, 2026) reproduce it, then split it into a timing part and a hindsight part. The timing part, the actual cost of when money moved, is 0.10% per year (source note, verbatim).
- DALBAR’s magnitudes are invalid (the research note §1A). Use DALBAR, if at all, for ordinal category rankings only. Do not quote DALBAR’s gap in SDC material.
- Vanguard’s own framing supports bf(path): coaching value “should not be viewed as an annual value-add but is likely to be intermittent. Some of the best opportunities to add value occur during periods of market duress or euphoria” (Putting a value on your value, July 2022). Its figure of 100–200 bps is Vanguard’s estimate of what coaching can add, not a measured investor gap.
- Two abandonment modes, one framework:
- Sell after the crash (drawdown pain), common with buy-and-hold.
- Quit the rule after years of lagging (lag pain), common with timing rules.
- Any claim that “the rule helps behaviourally” must show it lowers the first more than it raises the second. Study A (§5b) tests exactly this.
4.3 A non-wealth objective: change the objective, not just bf
Section titled “4.3 A non-wealth objective: change the objective, not just bf”- For someone with “enough”, terminal wealth is the wrong target. Better targets:
- the certainty equivalent at a stated risk aversion;
- the Martin ratio (depth × time underwater);
- the probability of sustaining a real spending floor.
- The L11 Red Team (Codex pass 3) concluded that a liability-matched real spending floor plus a permanent, smaller equity sleeve usually beats open-ended waiting. The waiting rule had no reliable re-entry point: from the 2013-11 episode it never re-entered, and with a 4% real withdrawal from 2014, WAIT ended at 0.40x vs HOLD 2.48x.
- bf connection: a permanent lower allocation is a one-time decision, a constant bf. A waiting rule is a path decision, bf(path). Constant-bf decisions are easier to keep. The household spending-floor test (lab task) is the place to quantify this.
- Evidence gap: no source located measures abandonment rates by investor objective. Malmendier & Nagel (QJE 2011; SCF data 1960–2007) show that lifetime experience shifts risk-taking and participation. That is an input to bf, not a measurement of it.
4.4 Leveraged, with or without margin: the tolls are mechanical
Section titled “4.4 Leveraged, with or without margin: the tolls are mechanical”- Margin replaces bf with a contract. The broker sells whatever the investor feels. The tolls are:
- (a) the probability of a call, given the loan-to-value and the decline path;
- (b) the forced-sale magnification: selling at the low locks the loss and forfeits the recovery;
- (c) carry: interest compounds whether markets rise or not;
- (d) emotional strain: the one within-person account (CBC) had the distress coming from the borrowing itself, not the position size;
- (e) the flywheel: forced sellers push prices down for everyone else.
- The flywheel is a documented mechanism, not a metaphor:
- Brunnermeier & Pedersen, “Market Liquidity and Funding Liquidity”, RFS 22(6), 2009: margins can be destabilizing, and market and funding liquidity reinforce each other into spirals.
- Bian, Da, He, Lou, Shue & Zhou, “The Drivers and Implications of Retail Margin Trading” (draft Aug 2025, Chinese account data for 2015): accounts very close to a margin call sell an additional 23.95% of their holdings.
- This is the direct link to SDC’s help-others mission: responsible deleveraging protects other investors too.
- With a no-margin-call term loan, bf returns as a calendar-trigger risk: can the payments be met from income? The research note’s seven-point design specification (§6F there) is the SDC standard to apply.
- Study B (§6) quantifies (a)–(c) and the value of pre-decline deleveraging. Routed to SDC/Risks (see §8).
5. The timing formula (sd-math terms, for sdc-market-timing-lab to implement)
Section titled “5. The timing formula (sd-math terms, for sdc-market-timing-lab to implement)”5.1 A timing rule: EV with bf(path)
Section titled “5.1 A timing rule: EV with bf(path)”E[W_S] = E over paths and simulated investors of W(the policy actually followed | S, hazard model)
The hazard model, as pre-registered, is applied identically to every strategy:
- Drawdown pain: p_dd × clip((DD − 15%)/25%, 0, 1).
- Lag pain: p_lag × clip((months behind buy-and-hold − 12)/24, 0, 1).
- Fallbacks:
- panic goes to cash, and re-enters buy-and-hold only after the market is 20% above the exit level;
- quitting the rule means switching to buy-and-hold.
Realized bf is the share of months the plan was followed.
sd-math/lab mapping: a BehaviourModel(p_dd, p_lag, dd_floor=0.15, dd_full=0.40, lag_floor=12, lag_full=36, panic_reentry=0.20) passed to the strategy simulator. Reference implementation: cape-timing/scripts/run_bf_triggers.py simulate().
5b. Study A: trigger rules under bf(path) (pre-registered)
Section titled “5b. Study A: trigger rules under bf(path) (pre-registered)”Setup:
- US CRSP 1926-07 to 2026-06; 1,081 overlapping 10-year starts.
- Cuts are halvings (100% to 50% stocks), with a 0.1% trading cost.
- 400 simulated investors per start.
Full-period results:
| Rule | CAGR | Worst drawdown | Martin | Median 10-yr Martin | P(behind B&H at 10 yrs) |
|---|---|---|---|---|---|
| B0 Buy and hold | 10.33% | −84% | 0.324 | 0.77 | — |
| B1 Trend only | 10.20% | −62% | 0.465 | 0.76 | 61% |
| B2 Latched V+T (MT-1) | 10.97% | −62% | 0.516 | 0.98 | 35% |
| C1 Volatility-managed | 9.08% | −84% | 0.259 | 0.59 | 95% |
| C2 Drawdown brake | 10.64% | −62% | 0.465 | 0.70 | 51% |
| C3 Expected-Martin tilt | 10.00% | −84% | 0.318 | 0.74 | 20% |
| C4 Expected-Martin + latched trend | 10.34% | −84% | 0.331 | 0.77 | 18% |
Pre-registered verdict: no candidate beats MT-1. All four fail on full-period Martin. They won 0, 1, 3 and 3 of the 10 decade blocks.
C3 and C4’s signal needs 30 years of completed windows, so it first exists in 1966-06. They could not act in 1929. A post-hoc comparison on 1966-06 onward (labelled as such, results/bf_triggers_posthoc.json) has:
- trend-only best (Martin 0.685, worst drawdown −30%);
- then B2 (0.624), C2 (0.610) and C4 (0.536).
So MT-1’s edge over plain trend-following comes mostly from 1929, which matches Red Team pass 1.
Under bf(path) (mean 10-year wealth relative to cash; realized bf in brackets). This is the rerun after Red Team pass 4, which fixed two things:
- trading cost is now charged inside the simulation;
- a buy-and-hold investor who re-enters after panicking now counts as back on plan.
| Panic hazard at −40% (p_dd) → | 0 | 3%/mo | 10%/mo | 25%/mo |
|---|---|---|---|---|
| B0 Buy and hold | 2.19 (1.00) | 2.12 (0.95) | 2.03 (0.88) | 1.93 (0.80) |
| B2 MT-1, no lag-quitting | 2.23 (1.00) | 2.17 (0.91) | 2.09 (0.81) | 2.00 (0.71) |
| B2 MT-1, lag-quit 10%/mo | 2.24 (0.83) | 2.17 (0.77) | 2.08 (0.69) | 1.99 (0.61) |
| B1 Trend only, no lag-quitting | 2.05 (1.00) | 2.02 (0.93) | 1.98 (0.85) | 1.92 (0.76) |
What it shows:
- MT-1’s behavioural value shows up in wealth, not in realized bf.
- Its lead over buy-and-hold widens only slightly as panic rises: +0.04 at no panic, +0.07 at the highest hazard.
- The implied behaviour drag falls from −1.27%/yr (buy-and-hold) to −1.10%/yr (MT-1).
- Trend-only halves the drag (−0.64%/yr), but gives up return.
- Realized bf is lower for MT-1 than for buy-and-hold. For MT-1, a panicked investor re-enters buy-and-hold, which is off-plan. For buy-and-hold, the same investor re-enters the plan itself. So realized bf is not comparable across strategies with different fallbacks. Compare the objective, not bf.
- (The first draft said MT-1 raised realized bf by 3–5 points. That came from the miscount Codex found, and is withdrawn.)
- Lag-quitting lowers MT-1’s realized bf to 0.61–0.83 yet barely moves its wealth. The fallback (buy-and-hold) is cheap. This is §2.1’s point: model the fallback, not just bf.
- The robust gain is in the tail and in crash decades.
- With no abandonment, MT-1’s 10th-percentile 10-year wealth is 1.06x cash, against 0.85x for buy-and-hold.
- By decade block at the illustrative (3%, 3%) cell, MT-1’s lead is concentrated in starts in the 1920s (1.26 vs 0.81) and 2000s (2.25 vs 1.86). It trails in the 2010s (2.92 vs 3.21).
- The published figures are not a calibration (withdrawn as a “calibration anchor” after Red Team pass 4).
- Buy-and-hold’s simulated drag at p_dd = 1%/month is −0.12%/yr, and at 25%/month it is −1.27%/yr. These sit near Fulkerson et al.’s 0.10%/yr and Morningstar’s ~1.2-point gap.
- But those measure aggregate fund-flow timing, not panic exits after drawdowns. The similarity identifies nothing. p_dd and p_lag remain uncalibrated.
- The incumbent was not untouched. MT-1 was chosen earlier on this same historical record, so “no candidate beats MT-1” is weak evidence that no better rule exists.
- C1 was specified narrowly: it scales by inverse volatility, capped at 100% stocks, whereas Moreira & Muir scale by inverse variance and can lever up. So C1’s result does not refute volatility management in general.
- Martin-ratio triggers. An ECY-conditioned expected-Martin trigger did not improve on CAPE-armed MT-1. “Or better” is more likely to come from changing the objective (§4.3: certainty equivalent, spending floor, time behind) than from another trigger.
- Time behind buy-and-hold should be a first-class reported metric, because it is what drives quitting.
6. Study B: the margin and leverage domain (pre-registered)
Section titled “6. Study B: the margin and leverage domain (pre-registered)”Setup:
- Per $1 of own money; US CRSP 1926-07 to 2026-06; 1,081 ten-year starts.
- Interest is capitalized at the T-bill rate plus a spread.
- A call is checked at month-end; when it fires, the investor is force-sold to equity/value = m + 5%.
Selected cells:
| Loan-to-value / spread / maintenance | P(call) | P(call) with pre-decline deleveraging | Median terminal equity: no-call loan / margin / margin + delever | 10th percentile: unlevered / no-call / margin / margin + delever |
|---|---|---|---|---|
| 20% / T-bill + 4% / 30% | 5% | 0% | 3.06 / 3.06 / 3.11 | 1.41 / 1.15 / 1.12 / 1.17 |
| 33% / T-bill + 4% / 30% | 16% | 7% | 3.23 / 3.23 / 3.34 | 1.41 / 0.93 / 0.59 / 0.97 |
| 50% / T-bill + 1.5% / 30% | 32% | 30% | 4.09 / 4.01 / 4.35 | 1.41 / 0.95 / 0.48 / 0.69 |
| 50% / T-bill + 4% / 30% | 38% | 31% | 3.70 / 3.57 / 3.78 | 1.41 / 0.38 / 0.32 / 0.51 |
Unlevered buy-and-hold: median 2.88, 10th percentile 1.41. All 12 cells are in results/margin.json.
What it shows:
- Margin leaves the median alone and wrecks the tail. On called starts at a 33% loan-to-value, the forced sale slightly reduces the trough (median trough equity 0.18 on margin vs 0.14 without calls) but locks the loss in: median terminal equity is 0.48 vs 0.82. The toll is the forgone recovery, not a deeper trough. At a 33% loan-to-value with bank-broker pricing, the 10th percentile falls from 0.93 (no-call loan) to 0.59 (margin).
- On starts that were called, the median forced-sale magnification is 1.4x to 7x across cells. It is dominated by 1929–32: at a 33% loan-to-value with bank pricing it is 4.8x for calls first made in 1929–32, and about 1.3x for 1973–74 and 2008–09.
- Every called margin start ended below unlevered buy-and-hold at 33% loan-to-value.
- Carry helps create calls.
- At a 33% loan-to-value, 79 of the 178 called starts had their first call in 1973–74.
- Codex’s decomposition (pass 4) splits it:
- with interest set to zero there, there are 0 first calls;
- on a flat-price path with the actual interest, there are 29.
- So carry was necessary for these calls, and can trigger calls on its own. But the historical 79 came from carry and the decline together.
- Carry is not only a cost; it erodes the cushion that stands between a decline and a forced exit.
- Pre-decline deleveraging (MT-1 used as the trigger to repay the loan) does two things at a 33% loan-to-value:
- it cuts P(call) from 16% to 7%;
- it restores the 10th percentile from 0.59 to 0.97, while raising the median (3.23 to 3.34).
- At a 50% loan-to-value it helps much less (38% to 31%). The remaining calls come mainly from re-levering to the full 50% after a trend recovery, just before the next leg down: first calls in 1933, 1937, 1970 and 1974, plus 1987’s fast drop. At a 33% loan-to-value, every remaining call is in 1974. A re-lever rule (for example, re-borrow in steps, or to a lower loan-to-value) is an untested design lever.
- This supports pre-decline-deleveraging as a leverage-domain use of MT-1, and a hard ceiling on loan-to-value, not the trigger, as the primary control.
- Month-end checks understate calls. On Nikkei daily data, daily-checked call frequency is 1.07x to 1.26x the month-end-checked figure. This shows that month-end checking misses intra-month first touches. It is not a correction factor: it comes from a Japanese price-only series with a fixed barrier, no interest and no remedial trades. Read the US figures as understated by an unknown amount, not as lower bounds with a known gap.
- Episodes, not trials. The first calls cluster in 6 to 46 distinct months (by cell), concentrated in 1929–32, 1937–38, 1973–74, 2000–02 and 2008–09. The effective sample is a handful of episodes.
- These are stress scenarios, not client probabilities. They come from overlapping starts clustered in a handful of episodes. A Canadian account needs its actual inputs:
- the security’s margin eligibility (CIRO’s 30% applies only to reduced-margin-list securities);
- the dealer’s house requirement;
- the contract’s liquidation terms;
- cash-flow capacity and taxes. The “no-margin-call loan with capitalized interest and no payments” used as counterfactual (ii) is unverified as an available Canadian product; real no-call term loans require payments.
- Not modelled:
- the flywheel (a single investor is a price-taker here);
- margin-rate changes by the broker;
- house margin above the regulatory minimum;
- the tax treatment of forced gains and losses.
- Wipe-outs were 0% at month-end sampling. Daily data and gaps would not make them zero.
7. Evidence register
Section titled “7. Evidence register”| Claim | Source | Status |
|---|---|---|
| Investor timing costs 0.10%/yr (the hindsight component removed) | Fulkerson, Jordan, Riley & Yan, FAJ 2026, doi 10.1080/0015198X.2026.2657253 (source note, verbatim) | Verified at the note |
| Morningstar gap −1.1 (to 2023), −1.2 (to 2024); allocation funds capture ~97% | Mind the Gap 2024, 2025 (via the research note) | Via the research note |
| DALBAR’s gap method is invalid | Research note §1A; Advisor Perspectives critiques | Via the research note |
| 86% made no change 4 months later | Thaler & Benartzi, SMarT, JPE (source PDF line 222) | Verified |
| Amortization-wealth elasticity 0.993 | Bernstein & Koudijs (FDIC WP 2020 version) | Verified at the note |
| Near-call accounts sell an additional 23.95% | Bian, Da, He, Lou, Shue & Zhou, draft Aug 2025 (raw PDF) | Verified |
| Margin users’ monthly turnover 15.2% vs 6.9% | Leveraging overconfidence (via the research note) | Via the research note |
| Coaching adds 100–200 bps; the value is intermittent, in duress or euphoria | Vanguard, Putting a value on your value, July 2022 | Verified (PDF). ⚠ That PDF is marked “For institutional and sophisticated investors only. Not for public distribution.” Do not quote it in public material; find a public Vanguard source first |
| Experienced returns shape risk-taking and participation | Malmendier & Nagel, QJE 126(1):373–416, 2011 | Verified (bibliographic + abstract) |
| Margin spirals | Brunnermeier & Pedersen, RFS 22(6):2201–2238, 2009 | Verified (bibliographic + abstract) |
| Reg T initial margin 50%; FINRA 4210(c)(1) maintenance 25% | finra.org margin page; Rule 4210 text | Verified |
| CIRO reduced-margin list: 30% client requirement | SDC/Strategy/Research/Investment-Debt-Providers.md [21] | Via the vault |
| RRSP refund: 57% treat themselves / 47% reinvest | IG Wealth Management survey, reported secondhand | Unverified: primary not found |
| 85% of default target-date investors fully invested 6 years later | EBRI/ICI (via the research note §2D) | Via the research note |
| NB leverage sweep: 68% of aggressive-leverage cases in a loss | FAIR Canada letters citing the NBSC 2010 sweep (via the research note §6B) | Via the research note; primary report unrecoverable |
| Leverage distress comes from the borrowing itself, not the position size | CBC, a single advisor’s account (via the research note §6C) | Anecdote |
| Prior art: implementation shortfall; behavioural life-cycle hypothesis (Shefrin & Thaler); commitment savings (Ashraf, Karlan & Yin, SEED); defaults (Madrian & Shea) | Named by Codex pass 4, with links in redteam/pass4_codex_bf/REVIEW.md | Unverified here: not re-read, and years not checked |
| Lump sum beats dollar-cost averaging ~2/3 of the time | — | Unverified; not used |
8. Routing and relations
Section titled “8. Routing and relations”- SDC/Risks: the Domain 4 analysis (Study B, and pre-decline deleveraging as a leverage-domain use of MT-1) was filed in
SDC/Risks/Inbox.mdon 2026-09-29. It is not signed off. - sdc-behavioural-solutions-debt decides the permissible intervention strength per risk tier. bf is the quantity those interventions move. The two should cross-link when that task runs.
Core/Processes/Behavioural-Solutions.mdis the evidence SSOT for “how to get people to act”. §4.1 here applies it; it does not restate it.- leverage_holding-period_behavioral_research is the evidence backbone for Domains 2 and 4.
- Downstream: Market-timing-report-rewrite (message) and sdc-market-timing-lab (implementation: §5 mappings,
BehaviourModel, time-behind as a metric).
10. Red Team pass 4 (OpenAI GPT-6 Sol via Codex, 2026-09-29): triage
Section titled “10. Red Team pass 4 (OpenAI GPT-6 Sol via Codex, 2026-09-29): triage”The review and its audit are in cape-timing/redteam/pass4_codex_bf/. It reproduced:
- B0/B1/B2 Martin ratios and drawdowns exactly;
- the Study B key cell (178/1,081 calls, 10th percentile 0.928 → 0.592, delever 7.40%);
- every worked-formula number.
| # | Finding | Verdict | Action |
|---|---|---|---|
| 1 | A scalar bf can reverse a decision: selection, covariance, utility vs wealth, double counting; prior art exists | Accepted | §2.1: conditions and prior art added; novelty marked unverified |
| 2 | The debt “0.98” assumes the invest-now contribution has adherence 1 | Accepted: it reverses the headline | §5.2 rewritten. At matched adherence (0.8), debt-first leads by $1,666; the lesson is to automate the freed flow |
| 3 | Margin results are stress scenarios, not Canadian probabilities; the 1973–74 carry attribution is too strong; Nikkei is not a correction factor; the no-call counterfactual product is unverified | Accepted | §6 reworded; the fee-inclusive sale target is fixed in code (results unchanged) |
| 4 | MT-1 was historically selected; C1 was specified narrowly | Accepted | §5b item 5 |
| 5 | The simulation omitted trading cost; B0 re-entrants were miscounted as off-plan; the “calibration anchor” is apples to oranges | Accepted: two code fixes and a rerun | MT-1’s “higher realized bf” claim is withdrawn; the calibration claim is withdrawn |
| 6 | Pre-registered outputs were missing: margin extra trough loss, behavioural results by episode | Accepted | Both added (median_trough_*, bf_central_by_block) |
| — | Decisive next test: a prospective, randomized Canadian field trial of automating the freed debt payment (3,000 households) | Noted for the CEO | A real-world, measured bf is the credibility gap. It needs partner firms; see Next Steps |
Appendix A: illustrations outside market timing
Section titled “Appendix A: illustrations outside market timing”Talbot, round 2: “These were simply examples of how bf can be the most important factor. They are not related to this market timing research.” They are kept here for reference only, and are not part of the timing work. The Red Team corrections to A.2 still apply.
A.1 The RRSP refund (Talbot’s formula, verbatim, then flagged extensions)
Section titled “A.1 The RRSP refund (Talbot’s formula, verbatim, then flagged extensions)”Talbot’s formula:
FV = PV × (1+i)^n + bf × (PV × TaxRate × (1+i)^n)
Worked example (PV $1,000, TaxRate 40%, i 6%, n 25; (1.06)^25 = 4.292):
| bf | FV | What the refund added |
|---|---|---|
| 0 (spent) | $4,292 | nothing |
| 0.47 (unverified survey share) | $5,099 | +$807 |
| 1 (all invested) | $6,009 | +$1,717 |
Assumptions the formula makes implicitly. These are flagged for Talbot to confirm, not silent edits:
- Refund lag: the refund arrives months after the contribution; the formula compounds it from day 0.
- Where it is invested:
- Into a TFSA: the term stays as written, and grows tax-free.
- Into a taxable account: its growth is taxed.
- Back into the RRSP: it earns its own refund. At bf applied every round, the series is 1 + bf·t + (bf·t)² + …, so FV_RRSP = PV × (1+i)^n / (1 − bf × TaxRate), before withdrawal tax. At bf = 1 this is exactly the gross-up refund ($7,153 above; bf = 0.47 gives $5,286).
- Withdrawal tax: the RRSP balance is taxed on the way out at the retirement rate t_r: multiply the RRSP terms by (1 − t_r).
sd-math mapping. strategies/rrsp.py already has the two endpoints:
spend_refundis bf = 0.reinvest_refundis bf = 1: its cash flows areat_annual_payment × (1 + tax_rate).gross_up_refundis the recursive bf = 1 case:at_annual_payment / (1 − tax_rate).
Proposed: one new field, RRSPCatchUpProfile.refund_invest_ratio: float = 1.0 (bf).
- Reinvest cash flows become
at_annual_payment × (1 + bf × tax_rate). - Gross-up becomes
at_annual_payment / (1 − bf × tax_rate). - bf = 0 and bf = 1 must reproduce today’s outputs exactly; that is the regression test.
ReturnProfile.distrib_reinvest_ratio is already a bf, on distributions. Name the new field to match it.
A.2 Invest vs pay down debt: the freed-cash-flow share
Section titled “A.2 Invest vs pay down debt: the freed-cash-flow share”Setup:
- A $400/month budget for 20 years.
- $15,000 of non-deductible debt at 8%.
- Investments earn 6%.
The two options:
- Debt first: all $400 goes to the debt until it is gone (44 months); then bf × $400 is invested. FV_debt-first = bf × P × [((1+i)^(N−T) − 1)/i]
- Invest now: the minimum payment ($182, a 10-year amortization) goes to the debt; the other $218 is invested from month 1. After year 10, bf × $182 of freed cash is added.
Results:
| bf (share of freed cash flow invested) | Debt first | Invest now |
|---|---|---|
| 1.0 (the robot assumption) | $132,637 | $130,554 |
| 0.8 | $106,109 | ~$124,590 |
| 0.5 | $66,318 | $115,641 |
The answer depends on the adherence of both designs, not on bf alone (corrected after Red Team pass 4).
Give the invest-now monthly contribution its own adherence a:
- invest-now = a × $100,728 + bf × $29,826;
- debt-first = bf × $132,637.
What that implies:
- If invest-now is automated (a = 1) and the freed flow is not, debt-first needs bf ≥ 0.98 to win. The first draft of this note stopped at this case.
- If both are equally manual (a = bf = 0.8), debt-first leads by $1,666. The robot answer (pay the 8% debt first) survives.
- With a = 0.8, the break-even bf is 0.78.
The behavioural lesson is therefore not “invest instead of repaying debt”. Two lessons hold:
- Whatever the plan is, automate the freed cash flow at the moment the debt disappears. That is the “arrives done” structure (§4.1). A comparison of two plans with different automation is a comparison of two designs, not of debt vs investing.
- A comparison that assumes bf = 1 for one side and not the other is rigged. Both adherence rates must be stated, and ideally measured.
There is a further caution: the 8% debt cost is contractual, while the 6% return is uncertain. A client-first comparison also needs loss capacity and after-tax outcomes.
sd-math mapping: there is no debt-vs-invest strategy yet. Proposed: a freed_cashflow_invest_ratio input (bf) on the future debt-vs-invest strategy. Its default should be sourced or asked, never 1.0 silently.
- bf — behavioural factor (working name): the share of a strategy actually implemented, 0–1.
- B&H — buy and hold.
- CAGR — compound annual growth rate.
- CE — certainty equivalent: the sure amount an investor values the same as a risky outcome.
- ECY — Excess CAPE Yield: 1/CAPE minus the real 10-year Treasury yield.
- CAPE — Cyclically Adjusted Price-to-Earnings ratio (Shiller).
- EV — expected value.
- FAJ / JPE / QJE / RFS — Financial Analysts Journal / Journal of Political Economy / Quarterly Journal of Economics / Review of Financial Studies.
- Martin ratio — annual excess return ÷ Ulcer Index.
- Ulcer Index — the root-mean-square drawdown, counting both depth and time underwater.
- MT-1 — the latched Valuation + Trend Rule (valuation-trend-rule).
- p_dd / p_lag — monthly abandonment hazard at full drawdown pain / full lag pain.
- Realized bf — the share of months the plan was actually followed.
- Loan-to-value — the loan ÷ the portfolio value.
- Maintenance margin (m) — the minimum equity ÷ value before a margin call.
- Reg T — US Federal Reserve Regulation T.
- FINRA — Financial Industry Regulatory Authority (US).
- CIRO — Canadian Investment Regulatory Organization.
- SMarT — Save More Tomorrow.
- EBRI/ICI — Employee Benefit Research Institute / Investment Company Institute.
- SCF — Survey of Consumer Finances.
- TFSA / RRSP — Tax-Free Savings Account / Registered Retirement Savings Plan.