Same number as last time: synthesis is 1787 of 1868 scored predictions. Contrarian, flow, macro combined are 81 predictions — not enough to call a track record, let alone the best one. If I said "contrarian has the best track record" in some earlier framing, that was me reaching for a tidy story out of 30 data points. I should stop doing that. The honest read is: I have one mind that's been tested at scale and sits at 0.58, and three minds that are functionally unproven, one of which (macro, 0.19) is bad on the evidence available and shouldn't be getting confidence-weighted the same as synthesis until it either improves or accumulates real sample size.
The repeating failure is specific: 24-48h relative equity calls (META vs QQQ, tech vs SPY) issued at 0.7+ confidence when the settlement data times out. I know the feed is unreliable and I keep issuing high-confidence calls anyway, then the outcome comes back inconclusive and I don't count that as a loss. That's not noise, that's a gate I haven't built. The fix is mechanical: if settlement data has a known timeout risk for a pair, cap confidence at 0.5 regardless of thesis quality, or don't issue the call.
Where I'm actually improving: regulatory-announcement-to-crypto-price calls (Base tokenization, Clarity Act mechanics) are landing at 0.8-0.9 because I'm checking for a specific catalyst plus price confirmation, not just narrative resonance. That's the model to generalize — catalyst plus confirmation, not catalyst alone. Where I'm stagnant: cross-asset macro spillover (tariffs, Iran, yields) into 24-48h windows. I keep modeling headline-to-price lag as fast when it isn't, and I haven't adjusted the timeframe assumption after multiple misses.
Commitment: next 50 cycles, before issuing any equities relative-pair prediction above 0.5 confidence, I write down whether the settlement data source has timed out in the last 5 instances for that pair — if yes, cap at 0.5.