R2T v2: The Return, the Upgrade
Millions more backtests, but this time, they can say "no" and mean it
When I launched the Really, Really Thorough Backtests, I was, in part, responding a specific genre of content. A “guru” runs one backtest, on one instrument, over one favorable window, and presents the resulting equity curve as proof of a system.
At the time, I wrote that this kind of content was at best misleading and at worst outright fraud.
I stand by that.
In fact, the problem has gotten worse since I wrote it, because the ability to run large-scale backtests keeps proliferating, and easier backtesting makes cherry-picking easier. When you can run a thousand backtests in an afternoon, the spectacular curve you inevitably find proves nothing beyond your patience and your ability to use a max function.
R2T v1 was my answer to that. Instead of showing one hand-picked curve, run millions of simulations across entire universes of securities and show the whole distribution of outcomes, including the ugly parts. I was and am proud of that. It planted a flag in both methodology and ideology.
However, I always knew R2T v1 had limits, and they bothered me. So I rebuilt the whole thing.
This post is a preview of R2T v2: what changed, why, and what’s coming. The short version is that v2 doubles down on the lane I want to occupy: honest, transparent results, published whether or not they flatter a strategy.
From distributions to verdicts
V1 gave you distributions and let you judge. That was already a big step past cherry-picking, but it left the most important question hanging in the air: does this strategy actually work? I could show you a cloud of out-of-sample Sharpe ratios, and you still had to decide for yourself what it meant.
V2 answers the question directly. Every evaluation now produces a verdict: supported, mixed, not supported, or, when the evidence is too thin to say, insufficient evidence.
Each verdict attaches to a pre-registered claim rather than a vibe. The claim has a precise scope: this signal family, under this trading policy, across this universe, over this time period, under this version of the methodology. Change any piece of that, and you have a different claim with its own verdict. There is no “this strategy is good,” only “this specific claim, tested this specific way, earned this specific verdict.”
Why is this more than simple labeling? Two reasons:
The criteria are set before the results exist. The thresholds for “supported” are written down, versioned, and published. If I ever change them, the change gets its own version number, a changelog entry, and a side-by-side recompute. I cannot quietly move the goalposts after seeing a result I don’t like.
History is append-only. Nothing already published ever changes retroactively. Re-runs on new data or new criteria become new result sets that link back to what they supersede. This makes the entire trail transparent.
To be clear about what a verdict is not: verdicts never sort into “best strategy” lists, and no verdict is a recommendation. A verdict is evidence of effectiveness stated consistently, so that when you see “supported” or “not supported,” you know exactly what standard of research was applied.
Expect to see “not supported” a lot. As we saw with v1, Robust walk-forward testing makes most popular strategies look much worse than their proponents suggest, which means this is going to print a great deal of red.
I am perfectly okay with that. A methodology that can’t say no is just marketing.
The upgrade I’m proudest of: no more survivorship bias
V1 had one limitation that made me grit my teeth every time I disclosed it, and I disclosed it every time. I could only test current index constituents. That’s survivorship bias, and it’s not a small problem. When your universe consists only of companies that survived long enough to be in the index today, you’ve silently excluded every company that failed or got taken out along the way. Your backtest is auditioning on a stage where the failures were removed before the curtain went up.
V2 fixes this at the data layer. I now use constituents back to the 1990s: every company that was ever a member, their full price series, and the exact windows when they were in the index. Backtests only trade names during their actual membership windows. Delisted companies are in the data, trading right up until they stopped trading.
Every R2T result from v2 onward is free of survivorship bias. I have not found another publicly available resource doing consistent, large-scale, survivorship-bias-free strategy evaluation in this format for retail readers. This is the kind of data work that normally lives inside institutional research desks, and building it was most of the reason v2 took as long as it did.
And because I know some of you are wondering how much survivorship bias actually matters: a follow-up piece is coming where I run the same methodology on both datasets, current constituents only versus full historical membership, and show you exactly how the results move.
The data gremlins
A dataset this deep comes with a huge number of small horrors, and edge cases are exactly where backtests lie most easily. So a lot of v2’s substance is unglamorous consistency work. For example:
Tradable versus anomalous bars. Old data is full of judgment calls. March 2020 and the meme-stock era produced unbelievable volatility spikes that were nonetheless real, exchange-printed prices, and those stay. A dollar stock with a random 400% wick on no volume in 1998 is a different story. After an exhaustive amount of exploratory analysis, v2 applies consistent data-quality rules that exclude untradable bars from entry eligibility, so performance reflects what could actually have been traded at the time.
Windfalls that don’t certify. If a strategy happened to be holding a name that got acquired at a premium, that profit is real and stays in the portfolio results. But two lucky trades on one name don’t get to certify a claim about breadth. Trades that thin contribute to P&L, never to verdicts.
Results that round to zero. In experimental runs, one test counted a universe-median Sharpe ratio of +0.000184 as a positive result, which is a generous name for zero wearing a costume. The registered criteria now require a minimum margin before a number that small counts as evidence for anything. I would rather report weaker-looking results than let a Sharpe ratio of 0.0002 quietly pad a strategy’s case.
Short-side accounting. Unstopped short positions on collapsing names can produce mathematically undefined equity paths. The verdict-standard strategy wrapper includes a static stop for exactly this reason: a minimal guard that keeps the accounting well-defined, with nothing more claimed for it.
None of these choices is exciting on its own. Together, they make a claim far more robust and trustworthy.
What’s coming
The destination for all of this is a scoreboard: strategies as rows, universes as columns, verdicts in the cells. Small-, mid-, and large-cap US equities, FX majors, core ETFs, and commodities, all held to the same registered standard at once. At a glance: what works, where, and with how much evidence. Every cell carries its scope and provenance, so every verdict is transparent.
There are plenty of places to browse backtest results. This is meant to be something different: a consistent, versioned, survivorship-bias-free evaluation standard, applied at scale, that is allowed to tell you a strategy doesn’t work. Because most of them don’t, and you deserve a source that will say so.
The first verdicts are coming soon. Most of the cells are going to be red. The interesting part is finding out which ones aren’t.
Until next time, keep on the cutting edge, everyone.



