FREE4/4+0ACC--

Twenty Backtests - selection bias practice

← Statistics
STATS · TWENTY BACKTESTS
LEVEL 1/4
3:00
SCORE 0
20 TESTED
THE BEST CURVE IN THE BOOK
--
EVERY SHARPE IN THE BOOK
0.00NO-EDGE BEST 1.871.00
SHARPE STD ERROR
1.00
AT 252 DAYS
BEST OF 20 IF NOTHING IS REAL
?
BEST OBSERVED
--
GAP
?
COMPUTE IT
LEVEL 1 · 20 STRATEGIES · 1Y EACH

Which one do you fund?

Twenty analysts, twenty strategies, one year of history each. Pick the one to fund.

THE SELECTION PROBLEM

With 20 candidates and 252 days each, the best Sharpe you would see from pure noise is about 1.87. Anything at or below that is not evidence.


SCORE0
LEVELS DONE0/4
LEVEL LOG0 CALLS

First book of the session.

CLICK A STRATEGY TO FUND IT · N FUND NONE · B BENCHMARK3 LEVELS TO GO

About Twenty Backtests

Decide which of a book of backtested strategies - if any - deserves funding, once you account for how many were tried to find the winner.

Each of four levels hands you a book of backtested strategies - equity curves, in-sample Sharpe ratios, total returns, and max drawdowns - and one question: which one do you fund? The levels vary the three things that decide the answer: how many strategies were tried (20 or 50), how much history each has (1, 2, or 5 years of daily returns), and whether any strategy genuinely has an edge. In two of the levels every single strategy is pure noise, and the correct answer is to fund nothing.

Each strategy's daily returns are drawn at 1% daily volatility with a drift implied by its true annualised Sharpe, which is zero for all but at most one hidden strategy per level. The winners that look seductive got there by being the best of many. Each level runs on a 3-minute clock; letting it expire counts as declining to fund.

Before deciding, you can press one button to compute the no-edge benchmark: the Sharpe standard error at this history length, the expected best-of-N Sharpe if every strategy were pure noise, the best Sharpe actually observed, and the gap between them in standard-error units, with a three-band verdict (clearly beyond luck / suggestive / in line with luck). After you choose, the level reveals the truth and shows your pick's out-of-sample performance.

Why quant interviews test this

Multiple testing and backtest overfitting are core interview territory for quant research seats and increasingly for risk and allocator roles. Classic prompts: 'An analyst shows you a Sharpe 2 backtest - what do you ask?' (how much data, how many variants were tried, what does out-of-sample look like), 'What is the standard error of a Sharpe ratio?' (roughly sqrt(252/T) annualised from daily data), and 'What does the best of N random strategies look like?' (an order-statistics question - the expected maximum grows with N even when every strategy is worthless).

The deeper skill being tested is selection-aware thinking: any time you see a winner - a strategy, a fund, a published result - the first question is how big the pool it was selected from was. Interviewers use this to separate candidates who evaluate evidence from candidates who evaluate performance.

The Twenty Backtests guide covers how scoring works, the strategy that wins, a worked example and the mistakes most players make.

More Statistics games