← VauntInsights

13

The Track Record Problem

Why Most Investment Performance Records Are Meaningless — and the Five Questions That Separate Real Evidence from Expensive Decoration

Figures as of mid-2026

In 1994, a hedge fund called Long-Term Capital Management raised $1.25 billion — then the largest launch in hedge fund history — on the strength of a track record that didn't exist. LTCM had been operating for less than a year. What it had was a team: two Nobel Prize winners in economics, a former vice chairman of the Federal Reserve, and some of the most celebrated quantitative traders of the era. Four years later, LTCM lost $4.6 billion in five weeks and required a Federal Reserve-orchestrated bailout to prevent cascading financial system failure[41].

LTCM is the extreme case. But the failure mode it represents — the substitution of credentials, story, and presentation for genuine, honest, independently verifiable evidence of investment skill — is not extreme at all. It is the norm. The financial industry has spent decades perfecting the art of presenting the appearance of a track record while carefully avoiding the substance of one.

The Six Ways a Track Record Can Lie Without Lying

Incubation bias — the most common and the least discussed. An asset manager runs ten different strategies internally, in paper or with small seed capital, for two years. Three of them perform well. Seven do not. The manager launches the three good ones publicly. The seven failures are never mentioned. The investor sees the three winners; the seven failures have been edited out.

Backtest overfitting — the systematic strategy's version of incubation bias. A quantitative researcher tests thousands of parameter combinations until they find one that looks extraordinary on historical data. With enough parameters and enough history, some combination will fit the noise perfectly. Research by Bailey and Lopez de Prado has quantified this rigorously: the more combinations tested, the higher the minimum required Sharpe ratio of the strategy to reflect genuine edge rather than overfitting[42]. Most published backtests don't come close to clearing this bar[45].

Benchmark selection — choosing a benchmark that makes the strategy look good rather than one that accurately reflects the strategy's risk exposure. A strategy that holds concentrated technology stocks and beats the S&P 500 looks impressive until you compare it to the Nasdaq, which is its actual relevant benchmark.

Gross-of-fees reporting — presenting returns before management fees, performance fees, and transaction costs are subtracted. The Global Investment Performance Standards (GIPS), developed by the CFA Institute, require net-of-fees presentation as the primary disclosure[43]. Most retail investment products are not held to this standard.

Short windows and cherry-picked periods — reporting performance over the window that starts just before a strong run and ends just after, omitting the periods of underperformance before and after. This is legal. It is common. It is presented without apology[44].

Risk-adjusted omission — reporting absolute returns without the risk metrics that give them context. The Sharpe ratio — return per unit of risk — is the basic tool for comparing strategies with different risk profiles, and its absence from a performance presentation is a reliable signal that the risk profile does not reflect well on the strategy[46].

How Long Does It Take to Know If a Strategy Is Real?

This is the question that the financial industry has the greatest incentive to avoid, because the honest answer is deeply inconvenient for anyone trying to sell a track record.

Research by Bailey and Lopez de Prado suggests that for a strategy with a Sharpe ratio of 1 — considered a good strategy in systematic investing — approximately three years of monthly data is required to reject the null hypothesis of luck at the 95% confidence level. For a Sharpe ratio of 0.5, the required period extends to more than ten years[42].

In practice, this means that a one-year track record — which is what most investment products present when launching, and which most media coverage treats as meaningful — is statistically insufficient to distinguish a genuinely skilled strategy from one that got lucky. A three-year track record, spanning at least one significant market stress event, begins to provide useful information. A ten-year track record, spanning multiple cycles, is the minimum required to have high confidence in the distinction between skill and luck for lower-Sharpe strategies.

This does not mean a short track record is worthless — it means its meaning is different from what it appears to be. A one-year live track record is evidence that the strategy works operationally, that the execution is clean, that the stated methodology is actually implemented. It is not evidence that the strategy has genuine, durable edge.

The Five Questions That Separate Evidence from Decoration

One: Does the track record include everything, or only the survivors? A manager who ran ten strategies and publishes three is showing you selection bias, not skill.

Two: Is the performance net of all fees, costs, and realistic transaction costs? Any performance figure that cannot be confirmed as net of these costs should be treated as uninformative.

Three: What is the appropriate benchmark for this strategy's actual risk exposures? A strategy should be compared to a benchmark that reflects the risks it takes, not the benchmark that makes it look best.

Four: Has the strategy been tested on data it didn't get to see during development? If a backtest uses the same data for both development and evaluation, it has proven nothing.

Five: Is the historical drawdown presented honestly, and does the strategy have a coherent explanation for surviving it? The maximum drawdown should be the first number discussed, not the last.

How Vaunt Thinks About This

This article is the standard we hold ourselves to, and publishing it is an act of transparency we think most investment platforms would prefer to avoid. Our own track record is, at the time of writing, short — first live capital deployed in July 2026. What we do have is: a testing process that addressed survivorship bias (we killed more strategies than we kept), out-of-sample validation (OOS Sharpe exceeded in-sample Sharpe on the dip strategy), honest cost modelling (real order-book slippage at actual position sizes), and a maximum drawdown figure — 74% on the dip strategy in its worst historical period — that appears in the first paragraph of every relevant document, not the last. As the live data accumulates, every result — including the months that underperform — will be published here. The track record that earns trust is not the one with the best numbers. It is the one that never edits anything out.

Educational purposes only — not financial advice.