11
Can AI Actually Pick Stocks?
A Rigorous Look at the Evidence — and the One Question No One Running These Portfolios Bothers to Answer
Figures as of mid-2026
Related — See Vaunt's live AI Equity Research →
Our public, daily-updated three-model portfolio experiment and sector-rotation picks vs SPY.
In February 2026, a financial media company published results showing that a portfolio managed by ChatGPT had returned 69% over the preceding twelve months, while a portfolio managed by Grok had returned 42% over the same period. The S&P 500 returned approximately 12% in the same window. The story spread widely. The implication was clear: AI had beaten the market by a margin that would make most professional fund managers envious[29].
What the coverage did not include — in any outlet that republished the story — was the simplest possible follow-up question: what would a basket of sector ETFs, weighted to match the implied sector exposures of these AI portfolios, have returned over the same period? The question was not asked because the answer would have complicated the narrative.
The Sector Concentration Problem
Both the ChatGPT and Grok portfolios in the referenced experiment were heavily weighted toward technology and semiconductors — Nvidia, Vertiv, Broadcom, AMD, and similar names. Large language models, trained predominantly on financial media, analyst reports, and earnings call transcripts from the preceding years, have absorbed the conventional wisdom that artificial intelligence is the defining investment theme of the current period. When asked to pick stocks, they pick what the financial press has been discussing most intensively.
MarketScreener, in its own coverage of the experiment, noted with admirable directness: "It is hardly surprising to see these portfolios outperform the US benchmark, given that tech and AI optimism are precisely what drove performance." What the coverage did not do was quantify the observation — to ask whether the AI's stock selection added anything above the sector bet it had implicitly made.
This question has a name in quantitative finance. It is the attribution analysis — the decomposition of a portfolio's return into the component generated by factor exposure (being in the right sectors at the right time) and the component generated by stock selection (picking the right companies within those sectors)[33]. A fund that generates 20% returns in a year where growth stocks return 22% has not beaten the market — it has underperformed its factor exposure.
When you apply the attribution test to the AI portfolios: iShares Semiconductor ETF (SOXX) returned approximately 60% over the twelve months ending February 2026. A simple portfolio of 60% SOXX and 40% QQQ — matching the rough sector exposure of the ChatGPT portfolio — would have returned approximately the same 69%. The AI did not beat the market. It expressed a sector view that happened to be correct.
The Academic Evidence on Systematic Stock Selection
The academic literature on systematic stock-picking strategies is extensive, rigorous, and unambiguous on a single point: durable alpha — excess returns that cannot be explained by exposure to known risk factors — is rare, hard to identify, and harder still to maintain. The Fama-French five-factor model, refined over decades of research, explains the bulk of cross-sectional stock return variation through exposure to market, size, value, profitability, and investment factors[33]. Momentum, documented systematically by Jegadeesh and Titman in 1993, adds a sixth widely-accepted factor[30]. The implication is that most strategies that appear to generate excess returns are, on examination, just expressing systematic exposure to one or more of these factors.
Research on alternative data — the systematic use of non-traditional information sources like satellite imagery, credit card transactions, and web-scraped data — has documented that information advantages exist but are short-lived and competitive rapidly[34]. The structural problem AI faces in stock selection is the same problem every investor faces: in a market where millions of participants are simultaneously analysing the same publicly available information, finding non-consensus insights that are also correct is extraordinarily hard.
The Test That Reveals the Truth
The honest test of any stock-picking system — AI-powered or otherwise — has three components. The first is the attribution question: how does the portfolio perform after subtracting the return that would have come from a simple sector or factor benchmark with equivalent exposures? Performance that disappears after factor adjustment was never alpha.
The second is the out-of-sample question: does the performance persist across market environments that are meaningfully different from the training period? A technology-heavy AI portfolio that outperforms in a technology bull market is not informative about AI stock-picking ability. The same portfolio in a regime where technology is out of favour is more informative.
The third is the statistical significance question: how many picks, over how many periods, are required to distinguish systematic skill from systematic luck? Research by David Bailey and Marcos Lopez de Prado on backtest overfitting suggests that even a few years of live track record is often insufficient to reject the null hypothesis of luck[32]. For an AI portfolio selecting 20 stocks per month, the genuine signal becomes visible only over multiple years and multiple market regimes[31] — not one compelling twelve-month window.
None of this means AI stock-picking is without value. It means the value requires rigorous testing to verify — the same testing framework that applies to any investment strategy, regardless of how impressive its underlying technology sounds.
How Vaunt Thinks About This
Vaunt's Research feature publishes monthly sector picks from multiple independent AI agents because we believe AI has a genuine, if modest, advantage in synthesising non-obvious connections between economic trends and company positioning. We also believe the only honest way to find out whether that advantage is real is to track every pick against its sector ETF benchmark — not just the winners — and publish the results in full, including the periods where buying the ETF would have been the smarter choice. The agent agreement signal (when multiple models independently arrive at the same company) is what we weight most heavily, precisely because it reduces the risk that any single model's sector bias is driving the result.
Educational purposes only — not financial advice.