DriversRecommendedOutdated drivers can make a good PC feel brokenScan driver issues before chasing fixes manually.Scan NowOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run Scan×
Skip to content
The Finance Base
The Money Desk · Blog
Re:

Study: GPT-4 May Improve Some Stock-Picking Strategies—But It Hasn’t Been Proven to Beat the Market

GPT-4 research is promising but conditional: historical signals and backtests are not proof that ChatGPT can reliably beat the market for retail investors.
From TheFinanceBase Team6 min to read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Short answer: Several academic studies found that GPT-4 generated useful investment signals in controlled or historical tests. None proves that an ordinary investor can reliably make more money by following ChatGPT’s recommendations in live markets. The results depend on the model version, data, prompt, market, portfolio rules, benchmark, trading costs and risk.

Which “GPT-4 study” does the headline mean?

There is no single experiment showing that GPT-4 is a dependable money-making machine. Several papers tested different tasks, from interpreting news to ranking stocks. Calling all of them “GPT-4 picked stocks” hides important differences.

Study What GPT-4 did Reported result Test type and main limitation
Pelster and Val Rated investment opportunities using internet information Attractiveness ratings were positively associated with later earnings announcements and stock returns; the authors described a positive-return strategy Specific live-information experiment and portfolio methodology; not evidence that random consumer prompts work similarly
Lopez-Lira and Tang Classified financial-news headlines as positive, negative or neutral for future performance Scores predicted some subsequent price drift, especially after news and in smaller stocks; reported strategy returns declined as LLM adoption increased News-signal test, not complete financial advice; small-stock spreads and liquidity can erase theoretical gains
LoGrasso Selected stocks using information constrained to historical decision dates Reported approximately 1% average monthly alpha for selected two-year holding periods beginning July 1 in each year from 1985 through 2021 Retrospective simulation, not an audited live record; sensitive to prompts, universe, dates and portfolio construction
MarketSenseAI Combined GPT-4 with financial statements, prices, news, macroeconomic data, APIs and ranking rules Reported superior total and risk-adjusted returns for some GPT-based ranking strategies Engineered research system, not an unassisted chatbot; relatively short evaluation and stated methodological assumptions
Risk-appetite study Constructed portfolios for different investor risk appetites and markets GPT-4o performed best in the tested U.S. setup, while GPT-4 performed best in the tested European setup Shows that model, geography and risk profile materially change outcomes

The Lopez-Lira and Tang work is also presented in a UCLA version at this link. The published version of the LoGrasso paper is available through Modern Finance, and the MarketSenseAI preprint is at arXiv.

What GPT-4 actually did in these experiments

“AI stock picking” can describe very different activities:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Scoring a company’s attractiveness.
  • Extracting facts from filings or earnings releases.
  • Classifying the tone of a news headline.
  • Ranking securities for a hypothetical portfolio.
  • Generating explanations alongside a separate quantitative strategy.

In most studies, GPT-4 did not determine a person’s asset allocation, tax strategy, position size, leverage, stop-loss rules or order execution by itself. The strongest systems supplied structured data and added software for selection, rebalancing and testing.

Backtest, retrospective test or live trading?

This distinction should come before any return figure. A backtest simulates rules on historical data. A retrospective test asks a model to recreate decisions using information from an earlier date. A live experiment produces assessments as information arrives, but trades may still be hypothetical. Paper trading uses simulated orders, while live trading uses real money and actual fills.

Most favorable GPT-4 findings are simulations, controlled experiments or hypothetical strategies—not a long-running, independently audited record showing that retail investors earned superior net returns.

Rank #2

What does “make more money” mean?

A higher headline return is not enough. A fair comparison asks whether the strategy delivered:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • More dollars or merely a higher percentage before costs.
  • A higher return than a comparable S&P 500 or total-market fund.
  • Better risk-adjusted performance than a risk-matched alternative.
  • Lower drawdowns, volatility and concentration.
  • Better after-tax results for the investor’s account and country.

Reported alpha is a statistical estimate relative to a benchmark or model. It is not a guaranteed return, and it does not show what one investor could capture after delays and execution.

Why GPT-4 might help

Language models can process large volumes of text quickly, apply a consistent classification prompt and turn filings or news into structured fields for a screening process. They can also expose assumptions in an investment thesis, compare user-supplied information and generate questions for further research.

Those capabilities may explain why some experiments found predictive signals. They do not establish that the model understands a company’s future cash flows or can forecast prices reliably.

Why a promising result may disappear

Information timing and data leakage

A historical decision is invalid if the model, data feed or researcher used a revised filing, later news, today’s surviving-company list or any other information unavailable at that time.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Overfitting and multiple testing

Trying many prompts, stocks, holding periods, model versions and portfolio rules makes an impressive result more likely to occur by chance. A genuinely untouched out-of-sample period is essential.

Costs and liquidity

Commissions, bid-ask spreads, slippage, market impact, taxes, data fees, API charges and rebalancing can consume a small edge. This is especially serious when signals favor small or thinly traded stocks.

Changing competition

Lopez-Lira and Tang reported declining strategy performance as LLM adoption increased. If many traders use the same information-processing signal, competition can reduce or eliminate it.

Model drift and prompt sensitivity

Different wording, data formatting, sampling settings, ticker symbols or model releases can produce different recommendations. Results from GPT-4 in 2023 or 2024 should not automatically transfer to GPT-4o, GPT-5.x or future systems.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Hidden factor exposure

An apparent AI advantage may simply be exposure to momentum, growth, profitability, large-cap technology or another conventional factor that happened to perform well during the sample.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

GPT-4’s practical failure modes

  • Hallucinations: OpenAI warns that GPT-4 can produce inaccurate information. Verify every material figure, date and citation against filings or authoritative market data at OpenAI’s GPT-4 research page.
  • Stale or incomplete data: A model may not have current prices, corporate actions or breaking news unless connected to a reliable source.
  • Overconfidence: Fluent explanations can make weak evidence sound certain.
  • Uncontrolled risk: A text model is not a substitute for controls on margin, options, short positions, leveraged ETFs or concentrated holdings.

How to evaluate an AI trading claim

  1. Identify the exact model and dates. Record the model version, prompt, inputs and decision timestamps.
  2. Establish reproducibility. Another person should be able to use the same data, universe and rules.
  3. Demand out-of-sample evidence. The test period must not have been used to design the strategy.
  4. Calculate net returns. Include spreads, slippage, commissions, taxes, data, API and subscription costs.
  5. Use a credible benchmark. Compare with a similar-risk index fund, factor strategy or buy-and-hold portfolio.
  6. Measure risk. Review maximum drawdown, volatility, Sharpe and Sortino ratios, turnover, concentration and factor exposure.
  7. Test stability. Small changes in prompt, model, news order or data format should not radically change the portfolio.
  8. Keep human approval. Do not let generated text change risk limits or place trades automatically without explicit controls.

Safer ways to use ChatGPT for investing research

More defensible uses include summarizing a filing, explaining an unfamiliar term, comparing companies with data you provide, stress-testing an investment thesis, identifying valuation assumptions, creating a risk checklist, translating a strategy into testable rules and reviewing backtest code.

OpenAI’s current personal-finance description says ChatGPT can help users understand financial information and investment risks, but is not a replacement for professional advice: see the product announcement. The newer product description is not evidence that current ChatGPT has the same behavior as the GPT-4 experiments.

Costs, tools and third-party services

ChatGPT may be useful as a research interface, while an API can be part of a larger pipeline combining filings, market data, alerts, portfolio rules and human approval. Current API information is listed at OpenAI’s pricing documentation. The original GPT-4 announcement listed launch pricing of $0.03 per 1,000 prompt tokens and $0.06 per 1,000 completion tokens; those figures are historical, not current prices.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A serious implementation still needs reliable corporate-action data, survivorship-bias-free history, a backtester, paper trading, a regulated broker, order controls and audit logs. FINRA warns investors about unregistered or unlicensed platforms claiming to use AI for investment advice at its generative-AI guidance. Buying an AI subscription or connecting a brokerage does not create a validated edge.

Verdict

GPT-4 has shown potential as a component of stock research and, in several specified experiments, produced signals associated with later returns. The evidence does not show that ChatGPT can reliably beat a low-cost index fund or make ordinary investors richer after risk, taxes and trading costs. Treat the model as an assistant for organizing and challenging research—not as an autonomous trading adviser.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More post from the Money Desk

  1. The Money DeskBlogTheFinanceBase07 MAR 2625 minWhat Is a 457 Plan?
  2. The Money DeskBlogTheFinanceBase07 MAR 2621 minTime Value of Money: What It Is and How It Works
  3. The Money DeskBlogTheFinanceBase07 MAR 2627 minAre You Living in One of These Top 10 Most Expensive Cities to Retire?
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.