Short answer: Several academic studies found that GPT-4 generated useful investment signals in controlled or historical tests. None proves that an ordinary investor can reliably make more money by following ChatGPT’s recommendations in live markets. The results depend on the model version, data, prompt, market, portfolio rules, benchmark, trading costs and risk.
Which “GPT-4 study” does the headline mean?
There is no single experiment showing that GPT-4 is a dependable money-making machine. Several papers tested different tasks, from interpreting news to ranking stocks. Calling all of them “GPT-4 picked stocks” hides important differences.
| Study | What GPT-4 did | Reported result | Test type and main limitation |
|---|---|---|---|
| Pelster and Val | Rated investment opportunities using internet information | Attractiveness ratings were positively associated with later earnings announcements and stock returns; the authors described a positive-return strategy | Specific live-information experiment and portfolio methodology; not evidence that random consumer prompts work similarly |
| Lopez-Lira and Tang | Classified financial-news headlines as positive, negative or neutral for future performance | Scores predicted some subsequent price drift, especially after news and in smaller stocks; reported strategy returns declined as LLM adoption increased | News-signal test, not complete financial advice; small-stock spreads and liquidity can erase theoretical gains |
| LoGrasso | Selected stocks using information constrained to historical decision dates | Reported approximately 1% average monthly alpha for selected two-year holding periods beginning July 1 in each year from 1985 through 2021 | Retrospective simulation, not an audited live record; sensitive to prompts, universe, dates and portfolio construction |
| MarketSenseAI | Combined GPT-4 with financial statements, prices, news, macroeconomic data, APIs and ranking rules | Reported superior total and risk-adjusted returns for some GPT-based ranking strategies | Engineered research system, not an unassisted chatbot; relatively short evaluation and stated methodological assumptions |
| Risk-appetite study | Constructed portfolios for different investor risk appetites and markets | GPT-4o performed best in the tested U.S. setup, while GPT-4 performed best in the tested European setup | Shows that model, geography and risk profile materially change outcomes |
The Lopez-Lira and Tang work is also presented in a UCLA version at this link. The published version of the LoGrasso paper is available through Modern Finance, and the MarketSenseAI preprint is at arXiv.
What GPT-4 actually did in these experiments
“AI stock picking” can describe very different activities:
#1 Best Overall
- Scoring a company’s attractiveness.
- Extracting facts from filings or earnings releases.
- Classifying the tone of a news headline.
- Ranking securities for a hypothetical portfolio.
- Generating explanations alongside a separate quantitative strategy.
In most studies, GPT-4 did not determine a person’s asset allocation, tax strategy, position size, leverage, stop-loss rules or order execution by itself. The strongest systems supplied structured data and added software for selection, rebalancing and testing.
Backtest, retrospective test or live trading?
This distinction should come before any return figure. A backtest simulates rules on historical data. A retrospective test asks a model to recreate decisions using information from an earlier date. A live experiment produces assessments as information arrives, but trades may still be hypothetical. Paper trading uses simulated orders, while live trading uses real money and actual fills.
Most favorable GPT-4 findings are simulations, controlled experiments or hypothetical strategies—not a long-running, independently audited record showing that retail investors earned superior net returns.
Rank #2
- Comes with secure packaging
- Easy to read text
- It can be a gift option
What does “make more money” mean?
A higher headline return is not enough. A fair comparison asks whether the strategy delivered:
Recommended Free Tools
- More dollars or merely a higher percentage before costs.
- A higher return than a comparable S&P 500 or total-market fund.
- Better risk-adjusted performance than a risk-matched alternative.
- Lower drawdowns, volatility and concentration.
- Better after-tax results for the investor’s account and country.
Reported alpha is a statistical estimate relative to a benchmark or model. It is not a guaranteed return, and it does not show what one investor could capture after delays and execution.
Why GPT-4 might help
Language models can process large volumes of text quickly, apply a consistent classification prompt and turn filings or news into structured fields for a screening process. They can also expose assumptions in an investment thesis, compare user-supplied information and generate questions for further research.
Rank #3
Those capabilities may explain why some experiments found predictive signals. They do not establish that the model understands a company’s future cash flows or can forecast prices reliably.
Why a promising result may disappear
Information timing and data leakage
A historical decision is invalid if the model, data feed or researcher used a revised filing, later news, today’s surviving-company list or any other information unavailable at that time.
Overfitting and multiple testing
Trying many prompts, stocks, holding periods, model versions and portfolio rules makes an impressive result more likely to occur by chance. A genuinely untouched out-of-sample period is essential.
Rank #4
Costs and liquidity
Commissions, bid-ask spreads, slippage, market impact, taxes, data fees, API charges and rebalancing can consume a small edge. This is especially serious when signals favor small or thinly traded stocks.
Changing competition
Lopez-Lira and Tang reported declining strategy performance as LLM adoption increased. If many traders use the same information-processing signal, competition can reduce or eliminate it.
Model drift and prompt sensitivity
Different wording, data formatting, sampling settings, ticker symbols or model releases can produce different recommendations. Results from GPT-4 in 2023 or 2024 should not automatically transfer to GPT-4o, GPT-5.x or future systems.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Hidden factor exposure
An apparent AI advantage may simply be exposure to momentum, growth, profitability, large-cap technology or another conventional factor that happened to perform well during the sample.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.GPT-4’s practical failure modes
- Hallucinations: OpenAI warns that GPT-4 can produce inaccurate information. Verify every material figure, date and citation against filings or authoritative market data at OpenAI’s GPT-4 research page.
- Stale or incomplete data: A model may not have current prices, corporate actions or breaking news unless connected to a reliable source.
- Overconfidence: Fluent explanations can make weak evidence sound certain.
- Uncontrolled risk: A text model is not a substitute for controls on margin, options, short positions, leveraged ETFs or concentrated holdings.
How to evaluate an AI trading claim
- Identify the exact model and dates. Record the model version, prompt, inputs and decision timestamps.
- Establish reproducibility. Another person should be able to use the same data, universe and rules.
- Demand out-of-sample evidence. The test period must not have been used to design the strategy.
- Calculate net returns. Include spreads, slippage, commissions, taxes, data, API and subscription costs.
- Use a credible benchmark. Compare with a similar-risk index fund, factor strategy or buy-and-hold portfolio.
- Measure risk. Review maximum drawdown, volatility, Sharpe and Sortino ratios, turnover, concentration and factor exposure.
- Test stability. Small changes in prompt, model, news order or data format should not radically change the portfolio.
- Keep human approval. Do not let generated text change risk limits or place trades automatically without explicit controls.
Safer ways to use ChatGPT for investing research
More defensible uses include summarizing a filing, explaining an unfamiliar term, comparing companies with data you provide, stress-testing an investment thesis, identifying valuation assumptions, creating a risk checklist, translating a strategy into testable rules and reviewing backtest code.
OpenAI’s current personal-finance description says ChatGPT can help users understand financial information and investment risks, but is not a replacement for professional advice: see the product announcement. The newer product description is not evidence that current ChatGPT has the same behavior as the GPT-4 experiments.
Costs, tools and third-party services
ChatGPT may be useful as a research interface, while an API can be part of a larger pipeline combining filings, market data, alerts, portfolio rules and human approval. Current API information is listed at OpenAI’s pricing documentation. The original GPT-4 announcement listed launch pricing of $0.03 per 1,000 prompt tokens and $0.06 per 1,000 completion tokens; those figures are historical, not current prices.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
A serious implementation still needs reliable corporate-action data, survivorship-bias-free history, a backtester, paper trading, a regulated broker, order controls and audit logs. FINRA warns investors about unregistered or unlicensed platforms claiming to use AI for investment advice at its generative-AI guidance. Buying an AI subscription or connecting a brokerage does not create a validated edge.
Verdict
GPT-4 has shown potential as a component of stock research and, in several specified experiments, produced signals associated with later returns. The evidence does not show that ChatGPT can reliably beat a low-cost index fund or make ordinary investors richer after risk, taxes and trading costs. Treat the model as an assistant for organizing and challenging research—not as an autonomous trading adviser.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




