The simplest way to build a reusable Yahoo Finance dataset in Python is to download data with the unofficial yfinance library, normalize it into one row per ticker and date, validate the results, and save both the data and its retrieval settings. The workflow below covers historical prices, dividends, splits, returns, fundamentals, common failures, and the point at which a licensed data provider becomes a better choice.
What you will build
The recommended canonical table is long (tidy) format:
| date | ticker | open | high | low | close | volume | dividends | stock_splits |
|---|---|---|---|---|---|---|---|---|
| 2025-01-02 | AAPL | … | … | … | … | … | 0 | 0 |
This layout is easy to filter, group, load into SQL or Parquet, and use in machine-learning pipelines. A download of several symbols normally arrives with pandas MultiIndex columns; that is useful for analysis but awkward for interchange, so the tutorial converts it.
Important limits before you start
yfinance is an open-source, unofficial way to access Yahoo Finance data. It is not a guaranteed commercial Yahoo API, and its documentation directs users to Yahoo’s terms for permitted use. Treat the basic workflow as suitable primarily for personal, educational, and research projects; public display, redistribution, and commercial products require a separate licensing review. See the project’s status and warning at the yfinance documentation.
#1 Best Overall
Coverage is not guaranteed to be complete or permanent for every symbol. Yahoo can change endpoints, return temporary errors, revise corporate-action adjustments, or rate-limit repeated requests. Intraday history is especially restricted: the current download reference says intraday data cannot extend beyond the most recent 60 days.
Install an isolated Python environment
- Create an environment:
python -m venv .venv. - Activate it on Windows PowerShell:
.venvScriptsActivate.ps1; on macOS or Linux:source .venv/bin/activate. - Install the packages:
python -m pip install --upgrade pip, thenpython -m pip install yfinance pandas pyarrow.
The last package enables Parquet output. The project’s installation page is here.
Download one stock
import yfinance as yf
ticker = yf.Ticker("AAPL")
prices = ticker.history(
start="2015-01-01",
end="2026-01-01",
interval="1d",
auto_adjust=True,
actions=True,
)
print(prices.head())
print(prices.dtypes)
start is inclusive and end is exclusive, so this request does not include January 1, 2026. Ticker.history() also supports pre/post-market selection, rounding, timeouts, and other controls documented at the price-history reference.
Download several tickers
import yfinance as yf
tickers = ["AAPL", "MSFT", "GOOG", "AMZN"]
prices = yf.download(
tickers=tickers,
start="2015-01-01",
end="2026-01-01",
interval="1d",
auto_adjust=True,
actions=True,
group_by="column",
multi_level_index=True,
threads=True,
progress=False,
)
print(prices.head())
print(prices.columns)
print(prices.columns.nlevels)
Set the important options explicitly rather than copying an old tutorial. The current API reference lists auto_adjust=True, actions=False, and multi_level_index=True as defaults, while older examples often assume different behavior. The function accepts a string or a list of symbols and can use threads for multiple downloads.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesAdjusted prices, quoted prices, and corporate actions
auto_adjust=True
OHLC prices are adjusted for corporate actions according to the library’s current behavior. For historical performance and return calculations, the resulting Close should generally be treated as the adjusted closing series.
Rank #2
auto_adjust=False
You receive raw-style OHLC fields and, depending on the returned schema and version, an additional adjusted-close field. This is preferable when you need the price that was quoted at the time rather than a performance-adjusted series.
actions=True
Yahoo returns dividend and stock-split events alongside prices. A dividend is a cash distribution; a split changes share count and per-share price. Neither event should be casually added to closing prices, and a split is not itself an economic gain or loss. The relevant parameters are documented at download and price history.
Convert MultiIndex output to long format
Inspect the actual structure first because pandas and yfinance versions can differ:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
print(prices.columns)
print(prices.columns.nlevels)
For common two-level output (field, ticker), this produces one row per ticker and date:
import pandas as pd
long_prices = (
prices
.rename_axis(index="date", columns=["field", "ticker"])
.stack(level="ticker", future_stack=True)
.reset_index()
)
long_prices.columns = [
str(column).lower().replace(" ", "_")
for column in long_prices.columns
]
print(long_prices.head())
If your pandas version does not support future_stack, use:
long_prices = prices.stack(level=1, dropna=False).reset_index()
long_prices.columns.name = None
A single-ticker download may already have ordinary columns:
single = yf.download(
"AAPL", start="2015-01-01", end="2026-01-01",
auto_adjust=True, actions=True, progress=False
).reset_index()
single["ticker"] = "AAPL"
single.columns = [str(c).lower().replace(" ", "_") for c in single.columns]
The project’s MultiIndex and reshaping examples are maintained in its README.
Validate before analysis
Reject empty results
if prices is None or prices.empty:
raise ValueError("No data was returned. Check the ticker, dates, interval, or connection.")
Check columns and duplicate observations
required = {"open", "high", "low", "close", "volume"}
missing = required - set(long_prices.columns)
if missing:
raise ValueError(f"Missing columns: {sorted(missing)}")
duplicates = long_prices.duplicated(["date", "ticker"]).sum()
if duplicates:
raise ValueError(f"Found {duplicates} duplicate date/ticker rows.")
Check OHLC relationships
bad_rows = long_prices[
(long_prices["high"] < long_prices["low"]) |
(long_prices["high"] < long_prices["open"]) |
(long_prices["high"] < long_prices["close"]) |
(long_prices["low"] > long_prices["open"]) |
(long_prices["low"] > long_prices["close"])
]
print(bad_rows)
Review coverage, nulls, and time zones
long_prices["date"] = pd.to_datetime(long_prices["date"], utc=True)
long_prices = long_prices.sort_values(["ticker", "date"])
coverage = long_prices.groupby("ticker").agg(
first_date=("date", "min"),
last_date=("date", "max"),
rows=("date", "size"),
missing_close=("close", lambda s: s.isna().sum()),
)
print(coverage)
Exchange calendars differ, and timestamp semantics are not identical across markets. The download reference documents interval-dependent time-zone handling and the ignore_tz option.
Calculate returns safely
Use one consistent adjustment policy. With adjusted closes:
long_prices["return_1d"] = (
long_prices.sort_values(["ticker", "date"])
.groupby("ticker")["close"].pct_change()
)
For log returns:
import numpy as np
long_prices["log_return_1d"] = (
long_prices.sort_values(["ticker", "date"])
.groupby("ticker")["close"]
.transform(lambda s: np.log(s).diff())
)
Do not calculate returns from raw closes across an unhandled split. A serious backtest also needs delisted securities, survivorship-bias controls, corporate-action timing, trading calendars, transaction costs, and realistic execution assumptions.
Save CSV, Parquet, and metadata
long_prices.to_csv("yahoo_prices.csv", index=False)
long_prices.to_parquet("yahoo_prices.parquet", index=False)
CSV is easiest to inspect and share. Parquet usually preserves types better and is more efficient for larger datasets. Record how the file was produced:
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →from importlib.metadata import version
from datetime import datetime, timezone
import json
metadata = {
"source": "Yahoo Finance via yfinance",
"downloaded_at_utc": datetime.now(timezone.utc).isoformat(),
"tickers": tickers,
"start": "2015-01-01",
"end_exclusive": "2026-01-01",
"interval": "1d",
"auto_adjust": True,
"actions": True,
"yfinance_version": version("yfinance"),
}
with open("dataset_metadata.json", "w", encoding="utf-8") as file:
json.dump(metadata, file, indent=2)
Dividends, splits, and fundamentals
Keep event data conceptually separate from price observations:
stock = yf.Ticker("MSFT")
dividends = stock.dividends
splits = stock.splits
income_statement = stock.income_stmt
quarterly_income = stock.quarterly_income_stmt
balance_sheet = stock.balance_sheet
cash_flow = stock.cashflow
yfinance also exposes holders, shares, options, analyst information, and other ticker-level features. Fundamentals are periodic and can have reporting dates, filing dates, restatements, and publication lags. For a backtest, use a value only after it was publicly available; do not join a statement to daily prices by calendar date alone.
Intraday data is a different problem
The current reference lists intervals including 1m, 2m, 5m, 15m, 30m, 60m, 90m, 1h, 1d, 5d, 1wk, 1mo, and 3mo. Intraday history cannot extend beyond 60 days:
intraday = yf.download(
"AAPL",
period="30d",
interval="5m",
prepost=False,
auto_adjust=False,
progress=False,
)
Do not use period="max" with a minute interval. Long historical intraday work generally requires a provider that explicitly licenses that history.
Recommended Free Tools
Best Value
Retries, caching, and reproducibility
Temporary network failures, rate limits, renamed symbols, exchange suffix errors such as 7203.T or VOD.L, and changing Yahoo responses are normal failure modes. Retries should recover transient failures, not evade provider restrictions:
import time
import yfinance as yf
def download_with_retries(ticker, attempts=3, pause_seconds=5):
last_error = None
for attempt in range(attempts):
try:
result = yf.download(
ticker, period="max", interval="1d",
auto_adjust=True, actions=True, progress=False,
timeout=30,
)
if result is not None and not result.empty:
return result
except Exception as error:
last_error = error
if attempt < attempts - 1:
time.sleep(pause_seconds * (attempt + 1))
if last_error:
raise RuntimeError(f"Download failed for {ticker}") from last_error
raise RuntimeError(f"No data returned for {ticker}")
- Cache raw responses or normalized files.
- Download large symbol lists in batches.
- Record failures and rerun only failed symbols.
- Pin the yfinance version in your requirements file.
- Check row counts and date coverage before replacing an existing dataset.
The yfinance project specifically recommends caching and rate limiting to reduce triggering Yahoo's limiter or blocker: project repository.
When another provider is safer
| Provider | Consider it when | Trade-off |
|---|---|---|
| Yahoo Finance via yfinance | Personal learning, daily research, prototypes, and small-to-moderate datasets | Unofficial access, possible rate limits, changing endpoints, and licensing uncertainty |
| Alpha Vantage | You want an official API-key workflow, JSON endpoints, adjusted series, or indicators | You manage keys, limits, endpoint-specific formats; commercial use requires contacting sales |
| Tiingo | You want paid EOD access and clearer usage tiers | Pricing and internal-use terms vary; confirm display or redistribution rights |
| Massive | You need deeper US equity history, aggregates, quotes, trades, or intraday access | Higher cost and business licensing considerations |
| Nasdaq Data Link | You need economic, alternative, or specialist datasets | Coverage, update frequency, pricing, and license vary by dataset |
Vendor pricing changes. On August 18, 2026, the listed individual Massive plans were Basic $0/month, Starter $29/month, Developer $79/month, and Advanced $199/month at Massive pricing. Tiingo displayed Starter at $0/month and Power at $30/month for individuals, with separate internal-commercial pricing information, at Tiingo pricing. Recheck current terms before buying.
Alpha Vantage documents its endpoints and commercial-use note at its official documentation. Nasdaq Data Link explains its API tools and free-versus-premium distinction at its documentation and getting-started guide.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Quick Recap
Final checklist
- Choose adjusted or unadjusted prices for a stated analytical purpose.
- Set dates, interval, actions, and adjustment options explicitly.
- Normalize multi-ticker output to one ticker/date row.
- Validate empty results, columns, duplicates, OHLC relationships, nulls, and coverage.
- Save CSV or Parquet together with retrieval metadata and package version.
- Cache downloads and limit retries.
- Keep fundamentals and corporate-action events logically separate from daily prices.
- Review licensing before public display, redistribution, or commercial use.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




