October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
The Finance Base
The Money Desk · Blog
Re:

How to Create a Financial Dataset with Yahoo Finance and Python

A practical yfinance guide for downloading historical Yahoo Finance data, converting it to tidy format, validating it, calculating returns, and saving reproducible CSV or Parquet files.
From TheFinanceBase Team7 min to read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The simplest way to build a reusable Yahoo Finance dataset in Python is to download data with the unofficial yfinance library, normalize it into one row per ticker and date, validate the results, and save both the data and its retrieval settings. The workflow below covers historical prices, dividends, splits, returns, fundamentals, common failures, and the point at which a licensed data provider becomes a better choice.

What you will build

The recommended canonical table is long (tidy) format:

date ticker open high low close volume dividends stock_splits
2025-01-02 AAPL … … … … … 0 0

This layout is easy to filter, group, load into SQL or Parquet, and use in machine-learning pipelines. A download of several symbols normally arrives with pandas MultiIndex columns; that is useful for analysis but awkward for interchange, so the tutorial converts it.

Important limits before you start

yfinance is an open-source, unofficial way to access Yahoo Finance data. It is not a guaranteed commercial Yahoo API, and its documentation directs users to Yahoo’s terms for permitted use. Treat the basic workflow as suitable primarily for personal, educational, and research projects; public display, redistribution, and commercial products require a separate licensing review. See the project’s status and warning at the yfinance documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Coverage is not guaranteed to be complete or permanent for every symbol. Yahoo can change endpoints, return temporary errors, revise corporate-action adjustments, or rate-limit repeated requests. Intraday history is especially restricted: the current download reference says intraday data cannot extend beyond the most recent 60 days.

Install an isolated Python environment

  1. Create an environment: python -m venv .venv.
  2. Activate it on Windows PowerShell: .venvScriptsActivate.ps1; on macOS or Linux: source .venv/bin/activate.
  3. Install the packages: python -m pip install --upgrade pip, then python -m pip install yfinance pandas pyarrow.

The last package enables Parquet output. The project’s installation page is here.

Download one stock

import yfinance as yf

ticker = yf.Ticker("AAPL")
prices = ticker.history(
    start="2015-01-01",
    end="2026-01-01",
    interval="1d",
    auto_adjust=True,
    actions=True,
)

print(prices.head())
print(prices.dtypes)

start is inclusive and end is exclusive, so this request does not include January 1, 2026. Ticker.history() also supports pre/post-market selection, rounding, timeouts, and other controls documented at the price-history reference.

Download several tickers

import yfinance as yf

tickers = ["AAPL", "MSFT", "GOOG", "AMZN"]
prices = yf.download(
    tickers=tickers,
    start="2015-01-01",
    end="2026-01-01",
    interval="1d",
    auto_adjust=True,
    actions=True,
    group_by="column",
    multi_level_index=True,
    threads=True,
    progress=False,
)

print(prices.head())
print(prices.columns)
print(prices.columns.nlevels)

Set the important options explicitly rather than copying an old tutorial. The current API reference lists auto_adjust=True, actions=False, and multi_level_index=True as defaults, while older examples often assume different behavior. The function accepts a string or a list of symbols and can use threads for multiple downloads.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Adjusted prices, quoted prices, and corporate actions

auto_adjust=True

OHLC prices are adjusted for corporate actions according to the library’s current behavior. For historical performance and return calculations, the resulting Close should generally be treated as the adjusted closing series.

auto_adjust=False

You receive raw-style OHLC fields and, depending on the returned schema and version, an additional adjusted-close field. This is preferable when you need the price that was quoted at the time rather than a performance-adjusted series.

actions=True

Yahoo returns dividend and stock-split events alongside prices. A dividend is a cash distribution; a split changes share count and per-share price. Neither event should be casually added to closing prices, and a split is not itself an economic gain or loss. The relevant parameters are documented at download and price history.

Convert MultiIndex output to long format

Inspect the actual structure first because pandas and yfinance versions can differ:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
print(prices.columns)
print(prices.columns.nlevels)

For common two-level output (field, ticker), this produces one row per ticker and date:

import pandas as pd

long_prices = (
    prices
    .rename_axis(index="date", columns=["field", "ticker"])
    .stack(level="ticker", future_stack=True)
    .reset_index()
)
long_prices.columns = [
    str(column).lower().replace(" ", "_")
    for column in long_prices.columns
]
print(long_prices.head())

If your pandas version does not support future_stack, use:

long_prices = prices.stack(level=1, dropna=False).reset_index()
long_prices.columns.name = None

A single-ticker download may already have ordinary columns:

single = yf.download(
    "AAPL", start="2015-01-01", end="2026-01-01",
    auto_adjust=True, actions=True, progress=False
).reset_index()
single["ticker"] = "AAPL"
single.columns = [str(c).lower().replace(" ", "_") for c in single.columns]

The project’s MultiIndex and reshaping examples are maintained in its README.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validate before analysis

Reject empty results

if prices is None or prices.empty:
    raise ValueError("No data was returned. Check the ticker, dates, interval, or connection.")

Check columns and duplicate observations

required = {"open", "high", "low", "close", "volume"}
missing = required - set(long_prices.columns)
if missing:
    raise ValueError(f"Missing columns: {sorted(missing)}")

duplicates = long_prices.duplicated(["date", "ticker"]).sum()
if duplicates:
    raise ValueError(f"Found {duplicates} duplicate date/ticker rows.")

Check OHLC relationships

bad_rows = long_prices[
    (long_prices["high"] < long_prices["low"]) |
    (long_prices["high"] < long_prices["open"]) |
    (long_prices["high"] < long_prices["close"]) |
    (long_prices["low"] > long_prices["open"]) |
    (long_prices["low"] > long_prices["close"])
]
print(bad_rows)

Review coverage, nulls, and time zones

long_prices["date"] = pd.to_datetime(long_prices["date"], utc=True)
long_prices = long_prices.sort_values(["ticker", "date"])
coverage = long_prices.groupby("ticker").agg(
    first_date=("date", "min"),
    last_date=("date", "max"),
    rows=("date", "size"),
    missing_close=("close", lambda s: s.isna().sum()),
)
print(coverage)

Exchange calendars differ, and timestamp semantics are not identical across markets. The download reference documents interval-dependent time-zone handling and the ignore_tz option.

Calculate returns safely

Use one consistent adjustment policy. With adjusted closes:

long_prices["return_1d"] = (
    long_prices.sort_values(["ticker", "date"])
    .groupby("ticker")["close"].pct_change()
)

For log returns:

import numpy as np
long_prices["log_return_1d"] = (
    long_prices.sort_values(["ticker", "date"])
    .groupby("ticker")["close"]
    .transform(lambda s: np.log(s).diff())
)

Do not calculate returns from raw closes across an unhandled split. A serious backtest also needs delisted securities, survivorship-bias controls, corporate-action timing, trading calendars, transaction costs, and realistic execution assumptions.

Save CSV, Parquet, and metadata

long_prices.to_csv("yahoo_prices.csv", index=False)
long_prices.to_parquet("yahoo_prices.parquet", index=False)

CSV is easiest to inspect and share. Parquet usually preserves types better and is more efficient for larger datasets. Record how the file was produced:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from importlib.metadata import version
from datetime import datetime, timezone
import json

metadata = {
    "source": "Yahoo Finance via yfinance",
    "downloaded_at_utc": datetime.now(timezone.utc).isoformat(),
    "tickers": tickers,
    "start": "2015-01-01",
    "end_exclusive": "2026-01-01",
    "interval": "1d",
    "auto_adjust": True,
    "actions": True,
    "yfinance_version": version("yfinance"),
}
with open("dataset_metadata.json", "w", encoding="utf-8") as file:
    json.dump(metadata, file, indent=2)

Dividends, splits, and fundamentals

Keep event data conceptually separate from price observations:

stock = yf.Ticker("MSFT")
dividends = stock.dividends
splits = stock.splits

income_statement = stock.income_stmt
quarterly_income = stock.quarterly_income_stmt
balance_sheet = stock.balance_sheet
cash_flow = stock.cashflow

yfinance also exposes holders, shares, options, analyst information, and other ticker-level features. Fundamentals are periodic and can have reporting dates, filing dates, restatements, and publication lags. For a backtest, use a value only after it was publicly available; do not join a statement to daily prices by calendar date alone.

Intraday data is a different problem

The current reference lists intervals including 1m, 2m, 5m, 15m, 30m, 60m, 90m, 1h, 1d, 5d, 1wk, 1mo, and 3mo. Intraday history cannot extend beyond 60 days:

intraday = yf.download(
    "AAPL",
    period="30d",
    interval="5m",
    prepost=False,
    auto_adjust=False,
    progress=False,
)

Do not use period="max" with a minute interval. Long historical intraday work generally requires a provider that explicitly licenses that history.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Retries, caching, and reproducibility

Temporary network failures, rate limits, renamed symbols, exchange suffix errors such as 7203.T or VOD.L, and changing Yahoo responses are normal failure modes. Retries should recover transient failures, not evade provider restrictions:

import time
import yfinance as yf

def download_with_retries(ticker, attempts=3, pause_seconds=5):
    last_error = None
    for attempt in range(attempts):
        try:
            result = yf.download(
                ticker, period="max", interval="1d",
                auto_adjust=True, actions=True, progress=False,
                timeout=30,
            )
            if result is not None and not result.empty:
                return result
        except Exception as error:
            last_error = error
        if attempt < attempts - 1:
            time.sleep(pause_seconds * (attempt + 1))
    if last_error:
        raise RuntimeError(f"Download failed for {ticker}") from last_error
    raise RuntimeError(f"No data returned for {ticker}")
  • Cache raw responses or normalized files.
  • Download large symbol lists in batches.
  • Record failures and rerun only failed symbols.
  • Pin the yfinance version in your requirements file.
  • Check row counts and date coverage before replacing an existing dataset.

The yfinance project specifically recommends caching and rate limiting to reduce triggering Yahoo's limiter or blocker: project repository.

When another provider is safer

Provider Consider it when Trade-off
Yahoo Finance via yfinance Personal learning, daily research, prototypes, and small-to-moderate datasets Unofficial access, possible rate limits, changing endpoints, and licensing uncertainty
Alpha Vantage You want an official API-key workflow, JSON endpoints, adjusted series, or indicators You manage keys, limits, endpoint-specific formats; commercial use requires contacting sales
Tiingo You want paid EOD access and clearer usage tiers Pricing and internal-use terms vary; confirm display or redistribution rights
Massive You need deeper US equity history, aggregates, quotes, trades, or intraday access Higher cost and business licensing considerations
Nasdaq Data Link You need economic, alternative, or specialist datasets Coverage, update frequency, pricing, and license vary by dataset

Vendor pricing changes. On August 18, 2026, the listed individual Massive plans were Basic $0/month, Starter $29/month, Developer $79/month, and Advanced $199/month at Massive pricing. Tiingo displayed Starter at $0/month and Power at $30/month for individuals, with separate internal-commercial pricing information, at Tiingo pricing. Recheck current terms before buying.

Alpha Vantage documents its endpoints and commercial-use note at its official documentation. Nasdaq Data Link explains its API tools and free-versus-premium distinction at its documentation and getting-started guide.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Final checklist

  • Choose adjusted or unadjusted prices for a stated analytical purpose.
  • Set dates, interval, actions, and adjustment options explicitly.
  • Normalize multi-ticker output to one ticker/date row.
  • Validate empty results, columns, duplicates, OHLC relationships, nulls, and coverage.
  • Save CSV or Parquet together with retrieval metadata and package version.
  • Cache downloads and limit retries.
  • Keep fundamentals and corporate-action events logically separate from daily prices.
  • Review licensing before public display, redistribution, or commercial use.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More post from the Money Desk

  1. The Money DeskBlogTheFinanceBase07 MAR 2625 minWhat Is a 457 Plan?
  2. The Money DeskBlogTheFinanceBase07 MAR 2621 minTime Value of Money: What It Is and How It Works
  3. The Money DeskBlogTheFinanceBase07 MAR 2627 minAre You Living in One of These Top 10 Most Expensive Cities to Retire?
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.