Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversFall ResetAmazon USFall reset deals: check better picks before checkoutAmazon US: today's deals, useful picks and quick comparisons.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix Now×
Skip to content
Blog

Data Science for Portfolio Optimization: Markowitz Mean-Variance Theory

By TheFinanceBase Team10 min read

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Markowitz mean-variance theory turns estimates of asset returns and co-movement into portfolio weights. Its central insight is that a portfolio’s risk depends not only on each holding’s volatility but also on how holdings move together. The framework is useful for understanding diversification and building allocation tools—but its output is only as dependable as its estimates, constraints, and testing.

What Markowitz optimization does

Portfolio optimization asks: given a set of investable assets, estimates of their returns and covariance, and rules about what can be held, which allocation best meets a stated objective? It is an allocation method, not a security-selection oracle. The answer is optimal only for the inputs, objective, and constraints specified.

Harry Markowitz formalized the trade-off between expected return and portfolio risk in “Portfolio Selection,” published in The Journal of Finance in 1952. Read the paper. Modern portfolio theory is the broader framework; mean-variance optimization is one portfolio-construction method within it. The Capital Asset Pricing Model is a related, later theory about expected asset returns, not another name for the optimizer.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Return, variance, and covariance

Let w be a vector of portfolio weights, μ the vector of expected asset returns, and Σ the covariance matrix of returns. The portfolio’s estimated expected return is:

E(Rp) = wTμ

Its estimated variance is:

σp2 = wTΣw

Volatility is the square root of variance. Covariance captures whether two assets tend to move together; correlation is covariance scaled by each asset’s volatility. Because portfolio variance includes the relationships between holdings, assets with imperfectly correlated returns can reduce portfolio risk. Counting holdings alone does not establish diversification: exposures, correlations, and contributions to risk matter too.

Common portfolio objectives

  • Global minimum variance: the feasible portfolio with the lowest estimated variance.
  • Target return: the lowest-variance portfolio that meets a specified estimated return.
  • Target risk: the highest estimated return subject to a volatility limit.
  • Maximum Sharpe ratio: the portfolio with the highest estimated excess return per unit of estimated volatility. The ratio is S = (E(Rp) − Rf)/σp, where the risk-free rate must use a currency and period consistent with the returns.

These choices answer different questions; a maximum-Sharpe allocation is not automatically preferable to a minimum-variance allocation.

The efficient frontier

For a long-only portfolio with a target expected return, a standard formulation is:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Minimize   wTΣw
Subject to   wTμ ≥ μ*,   1Tw = 1,   wi ≥ 0

Here, μ* is the target return; the weight-sum condition means the portfolio is fully invested, and nonnegative weights prohibit short positions. Varying the target traces the efficient frontier: the boundary of feasible portfolios that offer the highest estimated return for a given estimated risk, or the lowest estimated risk for a given return. Under common constraints, this is a convex quadratic optimization problem. PyPortfolioOpt’s guide explains the standard formulation and frontier.

On a chart, annualized volatility is usually on the horizontal axis and annualized expected return on the vertical axis. The global minimum-variance portfolio is the lowest-risk point; the maximum-Sharpe portfolio is often called the tangency portfolio when a risk-free asset is part of the analysis. An equal-weight portfolio is a useful reference point. All plotted results are estimates, not promises about realized returns or future dominance.

Build the inputs before optimizing

The solver does not discover expected returns. It converts supplied assumptions into weights. Treat each input as an estimate or model choice, not a known fact.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose a defensible dataset

At minimum, define the asset universe, price or total-return history, estimation window, return frequency, rebalance schedule, constraints, and—if using the Sharpe ratio—a risk-free-rate assumption. Use total-return data or appropriately adjusted prices so dividends, distributions, and splits are handled consistently; raw closing prices can misstate returns. Record the data source and the date on which each observation would have been available.

Missing observations, assets that began trading at different times, and asynchronous market closes can distort comparisons. Decide how to handle them before estimating inputs, and do not silently use future information. For backtests, the universe must reflect what could actually have been selected at the time, including assets that later disappeared.

Estimate expected returns

A simple arithmetic historical estimate for asset i is:

μ̂i = (1/T) Σt=1T ri,t

For periodic returns, a common approximation annualizes the arithmetic mean as m × μ̂periodic, where m is the number of periods per year. A geometric return describes compounded historical growth, but it is not interchangeable with the arithmetic mean in every optimization setup. Alternatives include factor models, analyst forecasts, dividend-growth assumptions, equilibrium-implied returns, or Black-Litterman estimates. Each adds assumptions; none removes forecast uncertainty.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Estimate covariance

The sample covariance between assets i and j is:

Σ̂ij = [1/(T−1)] Σt=1T(ri,t − r̄i)(rj,t − r̄j)

For regular periodic returns, annualized covariance is commonly approximated by m × Σ̂periodic. The return frequency, annualization convention, and risk-free-rate frequency must be compatible. A short history, many assets relative to observations, highly correlated holdings, or changing market relationships can make the estimate noisy or ill-conditioned. Covariance matrices used in quadratic optimization also need to be positive semidefinite. Shrinkage methods blend the noisy sample estimate with a more structured estimate to improve stability; they do not guarantee better realized performance. PyPortfolioOpt documents shrinkage-based risk models.

Python: a long-only baseline with PyPortfolioOpt

PyPortfolioOpt provides standard return and risk estimators and efficient-frontier methods. Its documentation lists the library’s supported workflows; check it against the version installed because APIs can change. Installation and functionality.

import pandas as pd
from pypfopt import expected_returns, risk_models
from pypfopt.efficient_frontier import EfficientFrontier

# CSV contains adjusted prices: one asset per column, dates as the index.
prices = pd.read_csv(
    "adjusted_prices.csv",
    index_col=0,
    parse_dates=True
)

# Library estimators use price history to produce annualized inputs.
mu = expected_returns.mean_historical_return(prices)
S = risk_models.sample_cov(prices)

# Long-only weights, with no holding above 30%.
ef = EfficientFrontier(mu, S, weight_bounds=(0, 0.30))
weights = ef.max_sharpe(risk_free_rate=0.02)

cleaned_weights = ef.clean_weights()
performance = ef.portfolio_performance(
    verbose=True,
    risk_free_rate=0.02
)
print(cleaned_weights)

The 0.30 cap is an example constraint, not a recommended allocation. Likewise, the 0.02 risk-free-rate input is only an example: choose a rate that matches the portfolio’s currency and evaluation period. For a different objective, use ef.min_volatility() or ef.efficient_return(target_return=...). The user guide covers these methods and weight bounds.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Regularize the weights

An L2 penalty discourages extreme weights and can make a solution less sensitive to noisy inputs. For example:

from pypfopt import objective_functions

ef = EfficientFrontier(mu, S, weight_bounds=(0, 0.30))
ef.add_objective(objective_functions.L2_reg, gamma=0.1)
weights = ef.min_volatility()

The penalty strength is a tuning choice. Select it using training and validation data, not by repeatedly checking the final test period.

Model turnover costs

To discourage trading away from a previous allocation, a documented transaction-cost objective can be added:

from pypfopt import objective_functions

previous_weights = {ticker: 0.10 for ticker in prices.columns}
ef = EfficientFrontier(mu, S, weight_bounds=(0, 0.30))
ef.add_objective(
    objective_functions.transaction_cost,
    w_prev=previous_weights,
    k=0.001
)
weights = ef.min_volatility()

The previous weights and cost parameter must reflect the actual portfolio and a defensible cost model; the code values are illustrative. PyPortfolioOpt’s mean-variance documentation covers constraints, regularization, and transaction-cost objectives.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Model constraints and execution reality

Constraints are part of the investment decision, not cosmetic adjustments after the solver runs. Common choices include:

  • Long-only and position bounds: require 0 ≤ wi ≤ wi,max.
  • Sector or asset-class limits: keep the sum of weights in group g between specified lower and upper bounds.
  • Turnover limits: constrain Σi|wi − wi,prev| to a chosen threshold.
  • Leverage and gross exposure: control borrowing and total absolute exposure for long-short portfolios.
  • Tracking error: limit benchmark-relative risk when the objective is benchmark-aware.
  • Liquidity: relate position size to volume, bid-ask spread, and expected market impact.
  • Cardinality or minimum holding size: control the number of positions or exclude tiny allocations; discrete holding rules can make the problem mixed-integer or otherwise nonconvex.

A practical objective can add a turnover penalty to estimated variance, such as wTΣw + λΣici|wi−wi,prev|, where the cost coefficients and penalty strength need calibration. Trading costs can include commissions, spreads, exchange and regulatory fees, slippage, market impact, borrow costs, and taxes. Zero commission does not mean zero implementation cost. Fractional-share availability, order failures, partial fills, and the delay between signal and execution also affect whether theoretical weights can be held.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why the clean mathematical answer can fail

Estimation error and unstable weights

Expected returns are especially difficult to estimate. Small changes in estimated means can produce large changes in maximum-return or maximum-Sharpe weights, while unstable correlations can move even a minimum-variance portfolio. A precise solver result can therefore reflect noise rather than a durable opportunity. Weight caps, long-only bounds, covariance shrinkage, L2 regularization, turnover penalties, bootstrap or resampling analysis, and minimum-variance objectives can reduce some instability, but no technique eliminates model risk.

Concentration, leverage, and ill-conditioning

Unconstrained solutions may contain large positive and negative positions or concentrate heavily in assets with favorable estimated statistics. Explicit bounds and leverage controls make the model more implementable. Too many assets for the available history, or assets that are nearly redundant, can make the covariance matrix ill-conditioned. Consider a narrower universe, a factor covariance model, shrinkage, or careful positive-semidefinite repair rather than trusting unstable numerical output.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Regimes, tails, and missing risks

Historical volatility and correlations can change during crises, inflation shocks, rate changes, or structural breaks. Variance is not a complete description of risk when returns are asymmetric or have fat tails. Mean-variance theory does not require every asset return to be perfectly normal, but its risk summary may be inadequate for an investor focused on severe losses, liabilities, cash flows, taxes, or illiquid assets. Stress scenarios and downside-focused measures can complement it.

Look-ahead bias and backtest overfitting

Leakage occurs when a test implicitly uses information unavailable at the decision date. Examples include estimating inputs with future data, using future index membership, omitting delisted assets, applying corporate-action information incorrectly, or choosing the lookback window after inspecting the test period. Repeatedly tuning the universe, rebalance frequency, bounds, objective, and cost assumptions on one history can make a backtest look better by chance. Keep training, validation, and final test periods distinct; reserve the final test for a limited, pre-specified evaluation.

Walk-forward validation: test the process, not just the weights

  1. Define the protocol: state the universe, data source and adjustment method, estimation window, signal date, execution date, rebalance frequency, missing-data treatment, and cost assumptions.
  2. Estimate only from past data: at each rebalance date, compute returns and covariance using the training window available then, and optimize under the intended constraints.
  3. Simulate the next holding interval: apply the weights only after the assumed execution point, account for costs and cash, and record what could actually have been traded.
  4. Roll forward: advance to the next rebalance date and repeat without using future observations in the estimates.
  5. Compare with simple alternatives: include equal weight, market-cap weight where relevant, minimum variance, a basic risk-parity portfolio, or a policy allocation.
  6. Report risk and implementation: include annualized return and volatility, Sharpe ratio, maximum drawdown, downside deviation, worst month or rolling period, turnover, cost drag, concentration, and weight stability. Break results out by market regime when appropriate.

Do not select a strategy solely because it has the highest in-sample Sharpe ratio. A frontier chart does not show turnover, taxes, liquidity, estimation uncertainty, or whether the portfolio survived a realistic out-of-sample protocol.

Alternatives and when to use them

Method Uses expected returns? Main strength Main limitation
Equal weight No Simple, transparent benchmark Ignores differences in risk and can create unintended concentration
Minimum variance Usually not directly Reduces reliance on return forecasts Still depends on covariance quality
Maximum Sharpe Yes Targets estimated excess return per unit of volatility Often sensitive to expected-return estimates
Risk parity No or limited Focuses on risk contributions May require leverage or produce low-return allocations in some settings
Black-Litterman Yes, structured Combines equilibrium-implied returns with investor views and confidence Adds assumptions about equilibrium and views
Hierarchical Risk Parity No traditional inverse-covariance allocation Uses a clustering and hierarchical diversification structure Less direct risk-return interpretation than a frontier target
Robust optimization Yes, with uncertainty sets Models parameter uncertainty explicitly Can be conservative and depends on uncertainty-set design
Downside or CVaR optimization Usually uses return assumptions or scenarios Focuses on downside variability or tail losses More complex and dependent on scenario design

Factor-based approaches can also constrain or target exposures such as value, momentum, quality, duration, or size, but the result depends on factor definitions and data quality. PyPortfolioOpt documents Black-Litterman, HRP, and other alternatives, as well as semivariance methods. See its alternative efficient-frontier methods.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use mean-variance optimization as an interpretable baseline when the universe is investable, the objective and constraints are explicit, and you can validate the entire process out of sample. If return forecasts are weak, start by comparing minimum variance and equal weight rather than treating maximum Sharpe as the default. If the central concern is tail loss, liabilities, tax-aware turnover, or illiquidity, use a method and data that model those concerns directly.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Written by TheFinanceBase Team

The Team behind TheFinanceBase.

Add your note

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.