Quantitative Methods · Reading 10

Least Squares Regression

CFA Level I · Quantitative Methods · Reading 10: Applications of Simple Linear Regression in Finance · about 1 h 1 min

What you'll learn

Module 10.1

Linear Regression Basics

This reading fits a simple linear regression by least squares, interprets its slope and intercept, and checks the model's assumptions with residual plots. It then measures fit with ANOVA, the standard error of estimate, and the F-test, tests the coefficients, builds prediction intervals, chooses among log-transformed functional forms, and estimates a stock's CAPM beta by regression.

LOS 10.a — Simple linear regression and the least squares criterion

Purpose and vocabulary

Simple linear regression explains the variation of one dependent variable with the variation of a single independent variable. Variation is the total squared deviation from the mean, ; it is related to, but not the same as, the variance (the variance divides the variation by ).

Dependent variable (Y)Independent variable (X)
RoleThe variable being explainedThe variable doing the explaining
Other namesExplained, endogenous, predicted variableExplanatory, exogenous, predicting variable
Example: using GDP growth to forecast stock returnsStock returnsGDP growth

A model with one dependent variable and one independent variable is a simple regression; a model with one dependent variable and more than one independent variable is a multiple regression (Level II).

The model and the regression line

Key concept

The hat marks an estimate. The residual for observation , , is the gap between the actual value and the value the fitted line predicts for that X.

  • Exam convention: the residual, the error term and the disturbance term are treated as the same thing, written .
  • Current practice: the error term is the unobservable deviation in the population model, and the residual is its sample estimate.

Least squares criterion

Of all the lines that could be drawn through the scatter plot, ordinary least squares (OLS) regression picks the one that minimizes the sum of squared errors (SSE), found by squaring each vertical gap between an observed value of the dependent variable and its fitted value, then adding the squares:

The resulting estimates are:

Key concept

The intercept formula shows that the regression line always passes through the point of means . The covariance and the variance share the divisor , which cancels, so from raw data the slope is .

Eight observations (X, Y) in percent: (-4, -2.9), (-2, -0.9), (-1, 0.1), (0, 0.2), (1, 1.8), (2, 1.7), (3, 2.6), (5, 4.7). The fitted regression line is Y-hat = 0.51 + 0.80X and passes through the point of means (0.50, 0.91). For the observation at X = 1, the vertical residual is 1.8 - 1.31 = 0.49.
Scatter plot with the least-squares regression line

Worked example (data in the figure)

Eight monthly observations give , , and (both in %²). The underlying sums are and , and gives the same slope.

For the observation at with : , so the residual is .

Interpreting the coefficients

  • Slope coefficient : how much the dependent variable is expected to move when the independent variable rises by one unit. In the example, a market excess return 1 percentage point higher goes with a stock excess return 0.80 point higher. When Y is a stock's excess return and X the market's, the slope is the stock's beta.
  • Intercept : the predicted value of Y when X equals zero (where the line crosses the vertical axis).
  • A large slope does not by itself mean a strong relationship; significance must be tested (t-test or confidence interval).

Types of data

Key concept

TypeDefinitionExample
Time seriesObservations of one variable at regular intervals over timeA company's quarterly EPS for five years
Cross-sectional dataObservations of many subjects at one point in timeToday's closing prices of 30 stocks in an index
Panel dataCross-sectional observations repeated over timeQuarterly EPS of 10 companies over five years

Common exam traps

  • The slope divides the covariance by the variance of X, the independent variable. Dividing by the variance of Y is a common error.
  • OLS minimizes the squared differences between actual and predicted values of Y. It does not compare coefficients or X values.
  • The slope is the change in Y per unit change in X. Reading it as the change in X per unit of Y reverses the roles.

Exam shortcuts

  • The regression line passes through , so once the slope is known the intercept is with no further sums.

Bottom line

  • Simple linear regression explains the variation of a dependent variable Y, , with the variation of a single independent variable X.
  • Ordinary least squares chooses the line that minimizes the sum of squared errors, .
  • The OLS estimates are and , so the fitted line passes through the point of means.
  • The slope is the expected change in Y for a one-unit rise in X, and the intercept is the predicted value of Y when X equals zero.
  • The residual is the gap between the actual and the fitted value; exam convention treats the residual, error term and disturbance term as the same thing, while current practice treats the residual as the sample estimate of the unobservable error term.
  • Time series data follow one variable over time, cross-sectional data cover many subjects at one point in time, and panel data repeat cross-sectional observations over time.

Quick check

Question 1Core

A credit analyst gathers data on 80 bond issuers at a single date to see whether an issuer's interest coverage ratio helps explain the yield spread on its bonds. In her simple linear regression, the dependent variable is:

Show answer and explanation

Correct answer: B

The dependent variable is the one whose variation the model explains, here the yield spread. The interest coverage ratio does the explaining, so it is the independent (explanatory) variable. Observations of many issuers at one date are cross-sectional data, which a simple linear regression can use as readily as time-series data.

Why the other options are wrong

  • A. An explanatory variable is the independent variable. Making the coverage ratio the dependent variable reverses the roles set by the analyst's question.
  • C. Data on many subjects at one point in time are cross-sectional data, and a regression can be run on them. Time-series data are not required.

Key takeaway Dependent variable (Y): explained, endogenous, predicted. Independent variable (X): explanatory, exogenous, predicting. The research question decides which is which.

Module 10.2

Analysis of Variance (ANOVA) and Goodness of Fit

LOS 10.b — Assumptions, residual analysis, ANOVA and goodness of fit

Assumptions of simple linear regression

Key concept

  1. Linearity: the relationship between X and Y is linear.
  2. Homoskedasticity: the residuals have the same variance for every observation.
  3. Independence: residuals are uncorrelated with one another (the X–Y pairs are independent).
  4. Normality: the residual term is normally distributed.

Violations show up in plots of the residuals against X (or against time). In a well-specified model the residuals scatter randomly around zero with no pattern.

ViolationWhat the residual plot showsName
Nonlinear relationshipA systematic pattern, e.g., residuals positive, then negative, then positive as X rises—
Residual variance not constantSpread of residuals widens (or narrows) with X, or changes over timeHeteroskedasticity
Residuals correlatedRecurring pattern, e.g., large errors every December (seasonality); the estimated variances of the coefficients are then wrongAutocorrelation
Residuals not normalSkewed or fat-tailed residuals; hypothesis tests may be unreliable in small samples (with a large sample, the central limit theorem means the estimates may still be valid)—
Two panels share a horizontal axis for the independent variable X from 0 to 10. Panel A: 21 observations at X = 0, 0.5, 1, ..., 10 lie on a curve that rises slowly at first and then steeply (Y = 2 at X = 0, about 7 at X = 5 and 22 at X = 10). The straight least-squares line through them is Y-hat = -1.17 + 2.00X. The points lie above the line for X below about 2, below it between about X = 2 and X = 8, and above it again for X above about 8. Panel B: the residuals from that line plotted against X form a U-shaped pattern: about +3.2 at X = 0, close to zero at X = 2, about -1.8 at X = 5, close to zero at X = 8 and about +3.2 at X = 10, instead of a random scatter around zero.
A straight line fitted to a curved relationship, and the residual pattern it leaves
Two panels of residuals plotted against the independent variable X (1 to 20), about 60 points each. Panel A: residuals scatter randomly within roughly -3 to +2.5 with the same spread at every X (homoskedastic). Panel B: residuals are tightly clustered near zero at low X and spread out steadily as X increases, reaching about -6 to +10 at the highest X (heteroskedastic).
Residual plots: homoskedastic vs. heteroskedastic

Outliers are observations far from the line. They can pull the estimated coefficients so that the line fits the other points poorly.

Analysis of variance (ANOVA)

Analysis of variance (ANOVA) splits the total variation of the dependent variable into an explained and an unexplained part:

  • Total sum of squares (SST) : total variation.
  • Sum of squares regression (SSR) : variation explained by the regression.
  • Sum of squared errors (SSE) : unexplained variation (also called the residual sum of squares).
Scatter of eight illustrative points with a fitted line (slope 0.71, intercept 1.26) and the mean of Y at 4.45. For the point at X = 8, Y = 7.9, the total deviation from the mean splits into an explained part (fitted value minus mean, summed over all points gives SSR) and an unexplained part (Y minus fitted value, summed gives SSE).
Total variation split into explained and unexplained parts (illustrative data)

The sample variance of Y is its total variation divided by its degrees of freedom, .

  • Mean square regression (MSR) ; with one independent variable (), MSR = SSR.
  • Mean squared error (MSE) in a simple regression.

Key concept

Source of variationDegrees of freedomSum of squaresMean sum of squares
Regression (explained)SSRMSR = SSR/1
Error (unexplained)SSEMSE = SSE/()
TotalSST

Goodness-of-fit measures

The standard error of estimate (SEE), also called the root mean squared error, is the standard deviation of the residuals; the lower it is, the better the fit.

Key concept

The coefficient of determination measures what fraction of the dependent variable's total variation the independent variable accounts for. In a simple regression , so the correlation coefficient is , with the sign of the slope. When only and SSR are known, SST SSR and SSE SST SSR, which then gives the MSE and the SEE.

F-test

The F-test compares the variation the regression explains, per degree of freedom, with the variation it leaves unexplained, per degree of freedom. A large ratio means the independent variable explains a lot relative to the noise.

The F-test asks whether the independent variables, as a group, explain the variation of Y; with several independent variables its null is that all slope coefficients are zero. With one independent variable, it tests vs. , the same hypothesis as the t-test of the slope. Reject if (a one-tailed test using the upper tail of the F-distribution). Rejecting means X and Y have a significant linear relationship.

Worked example

; SSR = 64 and SSE = 112, so SST = 176. MSE ; SEE ; ; (if the slope is positive); , well above the 5% critical value of about 4.2 for 1 and 28 degrees of freedom, so the slope is significant.

t-test of a regression coefficient

To test significance, set (two-tailed) and reject if . The test of the slope is equivalent to testing the correlation, . With one independent variable, the F-statistic equals the square of this t-statistic, so the F-test and the t-test of the slope always reach the same decision (in the worked example above, corresponds to ). The intercept can be tested the same way ().

Indicator variables

An indicator variable (dummy variable) takes the value 1 when a condition holds and 0 otherwise (e.g., 1 in recession quarters). Regressing Y on the indicator tests whether the condition matters: the slope equals the difference in the mean of Y between periods with and without the condition, and testing is equivalent to a difference-in-means test. The intercept is the mean of Y in periods without the condition (indicator = 0), and intercept plus slope is the mean in periods with it. Example. Regressing a company's quarterly earnings on a recession indicator tests whether earnings are cyclical; for cyclical earnings the slope is expected to be negative (lower earnings in recession quarters). If quarterly EPS averaged 1.40 in expansion quarters and 0.95 in recession quarters, the fitted line is , where in a recession quarter.

Common exam traps

  • . SSE/SST is the unexplained share (), and SSR/SSE is not a standard measure. In the worked example, is the unexplained share; is 0.364.
  • F is a ratio of mean squares, MSR/MSE. Dividing SSR by SSE ignores the degrees of freedom. In the worked example, instead of .
  • SEE divides SSE by , the degrees of freedom, before taking the square root. Dividing by understates it. In the worked example, instead of 2.
  • A negative slope gives a negative even though is positive. The standard error of the slope is a different statistic and is not a correlation.

Exam shortcuts

  • With one independent variable, for the test of against , so at the same significance level the F-test and the two-tailed slope t-test reach the same decision and only one needs computing.
  • When only and SSR are given, SST SSR and SSE SST SSR, which lead straight to the MSE and the SEE.
  • In a simple regression , taking the sign of the slope.

Bottom line

  • Simple linear regression assumes a linear relationship, homoskedastic residuals, residuals uncorrelated with one another, and normally distributed residuals.
  • Residual variance that widens or narrows with X or over time is heteroskedasticity, and residuals correlated with one another, such as a recurring seasonal pattern, show autocorrelation.
  • ANOVA splits total variation into explained and unexplained parts, .
  • In a simple regression , , and , which equals .
  • with 1 and degrees of freedom is a one-tailed test, and with one independent variable it tests , the same hypothesis as the slope's t-test.
  • Regressing Y on an indicator variable gives a slope equal to the difference in the mean of Y between periods with and without the condition, and an intercept equal to the mean when the indicator is 0.

Quick check

Question 2Core

An analyst regresses the monthly sales of 80 stores on each store's floor area. The figure below plots the regression residuals against floor area. The pattern most likely indicates that the regression's residual term:

Scatter of 80 residuals (vertical axis, monthly sales in $ thousands, from about -3.5 to +5) against store floor area (horizontal axis, 200 to 2,000 square meters). For small stores the residuals cluster tightly around zero (within about +/-1); as floor area increases, the residuals spread out more and more, reaching roughly -3.5 to +5 for the largest stores. The residuals are centered on zero with no curved pattern.
Residuals plotted against store floor area
Show answer and explanation

Correct answer: A

The residuals fan out as floor area increases: their spread (variance) grows with the independent variable. A residual variance that is not constant across observations is heteroskedasticity, a violation of the homoskedasticity assumption.

Why the other options are wrong

  • B. The plot shows residuals centered on zero with a changing spread; nothing in it points specifically to a nonnormal shape, and normality is judged from the shape of the residuals' distribution. How their spread changes with X is a separate question.
  • C. Non-independence (autocorrelation) shows up as residuals that are correlated with one another, such as a recurring seasonal pattern over time. A widening spread with X is a variance problem.

Key takeaway Residual spread widening with X = heteroskedasticity (non-constant residual variance).

Module 10.3

Predicted Values and Functional Forms of Regression

LOS 10.c — Predicted values, prediction intervals and functional forms

Predicted values

A predicted value of the dependent variable plugs a forecast of the independent variable, , into the estimated equation:

Example: and : . Include the intercept and keep the sign of each term.

Confidence (prediction) interval for a predicted value

Key concept

where is the standard error of the forecast. At Level I it is usually given; if not:

Key concept

The forecast is more precise (smaller ) when the SEE is smaller, the sample is larger, is closer to , and the independent variable has a greater variance.

Example (continued): , , so and the two-tailed 5% critical value is 2.060. The 95% interval is , i.e., 1.71 to 7.89.

Computing when it is not given (a separate regression): SEE = 2.0, , , , :

The standard error of the forecast is always larger than the SEE: the forecast carries both the residual uncertainty and the uncertainty in the estimated coefficients.

IntervalCentered onWidth driven by
Confidence interval for a coefficient (or )Standard error of that coefficient
Prediction interval for Y at Standard error of the forecast

Functional forms

If the relationship between X and Y is not linear, taking natural logarithms of the dependent variable, the independent variable or both can make it linear. For example, EPS growing at a roughly constant percentage rate over time traces an upward curve.

Key concept

ModelEquationSlope interpretationTypical use
Log-lin modelRelative (%) change in Y for an absolute change in XConstant growth rate over time
Lin-log modelAbsolute change in Y for a relative (%) change in XY responds to proportional changes in X
Log-log modelRelative change in Y for a relative change in X (an elasticity)Both variables in proportional terms

Example. In a log-lin regression of a company's annual EPS on a time index (1, 2, 3, …), a slope of 0.07 means EPS grows by roughly 7% a year (exactly ). It does not mean EPS rises by 0.07 currency units a year.

Choose the form by the nature of the variables and by comparing goodness-of-fit measures and residual plots.

  • Exam convention: a higher , a higher F-statistic and a lower SEE indicate the better-fitting functional form.
  • Current practice: and SEE are directly comparable only when the dependent variable is the same; a model of and a model of measure fit on different scales.
Left: illustrative EPS for years 1 to 12 rising from 1.50 to 3.24 along an upward curve, with a straight line fitted to it missing the curvature. Right: ln(EPS) against time lies on a straight line, ln(EPS) = 0.335 + 0.070t.
Log-lin model: EPS growing about 7% a year (illustrative)

Common exam traps

  • Prediction intervals use in a simple regression (two estimated parameters), not . In the example, 26 degrees of freedom give a critical value of 2.056 and an interval of 1.72 to 7.88 instead of 1.71 to 7.89.
  • A prediction interval uses the standard error of the forecast . The standard error of the slope belongs to the interval for the slope coefficient.
  • "Log" is attached to the variable that is logged: log-lin means Y is logged and X is linear.

Exam shortcuts

  • The standard error of the forecast is always larger than the SEE, so a prediction interval narrower than SEE can be ruled out.

Bottom line

  • A predicted value is , using the intercept and the sign of each term.
  • The prediction interval for Y is with degrees of freedom, where is the standard error of the forecast.
  • The forecast is more precise when the SEE is smaller, the sample is larger, is closer to and the independent variable has a greater variance.
  • A log-lin model ( on X) has a slope giving the relative change in Y for an absolute change in X, a lin-log model the absolute change in Y for a relative change in X, and a log-log model an elasticity.
  • Exam convention: a higher , a higher F-statistic and a lower SEE identify the better-fitting functional form; current practice: and SEE are directly comparable only when the dependent variable is the same.

Quick check

Question 3Core

An analyst regresses a fertilizer producer's monthly stock returns on monthly changes in a crop price index, using 32 observations. The estimated coefficients are in the table below. She expects next month's change in the index to be −0.015 and computes a standard error of the forecast of 0.028. Critical t-values with 30 degrees of freedom are 1.310 with 10% in one tail and 1.697 with 5% in one tail. The 90% prediction interval for next month's return is closest to:

Regression of monthly stock returns on crop price changes
CoefficientStandard error
Intercept0.00400.0035
Crop price change0.65000.2100
Show answer and explanation

Correct answer: A

A prediction interval is the predicted value plus or minus the critical t-value times the standard error of the forecast. A two-sided 90% interval leaves 5% in each tail, so with n − 2 = 30 degrees of freedom the critical value is 1.697.

  • Predicted return:
  • Half-width:
  • Interval: , that is, to

Calculator: 0.65 [×] 0.015 [=] gives 0.00975; 1.697 [×] 0.028 [=] gives 0.047516.

Why the other options are wrong

  • B. This interval uses 1.310, the critical value with 10% in one tail. A two-sided 90% interval splits the 10% between the two tails, so each tail holds 5%.
  • C. This interval is centered on , which drops the minus sign of the forecast change in the index.

Key takeaway Prediction interval , with n − 2 degrees of freedom and α/2 in each tail. Keep the sign of the forecast X when computing .

Module 10.4

Estimating CAPM Values With Simple Linear Regression

LOS 10.d — Estimating the CAPM with simple linear regression

The CAPM as a one-factor regression

The capital asset pricing model states

Because the only explanatory variable is the market risk premium, the CAPM is a one-factor risk model and can be estimated with simple linear regression, using past observations as proxies for expected values. Subtracting from both sides gives the estimable form:

Key concept

Key concept

Regression elementCAPM meaning
Dependent variableThe asset's excess return
Independent variableThe market risk premium
Slope The asset's estimated beta
Intercept Return not explained by market exposure (zero if the CAPM holds exactly)

Steps

  1. Collect data: the asset's returns; a security market index as the proxy for the market return; yields on short-term government debt as the proxy for the risk-free rate. At least 30 observations is the usual minimum (e.g., 1–3 years of weekly data or 3–5 years of monthly data).
  2. Compute the asset's excess returns and the market risk premium for each period.
  3. Run the least-squares regression of excess returns on the market risk premium.
  4. Use and an expected market risk premium to estimate the expected return: .

Testing and interpreting beta

Example: the slope is 1.25 with a standard error of 0.10 (large sample, so use z-values).

  • Beta above 1: the stock's excess returns amplify the market's; below 1: they are dampened.
  • Test : , so reject; beta is significantly different from zero.
  • 95% confidence interval: to .
  • Test (does the stock move one-for-one with the market?): , so reject; beta is significantly different from 1, and the estimate lies above it. This matches the interval above, which does not contain 1.
  • A one-sided question, such as whether beta exceeds 1, uses versus , the same statistic and a one-tailed critical value.
  • With and an expected market risk premium of 5%: .

Changing inputs: . With , a 0.4% rise in combined with a 0.4% fall in the market risk premium changes expected return by ; for a beta above 1 the same changes would lower it.

Reading a typical output

CoefficientStandard errort-statistic
Intercept0.050.080.63
Market risk premium (slope)1.250.1012.50

The slope row carries beta and its test; the intercept row is not significant here (t = 0.63), consistent with the CAPM's prediction of a zero intercept. The of such a regression tells how much of the asset's excess-return variation is systematic (market-driven); is the share that is unsystematic.

Choosing the sample

Longer samples give more observations and more precise estimates but risk including periods when the company's business mix, leverage and therefore beta were different. More frequent data (daily or weekly) add observations quickly but can be noisy for thinly traded assets. The choice depends on the purpose: a short-term speculator favors recent high-frequency data; a buy-and-hold investor favors monthly or quarterly data spanning at least one full business cycle.

Worked example: estimating and testing a CAPM beta

Sixty monthly observations of a stock's excess return and the market risk premium give a covariance of 0.00168 and a variance of the market risk premium of 0.00140. The regression reports a standard error of 0.16 for the slope, and with 60 observations the sample is large, so the two-tailed 5% critical value is 1.96. Estimate beta, test at the 5% level whether beta differs from 1, build the matching confidence interval for beta, and estimate the stock's expected return with a risk-free rate of 2.8% and an expected market risk premium of 5.5%.

Step 1. Beta is the slope coefficient: .

Step 2. Test against : . Because 1.25 is below 1.96, is not rejected, so beta is not significantly different from 1.

Step 3. The 95% confidence interval is , or 0.886 to 1.514. It contains 1, which agrees with Step 2.

Step 4. Expected return: .

Result. The estimated beta is 1.20, not significantly different from 1 at the 5% level, and the CAPM expected return is 9.4%.

Common exam traps

  • The dependent variable is the asset's excess return and the independent variable is the market risk premium. A regression of raw returns on raw returns is not the CAPM specification.
  • Beta is the slope coefficient. The intercept says nothing about beta, so a negative intercept does not make beta negative.
  • Confidence interval for beta: slope critical value standard error of the slope. The intercept's standard error plays no part. With the output above, using the intercept's 0.08 gives 1.093 to 1.407 instead of 1.054 to 1.446.

Exam shortcuts

  • A two-tailed 95% confidence interval for beta that excludes 1 means is rejected at the 5% level, and one that includes 1 means it is not, so either result gives the other.

Bottom line

  • The CAPM is estimated by regressing the asset's excess return on the market risk premium ; the slope is the asset's beta and the intercept is zero if the CAPM holds exactly.
  • A security market index proxies the market return, short-term government debt yields proxy the risk-free rate, and at least 30 observations is the usual minimum.
  • Beta is tested with , and its confidence interval is the slope critical value the slope's standard error.
  • The expected return is , and a change in inputs gives .
  • The of the CAPM regression is the share of the asset's excess-return variation that is systematic, and is the unsystematic share.
  • Longer samples give more precise estimates but may include periods with a different beta, and more frequent data can be noisy for thinly traded assets.

Quick check

Question 4Core

To estimate a stock's beta, an analyst regresses the stock's excess returns on the market risk premium and obtains an intercept of −0.15 and a slope of 0.72. Based on this model, the stock's beta is:

Show answer and explanation

Correct answer: B

In the regression , the slope coefficient is the estimate of the asset's beta. The slope is 0.72, so beta lies between 0 and 1: the stock's excess returns tend to move less than one-for-one with the market risk premium.

Why the other options are wrong

  • A. The negative number is the intercept, which is not beta. Beta is the slope, 0.72.
  • C. A beta above 1 would require a slope above 1; the estimated slope is 0.72.

Key takeaway Beta = slope of the excess-return regression; the intercept is irrelevant to beta.

This reading has 29 questions in the full bank. Practice all of them.

Key Takeaways