Multiple Regression for Cost Estimation, AMIS 3300 – Study Notes
offline

Source: Cost Estimation and Regression Discussion Problems (Turner)

Difficulty: Intermediate | Prerequisites: High-low method, comfort with Y = a + bX, basic idea of what a "good fit" means.

Big Picture

Regression analysis is the statistical upgrade to the high-low method. Instead of drawing a line through just two data points, regression fits the best possible line (or plane, in multiple regression) through all available observations. In a cost accounting context, it is used to estimate cost functions with greater precision and to measure how well the model explains cost behaviour. AMIS 3300 expects you to read and interpret regression output, not to run the software yourself. This topic typically appears after the high-low method and before cost-volume-profit analysis.


TL;DR

Multiple regression fits a cost equation using all data points and more than one cost driver. You read the coefficients from a computer printout, write the equation, and use R² and p-values to judge how reliable the model is. For AMIS 3300, the skill is interpretation, not computation.


Key Terms

Multiple regression

A statistical method that models a dependent variable (cost) as a linear function of two or more independent variables (cost drivers). The equation takes the form Y = a + b₁X₁ + b₂X₂ + ... + bₙXₙ.

In simple terms, it is the high-low method's smarter sibling: it uses all data points and can handle more than one cost driver at once.

R-squared (R², coefficient of determination)

The proportion of the variation in the dependent variable that is explained by the independent variables. Ranges from 0 to 1. An R² of 0.83 means 83% of the variation in cost is explained by the model.

In simple terms, it tells you how well the regression line fits the data. Higher is better, but it does not mean the model is "correct," only that it captures most of the pattern.

Standard error of the estimate (Sₑ)

A measure of the average distance between actual cost observations and the predicted values from the regression line. Smaller is better. Used to build prediction intervals around a cost estimate.

In simple terms, it is the typical size of the "miss" when you use the equation to predict a cost.

Intercept (a)

The estimated fixed cost, the value of Y when all independent variables equal zero. In the printout for Problem 3 this is 75.

Parameter estimate (coefficient, b)

The estimated change in the dependent variable for a one-unit change in that independent variable, holding the others constant. Setup hours has a coefficient of 13, meaning each additional setup hour adds an estimated $13 to overhead.

t-statistic (t for H₀: Parameter = 0)

Tests whether a particular coefficient is statistically different from zero. A larger absolute value means the variable is more likely a genuine cost driver rather than noise.

p-value (Pr > t)

The probability of observing a t-statistic at least as extreme as the one calculated, assuming the true coefficient is zero. A small p-value (typically below 0.05 or 0.10) means you reject the null hypothesis and conclude the variable is statistically significant.

In simple terms, a low p-value means the cost driver matters; a high p-value means the data cannot confirm it matters.

Degrees of freedom (df)

For multiple regression: df = n − 1 − p, where n is the number of observations and p is the number of independent variables. In Problem 3: df = 70 − 1 − 2 = 67. Degrees of freedom are used to look up t-values for confidence intervals.

Confidence / prediction interval

A range around a point estimate that accounts for uncertainty. Built using the t-value for the desired confidence level and the standard error: point estimate ± t(df, confidence level) × Sₑ.


Core Content

Reading the Regression Printout (Problem 3)

The printout from Problem 3 gives you everything needed to write and evaluate the model.

  • Dependent variable: Overhead costs (this is what the regression estimates).

  • Independent variables (cost drivers): Setup hours and number of parts.

  • Number of observations: 70 data points were used to build the model.

Writing the Multiple Regression Equation

Read the coefficients straight from the "Estimate" column:

  • Intercept = 75

  • Setup hours coefficient = 13

  • Number of parts coefficient = 50

The equation is:

Overhead = 75 + 13 × (Setup Hours) + 50 × (Number of Parts)

The intercept (75) represents estimated fixed overhead when both cost drivers are zero. Each additional setup hour adds $13 to estimated overhead. Each additional part adds $50.

Using the Equation: A Worked Prediction

Suppose a job requires 500 setup hours and 300 parts:

Overhead = 75 + 13(500) + 50(300) = 75 + 6,500 + 15,000 = $21,575

This is a point estimate, a single best guess. To express uncertainty, you build a prediction interval (see Formulas section).

Interpreting R-Squared

R² = 0.83 means that 83% of the variation in overhead costs across the 70 observations is explained by setup hours and number of parts together. The remaining 17% is unexplained variation, caused by factors not in the model or random noise.

R² does not mean the model is "83% accurate." It means the two cost drivers together capture most of the pattern in the data.

Interpreting the t-Statistics and p-Values

  • Setup hours: t = 5.10, p = 0.0001. Highly significant. There is very strong evidence that setup hours drive overhead costs.

  • Number of parts: t = 1.65, p = 0.0500. Borderline. At the 5% significance level, this variable barely clears the threshold. At 10% it is clearly significant. Whether you call it significant depends on the cutoff the problem specifies.

  • Intercept: t = 2.25, p = 0.0250. Significant at the 5% level. The fixed cost component is statistically distinguishable from zero.

Building a Prediction Interval

A prediction interval puts a range around the point estimate:

Predicted value ± t(df, confidence level) × Sₑ

For the example above (OH = $21,575) with a 90% confidence interval:

  • df = 70 − 1 − 2 = 67

  • Look up t(67, 90%) from a t-table (approximately 1.668 for a two-tailed 90% interval).

  • Sₑ = 50

  • Interval = 21,575 ± 1.668 × 50 = 21,575 ± 83.40

  • Lower bound: $21,491.60

  • Upper bound: $21,658.40

This means you are 90% confident that actual overhead for this job will fall within that range.

Degrees of Freedom

df = n − 1 − p, where n = number of observations and p = number of independent variables (parameters, not counting the intercept).

For Problem 3: df = 70 − 1 − 2 = 67. You need this number to look up the correct t-value for confidence intervals.


Formulas

Multiple regression equation:

Y = a + b₁X₁ + b₂X₂ + ... + bₙXₙ

R-squared:

R² = Explained variation / Total variation (given on the printout; you do not calculate it by hand in this course)

Degrees of freedom:

df = n − 1 − p

Where n = number of observations, p = number of independent variables.

Prediction interval:

Point estimate ± t(df, confidence level) × Sₑ

Where Sₑ is the standard error of the estimate and the t-value comes from a t-distribution table at the appropriate df and confidence level.


Real-World Applications

Companies use regression models to forecast overhead for budgeting and to set standard costs. A manufacturing firm might regress overhead on machine hours, number of setups, and batch size to build a more precise cost allocation system. Hospitals use similar models to predict patient costs based on diagnosis category, length of stay, and procedure count. The underlying logic is the same: identify the cost drivers and quantify how much each one contributes.


Common Misconceptions

  • Students often think R² = 0.83 means the model predicts costs with 83% accuracy. It does not. It means 83% of the variation in the dependent variable is captured by the independent variables. Prediction accuracy depends on the standard error and the specific input values.

  • Some students confuse the standard error of the estimate (Sₑ) with the standard error of a parameter. Sₑ measures overall prediction spread; parameter standard errors measure uncertainty about individual coefficients.

  • Students sometimes assume a high p-value means the variable has no effect. A high p-value means the data do not provide strong enough evidence to conclude there is an effect. That is a subtly different statement.

  • When computing degrees of freedom, students sometimes count the intercept as a parameter. In the formula df = n − 1 − p, p is only the number of independent variables, not including the intercept.


Why It Matters / Exam Flags

⚠️ You will be given a regression printout and asked to write the equation. Read the coefficients from the Estimate column and plug them straight in: Y = intercept + b₁X₁ + b₂X₂.

⚠️ Expect a question asking you to interpret R². The answer is always: "X% of the variation in [dependent variable] is explained by [independent variables]." Use the specific variable names from the problem.

⚠️ Know how to compute degrees of freedom (df = n − 1 − p) and use it with a t-table to build a prediction interval.

⚠️ Be prepared to assess whether individual variables are significant by checking their p-values against a given significance level (usually 0.05 or 0.10).

⚠️ The prediction interval formula (point estimate ± t × Sₑ) is a likely exam calculation. Have the steps memorised.


Quick Self-Test

  1. Fill in the blank: R² measures the proportion of ______ in the dependent variable explained by the independent variables.
    Answer: variation.

  1. True or False: A p-value of 0.0001 means the variable is not statistically significant.
    Answer: False. A p-value of 0.0001 is very small, indicating strong statistical significance.

  1. Fill in the blank: For Problem 3, the degrees of freedom are ______.
    Answer: 67 (calculated as 70 − 1 − 2).

  1. True or False: The intercept in a regression model always represents the fixed cost.
    Answer: True in a cost-estimation context. It is the estimated cost when all cost drivers are zero.

  1. True or False: Multiple regression uses only the two most extreme data points.
    Answer: False. That is the high-low method. Regression uses all observations.


Practice Q&A

Q: Write the multiple regression equation from Problem 3.

A: Overhead = 75 + 13 × (Setup Hours) + 50 × (Number of Parts).

Q: What does R² = 0.83 mean in the context of Problem 3?

A: 83% of the variation in overhead costs is explained by setup hours and the number of parts.

Q: How many degrees of freedom does the regression in Problem 3 have?

A: 67. Calculated as n − 1 − p = 70 − 1 − 2.

Q: Is the "number of parts" variable statistically significant at the 5% level?

A: It is borderline. Its p-value is 0.0500, which is exactly at the 5% threshold. Depending on the convention used, it may or may not be considered significant. At the 10% level it is clearly significant.

Q: A job requires 500 setup hours and 300 parts. What is the estimated overhead?

A: 75 + 13(500) + 50(300) = 75 + 6,500 + 15,000 = $21,575.

Q: Using the answer above, construct a 90% prediction interval. Assume t(67, 90%) ≈ 1.668.

A: 21,575 ± 1.668 × 50 = 21,575 ± 83.40. The interval is approximately $21,491.60 to $21,658.40.

Q: What is the standard error of the estimate, and what does it tell you?

A: Sₑ = 50. It means the regression's predictions typically deviate from actual overhead by about $50 on average.


Connections to Other Topics

Regression builds directly on the high-low method. If you understand the cost equation Y = a + bX, regression is simply a more rigorous way of estimating a and b (and adding more Xs). Later in the course, the cost equations you develop here feed into flexible budgeting, variance analysis, and activity-based costing, all of which require separating fixed from variable costs.

The statistical concepts (R², p-values, confidence intervals) also appear in business statistics courses. Strengthening your understanding here pays off beyond cost accounting.


Related Terms / Search Tags

Tags: multiple regression, linear regression, cost estimation, R-squared, coefficient of determination, standard error, p-value, t-statistic, degrees of freedom, prediction interval, confidence interval, overhead costs, cost drivers, setup hours, AMIS 3300, cost accounting, regression output, parameter estimates, Ohio State, Turner, managerial accounting, goodness of fit, statistical significance