Difficulty: Intermediate | Prerequisites: Basic algebra, familiarity with summation notation, introductory probability concepts.
This material sits at the core of an introductory applied statistics course. The formulas here connect descriptive statistics (summarising data) to inferential statistics (drawing conclusions from samples). You need to be comfortable with probability distributions before moving to hypothesis testing and confidence intervals later in the course. If terms like "random variable" or "probability distribution" feel unfamiliar, revisit your Chapter 1 and 2 notes first.
This formula sheet covers the essential statistical formulas for the MATH 601 midterm: how to compute the mean and standard deviation of discrete and binomial random variables, how geometric probabilities work, and how sampling distributions behave. It also lists the Excel functions you will need for descriptive statistics and probability calculations on the exam.
Discrete random variable
A variable that can take on a countable number of distinct values, each with an associated probability. Think of it as a list of possible outcomes (like the number of heads in 10 coin flips) where you can assign a probability to each one.
Probability distribution (discrete)
A table or function listing every possible value of a discrete random variable alongside its probability. In simple terms, it is the complete map of "what can happen" and "how likely is each outcome."
Mean of a discrete random variable (expected value, μ)
The weighted average of all possible values, where each value is weighted by its probability: μ = ∑[x · P(x)]. Think of it as the long-run average outcome if you repeated the experiment thousands of times.
Standard deviation of a discrete random variable (σ)
A measure of how spread out the values of a random variable are around the mean: σ = √[∑(x – μ)² · P(x)]. In simple terms, a larger standard deviation means the outcomes are more scattered and less predictable.
Binomial distribution
A probability distribution for the number of successes in a fixed number of independent trials, where each trial has the same probability of success. Think of it as the model for "how many times does the thing happen out of n attempts" (e.g. how many heads in 20 coin flips).
Binomial variable parameters (n and p)
n is the number of trials; p is the probability of success on any single trial. These two numbers fully define a binomial distribution.
Geometric probability
The probability that the first success occurs on the nth trial in a sequence of independent trials, each with success probability p. The formula is p · q^(n–1), where q = 1 – p. In simple terms, it answers "how likely is it that I fail (n–1) times and then succeed?"
Sampling distribution of the sample mean (x̄)
The probability distribution of the sample mean across all possible samples of a given size from a population. Think of it as: if you drew thousands of different samples and computed the mean of each, the sampling distribution describes the pattern of all those means.
Standard error (standard deviation of x̄)
The standard deviation of the sampling distribution, calculated as σ / √n. It measures how much the sample mean is expected to vary from sample to sample. Larger samples produce smaller standard errors, which is why bigger samples give more precise estimates.
=MEAN(array) computes the arithmetic average of a range of cells. Note: in most Excel versions the actual function name is =AVERAGE(array), but your formula sheet lists =MEAN. Check which your instructor expects.
=MEDIAN(array) returns the middle value when data is sorted. For an even number of values, it averages the two middle ones.
=STDEV.S(array) calculates the sample standard deviation (divides by n–1, not n). Use this when working with a sample rather than an entire population.
=MIN(array) and =MAX(array) return the smallest and largest values in the range, respectively.
=PERCENTILES.EXC(array, k) returns the k-th percentile using the exclusive method (k must be between 0 and 1, exclusive). For example, k = 0.25 gives the 25th percentile.
=QUARTILES.EXC(array, k) returns the k-th quartile (k = 1 for Q1, 2 for Q2/median, 3 for Q3) using the exclusive method.
=BINOM.DIST(x, n, p, FALSE) returns the probability of exactly x successes in n trials (the probability mass function). Use FALSE for "exactly x."
=BINOM.DIST(x, n, p, TRUE) returns the cumulative probability of x or fewer successes (the cumulative distribution function). Use TRUE for "x or fewer."
=NORM.DIST(x, mean, stdev, TRUE) returns the cumulative probability P(X ≤ x) for a normal distribution with the given mean and standard deviation.
=NORM.INV(probability, mean, stdev) returns the x-value for which P(X ≤ x) equals the given probability. This is the inverse of NORM.DIST. Use it to find cutoff values (e.g. "what score marks the top 10%?").
The mean (expected value) of a discrete random variable is the sum of each outcome multiplied by its probability:
\mu = \sum x \cdot P(x)The standard deviation measures spread around that mean:
\sigma = \sqrt{\sum (x - \mu)^2 \cdot P(x)}To compute these by hand: build a table with columns for x, P(x), x · P(x), (x – μ), (x – μ)², and (x – μ)² · P(x). Sum the last column, then take the square root.
The probability that the first success occurs on trial n:
P(X = n) = p \cdot q^{n-1}where p = probability of success and q = 1 – p (probability of failure). This models scenarios like "how many calls before I make a sale" or "how many attempts before the machine breaks."
For a binomial random variable with n trials and success probability p:
\mu = n \cdot p\sigma = \sqrt{n \cdot p \cdot (1 - p)}These are shortcuts. You do not need to build a full probability table when you know you are dealing with a binomial setting.
When sampling from a population with standard deviation σ, the standard error of the sample mean is:
\sigma_{\bar{x}} = \frac{\sigma}{\sqrt{n}}Notice the √n in the denominator. Quadrupling the sample size cuts the standard error in half, not by a quarter. This is a common source of exam mistakes.
Formula Name | Expression | When to Use |
|---|---|---|
Mean of discrete variable | μ = ∑[x · P(x)] | Given a probability distribution table |
SD of discrete variable | σ = √[∑(x – μ)² · P(x)] | Given a probability distribution table |
Geometric probability | P(X = n) = p · q^(n–1) | First success on trial n |
Binomial mean | μ = n · p | Number of successes in n trials |
Binomial SD | σ = √[n · p · (1 – p)] | Spread of successes in n trials |
Standard error of x̄ | σ/√n | Variability of sample means |
Binomial distributions are used in quality control: if a factory knows 2% of its products are defective, the binomial model predicts how many defectives to expect in a batch of 500. Geometric probability shows up in reliability engineering, modelling how many cycles a component survives before failure. The sampling distribution and standard error underpin every opinion poll you see reported with a margin of error.
Students often confuse the population standard deviation formula (dividing by n) with the sample standard deviation (dividing by n–1). On this exam, STDEV.S uses n–1 because you are working with samples.
Students frequently mix up BINOM.DIST with FALSE (exactly x successes) and TRUE (x or fewer successes). The cumulative version (TRUE) is far more common in applied problems like "what is the probability of at most 3 defectives?"
A common error with the standard error formula: students assume that doubling n halves the standard error. It does not. Because n sits under a square root, you must quadruple n to halve the standard error.
Students sometimes apply the geometric probability formula starting from n = 0 rather than n = 1. The first trial is n = 1; the exponent on q is (n–1).
⚠️ Know when to use BINOM.DIST with FALSE versus TRUE. Exam questions often hinge on whether the problem says "exactly" or "at most" / "no more than."
⚠️ The standard error formula σ/√n is the bridge to confidence intervals and hypothesis testing in the second half of the course. Expect at least one question requiring you to compute it.
⚠️ NORM.INV and NORM.DIST are inverses of each other. If a question gives you a probability and asks for the cutoff value, use NORM.INV. If it gives you a value and asks for the probability, use NORM.DIST.
⚠️ Geometric probability: the exam may phrase it as "the probability of the first success on the nth trial" or equivalently "the probability of (n–1) failures followed by a success." Recognise both framings.
True or false: the mean of a binomial distribution is always n · p, regardless of the shape of the distribution.
Fill in the blank: to halve the standard error of the sample mean, you must multiply the sample size by ____.
True or false: =BINOM.DIST(3, 10, 0.5, TRUE) gives the probability of exactly 3 successes.
Fill in the blank: in the geometric probability formula p · q^(n–1), the letter q stands for ____.
True or false: =NORM.INV takes a probability as input and returns an x-value.
Answers: 1. True. 2. Four. 3. False (it gives the probability of 3 or fewer successes). 4. 1 – p (the probability of failure). 5. True.
Q: A discrete random variable X has the following distribution: X = 1 with P = 0.3, X = 2 with P = 0.5, X = 3 with P = 0.2. What is the mean of X?
A: μ = (1)(0.3) + (2)(0.5) + (3)(0.2) = 0.3 + 1.0 + 0.6 = 1.9
Q: In a binomial experiment with n = 20 and p = 0.4, what are the mean and standard deviation?
A: Mean = 20 × 0.4 = 8. Standard deviation = √(20 × 0.4 × 0.6) = √4.8 ≈ 2.19.
Q: What Excel formula would you use to find the probability of getting at most 5 successes in 12 trials with p = 0.3?
A: =BINOM.DIST(5, 12, 0.3, TRUE). The TRUE flag gives the cumulative probability P(X ≤ 5).
Q: A population has σ = 16. If you take samples of size n = 64, what is the standard error of the sample mean?
A: σ/√n = 16/√64 = 16/8 = 2.
Q: What is the probability that the first success occurs on the 4th trial, given p = 0.25?
A: P = 0.25 × (0.75)^3 = 0.25 × 0.4219 ≈ 0.1055.
Q: You need to find the score that separates the top 5% of a normal distribution with mean 70 and standard deviation 10. Which Excel function do you use, and what do you enter?
A: =NORM.INV(0.95, 70, 10). You enter 0.95 because 95% of the distribution falls below the top-5% cutoff.
The sampling distribution and standard error formula connect directly to confidence intervals and hypothesis testing, which form the second half of most introductory statistics courses. The binomial distribution is a building block for understanding the normal approximation to the binomial (relevant when n is large). Geometric probability relates to the broader family of waiting-time distributions, including the negative binomial.
Expected value, weighted average, probability mass function, PMF, CDF, cumulative distribution function, binomial experiment, Bernoulli trials, geometric distribution, sampling variability, standard error, central limit theorem, CLT, NORM.INV, NORM.DIST, BINOM.DIST, STDEV.S, PERCENTILES.EXC, QUARTILES.EXC, discrete probability distribution, random variable, MATH 601, midterm formulas, statistics formula sheet