Difficulty: Intermediate | Prerequisites: Probability Fundamentals and Random Variables study notes
Big picture: Once you know what a random variable is, the next step is learning the named distributions that describe how outcomes are spread. This set of notes covers the main discrete and continuous distributions on the intro stats syllabus, then takes a close look at the normal distribution, which is the single most important distribution in the course. You will also meet the idea of approximating one distribution with another and checking normality with Q-Q plots.
Probability distributions are named models that describe how the outcomes of a random variable are spread. Discrete distributions (binomial, Poisson, geometric, and others) handle countable outcomes; continuous distributions (normal, exponential, beta, and others) handle outcomes on a scale. The normal distribution dominates intro stats because of its symmetric bell shape, the empirical rule, and the fact that many real-world measurements follow it closely enough to be useful.
Binomial distribution
Models the number of successes in n independent trials, each with the same probability p of success. Think of it as: "out of n attempts, how many times does the thing happen?"
Bernoulli distribution
A special case of the binomial with n = 1. One trial, two outcomes: success or failure.
Hypergeometric distribution
Models the number of successes when drawing from a finite population without replacement. The trials are dependent because each draw changes the pool. Think of it as the binomial's cousin for sampling without replacement.
Poisson distribution
Models the count of independently occurring events in a fixed interval of time or space. In simple terms, it answers "how many times does something happen per hour (or per kilometre, per page)?".
Geometric distribution
Models the number of independent trials needed to get the first success. Think of it as: "how long until it finally works?"
Negative binomial distribution
Models the number of independent trials needed to achieve the rth success. A generalisation of the geometric (which is the r = 1 case).
Normal distribution (Gaussian distribution)
A continuous, symmetric, bell-shaped distribution defined by its mean (mu) and standard deviation (sigma). It is the most widely used distribution in statistics.
Standard normal distribution
A normal distribution with mean 0 and standard deviation 1. Any normal variable can be converted to a standard normal using a z-score.
Z-score (standard score)
The number of standard deviations a value sits from the mean: z = (x - mu) / sigma. In simple terms, it tells you how unusual a value is.
Empirical rule (68-95-99.7 rule)
For a normal distribution, roughly 68% of values fall within 1 standard deviation of the mean, 95% within 2, and 99.7% within 3.
Continuity correction
An adjustment of plus or minus 0.5 applied when a discrete distribution is approximated by a continuous one. It accounts for the fact that a discrete value like 5 occupies the interval 4.5 to 5.5 on a continuous scale.
Normal Q-Q plot
A graphical tool that plots the quantiles of your sample data against the quantiles you would expect from a normal distribution. If the points fall roughly on a straight line, the data are approximately normal.
Uniform (discrete): Every outcome has the same probability. Rolling a fair die is the textbook example.
Bernoulli: A single trial with two outcomes. Success with probability p, failure with probability 1 - p.
Binomial: n independent Bernoulli trials. You count total successes. Parameters: n (number of trials), p (probability of success on each trial).
Hypergeometric: Like the binomial, but draws are without replacement from a finite population. The probability changes with each draw.
Poisson: Counts events occurring independently at a constant average rate (lambda) over a fixed interval. Useful for rare events: calls per hour, typos per page.
Geometric: Number of trials until the first success. Memoryless property: past failures do not change the probability of success on the next trial.
Negative binomial: Number of trials until the rth success. The geometric distribution is a special case with r = 1.
Uniform (continuous): Every value in some interval [a, b] is equally likely. The density is flat.
Normal: Symmetric bell curve defined by mean (mu) and standard deviation (sigma). The area under the entire curve equals 1.
Exponential: Models the time between events in a Poisson process. The density decreases at an exponential rate. Also memoryless.
Beta: Takes values between 0 and 1. Commonly used to model proportions or probabilities themselves.
Gamma: Generalises the exponential. Can be thought of as the sum of multiple independent exponential random variables.
t-distribution: Similar shape to the normal but with heavier tails. Used when the population standard deviation is unknown and the sample size is small.
The normal distribution is defined entirely by two parameters: the mean (mu, the centre) and the standard deviation (sigma, the spread). The curve is symmetric about the mean, so the mean, median, and mode are all equal.
The empirical rule provides quick probability benchmarks:
About 68% of values lie within 1 sigma of the mean.
About 95% within 2 sigma.
About 99.7% within 3 sigma.
To look up probabilities in a standard normal table, convert any normal variable X to Z using z = (x - mu) / sigma. The resulting Z follows a standard normal distribution with mean 0 and standard deviation 1.
When n is large enough that both np > 5 and n(1 - p) > 5, the binomial distribution is well approximated by a normal distribution with mean = np and standard deviation = sqrt(np(1 - p)).
Because you are approximating a discrete distribution with a continuous one, a continuity correction improves accuracy. For example, to find P(X >= 10) for a binomial, compute P(Y >= 9.5) using the normal approximation.
A Q-Q plot is a diagnostic tool. You order your sample data from smallest to largest and plot each value against the corresponding quantile of a standard normal distribution. If the points lie close to a straight line, normality is a reasonable assumption. Curvature or S-shapes indicate skewness or heavy tails.
Z = \frac{X - \mu}{\sigma}P(\mu - \sigma < X < \mu + \sigma) \approx 0.68P(\mu - 2\sigma < X < \mu + 2\sigma) \approx 0.95P(\mu - 3\sigma < X < \mu + 3\sigma) \approx 0.997Binomial-to-normal approximation conditions:
np > 5 \quad \text{and} \quad n(1-p) > 5Approximating mean and standard deviation:
\mu = np, \quad \sigma = \sqrt{np(1-p)}Continuity correction: when computing P(X >= k) for a discrete variable using a continuous approximation, use P(Y >= k - 0.5). When computing P(X <= k), use P(Y <= k + 0.5).
The binomial distribution is used in quality control (how many defective items in a batch?) and clinical trials (how many patients respond to treatment?). The Poisson distribution models events like website hits per minute or insurance claims per year. The normal distribution appears everywhere: standardised test scores, measurement errors in manufacturing, heights and weights in a population. The exponential distribution models the time between calls at a help desk or between arrivals at a queue.
Students often think the binomial and hypergeometric are interchangeable. The binomial assumes independent draws (with replacement or from an effectively infinite population); the hypergeometric does not.
A frequent error is applying the empirical rule to distributions that are not normal. The 68-95-99.7 percentages hold only for normal (or very nearly normal) data.
Students sometimes forget the continuity correction when approximating a binomial with a normal. Omitting it introduces a systematic error, especially for moderate sample sizes.
Many students assume that a Q-Q plot must be perfectly linear. Small deviations at the tails are common even in genuinely normal data; what matters is the overall pattern.
⚠️ You will be asked to identify which distribution fits a word problem. The key discriminators: independent trials with fixed n? Binomial. Dependent draws from a finite pool? Hypergeometric. Events per interval? Poisson.
⚠️ Know the empirical rule percentages (68, 95, 99.7) and be able to apply them to find probabilities for intervals stated in terms of sigma.
⚠️ z-score computation and table lookup are almost guaranteed exam material. Practise converting between raw scores and z-scores.
⚠️ The conditions for the normal approximation to the binomial (np > 5 and n(1-p) > 5) and the continuity correction are commonly tested.
⚠️ Interpreting a Q-Q plot may appear as a short-answer or multiple-choice question. Know what linearity implies and what curvature signals.
True or false: The Poisson distribution models the number of successes in a fixed number of independent trials.
False. That is the binomial. The Poisson models the count of events in a fixed interval of time or space.
Fill in the blank: About ______% of values in a normal distribution lie within 2 standard deviations of the mean.
95%.
True or false: A z-score of -1.5 means the value is 1.5 standard deviations below the mean.
True.
Fill in the blank: The normal approximation to the binomial requires both np > ______ and n(1-p) > ______.
5 and 5.
True or false: In a Q-Q plot, a straight line indicates the data are perfectly normally distributed.
False. A roughly straight line suggests approximate normality. Perfect linearity is not expected in real data.
Q: You flip a fair coin 3 times. List the sample space and identify the distribution of the number of heads.
A: The sample space is {HHH, HHT, HTH, THH, HTT, THT, TTH, TTT}. The number of heads follows a Binomial(n = 3, p = 0.5) distribution.
Q: A machine produces items with a 2% defect rate. In a batch of 100, what distribution models the number of defective items, and what are its parameters?
A: Binomial with n = 100 and p = 0.02. (If the batch is drawn without replacement from a very large production run, the binomial is a good approximation to the hypergeometric.)
Q: Scores on an exam are normally distributed with mean 72 and standard deviation 8. What percentage of students score between 64 and 80?
A: 64 = 72 - 8 (one sigma below) and 80 = 72 + 8 (one sigma above). By the empirical rule, about 68% of students score in this range.
Q: For the same exam, what is the z-score of a student who scored 88?
A: z = (88 - 72) / 8 = 2.0. The student scored 2 standard deviations above the mean.
Q: X follows Binomial(n = 50, p = 0.4). Can you use the normal approximation? If so, what mean and standard deviation do you use?
A: Check conditions: np = 20 > 5 and n(1-p) = 30 > 5, so yes. Use mean = 20 and standard deviation = sqrt(50 x 0.4 x 0.6) = sqrt(12) = approximately 3.46.
The distributions covered here connect directly to sampling distributions and the Central Limit Theorem, which explains why the normal distribution appears so often in practice. Hypothesis testing and confidence intervals both rely on knowing which distribution applies and how to compute probabilities under it. The Poisson distribution links to the exponential through the Poisson process, a connection that surfaces in queuing theory and reliability engineering.
Related Terms / Search Tags
binomial distribution, Bernoulli trial, hypergeometric, Poisson distribution, geometric distribution, negative binomial, normal distribution, Gaussian, bell curve, standard normal, z-score, z-table, empirical rule, 68-95-99.7, continuity correction, normal approximation, Q-Q plot, quantile-quantile plot, discrete distribution, continuous distribution, exponential distribution, beta distribution, gamma distribution, t-distribution, intro to statistics, Purdue STAT