Difficulty: Introductory to Intermediate | Prerequisites: Basic algebra, familiarity with integration
Probability and continuous random variables form the backbone of inferential statistics. Everything you will do later in the course, from hypothesis testing to confidence intervals to ANOVA, rests on understanding how probability works and how continuous distributions behave. This material bridges the gap between descriptive statistics (summarising data you already have) and making claims about populations you have not fully observed. You should already be comfortable with basic set notation, summation, and simple integration before diving in.
Probability quantifies uncertainty using rules for combining events (complement, addition, multiplication). Continuous random variables are described by probability density functions (PDFs), and you find probabilities by integrating the PDF over an interval. The Central Limit Theorem tells you that sample means follow an approximately normal distribution regardless of the population shape, which is why so many statistical tests work the way they do.
Probability
A number between 0 and 1 that measures how likely an event is to occur. Think of it as the long-run relative frequency if you repeated the experiment forever.
Sample space (S)
The set of all possible outcomes of a random experiment. In simple terms, it is everything that could happen.
Independent events
Two events A and B are independent if the occurrence of one does not change the probability of the other. Formally, P(A and B) = P(A) x P(B). Think of it as: knowing A happened tells you nothing new about B.
Complement (A')
The event that A does not occur. P(A') = 1 - P(A). In simple terms, the complement covers everything outside A.
Conditional probability
The probability of event A given that event B has already occurred, written P(A | B) = P(A and B) / P(B). Think of it as narrowing the sample space down to only the outcomes where B happened.
Addition rule
P(A or B) = P(A) + P(B) - P(A and B). You subtract the intersection to avoid counting it twice.
Probability density function (PDF)
A function f(x) that describes the relative likelihood of a continuous random variable taking a particular value. The total area under the curve equals 1, and the probability of any single exact value is 0. Think of it as a smooth curve whose area over an interval gives you the probability of landing in that interval.
Cumulative distribution function (CDF)
F(x) = P(X <= x), the probability that the random variable is less than or equal to x. It is the running total of the PDF from negative infinity up to x.
Expected value (mean), E(X)
The long-run average of a random variable. For a continuous RV, E(X) = the integral of x times f(x) over the entire range. Think of it as the balance point of the distribution.
Variance, Var(X)
A measure of how spread out the distribution is around its mean. Var(X) = E(X squared) - [E(X)] squared. In simple terms, larger variance means the values are more scattered.
Standard deviation
The square root of the variance. It lives in the same units as X, making it easier to interpret than variance.
Central Limit Theorem (CLT)
For a sufficiently large sample size n, the distribution of the sample mean is approximately normal, regardless of the shape of the population distribution. The mean of the sampling distribution equals the population mean, and its standard deviation equals sigma / square root of n. This is why normal-based methods work so broadly in statistics.
Independent events and the multiplication rule
Two events A and B are independent if and only if P(A and B) = P(A) x P(B).
If A and B are independent with positive probabilities, the following are always true:
P(A and B) < P(A), because you are multiplying P(A) by a fraction less than 1.
P(A and B) <= P(B), for the same reason.
P(A and B) = P(A) x P(B), by definition.
All three of those statements can be true simultaneously. Exam questions sometimes ask which "must be false" for independent events. None of the three above is false on its own.
Complement rule
P(A') = 1 - P(A).
Useful when it is easier to calculate "not A" than A directly. For example, P(X >= 2) = 1 - P(X < 2).
Addition rule (union of events)
P(A or B) = P(A) + P(B) - P(A and B).
If A and B are mutually exclusive (cannot both happen), the intersection is 0, so P(A or B) = P(A) + P(B).
Conditional probability and Venn diagrams
P(A | B) = P(A and B) / P(B).
In applied problems you often build a Venn diagram or a two-way table first, fill in the known probabilities, and then read off whatever the question asks.
Example from the exam material: 80% of Purdue students like Chinese food, 60% like American food, and the probability a student who does not like Chinese food also does not like American food is 0.30.
Let C = likes Chinese food, A = likes American food.
P(C) = 0.8, P(A) = 0.6, P(A' | C') = 0.30.
P(C') = 0.2, so P(A' and C') = P(A' | C') x P(C') = 0.30 x 0.20 = 0.06.
Students who do not like American food and do not like Chinese food: 6%.
P(A or C) = 1 - P(A' and C') = 1 - 0.06 = 0.94.
P(A and C) = P(A) + P(C) - P(A or C) = 0.6 + 0.8 - 0.94 = 0.46.
Properties of a valid PDF
f(x) >= 0 for all x.
The total area under f(x) over its entire range equals 1.
P(a < X < b) = the integral of f(x) from a to b.
P(X = any single value) = 0 for continuous variables.
Computing probabilities by integration
Given a PDF like f(x) = (3/80)(6 + 3x - x^2) for 0 < x < 4 and 0 otherwise:
To find P(X < 2), integrate f(x) from 0 to 2.
Set up the integral: (3/80) times the integral of (6 + 3x - x^2) dx from 0 to 2.
Evaluate the antiderivative: 6x + (3/2)x^2 - (1/3)x^3, then plug in the limits.
Result: (3/80)(12 + 6 - 8/3) = (3/80)(46/3) = 46/40 = 23/40 = 0.575.
Median of a continuous random variable
The median m satisfies P(X < m) = 0.50.
If P(X < 2) = 0.575, then since 0.575 > 0.50, the median is less than 2.
You do not need to recalculate if you already know P(X < 2). Just compare it to 0.50.
Expected value and variance
E(X) = integral of x times f(x) dx over the range.
E(X^2) = integral of x^2 times f(x) dx over the range.
Var(X) = E(X^2) - [E(X)]^2.
Example: if E(X) = 1.8 and E(X^2) = 4.32, then Var(X) = 4.32 - (1.8)^2 = 4.32 - 3.24 = 1.08.
What the CLT says
If you take random samples of size n from a population with mean mu and standard deviation sigma, then the sampling distribution of the sample mean (X-bar) is approximately normal when n is large enough (typically n >= 30).
Mean of sampling distribution: mu_X-bar = mu.
Standard deviation of sampling distribution (standard error): sigma_X-bar = sigma / sqrt(n).
Applying the CLT to find probabilities about sample means
Example: 50 stereo cartridges have a population mean weight of 1.8 g and standard deviation of 1.08 (from Var(X) = 1.08, so sigma = sqrt(1.08) approx 1.04). However, if the problem states sigma directly, use that.
The standard error is sigma / sqrt(50).
Convert to a z-score: z = (X-bar - mu) / (sigma / sqrt(n)).
Use the standard normal table to find the probability.
Example: P(X-bar > 2) = 1 - P(Z <= (2 - 1.8) / (sigma / sqrt(50))).
Note: if mu and sigma are not explicitly stated, you must show them in the z-score calculation. Even if you compute them from earlier parts of the problem, state them clearly.
Percentiles of the sampling distribution
To find the value x such that the sample mean is at or below it with probability p, use: x = mu + z_p times (sigma / sqrt(n)).
Example: the 90.66th percentile. P(Z <= b) = 0.9066 gives b = 1.32.
Then x = 1.8 + 1.32 times (sigma / sqrt(n)).
If sigma = 1.08 and n = 60: standard error = 1.08 / sqrt(60) approx 0.1342. So x = 1.8 + 1.32 x 0.1342 approx 1.977.
When to use the CLT vs. the exact distribution
If the population is normal, X-bar is exactly normal for any n.
If the population is not normal, use CLT only when n is large enough (rule of thumb: n >= 30).
For small n from a non-normal population, you cannot use the normal approximation.
P(A') = 1 - P(A)P(A \cup B) = P(A) + P(B) - P(A \cap B)P(A \mid B) = \frac{P(A \cap B)}{P(B)}E(X) = \int_{-\infty}^{\infty} x \, f(x) \, dxE(X^2) = \int_{-\infty}^{\infty} x^2 \, f(x) \, dx\text{Var}(X) = E(X^2) - [E(X)]^2\sigma = \sqrt{\text{Var}(X)}\text{Standard error} = \frac{\sigma}{\sqrt{n}}z = \frac{\bar{X} - \mu}{\sigma / \sqrt{n}}\text{Percentile value: } x = \mu + z_p \cdot \frac{\sigma}{\sqrt{n}}PDFs model physical quantities where outcomes are measured on a continuous scale: the weight of manufactured parts, the time until a device fails, the voltage across a circuit. Quality control engineers integrate the PDF of part weights to determine what fraction fall outside specification limits. The CLT is the reason polling organisations can estimate national opinion from a sample of 1,000 people.
Students often think P(X = 3) has a nonzero value for a continuous variable. It does not. Only intervals have positive probability.
Students confuse "independent" with "mutually exclusive." Mutually exclusive events cannot happen together (P(A and B) = 0), whereas independent events can happen together but do not influence each other.
A common mistake is applying the CLT to small samples from non-normal populations. The CLT requires a large enough n (usually n >= 30) when the underlying distribution is not normal.
Students sometimes forget to square E(X) when computing variance. The formula is E(X^2) minus [E(X)] squared, not E(X^2) minus E(X).
⚠️ Integration of the PDF is the single most common calculation for continuous RVs. Be comfortable setting up and evaluating integrals by hand.
⚠️ The median comparison shortcut: if P(X < c) > 0.50, the median is less than c. This saves time.
⚠️ For CLT problems, always state mu and sigma clearly before computing the z-score. Marks are lost when these values appear without explanation.
⚠️ Venn diagram problems (like the Chinese/American food question) appear regularly. Practise filling in all four regions systematically.
⚠️ Percentile questions require you to look up z from the normal table in reverse (given the area, find z). Know how to read the table both ways.
True or False: For a continuous random variable, P(X = 5) = 0. (True)
True or False: If A and B are independent, then A and B are mutually exclusive. (False: independent events can occur together)
Fill in the blank: Var(X) = E(X^2) - ______. ([E(X)]^2)
True or False: The Central Limit Theorem requires the population to be normally distributed. (False: it works for any population shape given large enough n)
Fill in the blank: The standard error of the sample mean equals sigma divided by ______. (sqrt(n))
Q: Given f(x) = (3/80)(6 + 3x - x^2) for 0 < x < 4, what is P(X < 2)?
A: Integrate from 0 to 2. The antiderivative is 6x + (3/2)x^2 - (1/3)x^3. Evaluating: (3/80)(12 + 6 - 8/3) = (3/80)(46/3) = 23/40 = 0.575.
Q: Using the PDF above, is the median greater than, less than, or equal to 2?
A: Less than 2, because P(X < 2) = 0.575 > 0.50, which means more than half the distribution lies below 2.
Q: If E(X) = 1.8 and E(X^2) = 4.32, what is Var(X)?
A: Var(X) = 4.32 - (1.8)^2 = 4.32 - 3.24 = 1.08.
Q: 80% of students like Chinese food, 60% like American food, and among those who do not like Chinese food, 30% also do not like American food. What percentage of students like at least one of the two?
A: P(A' and C') = 0.30 x 0.20 = 0.06. So P(A or C) = 1 - 0.06 = 0.94, meaning 94%.
Q: A sample of 50 cartridges is drawn from a population with mean 1.8 and known standard deviation. Using the CLT, what distribution does the sample mean follow?
A: The sample mean follows an approximately normal distribution with mean 1.8 and standard error sigma / sqrt(50).
Q: Find the 90.66th percentile of the sampling distribution if mu = 1.8, sigma = 1.08, and n = 60.
A: z = 1.32 (from the normal table). Standard error = 1.08 / sqrt(60) = 0.1342. Percentile value = 1.8 + 1.32 x 0.1342 = 1.977.
This material connects directly to hypothesis testing, because every test statistic is built on the idea of a sampling distribution (which the CLT provides). Confidence intervals use the standard error formula from this section to set the width of the interval. The Poisson and exponential distributions (covered in the next set of notes) are special cases of the continuous/discrete framework introduced here.
Probability rules, independent events, mutually exclusive, complement rule, addition rule, conditional probability, Bayes' theorem, Venn diagram, two-way table, probability density function, PDF, cumulative distribution function, CDF, expected value, mean, variance, standard deviation, integration, continuous random variable, Central Limit Theorem, CLT, sampling distribution, standard error, z-score, percentile, normal approximation, STAT 101, Purdue, introduction to statistics