Difficulty: Introductory to intermediate | Prerequisites: Part 1 of these notes (sample space, events, set operations, Venn diagrams).
Once you have the language of sample spaces and events (Part 1), the next step is assigning numbers to those events. This section covers the three ways probability gets assigned (subjective, empirical, theoretical), the core properties every probability must satisfy, and the two workhorse rules: the complement rule and the general addition rule. These rules are the tools you will reach for first in nearly every probability calculation for the rest of the course.
Probability values come from personal judgement (subjective), observed frequencies (empirical), or equally likely outcomes (theoretical). All probabilities sit between 0 and 1, the whole sample space has probability 1, and the empty event has probability 0. The complement rule and the addition rule let you compute new probabilities from ones you already know.
Subjective probability
A probability assigned based on personal judgement, experience, or intuition rather than data or symmetry. Think of it as "calculating the odds" from gut feeling. A sports pundit saying a team has a 60% chance of winning is using subjective probability.
Empirical probability (relative frequency)
Probability estimated by running the experiment many times and dividing the number of times event A occurs by the total number of trials. In simple terms, P(A) ≈ n/N, where n is how often A happened and N is the total number of trials. The approximation improves as N grows large.
Theoretical probability (classical / equally likely)
Probability calculated by counting: P(A) = (number of outcomes in A) / (total number of outcomes in S). This works only when all outcomes are equally likely. Think of it as the textbook coin-flip approach: 1 favourable outcome out of 2 total, so P(Heads) = 1/2.
Complement rule
For any event A, P(A') = 1 − P(A). The probability of "not A" is one minus the probability of A. In simple terms, if there is a 70% chance of rain, there is a 30% chance of no rain.
General addition rule
P(A ∪ B) = P(A) + P(B) − P(A ∩ B). You add the two probabilities and subtract the overlap to avoid counting it twice. Think of it as: OR = sum of each, minus the double-counted AND.
Disjoint addition rule
When A and B are mutually exclusive (A ∩ B = ∅), the general addition rule simplifies to P(A ∪ B) = P(A) + P(B), because the overlap is zero.
There are three recognised approaches to assigning a probability value.
Subjective: based on personal belief or expert opinion. No formula, no repeated trials. Common in risk assessment and forecasting.
Empirical (relative frequency): based on observed data.
P(A) = (number of times A occurs) / (total number of trials).
Requires a large sample for the estimate to be reliable. As the number of trials N approaches infinity, the empirical probability converges to the true probability (this is the law of large numbers in informal terms).
Theoretical (equally likely / classical): based on counting.
P(A) = (number of outcomes favourable to A) / (total number of outcomes in S).
Only valid when every outcome in S is equally likely. Rolling a fair die qualifies; drawing from a loaded deck does not.
Bayesian statistics uses probability as a measure of evidential support for a hypothesis. The core idea: start with a prior probability, then update it in light of new data to get a posterior probability. This connects to Bayes' theorem, covered in Part 3 of these notes.
Every valid probability assignment must satisfy these axioms.
For any event A: 0 ≤ P(A) ≤ 1.
If w is an individual outcome in event A, then P(A) = Σ P(wᵢ) for all outcomes wᵢ in A.
P(S) = 1. The probability of the entire sample space (something happens) is 1. This is sometimes called the completeness property.
P(∅) = 0. The probability of the empty event (nothing happens) is 0.
For any event A:
P(A') = 1 − P(A)
This is useful whenever computing P(A) directly is difficult but computing P(A') is easy, or vice versa. "At least one" problems are the classic use case: P(at least one) = 1 − P(none).
For any two events A and B:
P(A ∪ B) = P(A) + P(B) − P(A ∩ B)
The subtraction of P(A ∩ B) corrects for the fact that outcomes in both A and B would otherwise be counted twice.
Special case (disjoint events):
When A ∩ B = ∅, the formula reduces to:
P(A ∪ B) = P(A) + P(B)
Suppose 70% of customers add sugar to their coffee (S), 35% add milk (M), and 25% add both sugar and milk (S ∩ M).
a) P(S ∪ M) = P(S) + P(M) − P(S ∩ M) = 0.70 + 0.35 − 0.25 = 0.80.
So 80% of customers add sugar or milk (or both).
b) P(M' ∩ S') = P((S ∪ M)') = 1 − P(S ∪ M) = 1 − 0.80 = 0.20.
So 20% of customers add neither sugar nor milk. Notice how the complement rule and De Morgan's law work together here: "neither S nor M" is the complement of "S or M."
Given simple events 0, 1, 2, 3 with probabilities 0.60, 0.25, 0.10, 0.05 respectively, and events A = {0}, B = {0, 1}, C = {3}, D = {0, 1, 2, 3}:
P(A') = P(1 ∪ 2 ∪ 3) = 0.25 + 0.10 + 0.05 = 0.40.
P(B) = P(0 ∪ 1) = 0.60 + 0.25 = 0.85.
P(A ∩ C) = P(∅) = 0, because A and C share no outcomes.
P(D) = P(S) = 1, because D contains every outcome.
Complement rule: P(A') = 1 − P(A)
General addition rule: P(A ∪ B) = P(A) + P(B) − P(A ∩ B)
Disjoint addition rule: P(A ∪ B) = P(A) + P(B) [when A ∩ B = ∅]
Empirical probability: P(A) = n / N
Theoretical probability: P(A) = |A| / |S| [equally likely outcomes]
Insurance companies use empirical probability (claims data from millions of policyholders) to set premiums. Theoretical probability underpins casino mathematics: every game is designed so the house edge is a known, calculated number. The complement rule shows up in reliability engineering, where it is often easier to calculate the probability that a system fails than the probability it works, and then subtract from 1.
Students often forget to subtract P(A ∩ B) when using the addition rule, accidentally double-counting the overlap. If you add two probabilities and get a number above 1, this is almost certainly the mistake.
Applying theoretical probability when outcomes are not equally likely. You cannot use (favourable outcomes) / (total outcomes) on a loaded die or a biased coin.
Confusing P(A') with P(B). The complement is specific to one event: everything in S that is not in A. It is not some other event B.
Assuming the complement rule is trivial and skipping it. It is one of the most powerful shortcuts in the course, especially for "at least one" problems.
⚠️ The addition rule (with and without overlap) is tested heavily. Know when to subtract P(A ∩ B) and when you do not need to.
⚠️ Expect a question that gives a probability table (outcomes with assigned probabilities) and asks you to compute probabilities of various events by summing the right entries.
⚠️ "At least one" phrasing almost always signals the complement rule: P(at least one) = 1 − P(none).
⚠️ Know the three types of probability and be able to identify which type a given scenario uses.
True or False: P(A ∪ B) can never exceed 1.
Fill in the blank: If P(A) = 0.3, then P(A') = ______.
True or False: The empirical probability of an event is always exactly equal to the theoretical probability.
Fill in the blank: The addition rule subtracts P(A ∩ B) to correct for ______.
True or False: P(S) = 0.
Answers: 1. True (probability never exceeds 1). 2. 0.7. 3. False (empirical is an approximation that improves with more trials). 4. Double-counting (outcomes in both A and B). 5. False, P(S) = 1.
Q: In a class, 60% of students passed the midterm (M), 75% passed the final (F), and 50% passed both. What is the probability a randomly chosen student passed at least one exam?
A: P(M ∪ F) = P(M) + P(F) − P(M ∩ F) = 0.60 + 0.75 − 0.50 = 0.85.
Q: Using the same class, what is the probability a student passed neither exam?
A: P((M ∪ F)') = 1 − P(M ∪ F) = 1 − 0.85 = 0.15.
Q: A bag contains 3 red and 7 blue marbles. What is the theoretical probability of drawing a red marble?
A: P(red) = 3/10 = 0.30. This uses theoretical probability because each marble is equally likely to be drawn.
Q: Name the three types of probability and give one example of each.
A: Subjective (an analyst estimates a 40% chance of recession), empirical (after 1,000 coin flips, heads came up 503 times, so P(H) ≈ 0.503), theoretical (a fair six-sided die gives P(rolling a 3) = 1/6).
The complement and addition rules here are the foundation for conditional probability (Part 3), which adds the multiplication rule and Bayes' theorem. The idea of empirical probability connects forward to sampling distributions and the frequentist interpretation of confidence intervals. The properties of probability (axioms) are the formal basis for everything in the probability unit, including expected value and variance covered in later chapters.
probability rules, complement rule, addition rule, general addition rule, union probability, empirical probability, relative frequency, theoretical probability, classical probability, subjective probability, equally likely outcomes, Bayesian statistics, prior probability, properties of probability, axioms, STAT, PIV 203, Chapter 4, introduction to statistics, Purdue