Difficulty: Introductory | Prerequisites: Basic algebra, familiarity with set notation helpful but not required.
This material sits at the very start of any probability or statistics course. Before you can calculate anything, you need the language: what counts as an experiment, what a "sample space" is, how events relate to each other through set operations, and the three axioms that every probability function must satisfy. If you have missed the first few weeks, this is where to begin. Everything that follows in the course (distributions, hypothesis testing, regression) rests on these definitions.
Probability theory starts with a few building blocks: experiments produce outcomes, outcomes live in a sample space, and events are subsets of that space. Set operations (union, intersection, complement) let you combine and manipulate events, while three axioms pin down what "probability" means mathematically. The frequentist and Bayesian interpretations offer two different lenses for reading those numbers.
Random experiment
A process that produces an uncertain outcome, one you cannot predict with certainty before running it. Rolling a die, flipping a coin, drawing a card.
In simple terms: any repeatable activity where you do not know the result in advance.
Trial
A single execution of a random experiment.
Think of it as one go, one flip, one draw.
Sample space (S or Ω)
The complete set of every possible outcome an experiment can produce. For a six-sided die, S = {1, 2, 3, 4, 5, 6}.
In simple terms: the full menu of things that could happen.
Event
Any subset of the sample space. It groups outcomes that share a condition, e.g. "rolling an even number" = {2, 4, 6}.
Think of it as a filter applied to the sample space.
Simple event
An event containing exactly one outcome.
Compound event (complex event)
An event containing more than one outcome.
Union (A ∪ B)
The set of outcomes in A or B or both. On a Venn diagram, the entire shaded area of both circles.
In simple terms: "A or B (or both) happens."
Intersection (A ∩ B)
The set of outcomes common to both A and B. The overlapping region on a Venn diagram.
In simple terms: "A and B both happen."
Complement (A′)
All outcomes in the sample space that are not in A.
Think of it as "not A."
Subset (A ⊆ B)
Every outcome in A is also in B. A is entirely contained within B.
Mutually exclusive (disjoint) events
Events that share no outcomes: A ∩ B = ∅. If one occurs, the other cannot.
In simple terms: they can never happen at the same time.
Probability function P
A rule that assigns a number between 0 and 1 to each event, representing how likely that event is.
Frequentist interpretation
Probability is the long-run relative frequency of an event across many repetitions of the experiment.
Think of it as: "If I ran this experiment thousands of times, what fraction of the time would this event occur?"
Bayesian interpretation
Probability is a degree of belief or subjective confidence about an event, updated as new evidence arrives.
In simple terms: how strongly you believe something, given what you know so far.
Posterior probability
The updated probability of an event after incorporating new evidence, central to the Bayesian approach.
A random experiment is any process whose outcome is uncertain. Each time you run it, that single run is called a trial, and the result of a trial is an outcome.
The sample space (S or Ω) collects every outcome the experiment can produce.
Finite: rolling a die, S = {1, 2, 3, 4, 5, 6}.
Countably infinite: counting the number of emails you receive in a day (0, 1, 2, 3, ...).
Continuous: measuring the exact time a lightbulb lasts.
An event is any subset of the sample space. "Rolling an even number" is the event {2, 4, 6}.
A simple event has exactly one outcome, e.g. {3}.
A compound event has more than one outcome, e.g. {2, 4, 6}.
An event "occurs" when the actual outcome of the trial falls inside it.
The distinction between outcomes and events trips people up early on. An outcome is a single result; an event is a collection of results you have chosen to group together.
Because events are sets, every set operation you already know applies here. These operations let you build new events from existing ones.
Union (A ∪ B): all outcomes in A, in B, or in both. "At least one of A or B occurs."
On a Venn diagram, the entire shaded region covering both circles.
Intersection (A ∩ B): only the outcomes that belong to both A and B. "A and B both occur."
The overlapping lens shape on a Venn diagram.
Complement (A′): everything in the sample space that is not in A. "A does not occur."
Subset (A ⊆ B): every outcome of A is also an outcome of B. Whenever A occurs, B must also occur.
Disjoint / mutually exclusive: A ∩ B = ∅. The two events share no outcomes and cannot both happen on the same trial.
Venn diagrams are the standard visual tool for these relationships. Drawing two overlapping circles inside a rectangle (the sample space) and shading the relevant region makes union, intersection, and complement concrete.
A probability function P assigns a number to each event, subject to three axioms:
Non-negativity: P(E) ≥ 0 for every event E.
Normalisation: P(Ω) = 1. The sample space (the certain event) has probability 1.
Additivity: For mutually exclusive events A and B, P(A ∪ B) = P(A) + P(B).
From these three axioms, everything else follows:
The impossible event (the empty set ∅) has probability 0: P(∅) = 0.
Probabilities cannot exceed 1: 0 ≤ P(E) ≤ 1.
The complement rule, the addition rule, and every other probability identity can be derived from these axioms.
The axioms are worth memorising word-for-word. Exam questions often ask you to prove a property from first principles, which means starting here.
The axioms tell you how probability behaves, but not what probability means. Two major schools of thought fill that gap.
Frequentist interpretation: probability is the long-run relative frequency. Flip a fair coin thousands of times and about half will be heads; that ratio is what P(heads) = 0.5 means. Probabilities are fixed properties of the experiment, not updated with each new data point.
Bayesian interpretation: probability is a degree of belief. You start with a prior (your initial belief about how likely something is), observe data, and use Bayes' rule to compute a posterior (your updated belief). Probabilities change as evidence accumulates.
Aspect | Frequentist | Bayesian |
|---|---|---|
Definition | Long-run relative frequency | Degree of belief |
Updating | Not typically updated | Updated with new data via Bayes' rule |
Example | "50% of many coin flips land heads" | "I believe there is a 70% chance of rain tomorrow" |
Most introductory statistics courses lean frequentist, but the Bayesian perspective appears whenever you encounter Bayes' rule or posterior probabilities.
Set identities let you rewrite complex event expressions into simpler forms. They mirror the algebraic identities you already know for numbers.
Commutative: A ∪ B = B ∪ A and A ∩ B = B ∩ A. Order does not matter.
Associative: (A ∪ B) ∪ C = A ∪ (B ∪ C) and likewise for intersection. Grouping does not matter.
Distributive: A ∩ (B ∪ C) = (A ∩ B) ∪ (A ∩ C) and A ∪ (B ∩ C) = (A ∪ B) ∩ (A ∪ C). Intersection distributes over union and vice versa.
DeMorgan's Laws:
(A ∪ B)′ = A′ ∩ B′. "Not (A or B)" equals "not A and not B."
(A ∩ B)′ = A′ ∪ B′. "Not (A and B)" equals "not A or not B."
DeMorgan's Laws are particularly useful in probability problems where you need to find the probability of a complement of a union or intersection. They convert a complement of a union into an intersection of complements (and vice versa), which is often easier to compute.
Kolmogorov's three axioms
P(E) ≥ 0
P(Ω) = 1
If A ∩ B = ∅, then P(A ∪ B) = P(A) + P(B)
Complement
P(A′) = 1 − P(A)
General addition rule
P(A ∪ B) = P(A) + P(B) − P(A ∩ B)
For mutually exclusive events (A ∩ B = ∅), this simplifies to P(A ∪ B) = P(A) + P(B).
DeMorgan's Laws
(A ∪ B)′ = A′ ∩ B′
(A ∩ B)′ = A′ ∪ B′
Distributive laws
A ∩ (B ∪ C) = (A ∩ B) ∪ (A ∩ C)
A ∪ (B ∩ C) = (A ∪ B) ∩ (A ∪ C)
Quality control in manufacturing relies on probability axioms to decide how many items to inspect and what defect rate is acceptable.
Weather forecasting uses sample spaces (all possible atmospheric states) and assigns probabilities to events like rain, which is why you hear "30% chance of rain" rather than a yes-or-no prediction.
Students often confuse "outcomes" with "events." An outcome is a single result (e.g. rolling a 3). An event is a set of outcomes (e.g. rolling an odd number = {1, 3, 5}). Keep the two straight.
Students sometimes think that P(A ∪ B) = P(A) + P(B) always. It only holds when A and B are mutually exclusive. For overlapping events you must subtract the intersection: P(A ∪ B) = P(A) + P(B) − P(A ∩ B).
"Probability of zero" does not always mean "impossible" in continuous sample spaces. An event can have P = 0 and still be a possible outcome (e.g. the probability of landing on any exact real number in [0, 1]).
DeMorgan's Laws are frequently applied in the wrong direction. Remember: complementing a union turns it into an intersection of complements, and complementing an intersection turns it into a union of complements.
⚠️ Be able to state all three axioms from memory. Exam questions regularly ask you to prove a result "from the axioms."
⚠️ Know when to use the general addition rule versus the simplified (mutually exclusive) version. A common exam mistake is dropping the intersection term.
⚠️ DeMorgan's Laws appear in nearly every problem set. Practise applying them in both directions until it feels automatic.
⚠️ Be ready to identify whether a sample space is finite, countably infinite, or continuous, and to list the sample space for simple experiments like dice, coins, and cards.
True or false: The sample space for flipping two coins is {H, T, HH, HT, TH, TT}.
True or false: If A and B are mutually exclusive, then P(A ∪ B) = P(A) + P(B).
Fill in the blank: (A ∪ B)′ = A′ ___ B′. (Union or intersection?)
True or false: P(∅) = 1.
Fill in the blank: An event that contains exactly one outcome is called a ___ event.
Answers: 1. False (the sample space is {HH, HT, TH, TT}; H and T alone are not outcomes of a two-coin experiment). 2. True. 3. Intersection (∩). 4. False, P(∅) = 0. 5. Simple.
Q: A fair six-sided die is rolled. Let A = {1, 2, 3} and B = {2, 4, 6}. What is P(A ∪ B)?
A: A ∪ B = {1, 2, 3, 4, 6}. Each outcome has probability 1/6, so P(A ∪ B) = 5/6. Alternatively, P(A) = 3/6, P(B) = 3/6, P(A ∩ B) = P({2}) = 1/6, so P(A ∪ B) = 3/6 + 3/6 − 1/6 = 5/6.
Q: If P(A) = 0.4 and A and B are mutually exclusive with P(A ∪ B) = 0.7, what is P(B)?
A: Because A and B are mutually exclusive, P(A ∪ B) = P(A) + P(B). So P(B) = 0.7 − 0.4 = 0.3.
Q: Using DeMorgan's Law, rewrite (A ∪ B ∪ C)′.
A: (A ∪ B ∪ C)′ = A′ ∩ B′ ∩ C′. The complement of a union is the intersection of the individual complements.
Q: A sample space has five equally likely outcomes: S = {a, b, c, d, e}. Event E = {a, b}. What is P(E′)?
A: P(E) = 2/5, so P(E′) = 1 − 2/5 = 3/5.
Q: Explain in one sentence why P(Ω) = 1 must be true.
A: The sample space contains every possible outcome, so the outcome of any trial is guaranteed to fall within it, making its probability 1.
This material connects directly to conditional probability and Bayes' rule, which build on the addition and multiplication of events covered here. If you are comfortable with unions, intersections, and complements, conditional probability will feel like a natural extension.
Set identities (especially DeMorgan's Laws) reappear throughout discrete mathematics and computer science, particularly in logic and database query design.
The frequentist vs Bayesian distinction matters later in the course when you encounter confidence intervals (frequentist) versus credible intervals (Bayesian).
probability theory, set operations, sample space, event, outcome, trial, random experiment, union, intersection, complement, subset, mutually exclusive, disjoint events, Venn diagram, probability axioms, Kolmogorov axioms, non-negativity, additivity, normalisation, frequentist probability, Bayesian probability, posterior probability, DeMorgan's Laws, De Morgan's Laws, commutative property, associative property, distributive property, set identities, simple event, compound event, elementary event, probability function, STAT 101, introduction to statistics