Difficulty: Introductory | Prerequisites: None, this is foundational material
This section covers how to read and critique the two most common statistical graphics (histograms and boxplots), plus the probability foundations you need before touching inference. You should be comfortable identifying the right probability distribution for a scenario, distinguishing independent events from disjoint events, and knowing whether a problem is about a single observation from a population or an average from a sampling distribution. Get these right and the inference topics later in the course will make far more sense.
Histogram
A bar chart for continuous data, where adjacent bars represent frequency (or relative frequency) of observations falling into consecutive, non-overlapping intervals called bins. Think of it as a picture of how your data is spread out.
Bin (class interval)
Each interval on the horizontal axis of a histogram. The number and width of bins determine how much detail the histogram reveals. In simple terms, bins are the buckets you sort your data points into.
Boxplot (box-and-whisker plot)
A five-number summary drawn as a graphic: minimum, Q1, median, Q3, maximum, with outliers plotted individually. Think of it as a compact snapshot of a distribution's centre, spread, and skewness.
Outlier
An observation that falls more than 1.5 times the interquartile range (IQR) below Q1 or above Q3. In simple terms, it is a data point that sits unusually far from the bulk of the data.
Interquartile range (IQR)
Q3 minus Q1, the range covered by the middle 50% of the data. Think of it as the width of the box in a boxplot.
Probability
The long-run relative frequency of an event, or a measure of how likely an event is to occur, ranging from 0 (impossible) to 1 (certain). In simple terms, it is the fraction of times something would happen if you repeated the experiment forever.
Independence (of events)
Two events are independent if the occurrence of one does not change the probability of the other. Formally, P(A and B) = P(A) × P(B). Think of it as: knowing A happened tells you nothing new about B.
Disjoint (mutually exclusive)
Two events are disjoint if they cannot both occur at the same time; P(A and B) = 0. In simple terms, if one happens, the other definitely did not.
Binomial distribution
Models the number of successes in a fixed number of independent trials, each with the same probability of success. Think of it as the "coin-flip counter" distribution.
Poisson distribution
Models the number of events occurring in a fixed interval of time or space, when events happen independently at a constant average rate. Think of it as the distribution for counting rare or random occurrences.
Normal distribution
A symmetric, bell-shaped continuous distribution defined by its mean and standard deviation. In simple terms, the classic bell curve where most values cluster near the centre.
Exponential distribution
Models the time between successive events in a Poisson process. Think of it as the "waiting time" distribution.
Uniform distribution
Every outcome in the range is equally likely. In simple terms, no value is favoured over any other.
Sampling distribution
The probability distribution of a statistic (such as the sample mean) computed from all possible samples of a given size from a population. Think of it as: not the distribution of individual data points, but the distribution of an average (or proportion) across many repeated samples.
Population distribution
The distribution of values for the entire population of interest. In simple terms, what the data looks like if you could measure every single individual.
A histogram's usefulness depends heavily on bin width. Too few bins and you smooth over genuine features of the data (bimodality, gaps, clusters). Too many bins and every bar is short, noisy, and hard to read.
A good rule of thumb: start with roughly the square root of n (the sample size) as the number of bins, then adjust.
The shape of the distribution should be visible: you should be able to identify whether it is symmetric, left-skewed, right-skewed, uniform, or bimodal.
If the histogram looks like a picket fence (many alternating tall and short bars), you likely have too many bins.
If the histogram looks like a single block or two blocks, you likely have too few.
All bins must be the same width for the visual to be honest. Unequal bin widths distort the impression of frequency.
A boxplot displays the five-number summary plus any outliers.
The box spans from Q1 (25th percentile) to Q3 (75th percentile). The width of the box is the IQR.
The line inside the box is the median (50th percentile), not the mean.
The whiskers extend from each end of the box to the most extreme data point that is not an outlier. That means the whisker reaches at most 1.5 × IQR beyond Q1 or Q3.
Individual points beyond the whiskers are outliers, observations more than 1.5 × IQR from the nearer quartile.
A boxplot does not show the sample size or the shape of the distribution within the box. You cannot tell from a boxplot alone whether the data is bimodal.
Reading a boxplot quickly:
If the median line is closer to Q1, the data is right-skewed (the upper tail is stretched).
If the median line is closer to Q3, the data is left-skewed.
If the median sits roughly centred, the data is approximately symmetric.
Many outliers on one side suggest a heavy tail in that direction.
This is one of the most commonly tested distinctions in intro stats.
Independent events can (and often do) happen at the same time. What makes them independent is that knowing one occurred does not change the probability of the other. P(A | B) = P(A).
Disjoint events cannot happen at the same time. P(A and B) = 0. Knowing one occurred tells you the other definitely did not.
The critical point: disjoint events are NOT independent (unless one of them has probability zero). If A and B are disjoint and A happened, then P(B | A) = 0, which is different from P(B). Knowing A changed your information about B, so they are dependent.
The exam will present a scenario and ask which distribution fits. Here is how to decide:
Binomial: fixed number of trials (n), two outcomes per trial (success/fail), constant probability of success (p), trials independent. Classic example: number of heads in 20 coin flips.
Poisson: counting events in a fixed interval, events are independent and occur at a constant average rate (lambda). Classic example: number of emails received per hour.
Normal: continuous data, symmetric bell shape. Appropriate when the variable can take any real value and is influenced by many small, additive factors. Also the limiting distribution for sample means via the Central Limit Theorem.
Exponential: the time (or distance) between consecutive Poisson events. Memoryless: the probability of waiting another 5 minutes does not depend on how long you have already waited.
Uniform: every value in an interval [a, b] is equally likely. Classic example: a random number generator between 0 and 1.
The population distribution describes individual observations. If you measure the height of every adult in a country, that is the population distribution.
The sampling distribution describes a statistic (often the sample mean) computed from repeated samples of the same size n. It is always less variable than the population distribution.
As n grows, the sampling distribution of the mean becomes approximately normal regardless of the population shape (Central Limit Theorem), and its standard deviation shrinks by a factor of 1/sqrt(n).
A question about "what is the probability that a randomly selected person weighs more than 200 lbs" refers to the population distribution.
A question about "what is the probability that the average weight of a sample of 36 people exceeds 200 lbs" refers to the sampling distribution.
Outlier boundaries (boxplot)
\text{Lower fence} = Q_1 - 1.5 \times IQR\text{Upper fence} = Q_3 + 1.5 \times IQRAny observation below the lower fence or above the upper fence is flagged as an outlier.
Multiplication rule for independent events
P(A \cap B) = P(A) \times P(B)Addition rule for disjoint events
P(A \cup B) = P(A) + P(B)Addition rule (general)
P(A \cup B) = P(A) + P(B) - P(A \cap B)Standard deviation of the sampling distribution of the mean
\sigma_{\bar{x}} = \frac{\sigma}{\sqrt{n}}This is often called the standard error when sigma is estimated from the sample.
Students often think disjoint and independent mean the same thing. They do not. Disjoint events are actually dependent (unless one has probability zero), because knowing one occurred rules out the other.
Students sometimes assume a histogram with many bars is always better. It is not. Too many bins create noise that masks the true shape of the distribution.
Students confuse the median line in a boxplot with the mean. The line inside the box is always the median. The mean is not shown on a standard boxplot.
Students sometimes believe the whiskers always extend to the minimum and maximum. They extend only to the most extreme non-outlier values. Outliers sit beyond the whiskers as individual points.
⚠️ Expect a question asking you to distinguish independence from disjointness, possibly with a follow-up asking whether disjoint events can be independent.
⚠️ Scenario-based distribution questions are very common: you will be given a real-world setup and asked to name the distribution.
⚠️ Know the difference between a question about an individual observation (population distribution) and a question about a sample mean (sampling distribution). The standard deviation changes.
⚠️ Boxplot reading: be prepared to identify skewness, outliers, and the five-number summary from a diagram.
True or false: if two events are disjoint, they must be independent. (False)
Fill in the blank: an observation is an outlier on a boxplot if it falls more than ______ × IQR beyond Q1 or Q3. (1.5)
True or false: the sampling distribution of the mean is always less variable than the population distribution. (True)
Fill in the blank: the ______ distribution models the number of events in a fixed interval at a constant average rate. (Poisson)
True or false: a histogram with unequal bin widths can still be interpreted the same way as one with equal bin widths. (False)
Q: A factory records the number of defective items per batch. Each batch has 50 items and the probability of any single item being defective is 0.03, independent of other items. Which distribution models the number of defectives per batch?
A: Binomial, because there is a fixed number of independent trials (n = 50), each with two outcomes (defective or not) and a constant probability of success (p = 0.03).
Q: You are told that events A and B are disjoint with P(A) = 0.3 and P(B) = 0.4. Are A and B independent? Why or why not?
A: No. If A and B are disjoint, then P(A and B) = 0. For independence we need P(A and B) = P(A) × P(B) = 0.12. Since 0 is not equal to 0.12, the events are not independent.
Q: A boxplot shows the median line sitting very close to Q3, with a long lower whisker and two outlier points below the lower fence. Describe the shape of the distribution.
A: The distribution is left-skewed. The median near Q3 and the long lower whisker (plus low outliers) indicate a stretched lower tail.
Q: A researcher wants to know the probability that the average height of 49 randomly selected adults exceeds 175 cm. Should she use the population distribution or the sampling distribution?
A: The sampling distribution of the mean, because the question asks about the average of a sample (n = 49), not an individual observation.
Q: On a histogram of exam scores, a student uses 2 bins for 200 data points. What is wrong with this choice?
A: Two bins are far too few. The histogram will hide nearly all features of the distribution (skewness, modes, gaps). A starting point of roughly sqrt(200), which is about 14 bins, would be more appropriate.
The probability distributions covered here (binomial, Poisson, normal) are the foundation for the inference section of the course. Confidence intervals and hypothesis tests rely on knowing which distribution to use in a given setting. The normal distribution in particular reappears through the Central Limit Theorem, which justifies the use of z- and t-based inference for sample means even when the population is not perfectly normal.
histogram bins, bin width, class interval, frequency distribution, boxplot, box-and-whisker, five-number summary, Q1, Q3, IQR, interquartile range, median, outlier detection, 1.5 IQR rule, probability, independence, disjoint, mutually exclusive, binomial distribution, Poisson distribution, normal distribution, bell curve, exponential distribution, uniform distribution, sampling distribution, population distribution, Central Limit Theorem, standard error