Special Distributions and Descriptive Statistics, STAT 101 – Study Notes
offline

Difficulty: Intermediate | Prerequisites: Probability and Continuous Random Variables notes

Big Picture

After learning general probability rules and the idea of a continuous random variable, this section introduces the named distributions you will use most often: the exponential (time between events), the Poisson (count of events in a fixed period), and the binomial (number of successes in a fixed number of trials). Alongside these, you need to be able to read and interpret histograms, identify skewness, locate the median class, and detect outliers using the interquartile range. These are the building blocks for the hypothesis tests and confidence intervals that come next.


TL;DR

The exponential distribution models waiting times, the Poisson models event counts, and the binomial models success/failure counts. Descriptive statistics tools (histograms, IQR, skewness) help you summarise and screen data before running any formal test. Know the formulas, know how to read a histogram, and know the 1.5 x IQR rule for outliers.


Key Terms

Exponential distribution

A continuous distribution that models the time between events in a Poisson process. Its single parameter is the rate lambda (average number of events per unit time). The mean is 1/lambda. Think of it as the distribution you use whenever someone says "average waiting time."

Poisson distribution

A discrete distribution that counts the number of events occurring in a fixed interval of time or space, given a known average rate lambda. Think of it as the go-to when you are counting how many of something happen in a set period.

Binomial distribution

A discrete distribution that counts the number of successes in n independent trials, each with success probability p. Think of it as the coin-flip distribution, generalised.

Histogram

A bar chart for continuous or grouped data where each bar's height (or area) represents the frequency or relative frequency of values in that class. It is the first thing you look at to understand the shape of a dataset.

Skewness

A measure of asymmetry in a distribution. Left-skewed (negatively skewed) means the tail extends to the left. Right-skewed (positively skewed) means the tail extends to the right. In simple terms, the skew is named for the direction of the longer tail.

Interquartile range (IQR)

IQR = Q3 - Q1, the range of the middle 50% of the data. Used to measure spread and to detect outliers.

Outlier (by the 1.5 x IQR rule)

Any data point below Q1 - 1.5 x IQR or above Q3 + 1.5 x IQR. Think of it as a value that sits unusually far from the bulk of the data.

Median class

The class (bin) in a histogram that contains the median observation. For n data points, the median is the ((n+1)/2)th value when n is odd, or the average of the (n/2)th and (n/2 + 1)th values when n is even.


Core Content: Exponential Distribution

Setup and parameters

  • Used when waiting times between events follow a Poisson process.

  • Parameter: lambda = average rate of events per unit time.

  • Mean waiting time: 1/lambda.

  • Example: if phone calls arrive at an average of 1 every 0.26 hours, then lambda = 1/0.26 = 3.846 per hour.

Finding probabilities: CDF approach

  • The CDF is F(x) = 1 - e^(-lambda x) for x >= 0.

  • P(a < X < b) = F(b) - F(a) = (1 - e^(-lambda b)) - (1 - e^(-lambda a)) = e^(-lambda a) - e^(-lambda b).

  • Example: P(0.2 < X < 0.34) with lambda = 3.846. F(0.34) = 1 - e^(-3.846 x 0.34). F(0.2) = 1 - e^(-3.846 x 0.2). Subtract to get the answer.

Finding probabilities: PDF approach

  • The PDF is f(x) = lambda x e^(-lambda x) for x >= 0.

  • Integrate f(x) from a to b to get P(a < X < b).

  • Both approaches give the same result. The CDF is usually faster.

Applying the CLT to averages of exponential variables

  • If you have n independent phone lines each with exponential waiting times, the average waiting time across all n lines has approximately a normal distribution (by CLT) when n is large.

  • Mean of the average: mu = 0.26 (same as the individual mean).

  • Standard error: sigma / sqrt(n) = 0.26 / sqrt(100) = 0.026.

  • Then convert to z-scores and use the normal table.

  • Example: P(average between 0.2 and 0.34) with n = 100 lines. z1 = (0.2 - 0.26)/0.026, z2 = (0.34 - 0.26)/0.026. Use normal table.

Percentiles of the average

  • To find the top k% of average waiting times: find the z-value such that P(Z <= z) = 1 - k/100, then x = mu + z x (sigma/sqrt(n)).

  • Example: top 31.92% means P(Z <= b) = 1 - 0.3192 = 0.6808, so b = 0.47. Then x = 0.26 + 0.47 x 0.026 = 0.2722.


Core Content: Poisson Distribution

Setup and parameters

  • Counts the number of events in a fixed interval when events occur independently at a constant average rate.

  • Parameter: lambda = average number of events in the interval.

  • P(X = x) = (e^(-lambda) x lambda^x) / x! for x = 0, 1, 2, ...

  • Mean = lambda. Variance = lambda.

Adjusting lambda for different intervals

  • If the average rate is 3 per hour, then for a 2-hour window lambda = 6.

  • Always match the lambda to the interval the question asks about.

Example: emergency room arrivals

  • Average rate: 3 patients per hour. Between 1 pm and 3 pm = 2 hours, so lambda = 6.

  • P(X >= 2) = 1 - P(X = 0) - P(X = 1) = 1 - e^(-6) - 6e^(-6) = 1 - 0.0025 - 0.0149 = 0.983.

Linear combinations of Poisson and independent variables

  • If H is Poisson with Var(H) = 10 and E is independent with Var(E) = 3, then for the combination 3H - 2E:

  • Var(3H - 2E) = 9 Var(H) + 4 Var(E) = 9(10) + 4(3) = 102.

  • Standard deviation = sqrt(102) = 10.10.

  • Note: variance of a difference still adds the variances (because of the squaring of the coefficients).


Core Content: Binomial Distribution

Setup and parameters

  • n = fixed number of independent trials. p = probability of success on each trial.

  • P(X = x) = C(n, x) x p^x x (1 - p)^(n - x).

  • Mean = np. Variance = np(1 - p).

Example: ordering coffee at Starbucks

  • 20% of customers order a Grande. Out of the next 15 customers, S = number who order Grande. S is Binomial(n = 15, p = 0.2).

  • P(S >= 2) = 1 - P(S = 0) - P(S = 1).

  • P(S = 0) = C(15,0)(0.2)^0(0.8)^15.

  • P(S = 1) = C(15,1)(0.2)^1(0.8)^14.

  • Result: 1 - 0.035 - 0.132 = 0.833.

Normal approximation to the binomial

  • When n is large and the "success" and "failure" counts are both at least 5, use the normal with mean np and variance np(1 - p).

  • The variance of a binomial is np(1 - p). For the standard deviation of the count of "Large" orders out of 30 people: if M ~ Binomial(30, 0.40), Var(M) = 30 x 0.40 x 0.60 = 7.2.

Combining independent binomials

  • If S ~ Binomial(15, 0.2) and M ~ Binomial(30, 0.4) are independent, Var(3S - 2M) = 9 x Var(S) + 4 x Var(M).

  • Var(S) = 15(0.2)(0.8) = 2.4. Var(M) = 30(0.4)(0.6) = 7.2.

  • Var(3S - 2M) = 9(2.4) + 4(7.2) = 21.6 + 28.8 = 50.4. SD = sqrt(50.4) = 7.099.


Core Content: Histograms, Skewness, and Outlier Detection

Reading a histogram

  • The x-axis shows class intervals (bins). The y-axis shows frequency or relative frequency.

  • The shape tells you about the distribution: symmetric, left-skewed, or right-skewed.

  • Left-skewed: the left tail is longer. The peak is towards the right.

Identifying the median class

  • For n observations, the median is approximately the (n/2)th observation (or between n/2 and n/2 + 1).

  • Count cumulative frequencies from the left until you reach or pass n/2. That class contains the median.

  • Example: with 60 observations, the median is between the 30th and 31st values. If cumulative frequency reaches 30 in the 65 to 70 class, the median is in that class.

Outlier detection: the 1.5 x IQR rule

  • Compute Q1 (25th percentile) and Q3 (75th percentile).

  • IQR = Q3 - Q1.

  • Lower fence = Q1 - 1.5 x IQR.

  • Upper fence = Q3 + 1.5 x IQR.

  • Any observation below the lower fence or above the upper fence is an outlier.

Worked example with temperature data

  • Data: 71, 71, 75, 77, 78, 81, 82, 83, 83, 84, 85, 85, 86, 86, 86, 86, 87, 87, 87, 87, 89, 89, 90, 90, 90, 90, 90, 91, 91.

  • Q1 = 83, Q3 = 89.

  • IQR = 89 - 83 = 6.

  • Lower fence = 83 - 1.5(6) = 83 - 9 = 74.

  • Upper fence = 89 + 1.5(6) = 89 + 9 = 98.

  • Outliers: any values below 74. In this dataset, 71 and 71 are below 74, so there are two outliers at the lower end.


Formulas

\text{Exponential CDF: } F(x) = 1 - e^{-\lambda x}, \quad x \geq 0
\text{Exponential PDF: } f(x) = \lambda e^{-\lambda x}, \quad x \geq 0
\text{Exponential mean: } \mu = \frac{1}{\lambda}
\text{Poisson PMF: } P(X = x) = \frac{e^{-\lambda} \lambda^x}{x!}
\text{Binomial PMF: } P(X = x) = \binom{n}{x} p^x (1-p)^{n-x}
\text{Binomial mean: } \mu = np, \quad \text{Variance: } \sigma^2 = np(1-p)
\text{IQR} = Q_3 - Q_1
\text{Lower fence} = Q_1 - 1.5 \times \text{IQR}, \quad \text{Upper fence} = Q_3 + 1.5 \times \text{IQR}
\text{Var}(aX + bY) = a^2 \text{Var}(X) + b^2 \text{Var}(Y) \quad \text{(if X, Y independent)}

Real-World Applications

Exponential distributions are used in reliability engineering to model the lifespan of components and in queueing theory to model customer waiting times. Poisson distributions model rare events: insurance claims per month, server requests per second, radioactive decays per minute. The IQR rule is a standard first pass for flagging suspect data in any field, from clinical trials to financial auditing.


Common Misconceptions

  • Students often confuse the Poisson rate parameter with the exponential mean. If events arrive at rate lambda per hour, the Poisson uses lambda directly, but the exponential mean waiting time is 1/lambda.

  • A common error is subtracting variances in Var(aX - bY). The sign does not matter: it is a^2 Var(X) + b^2 Var(Y) when X and Y are independent, always addition.

  • Students sometimes label a histogram "right-skewed" because the tallest bars are on the right. Skewness is named for the tail, not the peak. If the long tail goes left, it is left-skewed.

  • The 1.5 x IQR rule identifies potential outliers, not definitive ones. In practice you investigate further, but on this exam treat it as the deciding criterion.


Why It Matters / Exam Flags

  • ⚠️ Exponential problems can be solved with either the CDF or by integrating the PDF. The CDF route is faster and less error-prone.

  • ⚠️ When the question says "average waiting time across n lines," that is a CLT problem, not a single-variable exponential problem. Switch to the normal distribution with standard error sigma/sqrt(n).

  • ⚠️ For Poisson, always check whether the time interval matches the stated rate. Adjust lambda before computing.

  • ⚠️ Outlier questions on the exam expect you to state the fences and then list which specific data values fall outside them.

  • ⚠️ Histogram questions often ask for the shape (skewness) and the median class in the same problem. Do both.


Quick Self-Test

  1. True or False: The mean of an exponential distribution with rate lambda = 4 is 4. (False: the mean is 1/4 = 0.25)

  1. True or False: For a Poisson random variable, the mean and variance are equal. (True)

  1. Fill in the blank: Var(3X - 2Y) = 9 Var(X) + ______ Var(Y), assuming independence. (4)

  1. True or False: In a left-skewed histogram, the longer tail points to the right. (False: the longer tail points to the left)

  1. Fill in the blank: A data point is an outlier if it falls below Q1 - ______ x IQR or above Q3 + ______ x IQR. (1.5, 1.5)


Practice Q&A

Q: Phone calls arrive at a rate of 1 every 0.26 hours. What is the probability that the waiting time for the next call is between 0.2 and 0.34 hours?

A: lambda = 1/0.26 = 3.846. P(0.2 < X < 0.34) = e^(-3.846 x 0.2) - e^(-3.846 x 0.34) = e^(-0.769) - e^(-1.308) = 0.463 - 0.270 = 0.193.

Q: An emergency room sees an average of 3 patients per hour. What is the probability that at least 2 arrive between 1 pm and 3 pm?

A: lambda = 3 x 2 = 6 for the 2-hour window. P(X >= 2) = 1 - P(X = 0) - P(X = 1) = 1 - e^(-6) - 6e^(-6) = 0.983.

Q: Given temperature data with Q1 = 83 and Q3 = 89, identify any outliers from the dataset.

A: IQR = 6. Lower fence = 83 - 9 = 74. Upper fence = 89 + 9 = 98. Any values below 74 or above 98 are outliers. The values 71 and 71 are below 74, so there are two outliers at the lower end.

Q: A histogram of automobile speeds is left-skewed. Where is the peak relative to the tail?

A: The peak (highest bars) is towards the right side, and the tail stretches out to the left.

Q: If S is Binomial(15, 0.2), what is P(S >= 2)?

A: P(S >= 2) = 1 - P(S = 0) - P(S = 1) = 1 - (0.8)^15 - 15(0.2)(0.8)^14 = 1 - 0.035 - 0.132 = 0.833.


Connections to Other Topics

The exponential and Poisson distributions are mathematically linked: if events follow a Poisson process, the time between them is exponential. Binomial proportions lead into confidence intervals for proportions and chi-squared tests later in the course. The descriptive statistics tools here (histograms, IQR) are the same ones you use to check assumptions before running a t-test or ANOVA.


Related Terms / Search Tags

Exponential distribution, Poisson distribution, binomial distribution, normal approximation to the binomial, lambda, rate parameter, waiting time, event count, histogram, left-skewed, right-skewed, symmetric, frequency, relative frequency, interquartile range, IQR, Q1, Q3, outlier, 1.5 IQR rule, fences, median class, cumulative frequency, linear combination of random variables, variance of a sum, variance of a difference, STAT 101, Purdue, introduction to statistics