Coinbase Data Scientist Statistics, Probability and Hypothesis Testing – Study Notes
offline

Difficulty: Intermediate | Prerequisites: Basic statistics (means, distributions, p-values), Part 1 (Interview Process) Date: 2025 | Source: Interview Query – Coinbase Data Science Interview Guide

Tags: Coinbase, statistics, hypothesis testing, p-value, chi-square test, t-test, type I error, type II error, significance level, time series, imputation, missing data, data scientist interview


Big Picture

Statistics and hypothesis testing form the backbone of data-driven decision-making at Coinbase. Whether you are evaluating a new feature's impact on user retention, comparing damage rates between shipping options, or assessing whether a month-over-month change in trading volume is meaningful, you need to choose the right test, set up hypotheses correctly, and interpret results in a business context. These questions appear frequently in the virtual interview rounds and are often framed around Coinbase-specific scenarios involving user behaviour, cryptocurrency markets, or product metrics.


TL;DR

Coinbase tests your ability to select appropriate statistical tests (chi-square, t-test, Mann-Whitney U), define null and alternative hypotheses, interpret p-values in business context, understand type I and type II errors, and handle missing data through imputation. Expect scenario-based questions that require you to connect statistical reasoning to product decisions.


Key Terms

P-value

The probability of observing results at least as extreme as the measured results, assuming the null hypothesis is true. In simple terms, a small p-value means the data is unlikely under the assumption that nothing interesting is happening, so you have reason to reject that assumption.

Null hypothesis (H₀)

The default assumption that there is no effect or no difference. You test against this. For example: "there is no difference in damage rates between parcel A and parcel B."

Alternative hypothesis (H₁ or Hₐ)

The claim you are trying to find evidence for. For example: "there is a difference in damage rates between parcel A and parcel B."

Significance level (alpha, α)

The threshold below which you reject the null hypothesis, typically set at 0.05. In simple terms, it is the false-alarm rate you are willing to tolerate.

Type I error (false positive)

Rejecting a null hypothesis that is true. You conclude there is an effect when there is none.

Type II error (false negative)

Failing to reject a null hypothesis that is false. You miss a real effect.

Statistical power

The probability of correctly rejecting a false null hypothesis (1 minus the type II error rate). Higher power means you are less likely to miss a real effect.

Chi-square test

A test for assessing whether observed categorical frequencies differ from expected frequencies. Used when comparing proportions across groups. Think of it as the go-to test for "is the distribution of outcomes different between these groups?"

Paired t-test

A test comparing the means of two related measurements (e.g., the same metric measured before and after a change). Used when the same subjects are measured twice.

Mann-Whitney U test

A non-parametric alternative to the t-test for comparing two independent groups when the data may not be normally distributed.

Wilcoxon signed-rank test

A non-parametric alternative to the paired t-test. Useful when paired differences are not normally distributed.

Imputation

The process of replacing missing data with substituted values. Methods range from simple (mean, median) to complex (predictive models).


Core Content

Comparing Two Groups – Parcel Damage Rates (Q9)

Scenario: You have two parcel types (A and B) with damage probabilities p = 0.4 and q = 0.6, from 200 shipments split evenly.

Choosing the right test:

  • The data is categorical (damaged vs. not damaged) across two independent groups.

  • A chi-square test of independence or Fisher's exact test is appropriate here.

Setting up the hypotheses:

  • H₀: There is no difference in the probability of package damage between parcel types A and B.

  • H₁: There is a difference in the probability of package damage between parcel types A and B.

Working through it:

  • 100 shipments per parcel type.

  • Parcel A: 40 damaged, 60 undamaged.

  • Parcel B: 60 damaged, 40 undamaged.

  • Compute the chi-square statistic from these observed vs. expected counts.

  • With these numbers, the difference (0.4 vs. 0.6) on 100 observations per group is likely to produce a p-value below 0.05, leading you to reject H₀.

Conclusion: There is a statistically significant difference, and parcel A (lower damage rate) is the better choice.

Month-over-Month Significance in Time Series (Q10)

Scenario: You have monthly data for five years and need to determine whether the difference between the current month and the previous month is significant.

Choosing the right test:

  • If comparing the same metric at two consecutive time points, consider a paired t-test (if differences are roughly normal) or a Wilcoxon signed-rank test (if not).

  • For independent samples at two time points, the Mann-Whitney U test is an option.

Process:

  • Calculate the difference between values for each pair of consecutive months.

  • Test whether the mean (or median) difference is significantly different from zero.

  • If the p-value is below your chosen significance level, the month-over-month change is statistically significant.

Practical note: In time series data, autocorrelation can violate independence assumptions. Be prepared to mention this as a caveat and discuss how you would address it (e.g., differencing, using time series-specific tests).

Interpreting Low P-values in User Behaviour (Q13)

Scenario: You find a low p-value linking a user action to an outcome on Coinbase.

Interpretation:

  • A low p-value means the observed association between the user action and the outcome is unlikely to have occurred by chance alone.

  • This does not prove causation. It signals a relationship worth investigating further.

Business application:

  • The finding could inform product recommendations, feature optimisation, or personalisation strategies.

  • Next steps would typically include running a controlled experiment (A/B test) to establish causality before making product changes.

Key caveat: With large datasets (common at Coinbase), even trivially small effects can produce low p-values. Always consider effect size alongside statistical significance.

Type I and Type II Errors (Q14)

Type I error (false positive):

  • You reject H₀ when it is true.

  • Example at Coinbase: concluding that a new feature increases user retention when it does not. The cost is wasted engineering resources and potentially confusing users.

Type II error (false negative):

  • You fail to reject H₀ when it is false.

  • Example at Coinbase: missing a genuinely effective feature improvement. The cost is lost revenue and user satisfaction.

Balancing the two:

  • Lowering α (e.g., from 0.05 to 0.01) reduces type I error risk but increases type II error risk.

  • Increasing sample size increases statistical power, reducing type II error risk without increasing type I error risk.

  • The right balance depends on the business context: if the cost of a false positive is high (e.g., rolling out a bad feature to millions of users), use a stricter α. If missing a real improvement is costly, prioritise power.

Imputation Methods for Missing Data (Q19)

Scenario: Coinbase user data is missing income or investment experience fields.

Common imputation methods:

  • Mean imputation: Replace missing values with the column mean. Simple, but sensitive to outliers and reduces variance in the data.

  • Median imputation: Replace with the median. More robust to outliers than mean imputation.

  • Mode imputation: Replace with the most frequent value. Suitable for categorical data.

  • Predictive imputation: Train a model (e.g., regression, k-nearest neighbours) to predict missing values from other features. Captures complex relationships but is computationally heavier.

  • Multiple imputation: Generate several plausible datasets with different imputed values and combine results. Accounts for the uncertainty introduced by imputation.

Impact on analysis:

  • Mean and median imputation can understate variability and bias standard errors downward.

  • Predictive imputation preserves relationships between variables better but adds modelling complexity.

  • The choice depends on the pattern of missingness (random vs. systematic), the proportion of missing data, and the goals of the analysis.


Formulas / Key Relationships

Chi-square test statistic: χ² = Σ (Observed - Expected)² / Expected

Type I error rate = α (the significance level you set)

Type II error rate = β

Statistical power = 1 - β

Relationship: For a fixed sample size, decreasing α increases β (and decreases power). Increasing sample size can decrease both error rates simultaneously.


Common Misconceptions

  • "A p-value of 0.03 means there is a 3% chance the null hypothesis is true." The p-value is the probability of the data given H₀, not the probability of H₀ given the data. These are different things.

  • "Statistical significance always means practical significance." With enough data, tiny, meaningless effects can be statistically significant. Always report effect sizes.

  • "Mean imputation is harmless because it preserves the mean." It artificially reduces variance and can distort correlations between variables.

  • "Type I and type II errors are equally important." Their relative importance depends entirely on the business context. At Coinbase, deploying a bad feature (type I) and missing a good one (type II) have very different costs.


Why It Matters / Exam Flags

⚠️ Know when to use a chi-square test vs. a t-test vs. a non-parametric alternative. The choice depends on whether your data is categorical or continuous, paired or independent, and whether normality assumptions hold.

⚠️ Be ready to set up null and alternative hypotheses for any scenario Coinbase gives you. State them clearly before jumping into the test.

⚠️ Interviewers will probe whether you understand the difference between statistical significance and practical significance, particularly for large-scale user data.

⚠️ Imputation questions test whether you understand the downstream consequences of your choice, not just the mechanics of filling in blanks.


Quick Self-Test

  1. True or False: A chi-square test is appropriate for comparing the means of two continuous variables.

  1. Fill in the blank: A type I error is also known as a _______ _______.

  1. True or False: Increasing your significance level (alpha) from 0.01 to 0.05 reduces the risk of a type II error.

  1. Fill in the blank: Mean imputation preserves the sample mean but artificially reduces the sample _______.

  1. True or False: A low p-value proves that the observed effect is large and practically meaningful.

Answers: 1. False (chi-square is for categorical data). 2. False positive. 3. True (a more lenient threshold makes it easier to reject H₀, reducing type II error). 4. Variance. 5. False (it shows statistical significance, not practical significance).


Practice Q&A

Q: You observe p = 0.4 and q = 0.6 for parcel damage rates from 200 total shipments (100 each). What test would you use, and what would you expect to conclude?

A: Use a chi-square test of independence (or Fisher's exact test). Set H₀ as "no difference in damage probability between parcels." With 40/100 vs. 60/100 damaged, the difference is substantial on this sample size, so you would likely reject H₀ at α = 0.05 and conclude parcel A has a significantly lower damage rate.

Q: How would you determine whether a month-over-month change in a time series metric is significant?

A: Calculate the difference between consecutive months, then use a paired t-test (if differences are approximately normal) or a Wilcoxon signed-rank test (if not). If the p-value falls below your chosen α, the change is statistically significant. Mention autocorrelation as a potential complication in time series data.

Q: A low p-value suggests a relationship between a Coinbase user action and increased retention. What should you do next?

A: Recognise that the low p-value indicates statistical significance but not necessarily causation. Examine the effect size to assess practical significance. If promising, design a controlled A/B test to establish whether the relationship is causal before making product changes.

Q: You are designing an analysis at Coinbase and need to balance type I and type II errors. What factors would you consider?

A: Consider the business cost of each error type. If a false positive means deploying a harmful feature, use a stricter α. If a false negative means missing a revenue-boosting improvement, prioritise statistical power by increasing sample size. The right balance is driven by the specific decision at stake.


Connections to Other Topics

The hypothesis testing concepts here underpin the A/B testing and model evaluation questions in Part 5 (Machine Learning). The imputation methods connect to the data-cleaning behavioural question in Part 2. SQL queries in Part 4 often feed the datasets that these statistical tests are run on, so understanding what the data looks like before it reaches your test is essential.


Related Terms / Search Tags

hypothesis testing, chi-square test, Fisher's exact test, paired t-test, Mann-Whitney U test, Wilcoxon signed-rank test, p-value interpretation, type I error, type II error, false positive, false negative, statistical power, significance level, alpha, imputation, mean imputation, median imputation, predictive imputation, multiple imputation, missing data, time series analysis, Coinbase statistics interview, data scientist statistics