Difficulty: Intermediate to Advanced | Prerequisites: Basic ML concepts (supervised learning, classification, regression), Parts 1 and 3 (Process and Statistics) Date: 2025 | Source: Interview Query – Coinbase Data Science Interview Guide
Tags: Coinbase, machine learning, bias-variance tradeoff, recommendation engine, collaborative filtering, content-based filtering, model evaluation, MAE, RMSE, precision, recall, F1-score, AUC-ROC, logistic function, softmax, PCA, dimensionality reduction, class imbalance, overfitting, underfitting, churn prediction, data scientist interview
Machine learning questions at Coinbase are applied, not theoretical. The interviewers want to see that you can choose the right algorithm for a business problem, evaluate models with appropriate metrics, and handle real-world complications like class imbalance and high-dimensional data. Coinbase's ML use cases include recommendation systems (for cryptocurrencies, features, and content), user churn prediction, fraud detection, delivery time estimation, and portfolio analysis. Questions in this area typically span the full lifecycle: problem framing, feature engineering, model selection, evaluation, and iteration.
Coinbase ML questions cover recommendation engines (collaborative and content-based filtering), the bias-variance tradeoff, model comparison using MAE and RMSE, precision vs. recall for churn prediction, class imbalance handling, PCA for dimensionality reduction, and the logistic and softmax functions. Every answer should connect the technical concept to a specific business outcome.
Bias-variance tradeoff
The fundamental tension between model simplicity and flexibility. In simple terms, a model that is too simple misses real patterns (high bias, underfitting), while a model that is too complex memorises noise (high variance, overfitting). The goal is to find the sweet spot.
Underfitting
When a model is too simple to capture the underlying structure of the data. Training and test performance are both poor.
Overfitting
When a model fits the training data too closely, including its noise. Training performance is strong, but test performance degrades.
Collaborative filtering
A recommendation approach based on user behaviour patterns. It recommends items liked by similar users. Think of it as "users who traded X also traded Y."
Content-based filtering
A recommendation approach based on item attributes. It recommends items similar to those the user has already interacted with. Think of it as "you traded a DeFi token, here are other DeFi tokens."
Mean Absolute Error (MAE)
The average of the absolute differences between predicted and actual values. Treats all errors equally regardless of size.
Root Mean Squared Error (RMSE)
The square root of the average of squared differences between predicted and actual values. Penalises large errors more heavily than MAE.
Precision
Of all the items the model predicted as positive, the proportion that are truly positive. In simple terms, "when the model says yes, how often is it right?"
Recall (sensitivity)
Of all the items that are truly positive, the proportion the model correctly identified. In simple terms, "of everything that should have been flagged, how much did the model catch?"
F1-score
The harmonic mean of precision and recall. Useful when you need a single metric that balances both, especially with imbalanced classes.
AUC-ROC
The area under the receiver operating characteristic curve. Measures a model's ability to distinguish between classes across all probability thresholds. Robust to class imbalance.
Class imbalance
When one class (e.g., "user churns") is far less frequent than another ("user stays"). Standard accuracy becomes misleading because a model that always predicts the majority class can score highly while being useless.
Principal Component Analysis (PCA)
A dimensionality reduction technique that transforms features into a smaller set of uncorrelated components, each capturing as much variance as possible. Think of it as compressing a dataset's information into fewer, more informative dimensions.
Logistic function (sigmoid)
A function that maps any real-valued number to a value between 0 and 1. Used in binary classification to convert model outputs into probabilities.
Softmax function
A generalisation of the logistic function for multi-class problems. It converts a vector of raw scores into a probability distribution across multiple classes that sum to 1.
Scenario: Build a job recommendation feed using LinkedIn profiles, job application history, and job search questionnaire answers.
Approach:
Data preprocessing: Extract features from profiles (skills, experience, industry), application history (which jobs were applied to, accepted, rejected), and questionnaire answers (preferences, salary range, location).
Feature engineering: Create user profile vectors and job representation vectors that encode the relevant attributes.
Algorithm selection:
Collaborative filtering: recommend jobs that similar users have applied to or been hired for.
Content-based filtering: recommend jobs whose attributes match the user's profile and stated preferences.
Hybrid approach: combine both for better coverage and personalisation.
Evaluation metrics: Precision, recall, and user engagement (click-through rate, application rate).
Iteration: Continuous monitoring and feedback loops. Track whether recommended jobs lead to applications and positive outcomes. Retrain periodically.
At Coinbase, this translates to: recommending cryptocurrencies, features, or educational content based on user trading history, portfolio composition, and stated preferences.
Scenario: Determine whether a new delivery time model outperforms the old one.
Process:
Define evaluation metrics: MAE or RMSE. Choose RMSE if large prediction errors are particularly costly.
Split data into training and testing sets.
Train both models on the same training data. Evaluate on the same test data.
Compare metrics directly.
Run a statistical test (paired t-test or Wilcoxon signed-rank test) on the per-observation errors to determine whether the improvement is statistically significant, not just numerically different.
Why the statistical test matters: A model might have a slightly lower RMSE by chance. The statistical test tells you whether the difference is reliable. This connects directly to the hypothesis testing concepts in Part 3.
Scenario: Building a loan approval model with features including age, occupation, zip code, height, number of children, and favourite colour.
High bias (underfitting):
A simple model (e.g., logistic regression with few features) may fail to capture the complex relationships between applicant features and loan outcomes.
Training error and test error are both high.
High variance (overfitting):
A complex model (e.g., a deep decision tree with no pruning) may memorise the training data, including irrelevant features like favourite colour and height.
Training error is low, but test error is high.
Finding the balance:
Tune hyperparameters (tree depth, regularisation strength).
Select appropriate features. Irrelevant features (favourite colour, height for loan approval) add noise and increase variance.
Use cross-validation to estimate out-of-sample performance.
Apply regularisation techniques (L1/L2) to penalise model complexity.
Interview tip: This question is partly about feature selection. Mentioning that favourite colour and height are unlikely to be predictive (and could introduce noise or bias) shows practical judgement.
Scenario: Recommend cryptocurrencies to Coinbase users using supervised learning.
Algorithm choices:
Decision trees, random forests, or gradient boosting models (e.g., XGBoost, LightGBM) handle complex, non-linear relationships well.
These algorithms can also provide feature importance, which is valuable for understanding what drives recommendations.
Class imbalance challenge:
Some cryptocurrencies are traded far more frequently than others. A model trained on raw data may overwhelmingly recommend popular coins and ignore niche ones.
Solutions for class imbalance:
Oversampling the minority class (e.g., SMOTE): creates synthetic examples of underrepresented cryptocurrencies.
Undersampling the majority class: reduces the number of examples from overrepresented coins.
Class weights / cost-sensitive learning: assigns higher misclassification costs to the minority class during training.
Evaluation metrics: Use AUC-ROC or F1-score instead of accuracy. Accuracy is misleading when classes are heavily imbalanced.
Scenario: Evaluating a model that predicts user churn on Coinbase.
Why accuracy alone is insufficient:
If only 5% of users churn, a model that always predicts "no churn" achieves 95% accuracy while being completely useless for its intended purpose.
Precision in this context:
Of all users the model flags as likely churners, what proportion will truly churn?
High precision means fewer false alarms and more efficient use of retention resources.
Low precision means wasting effort on users who were never going to leave.
Recall in this context:
Of all users who will truly churn, what proportion does the model catch?
High recall means fewer churners slip through undetected.
Low recall means losing users (and revenue) that could have been retained.
Balancing the two:
If retention interventions are cheap (e.g., a personalised email), prioritise recall to catch as many churners as possible.
If interventions are expensive or intrusive (e.g., large discounts), prioritise precision to avoid wasting resources on non-churners.
The F1-score provides a single metric that balances both.
Scenario: Analysing user portfolios with a high number of cryptocurrency features.
Process:
Standardise the features (zero mean, unit variance). PCA is sensitive to scale.
Apply PCA to identify orthogonal principal components that capture the maximum variance.
Select the number of components based on the proportion of variance explained (e.g., keep enough components to explain 90 or 95% of total variance).
Benefits:
Reduces multicollinearity between correlated features (e.g., Bitcoin and Ethereum prices often move together).
Lowers computational cost for downstream models.
Improves model interpretability by collapsing many features into fewer meaningful dimensions.
Interpretation:
The principal components themselves can reveal underlying portfolio structures, such as "DeFi-heavy portfolios" vs. "stablecoin-heavy portfolios," which inform product recommendations.
Problem: Determine whether two rectangles overlap given their corner coordinates.
Approach:
Find the bounding box of each rectangle by identifying the minimum and maximum x and y values from its four corner points.
Two rectangles do not overlap if one is entirely to the left, right, above, or below the other.
They overlap if and only if they overlap in both the x dimension and the y dimension simultaneously.
Edge cases to mention:
Unordered corner points (hence finding min/max rather than assuming order).
Rectangles that share an edge or a single corner (whether you count these as "overlapping" depends on the problem definition; ask the interviewer).
Logistic (sigmoid) function:
Maps any real number to (0, 1).
Used in binary classification: converts a model's raw output into a probability of belonging to the positive class.
Formula: σ(z) = 1 / (1 + e^(-z))
Softmax function:
Generalises the logistic function to multiple classes.
Takes a vector of raw scores and outputs a probability distribution where all values are in (0, 1) and sum to 1.
Used in multi-class classification.
Why they matter for logistic regression:
They turn raw model outputs (which can be any number) into interpretable probabilities.
Logistic regression uses the sigmoid for binary outcomes; multinomial logistic regression uses softmax for multi-class outcomes.
MAE: (1/n) × Σ |actual - predicted|
RMSE: √((1/n) × Σ (actual - predicted)²)
Precision: True Positives / (True Positives + False Positives)
Recall: True Positives / (True Positives + False Negatives)
F1-score: 2 × (Precision × Recall) / (Precision + Recall)
Sigmoid: σ(z) = 1 / (1 + e^(-z))
"High accuracy means the model is good." With imbalanced classes, accuracy is misleading. A churn model with 95% accuracy might simply be predicting "no churn" every time.
"PCA always improves model performance." PCA can help by reducing noise and multicollinearity, but if the discarded components contain meaningful signal, performance can drop. Always check.
"The bias-variance tradeoff is purely theoretical." It has direct practical implications. Choosing between a simple logistic regression and a gradient boosting model for loan approval is a bias-variance decision.
"Collaborative filtering and content-based filtering are interchangeable." They have different strengths. Collaborative filtering struggles with new items (cold start); content-based filtering struggles with serendipity (recommending only similar items).
⚠️ The bias-variance tradeoff question is one of the most commonly asked at Coinbase. Be ready to explain it with a concrete example, not just the textbook definition.
⚠️ For churn prediction, interviewers expect you to discuss precision and recall, not just accuracy. Know which to prioritise based on the business context.
⚠️ Class imbalance solutions (SMOTE, class weights, appropriate metrics) come up whenever the scenario involves rare events: fraud, churn, or niche cryptocurrency recommendations.
⚠️ Know the difference between MAE and RMSE and when to prefer each. RMSE penalises large errors more heavily, so it is preferred when big misses are especially costly.
⚠️ The logistic vs. softmax question tests foundational understanding. Be concise: sigmoid for binary, softmax for multi-class, both convert raw outputs to probabilities.
True or False: High bias in a model leads to overfitting.
Fill in the blank: Collaborative filtering recommends items based on the behaviour of _______ users.
True or False: RMSE penalises large errors more heavily than MAE.
Fill in the blank: The softmax function is a generalisation of the _______ function for multi-class classification.
True or False: PCA should be applied to unstandardised data.
Answers: 1. False (high bias leads to underfitting; high variance leads to overfitting). 2. Similar. 3. True. 4. Logistic (sigmoid). 5. False (standardise first so that features on different scales do not dominate the principal components).
Q: How would you determine whether a new delivery time model is better than the old one?
A: Define an evaluation metric (MAE or RMSE). Train both models on the same training set and evaluate on the same test set. Compare metric values and run a paired t-test or Wilcoxon signed-rank test on the per-observation errors to confirm the improvement is statistically significant.
Q: Explain the bias-variance tradeoff using the loan approval example.
A: A simple model (few features, low complexity) may underfit by missing real patterns between applicant characteristics and loan outcomes (high bias). A complex model with all features (including irrelevant ones like favourite colour) may overfit by memorising training noise (high variance). The optimal model balances complexity through feature selection, regularisation, and hyperparameter tuning, validated by cross-validation performance.
Q: Why is accuracy insufficient for evaluating a churn prediction model?
A: If only 5% of users churn, a model predicting "no churn" for everyone achieves 95% accuracy. Precision and recall are more informative: precision tells you how reliable the churn flags are, and recall tells you what proportion of true churners you catch. The right balance depends on the cost of false positives vs. false negatives.
Q: How would you handle class imbalance in a cryptocurrency recommendation system?
A: Options include oversampling the minority class (SMOTE), undersampling the majority class, using class weights or cost-sensitive learning in the algorithm, and evaluating with AUC-ROC or F1-score instead of accuracy.
Q: What is the difference between the logistic function and the softmax function?
A: The logistic (sigmoid) function maps a single value to (0, 1) and is used for binary classification. The softmax function maps a vector of values to a probability distribution that sums to 1 and is used for multi-class classification. Both convert raw model outputs into interpretable probabilities.
Model evaluation connects directly to the hypothesis testing in Part 3: determining whether one model is significantly better than another requires the same statistical tests (t-test, Wilcoxon) covered there. The feature engineering steps for recommendation engines rely on the SQL querying skills from Part 4. The churn prediction and recommendation scenarios draw on the behavioural questions from Part 2, where you may be asked to describe how you communicated model results to stakeholders.
machine learning interview, bias-variance tradeoff, underfitting, overfitting, recommendation engine, collaborative filtering, content-based filtering, hybrid recommendation, MAE, RMSE, model evaluation, precision, recall, F1-score, AUC-ROC, class imbalance, SMOTE, oversampling, undersampling, cost-sensitive learning, churn prediction, PCA, principal component analysis, dimensionality reduction, logistic function, sigmoid, softmax, logistic regression, multi-class classification, Coinbase ML interview, cryptocurrency recommendation, user portfolio analysis, feature engineering, cross-validation, regularisation, hyperparameter tuning