Following up on my previous SQL post:
This was the Python screen in the same Mixpanel DS phone-interview round, about 45 minutes.
There were four questions: two hands-on pandas tasks and two open-ended questions.
- Aggregate the users and events dataframes separately.
- Merge the two dataframes and check for duplicates. An easy one.
- Open-ended: Which features might predict retention? Explore the data yourself and give an answer.
- Follow-up: How do you convince someone these features really can predict retention?
My approach:
- Start with basic features: first-week event count, active days, plan_type, signup day of the week, and so on.
- After merging, use duplicated() to confirm user_id is unique.
- Modeling: train/test split, then inspect feature importance with a random forest.
- Build confidence: cross-validation, ablation (remove one feature and see how much the metric drops), and checks for data leakage. Features can only use information from before the retention window.
Basic code:
import pandas as pd
# 1. Aggregate per user
user_agg = events.groupby('user_id').agg(
event_cnt=('event_ts', 'count'),
active_days=('event_ts', lambda s: s.dt.normalize().nunique()),
first_event=('event_ts', 'min'),
).reset_index()
# 2. Merge and check duplicates
df = users.merge(user_agg, on='user_id', how='left')
assert df.duplicated(subset='user_id').sum() == 0, "user_id is not unique after the merge!"
df = df.drop_duplicates(subset='user_id')
# 3. Retention label: an event within 7 days after signup counts as retained
df['retained'] = (df['first_event'] < df['signup_date'] + pd.Timedelta(days=7)).astype(int)
# 4. Simple modeling to inspect feature importance
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
X = df[['event_cnt', 'active_days', 'plan_type']]
X = pd.get_dummies(X, columns=['plan_type'])
y = df['retained']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
clf = RandomForestClassifier(random_state=42).fit(X_train, y_train)
print("test accuracy:", clf.score(X_test, y_test))
print(pd.Series(clf.feature_importances_, index=X.columns).sort_values(ascending=False))
This round went well too. I've already scheduled the onsite and will come back with an update after it.
Discussion
Loading comments…