Mixpanel Data Scientist Interview Experience — Pandas Tasks and Open-Ended Retention Prediction

Mixpanel·Data Scientist·Sep 2026
Technical ScreenIn progressmedium

Following up on my previous SQL post:
This was the Python screen in the same Mixpanel DS phone-interview round, about 45 minutes.
There were four questions: two hands-on pandas tasks and two open-ended questions.

  1. Aggregate the users and events dataframes separately.
  2. Merge the two dataframes and check for duplicates. An easy one.
  3. Open-ended: Which features might predict retention? Explore the data yourself and give an answer.
  4. Follow-up: How do you convince someone these features really can predict retention?

My approach:

  • Start with basic features: first-week event count, active days, plan_type, signup day of the week, and so on.
  • After merging, use duplicated() to confirm user_id is unique.
  • Modeling: train/test split, then inspect feature importance with a random forest.
  • Build confidence: cross-validation, ablation (remove one feature and see how much the metric drops), and checks for data leakage. Features can only use information from before the retention window.

Basic code:

import pandas as pd
# 1. Aggregate per user
user_agg = events.groupby('user_id').agg(
event_cnt=('event_ts', 'count'),
active_days=('event_ts', lambda s: s.dt.normalize().nunique()),
first_event=('event_ts', 'min'),
).reset_index()
# 2. Merge and check duplicates
df = users.merge(user_agg, on='user_id', how='left')
assert df.duplicated(subset='user_id').sum() == 0, "user_id is not unique after the merge!"
df = df.drop_duplicates(subset='user_id')
# 3. Retention label: an event within 7 days after signup counts as retained
df['retained'] = (df['first_event'] < df['signup_date'] + pd.Timedelta(days=7)).astype(int)
# 4. Simple modeling to inspect feature importance
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
X = df[['event_cnt', 'active_days', 'plan_type']]
X = pd.get_dummies(X, columns=['plan_type'])
y = df['retained']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
clf = RandomForestClassifier(random_state=42).fit(X_train, y_train)
print("test accuracy:", clf.score(X_test, y_test))
print(pd.Series(clf.feature_importances_, index=X.columns).sort_values(ascending=False))

This round went well too. I've already scheduled the onsite and will come back with an update after it.

Published

Curated and edited by PracHub

Practice the questions from this interview

Discussion

Sign in to join the discussion. The author is notified of every comment.

Loading comments…

Interview at a glance

Company
Mixpanel
Role
Data Scientist
Rounds
Technical Screen
Outcome
In progress
Difficulty
medium
Interview date
Sep 2026
Questions from this interview
2 questions

Real Mixpanel interview experiences

First-hand reports from Mixpanel candidates — the rounds, the questions they were asked, and how it went.

All 6 Mixpanel interview experiences