Diagnose a metric drop in search time
Company: Google
Role: Data Scientist
Category: Analytics & Experimentation
Difficulty: medium
Interview Round: Onsite
Over the last 3 calendar months, the metric 'searching time per user per session' dropped by 35%. A teammate proposes modeling two distributions: T1 = time before first successful result and T2 = time before giving up. Critique this and design a robust analysis to find root cause.
- Precisely define the metric and population; handle multi‑tab sessions, background inactivity, and timeouts. Specify inclusion/exclusion rules.
- Explain biases from splitting into T1 and T2: selection bias (excluding no‑success sessions), right‑censoring, competing risks (success vs abandonment), and left‑truncation. How would you detect and correct them?
- Choose methods (e.g., survival/hazard models with censoring; mixture models) and show how to estimate and compare hazards across cohorts (device, locale, query type).
- Rule out non‑product causes: instrumentation changes, seasonality, traffic mix shifts, bot filtering, release flags. List concrete checks and guardrail metrics.
- Build a time‑series decomposition and change‑point analysis; specify covariates and counterfactual baselines.
- Propose a minimal experiment or holdout (e.g., rollback of ranking feature) with success criteria and expected directional outcomes.
Overview: This question evaluates a data scientist's competency in precise metric and population definition, time‑to‑event and survival/hazard modeling, time‑series decomposition and change‑point analysis, instrumentation and traffic forensics, and experimental design for root‑cause identification.
Community answers
Answer by SS
🧠 1. First: Critique the T1 / T2 idea
Your teammate suggests:
T1: time until first success
T2: time until user gives up
❌ Problem with this approach
It sounds intuitive, but it’s biased:
Selection bias:
T1 only uses sessions where users succeeded → ignores failures
Right-censoring:
Some sessions are still ongoing → we don’t observe final outcome
Competing risks:
User can either succeed OR abandon → not independent
Left-truncation:
If tracking starts late (e.g., page reload), we miss early behavior
👉 Conclusion:
T1/T2 split loses information and gives misleading results.
📏 2. Define the metric properly (very important)
Metric:
“Searching time per user per session = time from first query to either success or exit”
Define:
Start: first search query
End:
success (click meaningful result) OR
abandonment (exit / inactivity)
Handle edge cases:
Multi-tab sessions:
Treat each tab independently OR merge if within same session window
Background inactivity:
Remove idle time (e.g., no activity > 30 sec)
Timeouts:
Cap sessions (e.g., max 30 mins)
Inclusion rules:
Include sessions with ≥1 search
Exclude:
bot traffic
broken/incomplete logs
extremely long idle sessions
⚠️ 3. What could cause the 35% drop?
Before modeling, rule out non-product causes:
🔍 Data / logging checks
Tracking definition changed?
Event missing or delayed?
Timeout logic changed?
📊 Traffic mix changes
More mobile users?
More easy queries?
Different geographies?
📅 Seasonality
Holidays → simpler queries?
🤖 Bots / fil