Diagnose a metric drop in search time

Quick Overview

This question evaluates a data scientist's competency in precise metric and population definition, time‑to‑event and survival/hazard modeling, time‑series decomposition and change‑point analysis, instrumentation and traffic forensics, and experimental design for root‑cause identification.

Diagnose a metric drop in search time

Company: Google

Role: Data Scientist

Category: Analytics & Experimentation

Difficulty: medium

Interview Round: Onsite

Over the last 3 calendar months, the metric 'searching time per user per session' dropped by 35%. A teammate proposes modeling two distributions: T1 = time before first successful result and T2 = time before giving up. Critique this and design a robust analysis to find root cause. - Precisely define the metric and population; handle multi‑tab sessions, background inactivity, and timeouts. Specify inclusion/exclusion rules. - Explain biases from splitting into T1 and T2: selection bias (excluding no‑success sessions), right‑censoring, competing risks (success vs abandonment), and left‑truncation. How would you detect and correct them? - Choose methods (e.g., survival/hazard models with censoring; mixture models) and show how to estimate and compare hazards across cohorts (device, locale, query type). - Rule out non‑product causes: instrumentation changes, seasonality, traffic mix shifts, bot filtering, release flags. List concrete checks and guardrail metrics. - Build a time‑series decomposition and change‑point analysis; specify covariates and counterfactual baselines. - Propose a minimal experiment or holdout (e.g., rollback of ranking feature) with success criteria and expected directional outcomes.

Overview: This question evaluates a data scientist's competency in precise metric and population definition, time‑to‑event and survival/hazard modeling, time‑series decomposition and change‑point analysis, instrumentation and traffic forensics, and experimental design for root‑cause identification.

Community answers

Answer by SS

🧠 1. First: Critique the T1 / T2 idea Your teammate suggests: T1: time until first success T2: time until user gives up ❌ Problem with this approach It sounds intuitive, but it’s biased: Selection bias: T1 only uses sessions where users succeeded → ignores failures Right-censoring: Some sessions are still ongoing → we don’t observe final outcome Competing risks: User can either succeed OR abandon → not independent Left-truncation: If tracking starts late (e.g., page reload), we miss early behavior 👉 Conclusion: T1/T2 split loses information and gives misleading results. 📏 2. Define the metric properly (very important) Metric: “Searching time per user per session = time from first query to either success or exit” Define: Start: first search query End: success (click meaningful result) OR abandonment (exit / inactivity) Handle edge cases: Multi-tab sessions: Treat each tab independently OR merge if within same session window Background inactivity: Remove idle time (e.g., no activity > 30 sec) Timeouts: Cap sessions (e.g., max 30 mins) Inclusion rules: Include sessions with ≥1 search Exclude: bot traffic broken/incomplete logs extremely long idle sessions ⚠️ 3. What could cause the 35% drop? Before modeling, rule out non-product causes: 🔍 Data / logging checks Tracking definition changed? Event missing or delayed? Timeout logic changed? 📊 Traffic mix changes More mobile users? More easy queries? Different geographies? 📅 Seasonality Holidays → simpler queries? 🤖 Bots / fil
|Home/Analytics & Experimentation/Google
Google logo
Google
Oct 13, 2025
mediumData ScientistOnsiteAnalytics & Experimentation
7
0

Over the last 3 calendar months, the metric 'searching time per user per session' dropped by 35%. A teammate proposes modeling two distributions: T1 = time before first successful result and T2 = time before giving up. Critique this and design a robust analysis to find root cause.

  • Precisely define the metric and population; handle multi‑tab sessions, background inactivity, and timeouts. Specify inclusion/exclusion rules.
  • Explain biases from splitting into T1 and T2: selection bias (excluding no‑success sessions), right‑censoring, competing risks (success vs abandonment), and left‑truncation. How would you detect and correct them?
  • Choose methods (e.g., survival/hazard models with censoring; mixture models) and show how to estimate and compare hazards across cohorts (device, locale, query type).
  • Rule out non‑product causes: instrumentation changes, seasonality, traffic mix shifts, bot filtering, release flags. List concrete checks and guardrail metrics.
  • Build a time‑series decomposition and change‑point analysis; specify covariates and counterfactual baselines.
  • Propose a minimal experiment or holdout (e.g., rollback of ranking feature) with success criteria and expected directional outcomes.
Loading comments...