Evaluate a New Version in Simulation and Explain a Flat Live Result
Company: Waymo
Role: Data Scientist
Category: Analytics & Experimentation
Difficulty: medium
Interview Round: Technical Screen
# Evaluate a New Version in Simulation and Explain a Flat Live Result
A new version performs better in offline simulation but shows no detectable improvement after launch. You also have historical scenarios that can be replayed under both the old and new versions. Explain how you would compare the versions using these scenarios and investigate the disagreement with live results, especially when the outcomes of interest are rare events.
### What a Strong Answer Covers
- A paired scenario comparison with a clear metric and uncertainty at the appropriate independent unit.
- Scenario representativeness, simulation fidelity, and separation of tuning from evaluation.
- A diagnosis of a flat live result that separates low power from a true lack of benefit.
- Exposure, rollout, metric, and distribution checks before making a release conclusion.
```hint Use the same scenario twice
A scenario that is intrinsically difficult can produce correlated outcomes under both versions.
```
### Follow-up Questions
- What if the simulation scenarios deliberately oversample hazardous situations?
- Why is a nonsignificant live comparison not evidence that the versions are equivalent?
Overview: Compare model versions on paired simulation scenarios and diagnose why offline improvements may disappear in live rare-event evaluation.
Evaluate a New Version in Simulation and Explain a Flat Live Result
A new version performs better in offline simulation but shows no detectable improvement after launch. You also have historical scenarios that can be replayed under both the old and new versions. Explain how you would compare the versions using these scenarios and investigate the disagreement with live results, especially when the outcomes of interest are rare events.
What a Strong Answer Covers Guidance
A paired scenario comparison with a clear metric and uncertainty at the appropriate independent unit.
Scenario representativeness, simulation fidelity, and separation of tuning from evaluation.
A diagnosis of a flat live result that separates low power from a true lack of benefit.
Exposure, rollout, metric, and distribution checks before making a release conclusion.
Follow-up Questions Guidance
What if the simulation scenarios deliberately oversample hazardous situations?
Why is a nonsignificant live comparison not evidence that the versions are equivalent?