OpenX · Data Scientist
Updated · 2026-09-24

OpenX Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

A Data Scientist at OpenX works at the absolute frontier of programmatic advertising and high-throughput technology. OpenX operates one of the world's largest independent advertising exchanges, processing billions of transactions daily and generating massive volumes of real-time data. In this role, you are not just analyzing static datasets; you are building the intelligent algorithms that power real-time bidding, yield optimization, ad fraud detection, and traffic shaping.

Allocate prep to your weakest link rather than your favourite topic. Of the three things that usually gate the outcome (SQL that is correct under messy joins, sound reasoning about experiments, and structured framing of an open-ended problem), candidates tend to over-invest in modelling theory and under-invest in framing.

OpenX candidates report 3 rounds · ≈ 3-5 weeks. The stages below are what candidates describe, not a published process.

Reconstruct the funnel from bid request to conversionDefend attribution windows and model choiceSeparate invalid traffic from genuine performance shifts

33 min read

Practice 15 Data Scientist prompts
15Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

A Data Scientist at OpenX works at the absolute frontier of programmatic advertising and high-throughput technology. OpenX operates one of the world's largest independent advertising exchanges, processing billions of transactions daily and generating massive volumes of real-time data. In this role, you are not just analyzing static datasets; you are building the intelligent algorithms that power real-time bidding, yield optimization, ad fraud detection, and traffic shaping.

The impact of your work is immediate and highly visible. A minor optimization in a machine learning model can lead to significant revenue shifts for publishers and advertisers alike. Your primary challenge will be to balance statistical rigor with computational efficiency, ensuring that complex predictive models can execute within milliseconds to keep pace with the lightning-fast programmatic ecosystem.

To succeed as a Data Scientist or Staff Data Scientist at OpenX, you must possess a rare combination of deep mathematical intuition, strong software engineering fundamentals, and a product-oriented mindset. You will collaborate closely with platform engineers, product managers, and business leaders to turn massive-scale data into actionable marketplace intelligence.

01

Recruiter Screen

reported

Data Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.

What to demonstrate

  • Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
  • Whether you name what you have not done instead of stretching to cover every line of the posting
  • Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage

How to prepare

  • Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
  • Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
  • Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
PracHub interview research ↗
02

Technical Conversation

reported

This round decides whether someone can hand you a schema and a question and trust the number that comes back. Correctness under a clock is the bar, not clever syntax. The habit that separates strong from weak answers is checking the grain: after every join, know how many rows you expect and whether the count moved. Most wrong answers in this format are not wrong logic, they are a fan-out from a key that turned out not to be unique, or a filter applied before an aggregate when it belonged after. Say what you expect before you run it.

What to demonstrate

  • Whether your row counts survive each join, and whether you notice on your own when they do not
  • Deliberate handling of rows that fail to match, including whether the question needs an inner join or a left join with the non-matches kept and counted
  • Whether NULLs are treated on purpose, given that a NULL compares equal to nothing and that COUNT of a column skips it
  • Reaching a defensible answer inside the window instead of a refined one after it

How to prepare

  • Take a two-table schema, write a join that fans out on purpose, then fix it by collapsing the many-side to one row per key before joining. Repeat until the fix is reflex rather than recall.
  • Write a funnel as one query and print the distinct user count at each stage, then confirm each stage is a subset of the one above it rather than assuming it
  • Do a few timed runs in a plain text box with no autocomplete and no formatter, since assessment editors often have neither
PracHub interview research ↗
03

Hands-on Coding Assessment

reported

Before anything else, this round is a reading test. You are given a small schema and a question phrased in business language, and most of the difficulty sits in the gap between them. Who counts as an active user, does a refunded order still count as an order, is that date column an event time or a load time. Weak answers start typing immediately and compute something precise about the wrong population. Strong ones pin the definition in one sentence, name the column that encodes it, then write the query. On a timed assessment with nobody to tell, write the definition in a comment anyway.

What to demonstrate

  • Whether an ambiguous term becomes a specific column and filter before any computation happens
  • Whether you read the schema for keys and cardinality rather than only for column names
  • Whether the result answers the question at the grain it was asked at, per user or per session or per day

How to prepare

  • Take three metrics you already use and write down the exact filter and exact grain behind each, then practise stating one of them in a single sentence out loud
  • On a schema you have never seen, spend the first minute writing what one row of each table means and which key it is unique on, then predict which joins can duplicate rows
  • Rehearse a version where the definition changes halfway through, and edit the query you have instead of starting over
PracHub interview research ↗

PracHub editorial advice for the preparation topics above.

01

Randomising users into treatment and control while both arms draw from the same campaign budget.

The arms compete in the same auctions and against the same budget, so treatment winning more impressions directly starves control, and the measured gap includes that cannibalisation rather than only the ad effect. The test looks methodologically clean, the randomisation is genuinely valid, and the lift is still partly manufactured. This is an interference violation, not a randomisation failure, so checking balance on covariates will not catch it. The fixes are to randomise at a unit that contains the budget, such as a geographic market, or to give each arm its own budget and its own pacing, and then be explicit that you are now comparing two separately-budgeted campaigns.

02

Reallocating budget using last-click attribution.

Last-click assigns full credit to whichever touchpoint sits closest to the conversion in time, which structurally favours channels that harvest existing demand — retargeting an already-interested user, or catching a branded search — over channels that create demand in the first place. Optimising against it therefore moves money toward tactics that would have converted many of those users anyway, and the reported cost per action improves at the exact moment true incremental performance gets worse. The signature of this failure is a portfolio where every channel's attributed conversions sum to well above the advertiser's total conversion count. The counter is to treat attributed numbers as a budget-splitting convention and to source the actual reallocation decision from holdout-based incrementality.

03

Sizing estimates built on unnamed, unrevisable assumptions

Write each assumption as a named number you can change, then show the arithmetic so the interviewer can challenge one input instead of the whole answer. Finish by saying which assumption the result is most sensitive to, which matters more than the point estimate.

04

Generalising beyond the population the sample actually supports

State the frame the sample was drawn from and where it diverges from the population you want to talk about: time window, platform, geography, opt-in. If a group is excluded from the frame, either weight to known margins or narrow the claim rather than quietly extending it.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

12 technical prompts3 include a worked solution

Walk me through the bias-variance tradeoff. How does increasing model …

medium
statistics and probability

Walk me through the bias-variance tradeoff. How does increasing model complexity affect both bias and variance?

Approach
  1. Quantify uncertainty explicitly rather than reporting a point estimate alone.
  2. Write down the assumption the method needs before you use the method.
  3. Say what the estimate is of, and over what population it generalises.
Follow-up
  • How would you explain this result to someone who does not know statistics?
  • What sample size would you need to detect an effect half this size?

Describe a time when you had to explain a highly complex machine learn…

medium
machine learning and modelling

Describe a time when you had to explain a highly complex machine learning model to a non-technical product manager. How did you structure your communication?

Approach
  1. Pick an evaluation metric that matches the cost of each error type, not a default.
  2. Frame the prediction: the label, the moment of prediction, and the action it triggers.
  3. Set a baseline first, so any model has something honest to beat.
Follow-up
  • Where could label leakage enter this setup?
  • How would you choose the decision threshold, and who owns that choice?

How do you handle highly imbalanced datasets when training a binary cl…

medium
machine learning and modelling

How do you handle highly imbalanced datasets when training a binary classification model for ad click prediction?

Approach
  1. Say how the offline result would be validated online before it is trusted.
  2. Frame the prediction: the label, the moment of prediction, and the action it triggers.
  3. Pick an evaluation metric that matches the cost of each error type, not a default.
Follow-up
  • What would you monitor after launch to know the model is still valid?
  • How would you choose the decision threshold, and who owns that choice?

Sessionise a device event stream with a 30-minute gap

mediumWorked solution
sessionisationevent-streamsalgorithmspandas

events holds about 5 million unsorted rows: device_id (nullable, roughly 18% null), event_ts, event_type in {impression, click, landing_arrival, conversion}, line_item_id, value_usd (nullable). Define a session as a maximal run of events for one device_id in which consecutive events are at most 30 minutes apart. Assign a session_id to every sessionisable row and return a per-session summary with device_id, start_ts, end_ts, n_impressions, n_clicks, converted and total value_usd. Do not use a library sessionisation helper, and do not scan the frame more than once after sorting.

Approach
  1. Separate null device_id rows first and report their share. Null identity is a supply-source property, not random missingness, so grouping them under one null key would fabricate a single session spanning millions of unrelated events; excluding them means the summary describes identified traffic only, and that has to be said out loud.
  2. Sort once by (device_id, event_ts) with a deterministic tiebreak such as an event_type rank then a row id. Without a tiebreak, an impression and its click sharing a timestamp order differently between runs and the summary is not reproducible.
  3. Compute the boundary mask as (device_id != device_id.shift()) | (event_ts - event_ts.shift() > 30min) and take its cumsum as session_id. This is a grouped run-length in two vectorised operations: O(n log n) for the sort and O(n) afterwards.
  4. Create indicator columns (is_impression, is_click, is_conversion) before the groupby, then summarise with one named aggregation pass. Building each count from its own filtered groupby re-scans the frame per metric and violates the single-pass constraint.
  5. Assert the structural invariant rather than eyeballing it: within a device, sessions ordered by start_ts must be disjoint and increasing, every within-session consecutive gap at most 30 minutes, every between-session gap strictly greater.
Worked solution 25 min
  1. Split on device_id.notna(), record the null share, and keep the null rows aside rather than dropping them silently.
  2. Add an event_type rank column, sort by (device_id, event_ts, type_rank, row_id), and reset the index.
  3. boundary = (device_id != device_id.shift()) | (event_ts.diff() > pd.Timedelta('30min')); session_id = boundary.cumsum().
  4. Add is_impression/is_click/is_conversion indicators, then one groupby('session_id').agg giving device_id first, start_ts min, end_ts max, the three sums, converted as is_conversion.max() > 0, and value_usd sum.
  5. Run the invariant assertions per device on the summary before returning it.
EXPECTED RESULTA session_id covering every non-null-device row exactly once, and a summary frame with one row per session. The summed n_impressions, n_clicks and remaining event counts across sessions equal the sessionised row count, and the null-device rows are reported as an excluded count.
Follow-up
  • A conversion arrives three hours after the click and lands in its own session. Is that the right answer, and how does a session boundary differ from an attribution window?
  • One device shows negative gaps from a clock four hours off. What does your boundary mask do with it, and what would you prefer it did?
  • Make the 30 minutes a parameter. How would you choose it from the data rather than from convention?

For someone who can already write the query and train the model but stalls when asked what to measure or whether a change is worth making. Metric definition and case structure come first; the technical work is kept as maintenance rather than the centre of the week.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Metric anatomy
  • For three products you use daily, write one primary metric, two input metrics that plausibly move it, and one guardrail that would catch a cheap way of moving the primary at the cost of the product.
  • For one of them, specify the metric precisely enough that two analysts would return the same number: numerator, denominator, unit of observation, time window, and how returning and deleted accounts are treated.
  • Pick a ratio metric and write what happens to it when the denominator shrinks for reasons unrelated to the numerator, with a concrete example of that happening.

Deliverable: A one-page metric tree for one product, with the primary metric written as an unambiguous spec.

Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗
02Diagnosing a drop without guessing
  • Take the prompt "weekly active users fell 8 percent week over week" and write the segmentation plan before proposing any cause: platform, region, tenure cohort, acquisition channel, and whether the movement sits in the numerator or in a changed denominator.
  • List the instrumentation failures that manufacture fake drops (a client release that stopped firing an event, a bot filter change, a shifted date boundary or timezone) and write the query that rules out each one.
  • Rehearse stating the boring explanations first, seasonality and day-of-week composition, before reaching for a product cause.

Deliverable: A drop-diagnosis checklist short enough to recite from memory in under a minute.

Practice prompt ↗Practice prompt ↗
03Should we build it
  • Take a feature idea and write it as a bet: what you believe is true, what would have to be true for it to pay off, the metric that would confirm it, and the effect size that would justify the engineering cost.
  • Size the opportunity top-down and bottom-up, then reconcile the two numbers in writing instead of quoting whichever is friendlier.
  • Write the counter-metric that would make you kill the feature even if it wins on the primary metric.

Deliverable: A one-page product memo ending in a decision rather than a list of considerations.

Practice prompt ↗Practice prompt ↗
04The places aggregate numbers lie
  • Construct a Simpson's paradox numerically: two segments where the treatment wins within each segment yet loses overall, and identify the shift in segment weights that causes it.
  • Take a heavy right-tailed quantity such as revenue per user and write why the mean is the wrong summary, which percentile you would report instead, and what a moving mean with a stable median tells you.
  • Write your definition of a session for the product from day one, then name two real behaviours it misclassifies.

Deliverable: One page holding a worked Simpson's paradox table and a session definition with its two known failure cases.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05Technical maintenance, aimed at metrics
  • Solve four timed SQL prompts that all end in a ratio metric, so the question of grain stays live in every answer.
  • Compute a 95 percent confidence interval for a proportion on a small sample, and state why the normal approximation is unreliable when either np or n(1 minus p) falls below roughly 10, along with which interval you would use instead.
  • Take one metric from your day-one tree, write the query that computes it correctly, then write the query that computes it wrong in the most plausible way and explain how you would notice.

Deliverable: Four solved prompts plus a matched correct and plausible-wrong query for one metric.

Practice prompt ↗Practice prompt ↗
06Turning engineering work into data science stories
  • Write three project stories as situation, decision, trade-off, outcome, each carrying one number and one thing you got wrong.
  • For the story you will lead with, prepare an answer to "what would you do differently" that names a decision you made, not a constraint you were handed.
  • Practise the sentence that reframes a systems project as a question project: the question the work answered, ahead of the pipeline it shipped.

Deliverable: Three written stories with the lead story delivered aloud and timed under four minutes.

Practice prompt ↗Practice prompt ↗
07Mock case and gap list
  • Run a 40-minute mock case with someone playing a product manager who pushes back on your metric choice, and record it.
  • Listen back and mark every moment you proposed a solution before the success metric existed.
  • Rewrite those moments as the question you should have asked, and rehearse the first 90 seconds of the case until scoping comes before solving.

Deliverable: A recorded case plus a rewritten opening 90 seconds.

Practice prompt ↗Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

Half of this section is about translation. Be ready to describe how you explained a result to someone who did not want the method, only the implication, and what you did when the simplified version started being repeated in a way that overstated it. Correcting your own simplification is a strong beat.

Tell me about a project where the data was extremely messy or incomple…

medium
behavioural and stakeholder questions

Tell me about a project where the data was extremely messy or incomplete. How did you unblock yourself to deliver results?

Approach
  1. State the situation in two sentences and spend the rest on your reasoning.
  2. Quantify the outcome, including what you would not claim credit for.
  3. Pick a story where you drove the decision, not one where you observed it.
Follow-up
  • What did you decide not to do, and why?
  • How did you know the outcome was caused by your change?

Defending a null incrementality result against a headline ROAS

hard
influenceincrementalityattributionconflict

Your geo holdout shows a retargeting line item's incremental conversions per $1,000 of spend are indistinguishable from zero. The same line item reports last-click ROAS of 8.3 and is the headline number in a renewal deck presented in six days. The account team believes your test is wrong; the advertiser has never questioned the 8.3. Deliverable: the internal case you make, the specific evidence you bring, the limits you concede, and the smallest reversible experiment you propose if your reading is rejected.

Approach
  1. The interviewer is probing whether you can hold an unpopular finding without turning it into a credibility fight, so open by granting that both numbers are computed correctly: last-click measures which touchpoint was nearest in time, the holdout measures what changed because of the spend. Framing attribution as 'wrong' converts a methodological point into a turf argument you will lose on relationship grounds.
  2. State the null as a bound rather than as zero. Report the minimum detectable effect the 20-market design could resolve and say 'we can rule out incremental ROAS above this value'. A null without an MDE is not a finding and is trivially dismissed as an underpowered test.
  3. Bring converging evidence from tables you already have, and say what each piece can and cannot show. Credit sensitivity: recompute the same conversion cohort under a second attribution_model, first-touch or even credit across in-window touchpoints, holding cohort and windows fixed, and report what share of the line item's credited volume survives; that bounds how much of the 8.3 is a property of the crediting rule rather than of the spend, and it is a sensitivity result, not a lift estimate. Harvesting signature: the share of retargeting-credited conversions whose user already had an earlier impression, click or site session inside click_window_days, together with the distribution of conversion_ts minus click_ts, where a heavy mass inside a few minutes is consistent with an ad served into a session that was already converting.
  4. Quantify both error directions before recommending anything: the spend at risk if you are right and ignored, and the renewal revenue at risk if you are wrong and act. This is what separates a defensible position from an ideological one.
  5. Propose the reversible version rather than the shutdown: a 20% budget reduction in half the matched markets for six weeks, powered against the MDE you just computed, with the analysis window extended past the flight by the click window plus ingestion lag.
Follow-up
  • The account lead says telling the advertiser their favourite channel does nothing will cost us the renewal. How do you respond without either caving or escalating?
  • What specific evidence would change your mind about this line item?
  • How do you handle it if the budget-down test also comes back underpowered?

Scoping a one-line request that attribution is wrong

easy
scopingstakeholder-managementattribution

A sales lead forwards a one-line request: 'attribution is wrong for this advertiser, fix it.' You have read access to conversion_event, attribution_credit and ad_click for the account, and thirty minutes with the account manager. You may not contact the advertiser this week. Deliverable: a one-page scope naming the single question you will answer, the questions you are explicitly not answering, the data you need, and the decision the answer changes. Bring the three clarifying questions you would ask the account manager first.

Approach
  1. The interviewer is probing whether you convert a complaint into a comparison before you start work, so begin by writing the complaint as 'number A versus number B' and leave B blank until the account manager names it: attributed conversions against the advertiser's own order table, attributed CPA against last month, or our report against a second vendor's. Each has a different investigation.
  2. Run the two cheap arithmetic checks that resolve a large share of these before any modelling: confirm SUM(credit_fraction) per conversion_id equals 1.0 within one attribution_model and model_version, and measure the duplication rate by counting dedup_key values that appear with more than one distinct source in conversion_event.
  3. Measure identity coverage for the account: share of conversion_event rows carrying a non-null click_tracking_id and share carrying device_id. A low match rate means the channel is under-credited, which is the opposite complaint and needs the opposite fix.
  4. Decide the scope from what the answer will change. If the decision is budget reallocation, the honest scope is an incrementality read, not an attribution repair. If the decision is an invoice dispute, the scope is a reconciliation against the advertiser's order count.
  5. Write the page with an explicit out-of-scope list: model choice, window changes and anything requiring advertiser data you cannot get this week, each with the condition that would bring it back in scope.
Follow-up
  • The account manager says the advertiser just wants the numbers to match. What do you tell them is actually achievable, and why is exact agreement not one of the options?
  • How does your scope change if the discrepancy is 4% rather than 40%?
  • 01

    Tell me about a project where the data was extremely messy or incomplete. How did you unblock yourself to deliver results?

  • 02

    Your geo holdout shows a retargeting line item's incremental conversions per $1,000 of spend are indistinguishable from zero. The same line item reports last-click ROAS of 8.3 and is the headline number in a renewal deck presented in six days. The account team believes your test is wrong; the advertiser has never questioned the 8.3. Deliverable: the internal case you make, the specific evidence you bring, the limits you concede, and the smallest reversible experiment you propose if your reading is rejected.

  • 03

    A sales lead forwards a one-line request: 'attribution is wrong for this advertiser, fix it.' You have read access to conversion_event, attribution_credit and ad_click for the account, and thirty minutes with the account manager. You may not contact the advertiser this week. Deliverable: a one-page scope naming the single question you will answer, the questions you are explicitly not answering, the data you need, and the decision the answer changes. Bring the three clarifying questions you would ask the account manager first.

PracHub interview preparation framework ↗
Is this an official OpenX interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at OpenX. Rounds and questions reflect what candidates have reported, not a process OpenX has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
How much coding should I expect in the Data Scientist interview?

You should expect at least one dedicated coding round using a live sharing platform like codepad. The focus is on writing clean, optimal Python code to solve algorithmic and data manipulation problems, similar to medium-difficulty Leetcode challenges.

PracHub interview research ↗
Does OpenX require prior experience in ad-tech?

While prior ad-tech experience is a significant advantage, it is not strictly required. However, you should take the time to learn the basics of programmatic advertising, real-time bidding, and common ad-tech metrics before your interview.

PracHub interview research ↗
How does the team evaluate communication skills?

Communication is evaluated throughout the process, particularly in the behavioral round and the technical chat with the director. They look for your ability to explain complex machine learning choices simply and your capacity to collaborate across engineering and product teams.

PracHub interview research ↗
What is the work environment and culture like for the data science team?

The culture is highly collaborative, data-driven, and fast-paced. Because the company processes such massive volumes of data, there is a strong emphasis on engineering excellence, continuous learning, and practical problem-solving.

PracHub interview research ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.