As a Data Scientist at eBay, you operate at the intersection of massive-scale e-commerce, cutting-edge machine learning, and strategic product decision-making. Your work directly influences how millions of buyers and sellers connect across more than 190 global markets. Whether you are optimizing search algorithms, scaling interactive live shopping experiences, or securing the marketplace against fraudulent listings, your analyses and models shape the future of digital commerce for enthusiasts worldwide.
This position demands both deep technical execution and strategic business acumen. You will own end-to-end analytical initiatives—from scoping ambiguous product challenges and defining core metric taxonomies to architecting data pipelines and deploying production-grade machine learning models. You will collaborate closely with engineering, product management, and operations teams to drive measurable outcomes in high-impact domains such as buyer engagement, seller success, category expansion, and trust and safety.
Expect an environment where intellectual curiosity and data-driven storytelling are paramount. You will be trusted to cut through ambiguity, synthesize complex datasets into crisp executive narratives, and guide senior leadership on critical roadmap decisions. If you thrive in a fast-paced, high-visibility ecosystem where your insights directly power marketplace growth and community trust, a career at offers extraordinary scope and impact.
Online Assessment
reportedMuch of what gets scored here happens out loud while you type. Nobody can see your reasoning inside a half-written query, so five silent minutes read as being stuck even when they are not. State the plan in plain language first: which tables, what grain you are aggregating to, and the one filter that defines the population. Then write it. The narration doubles as insurance, because a wrong plan gets caught early and cheaply while a wrong query gets caught at the end with no time left to redo it. A timed statistics section, where one exists, is a separate test with its own clock.
What to demonstrate
- Whether the query you write matches the plan you just described
- What you do with a hint, meaning whether the correction gets absorbed or the first approach gets defended
- Whether you can debug your own wrong output by reading the result set and naming which part of the query produced the anomaly
How to prepare
- Solve three problems while screen-sharing into a recording, then watch it back and mark every stretch longer than thirty seconds where you said nothing
- Practise compressing the plan into one sentence before typing, then check afterwards whether the finished query actually matched it
- Time yourself on statistics questions that carry a business reading, such as what a confidence interval does and does not claim, rather than re-reading notes without a clock
Recruiter Screening
reportedData Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.
What to demonstrate
- Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
- Whether you name what you have not done instead of stretching to cover every line of the posting
- Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage
How to prepare
- Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
- Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
- Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
Technical Rounds
reportedBefore anything else, this round is a reading test. You are given a small schema and a question phrased in business language, and most of the difficulty sits in the gap between them. Who counts as an active user, does a refunded order still count as an order, is that date column an event time or a load time. Weak answers start typing immediately and compute something precise about the wrong population. Strong ones pin the definition in one sentence, name the column that encodes it, then write the query. On a timed assessment with nobody to tell, write the definition in a comment anyway.
What to demonstrate
- Whether an ambiguous term becomes a specific column and filter before any computation happens
- Whether you read the schema for keys and cardinality rather than only for column names
- Whether the result answers the question at the grain it was asked at, per user or per session or per day
How to prepare
- Take three metrics you already use and write down the exact filter and exact grain behind each, then practise stating one of them in a single sentence out loud
- On a schema you have never seen, spend the first minute writing what one row of each table means and which key it is unique on, then predict which joins can duplicate rows
- Rehearse a version where the definition changes halfway through, and edit the query you have instead of starting over
Final Loop
reportedWhere a loop ends with a senior leader, that conversation is rarely another skills test. The technical signal already exists by then, so the questions tend to open up: what you would look at first, where a metric you have heard about could mislead, what you would push back on. The decision being made is scope, which in practice means level and how much you would be trusted to own unsupervised. Treating it as a formality is the usual mistake. An open question late in the day is still being scored, and a vague answer reads as someone who has not run anything themselves.
What to demonstrate
- Whether your view of the business has anything specific behind it, given that you are working only from what is public and are expected to say so
- Whether the scope of work you describe owning matches the scope of the role, instead of sitting a level below it
- Whether you can disagree with something concrete and stay useful about it, rather than agreeing with everything said in the room
- Whether your questions are ones only this person could answer, as opposed to ones the recruiter already covered
How to prepare
- Build one view you could defend for two minutes using only public information: what the funnel probably looks like, which metric likely drives decisions, and where that metric could mislead. Being wrong for a stated reason survives this round; having no view does not
- Write down the largest piece of work you have owned from question to decision, who else touched it, and what you decided alone, then check that it reads at the level you are interviewing for
- Prepare one thing you would want changed if you joined and phrase it as a question rather than a verdict, so it opens a conversation instead of closing one
7 candidate reports. Individual accounts describe a particular role and hiring cycle.
eBay Software Engineer Interview Experience — An Onsite With No LeetCode
This was the most unusual interview I have had so far. There was not a single LeetCode question in the onsite. First round I wrote a shopping cart. There was not much algorithm work, and it felt as though the round was testing how I communicated with the interviewer. Second round I was asked basic questions about data structures and complexity, including very basic topics such as lists, hashes, a…
Read full experienceeBay Software Engineer Online Assessment Experience — Four CodeSignal Questions in 70 Minutes
View report detailseBay Software Engineer Interview Experience — AI-Coding Phone Screen, Rejected After Onsite Coding Round
In June, an HR person reached out about an MTS role, and we talked through some behavioral questions and my project experience. Phone screen: AI coding — basically they give you some code and have you spot design pattern issues (like something not being extensible), or point out a bad data type. On CodeSignal, the AI is really strong and basically does the work for you. You just need to talk thro…
Read full experienceeBay Software Engineer Interview Experience — Three Onsite Rounds With Two System Designs
Round 1 Given an array heights, where each element represents the height of a vertical line. Choose two lines to act as the walls of a container. Return the maximum amount of water the container can hold (max area). Given an array of integers temps representing daily temperatures, write a function to calculate, for each day, how many days you'd have to wait until a warmer temperature. The functio…
Read full experienceeBay Intern Data Analyst Interview Experience — Three SQL-Heavy Phone and Video Rounds
View report detailsPracHub editorial advice for the preparation topics above.
Reading incentive impact without a cell-level holdout
A bonus in one hour or one zone pulls provider hours and consumer orders from adjacent hours and zones rather than creating them, so a before-and-after read on the treated cell counts displaced volume as incremental and can show a positive result for a spend that produced nothing. Only a randomised holdout at the same granularity as the incentive, or a comparison against untreated cells that share the demand shock, separates increment from displacement. Always state incremental orders per incentive dollar, never total orders in treated cells.
Denominator drift in per-active-user metrics
Orders per active consumer falls when acquisition succeeds, because new cohorts transact less than tenured ones, so the metric penalises the thing the company is trying to do. A team that optimises it will quietly prefer weaker acquisition. Decompose into cohort size times cohort frequency, or hold the cohort fixed and read frequency by tenure bucket, before drawing any conclusion about engagement.
Interpreting a change before checking data quality and logging
Spend the first pass on row volume by day, null rates, duplicate keys, and whether the step change lands on a release or tracking-migration date. A discontinuity that coincides with a deploy is an instrumentation hypothesis before it is a behavioural one.
Building features from data that postdates the prediction time
Check every feature against the timestamp at which the model would actually score, and drop anything computed from a window that includes or follows the label event. For a forecasting use case, split train and test by time rather than at random, and split by entity when the same entity recurs.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
How would you build and evaluate a classification model to predict use…
How would you build and evaluate a classification model to predict user churn or buyer intent?
Approach
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
- Set a baseline first, so any model has something honest to beat.
Follow-up
- Where could label leakage enter this setup?
- What would you monitor after launch to know the model is still valid?
Sessionise a provider heartbeat stream into supply sessions
You are given pings: provider_id, market_id, event_at_utc, status in ('idle','en_route','engaged','offline'), one row per app heartbeat, nominally every 30 seconds but with gaps. Reconstruct the supply-session table. A session starts at the first non-offline ping and ends at an explicit offline ping, at a market change, or when the gap to the next ping exceeds 10 minutes. Produce session_id, provider_id, market_id, online_at, offline_at, online_seconds, engaged_seconds, en_route_seconds, idle_seconds and end_reason, with the three state components summing to online_seconds exactly in integer seconds.
Approach
- Sort by (provider_id, event_at_utc), then build a new_session boolean: first ping of a provider, previous status equals 'offline', market_id changed, or the gap from the previous ping exceeds 600 seconds. A cumsum over that boolean is the session key, and it removes any need for a per-provider Python loop.
- Attribute duration to intervals, not to pings: each ping owns the seconds until the next ping inside the same session, and the final ping owns a capped 30 seconds. Because the state seconds are the intervals themselves, they sum to online_seconds by construction rather than by a correction step.
- Encode the three terminations distinctly. Gap timeout ends at last_ping + 30s with end_reason 'app_background_timeout'; an explicit offline ping ends at that ping with 'manual_offline'; a session with no terminating event before the data ends is 'session_still_open' with offline_at NaT.
- On a market change, close the old session at its last ping in the old market and open the new one at the first ping in the new market; the seconds in between belong to neither session, and the output should say so rather than quietly padding one side.
- Aggregate with a single groupby on the session key, pivoting the per-interval state into the three second columns, then assert the sum identity and that consecutive sessions for one provider never overlap.
Follow-up
- A provider is engaged on a 40-minute order and the app backgrounds mid-order. What does your 10-minute rule do to that session, and what does it do to utilisation?
- Utilisation divides engaged by online. Which of your three end_reason cases biases it most, and in which direction?
Trailing 30-day prior-order counts without rolling or asof
You are given orders: order_id, consumer_id, completed_at_utc (tz-aware UTC), about two million rows, one row per completed order. For every order, compute how many completed orders the same consumer had in the 30 days before that order, counting the window as [t - 30 days, t) so the order itself and any exact-timestamp twin are excluded. You may not use groupby().rolling, merge_asof, or apply over groups. Return the input frame, in its original row order and index, with one added integer column prior_30d.
Approach
- Sort once by (consumer_id, completed_at_utc) while keeping the original index, and move to NumPy int64 nanoseconds; the whole problem is two searchsorted calls per group, and a Python loop over two million rows is what makes this fail on time rather than on logic.
- Find group boundaries with np.flatnonzero on a consumer_id change mask instead of iterating a groupby object, then slice the timestamp array per block.
- Within a block, prior_30d[i] = searchsorted(ts_block, t_i, 'left') - searchsorted(ts_block, t_i - 30 days, 'left'), which is exactly the half-open window and needs no special case for the first order.
- State the tie rule out loud: side='left' on the upper bound means simultaneous orders do not count each other, which is the defensible choice when the timestamp has second resolution.
- Scatter the result back through the sort permutation so the added column aligns with the caller's frame, and assert the index is unchanged before returning.
Worked solution 30 min
- order_idx = np.lexsort((ts_ns, consumer_ids)); ts = ts_ns[order_idx]; cid = consumer_ids[order_idx].
- starts = np.concatenate(([0], np.flatnonzero(cid[1:] != cid[:-1]) + 1, [len(cid)])).
- For each block, lo = np.searchsorted(block, block - 308640010**9, 'left'); hi = np.searchsorted(block, block, 'left'); out_block = hi - lo.
- Concatenate block results, then invert: result = np.empty(n, int); result[order_idx] = out_sorted.
- Attach as orders['prior_30d'] and assert the frame's index and row order are identical to the input.
Follow-up
- How does the implementation change if the count must be restricted to the same market?
- This column will feed a model scored at request time. What leakage would you check for, and which timestamp defines the cut-off?
Write a SQL query to calculate month-over-month user retention cohorts…
Write a SQL query to calculate month-over-month user retention cohorts for active marketplace buyers.
Approach
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- State the window function and its partition and ordering out loud before writing it.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
Follow-up
- How does the query change if the join becomes one-to-many?
- What breaks if events arrive late or out of order?
This category evaluates your ability to write efficient queries, manip…
This category evaluates your ability to write efficient queries, manipulate dataframes, and extract actionable insights from large-scale relational databases.
Approach
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
Follow-up
- What breaks if events arrive late or out of order?
- How would you verify this result without re-running the same query?
Approved providers who never matched and never went online
dim_user has user_id, side, provider_status, provider_approved_at_utc. fct_request has request_id, provider_id and matched_at_utc, where provider_id is NULL on every request that never matched. fct_supply_session has session_id, provider_id, online_at_utc. List providers whose provider_status is 'approved' and whose provider_approved_at_utc falls in a given calendar month, who have never appeared as fct_request.provider_id and who have no row in fct_supply_session. Return user_id, provider_approved_at_utc and days since approval, oldest first.
Approach
- Use NOT EXISTS for both exclusions, correlated on provider_id. NOT EXISTS is unaffected by NULLs in the subquery and lets the planner stop at the first matching row.
- If you reach for NOT IN, provider_id is nullable, so the subquery result contains at least one NULL, every comparison evaluates to UNKNOWN instead of TRUE, and the query returns zero rows. There is no error and no warning; an empty result set reads as good news.
- The correct alternatives are a LEFT JOIN with WHERE right_key IS NULL, or NOT IN with an explicit WHERE provider_id IS NOT NULL inside the subquery. Both work; NOT EXISTS is the one that survives someone later making a second column nullable.
- Filter the cohort on provider_approved_at_utc within the month and provider_status = 'approved', so suspended and deactivated accounts do not inflate what will be read as an onboarding failure.
- Compute days since approval as a date difference on the approval timestamp, and sanity check the result size against the cohort size before drawing any conclusion.
Worked solution 20 min
- Run SELECT COUNT(*) FROM fct_request WHERE provider_id IS NULL. A non-zero count is the proof that NOT IN is unusable against this column.
- Write the cohort CTE and count it, so the denominator is known before any filtering.
- Add the two NOT EXISTS clauses one at a time and record the row count after each; which one does most of the work is itself the finding.
- Order by provider_approved_at_utc ascending and add the days-since column.
Follow-up
- The list is 40 percent of the approval cohort. Is that an onboarding failure or a data problem, and which single query settles it?
- How would you separate providers who never went online from providers who went online and were never offered anything, given that fct_request only records the provider who actually matched?
These questions test your ability to define success metrics, evaluate …
These questions test your ability to define success metrics, evaluate product features, and diagnose unexpected business drops.
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- State what result would change your recommendation, so the answer is falsifiable.
- Restate the decision this analysis has to support, and who acts on the answer.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- How would you detect that the metric is being gamed rather than genuinely improving?
How would you investigate and diagnose a sudden 15 percent drop in dai…
How would you investigate and diagnose a sudden 15 percent drop in daily active sellers on the platform?
Approach
- Fix the population and the time window before naming any metric.
- State what result would change your recommendation, so the answer is falsifiable.
- Restate the decision this analysis has to support, and who acts on the answer.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- How would you detect that the metric is being gamed rather than genuinely improving?
How will you identify fraudulent listings and bad actors on eBay while…
How will you identify fraudulent listings and bad actors on eBay while minimizing false positives for legitimate sellers?
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- State what result would change your recommendation, so the answer is falsifiable.
- Fix the population and the time window before naming any metric.
Follow-up
- Which segment would you cut first, and what would that rule out?
- What would you do if the primary metric and the guardrail moved in opposite directions?
What framework would you use to evaluate the success of a major redesi…
What framework would you use to evaluate the success of a major redesign of the buyer search results page?
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Fix the population and the time window before naming any metric.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- Which segment would you cut first, and what would that rule out?
Interviewers use these scenarios to assess your understanding of causa…
Interviewers use these scenarios to assess your understanding of causal inference, experimental design, and statistical validity.
Approach
- Say whether units interfere with each other, and switch design if they do.
- Name the guardrails that would stop a launch even on a positive primary result.
- Name the randomisation unit first; it decides the variance and what the test can detect.
Follow-up
- How would you handle interference between treated and control units?
- What would you conclude if the result is positive but the test is underpowered?
Your experiment shows a statistically significant lift in clicks but z…
Your experiment shows a statistically significant lift in clicks but zero change in final purchases. How do you analyze what went wrong?
Approach
- Say whether units interfere with each other, and switch design if they do.
- Name the randomisation unit first; it decides the variance and what the test can detect.
- Name the guardrails that would stop a launch even on a positive primary result.
Follow-up
- What would you do if you could not randomise at all?
- What would you conclude if the result is positive but the test is underpowered?
Choose a randomisation unit for a dispatch radius change
Engineering wants to widen the dispatch radius so a request can be offered to providers farther away. The proposed test randomises consumers 50/50 and compares SLA fill rate from fct_request, using matched_at_utc minus requested_at_utc against the market SLA. Supply in a market is a shared pool. Explain what this design actually estimates, state the direction of the bias and the conditions under which it is worst, and propose a design that estimates the market-level effect. Name the randomisation unit, what it costs you, and the guardrail you would read alongside fill rate.
Approach
- Name the violated assumption precisely. SUTVA requires that one unit outcome not depend on another unit assignment. Here a treated request reaching a farther provider removes that provider from the pool available to a control request in the same market-minute, so control outcomes are a function of the treatment share.
- Sign the bias and locate it. Treatment gains partly at the expense of control, so the measured difference exceeds the market-level effect. The transfer approaches one-for-one as utilisation rises, because a tight pool has no slack, and shrinks toward zero at low utilisation. The design is therefore most wrong in exactly the peak hours the change is meant to fix.
- Point at the effect user-level randomisation cannot see at all. Longer pickups raise en_route_seconds in fct_supply_session, which cuts utilisation and earnings per online hour. At full rollout that can reverse the sign, and no consumer-level comparison contains the information needed to detect it.
- Pick the unit. Market by time block, a switchback, if the matching rule can be toggled per market-minute; market-level clusters if it cannot. Geohash zones look attractive because they give more units, but providers cross zone boundaries inside a single offer cycle, so zone assignment leaks at exactly the margin the change operates on.
- State the cost honestly. Effective sample becomes blocks or markets rather than requests, power falls by orders of magnitude, and between-market variance dominates. That is why hour-of-week stratification and pre-period covariate adjustment stop being optional refinements and become part of the design.
- Attach the guardrail: provider earnings per online hour, (gross_earnings_cents + incentive_earnings_cents) over online_seconds from fct_supply_session. Fill rate can always be bought with longer pickups that raise idle and en-route time and quietly erode the supply base that produced the fill rate.
Worked solution 25 min
- Write the interference mechanism as one sentence about the shared pool, then state the sign of the resulting bias.
- Show the bias is increasing in utilisation by contrasting a slack hour, where a reassigned provider costs no one a match, with a saturated hour, where it costs exactly one.
- Choose market-by-hour switchback as the primary design, with market-level clusters as the fallback when per-minute toggling is impossible.
- Quantify what the choice costs: replace a sample of millions of requests with roughly one thousand blocks, and recompute the MDE on that basis.
- Name the supply guardrail and the decision rule that pairs it with fill rate.
Follow-up
- The team proposes randomising providers instead. Does that remove the interference?
- How would you detect interference empirically inside the consumer-level test that has already run?
- What changes if the wider radius only fires when surge_multiplier exceeds 1.5?
Requests fell four percent on one local day only
Daily requests in one market fell 4% against the previous day and recovered the day after. The dashboard groups fct_request by DATE(requested_at_local), and dim_market.timezone holds the market's IANA zone. Before anyone opens an incident, rule out calendar and clock causes. Deliverable: the ordered checks you run, the arithmetic for the largest mechanical cause you find, and the residual threshold at which you would still escalate.
Approach
- Compare like weekday to like weekday. Marketplace demand has a strong weekly cycle, so a day-over-day comparison that crosses a weekend boundary is not a signal at all. Use the same weekday for the trailing four to eight weeks as the baseline.
- Read the day's length from the zone, not from the row labels. requested_at_local is a naive wall-clock timestamp, so on an autumn fall-back day the repeated hour lands on the same label twice: COUNT(DISTINCT date_trunc('hour', requested_at_local)) returns 24 for a 25-hour day, never 25. That count can only ever catch spring-forward, and even then only in a market busy enough that every surviving hour has at least one request. Compute the length from dim_market.timezone instead - ((d + 1)::timestamp AT TIME ZONE m.timezone) - (d::timestamp AT TIME ZONE m.timezone) returns 23, 24 or 25 hours for local date d, does not depend on volume, and still returns 23 or 25 in the zones whose transition falls at local midnight, where the non-existent instant resolves forward.
- Size the clock effect properly. The shortfall is not one hour in 24, which would be 4.2%, but the share of daily requests that normally falls in the hour that disappeared, read off the trailing same-weekday hourly profile. In most zones the transition lands overnight, so the true mechanical effect is usually well under one percent. A fall-back day runs the other way, and its data-side fingerprint is the duplicated label carrying roughly twice its usual share - which is the only way that case shows up at all if you are counting distinct local hours.
- Check the local calendar for the date: public holidays, school terms, large scheduled events, severe weather. These move the whole day rather than a single hour, which is how you tell them apart from a clock change.
- Escalate only the residual. Subtract the weekday effect and the clock effect, then express what is left in units of the same-weekday standard deviation over the trailing eight occurrences, and give the threshold you are using.
Follow-up
- The same 4% drop lands on the same local date in five markets across three timezones. What do you check first now?
- The product dashboard groups by local date and the finance report by UTC date. What reconciles the two?
Roughly 90 minutes a night on weekdays with one longer weekend block. The plan deliberately cuts scope rather than compressing everything, on the assumption that finishing one thing a night beats half-starting four.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Fix the scope and set a baseline
- Read the role description and write the three things the loop will almost certainly test, then write an explicit not-doing list for everything else and keep it visible all week.
- Take one 20-minute SQL prompt and one 10-minute metric question cold, and write the single sentence that says what blocked each attempt, since that sentence is what decides which two topics get the most evenings.
- Set the week's one rule: one problem finished to completion every night, including the night you only have 40 minutes.
Deliverable: A one-page scope with an explicit not-doing list and two cold attempts, each carrying one sentence on what blocked it.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02One query pattern, written three times
- Choose the single pattern most likely to appear (a cohort retention grid, or a funnel counted by user) and write it three times from a blank file rather than editing the previous attempt.
- On the third attempt, write the grain of every CTE as a comment before writing its body.
- Stop at 90 minutes even if the third version is imperfect, and write the one thing you would fix with another hour.
Deliverable: Three independent versions of the same query plus a note on what changed between them.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Only the statistics you will be asked to defend
- Write, in under 200 words, how you would decide whether a difference between two groups is real: the test, its assumptions, and what you would switch to when an assumption fails.
- Compute a 95 percent confidence interval for a difference in proportions by hand on realistic numbers, then write in one sentence what changes if the two samples are paired rather than independent.
- Write your answer to "what does a p-value mean", check it against a definition, and delete the version that describes it as the probability the hypothesis is true.
Deliverable: A 200-word written answer and one hand-computed interval you can reproduce under pressure.
Practice prompt ↗Practice prompt ↗Practice prompt ↗04One case, and the assumptions holding it up
- Answer one product case aloud in 20 minutes with a recording running, then listen back with a pen and mark every claim you asserted without saying what it rested on: an assumed user behaviour, an assumed data source, an assumed baseline rate, an assumed grain.
- Pick the three assumptions the recommendation actually depends on, write how you would check each one against data, and say which one being wrong would flip the recommendation rather than merely weaken it.
- Write the four-step structure you used onto a card small enough to hold in working memory when you are nervous.
Deliverable: One recording, three load-bearing assumptions each with a written check, and a four-step structure card.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Your own work, timed
- Write a 90-second version and a four-minute version of your main project, and time both out loud rather than reading them.
- Prepare answers to the two follow-ups that always come: what you would do differently, and how you knew it worked.
- Put one number in the first sentence and be able to say exactly where that number came from and what it excludes.
Deliverable: Two timed narratives with one defensible number in the opening line.
Practice prompt ↗Practice prompt ↗06The one full rehearsal, in a longer weekend block
- Run a 60-minute mock covering query work, a case and a behavioural question in a single sitting with no breaks, because sustained attention is the thing evenings have not trained.
- Immediately afterwards, and before hearing any feedback, write the three moments you lost the thread.
- Spend the rest of the block only on those three moments, and on nothing you merely feel shaky about.
Deliverable: Mock notes naming three failure moments with a specific fix written under each.
Practice prompt ↗Practice prompt ↗07Taper
- Write the 20-minute warm-up you will actually do on the morning of the interview: one query you can already write from a blank file, one metric you can define out loud, and nothing you have never seen before.
- Re-read only your own notes from this week, and open no new material.
- Write down the logistics: the tool you will be asked to work in, whether lookups are allowed, and the sentence you will use when you do not know something.
Deliverable: A one-page card holding the case structure, the project numbers, and the logistics.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
An answer without a quantity is hard to interrogate, so interviewers keep probing until they find one. Come with the baseline, the change, the window it was measured over, and how confident you were. If the effect never got measured, say so and say what you would have measured. Fabricated precision is worse than an honest gap.
Tell me about a time that you’ve delivered high-impact analytics as pa…
Tell me about a time that you’ve delivered high-impact analytics as part of a complex cross-functional project.
Approach
- Close with what you would do differently, concretely.
- Quantify the outcome, including what you would not claim credit for.
- Pick a story where you drove the decision, not one where you observed it.
Follow-up
- What would you do differently if you ran that project again?
- What did you decide not to do, and why?
Handle a request to re-cut a test after the readout
A four-week consumer-credit test reads flat on completed orders per active consumer and negative on contribution margin per completed order, which was the pre-registered guardrail. The sponsor asks for three re-cuts: on gross bookings instead of margin, on a seven-day window instead of four weeks, and excluding one market that 'had an outage'. One of the three is defensible under conditions. Deliverable: which you run, which you decline, the words you use to decline, and what appears in the written readout about all three requests.
Approach
- Sort the three requests by one test: could this have been specified before anyone saw the result, and is it symmetric across arms. That test, not the sponsor's seniority, decides what you run.
- Decline the gross-bookings switch on mechanism rather than on process: the credit operates by spending incentive dollars, and gross bookings excludes incentive spend by construction, so it cannot see the cost the guardrail exists to catch.
- Decline the seven-day window because the credit's payback horizon is longer than the window, so a short read measures the redemption spike rather than the behaviour change, and because the window was chosen after the four-week result was known.
- Run the outage exclusion only under stated conditions: the outage is visible in a metric nobody selected, such as requests per market-hour in fct_request, it hit both arms in the same proportion, and it is timestamped independently of this test. Report it as a sensitivity beside the primary, never as a replacement.
- Put all three requests in the readout with their status and reasoning, which makes the selection visible and removes the incentive to ask again quietly; then give the sponsor a real path forward: the incentive level at which the credit would break even on contribution margin, and a powered follow-up if that level is reachable.
Follow-up
- The sponsor says the guardrail was the wrong metric all along - how do you respond?
- What if the outage is real but hit only the treatment arm?
- How would you have pre-registered exclusions so that this conversation never happened?
Explain a switchback confidence interval to a non-technical executive
A switchback test of a dispatch-radius change ran 1,152 market-hour blocks across six markets. SLA fill rate moved +1.8 percentage points, 95% interval [-0.4, +4.0], variance clustered at the block. Those markets serve about 250,000 eligible requests a week at 88% fill and 93% completion. An executive with no statistics background wants a ship-or-wait answer inside a five-minute update. Deliverable: the two-minute spoken explanation, your recommendation, and the single condition that would change it. You may not use the words significant, p-value, or confidence interval.
Approach
- Open with the decision and the recommendation, then justify; an executive who hears the caveat first stops listening before the ask arrives.
- Translate both interval bounds into the unit the executive already manages: eligible requests times percentage points times completion rate gives weekly completed orders, so the range becomes 'between about 1,000 fewer and about 9,300 more completed orders a week, best single guess about 4,200 more'.
- Say plainly what the range does and does not rule out: it does not rule out a small loss, and it is wide because the test has 1,152 effective units, not 250,000 consumers. Block-level randomisation is the reason the sample is small, and it is the reason the number is trustworthy at market level.
- Price the two errors against each other: a reversible dispatch parameter with a bounded downside is cheap to ship and cheap to revert, so the decision rule is not 'is the effect proven' but 'is the worst case affordable and detectable'.
- End with the one condition that flips you: name the monitoring metric (provider utilisation and idle time, since a wider radius can raise fill by burning provider hours) and the threshold at which you revert.
Follow-up
- How many more weeks of blocks would it take to halve the width of that range, and is that worth the delay?
- The executive asks 'so is it real or not' - what do you say without reaching for statistical vocabulary?
- What would you monitor post-ship that the experiment itself could not measure?
- 01
Tell me about a time that you’ve delivered high-impact analytics as part of a complex cross-functional project.
- 02
A four-week consumer-credit test reads flat on completed orders per active consumer and negative on contribution margin per completed order, which was the pre-registered guardrail. The sponsor asks for three re-cuts: on gross bookings instead of margin, on a seven-day window instead of four weeks, and excluding one market that 'had an outage'. One of the three is defensible under conditions. Deliverable: which you run, which you decline, the words you use to decline, and what appears in the written readout about all three requests.
- 03
A switchback test of a dispatch-radius change ran 1,152 market-hour blocks across six markets. SLA fill rate moved +1.8 percentage points, 95% interval [-0.4, +4.0], variance clustered at the block. Those markets serve about 250,000 eligible requests a week at 88% fill and 93% completion. An executive with no statistics background wants a ship-or-wait answer inside a five-minute update. Deliverable: the two-minute spoken explanation, your recommendation, and the single condition that would change it. You may not use the words significant, p-value, or confidence interval.
Is this an official eBay interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at eBay. Rounds and questions reflect what candidates have reported, not a process eBay has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult is the interview process, and how much preparation time is recommended?
The interview process is rigorous and technical, requiring solid preparation across SQL, statistics, experimentation, and product sense. Most candidates dedicate 4 to 6 weeks of focused study, practicing coding problems and structuring mock product cases before their loops.
PracHub interview research ↗What is the best way to stand out during the product sense and case study rounds?
Structure your answers clearly by starting with clarifying questions, defining key metrics, and breaking the problem down into logical components. Always connect your analytical recommendations back to user impact and business value for eBay.
PracHub interview research ↗Are machine learning questions asked across all Data Scientist teams?
Not necessarily. While machine learning knowledge is essential for specialized teams focusing on AI, automation, or recommendation engines, many product analytics roles place a much heavier emphasis on SQL, experimentation, and causal inference.
PracHub interview research ↗What is the typical timeline from initial recruiter screen to final offer?
The timeline can vary depending on team urgency and scheduling alignment, but a standard loop typically spans 3 to 5 weeks from the initial online coding assessment through the final round of interviews.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22