The reported question set mixes five kinds of work: diagnosing and designing product metrics, SQL with window functions, A/B testing and statistics, machine learning from classification pipelines to NLP, and behavioral prompts about methodology and disagreement. Because the role is described as sitting between data and product strategy, the product-metric questions are not a side topic. Several reported questions are phrased as business situations, such as a sudden drop in a core metric or the success measure for an insurance feature, so start your answer from the decision the metric has to support.
For SQL, the reported examples are the third highest salary in a table, running totals, cohort analysis, and, among other reported examples, the top 3 users by spend per month. All of them come down to choosing the right window function and frame: DENSE_RANK versus ROW_NUMBER when values tie, a frame with a unique tiebreaker in ORDER BY for a running sum, and a first-activity date computed per user to assign cohorts. For experimentation, be ready for sample-size calculation, one-tailed versus two-tailed tests, seasonality and outliers. Statistical significance, p-values and confidence intervals are listed as foundations, and one reported prompt asks you to explain significance to a non-technical stakeholder.
The machine learning side is broad. Reported material names NLP, deep learning and Python-based data manipulation, a multi-class classification problem from start to finish, traditional models versus Transformer-based architectures for NLP, and the practical use of Open CV in past projects. Scikit-learn, TensorFlow and PyTorch appear in the listed skills. Prepare your own past projects in enough detail to say which library you used, why, and what the result was in business terms. Candidates are also told to be ready to quantify the impact of resume projects.
Reported tips add a few habits worth rehearsing. Use STAR for behavioral answers. Open a case by asking what the business wants from the analysis, agree which metrics will show it, and list the unusual cases that could distort them before you choose a technique. Expect your assumptions to be questioned and treat that as a joint working session rather than an attack. For logic puzzles, candidates are told the reasoning steps matter more than the mechanics of the puzzle, so practise thinking aloud.
Initial Screening
reportedCandidates describe the Initial Screening as the opening stage, where fit for the role is assessed. Little more is reported about its format, so prepare a short summary of your background and map yourself against the listed skills: Python, advanced SQL, a machine learning framework such as scikit-learn, TensorFlow or PyTorch, and A/B testing. Reports also suggest that freshers lean on logical reasoning and mathematical foundations, while experienced hires lean on influence on product outcomes. That advice is given for the role overall, not for this stage, but it tells you how to frame your background. Choose which framing is yours and have one project ready where you can state a measured business result.
What to demonstrate
- Fit for the role, which is how candidates describe the purpose of this stage
- How clearly you can summarise your background and your reasons for wanting a product-focused data science role, which a screening conversation inherently covers
How to prepare
- Write a short spoken summary of your path that names the stack you have used and ends with one project result expressed as a number
- Mark each listed skill as used in production, used in coursework or a side project, or not used, and have a one-sentence answer for each gap
- Pick your framing: freshers lean on reasoning and maths, experienced hires on influence over product outcomes, and prepare one project that supports it
- Ask the recruiter about the format of the later rounds, including which parts are remote and which are in person
Technical Assessments
reportedCandidates report that the Technical Assessments cover logical reasoning and machine learning. Across the loop, reports mention both remote assessments and, in many cases, in-office meetings, so ask the recruiter which applies to this stage. Candidates also list SQL, statistical knowledge and deep learning among the topics tested without tying them to one round, so keep all of them warm for this stage. Prepare to write working code and to state your reasoning as you go. For logic questions, name each step you take and avoid getting lost in puzzle mechanics. For the machine learning topics, practise a full pass from problem framing through model choice to evaluation.
What to demonstrate
- Logical reasoning, shown through the steps you state while solving
- Machine learning judgement: framing, model choice and evaluation for problems such as multi-class classification
- If the assessment includes SQL, Python or statistics (all listed for the role but not tied to a round), code and queries that return the right result and statistical reasoning you can explain
How to prepare
- Build a multi-class classifier in scikit-learn end to end: baseline, pipeline, cross-validation, macro-averaged metric and confusion matrix
- Write a one-page comparison of TF-IDF with a linear model versus a fine-tuned Transformer for a text classification task, covering data size, latency, cost and interpretability
- Keep SQL warm even though no report ties it to this stage: write a ranking query and a running total from a blank file and test both on tied values
- Solve two logic puzzles aloud, naming each deduction before you make it
Behavioral Assessments
reportedCandidates describe the Behavioral Assessments as the stage that evaluates soft skills and cultural fit. The reported behavioral questions, which are not tied to a specific round, include defending a technical methodology to a skeptical stakeholder, pivoting after unexpected data findings, disagreeing with colleagues over how to read model results, and mentoring a colleague or contributing to a team's growth. Reports also advise preparing to discuss how you take feedback and navigate disagreement, and to use STAR. Prepare a distinct story for each prompt so each one shows a different decision of yours.
What to demonstrate
- Soft skills and cultural fit, which is how candidates describe this round
- How you handle feedback and navigate disagreement, which reports advise preparing to discuss
- How you contribute to a team, the other area reports name for this kind of discussion
How to prepare
- Write four STAR stories, one for each reported prompt, and give each a result with a number and its source
- For the methodology-defence story, write down the specific objection the stakeholder raised and the evidence that answered it
- For the pivot story, state what the unexpected finding was, what you changed in the approach and what happened afterwards
- Rehearse each story aloud: keep the situation to two sentences and spend the rest on what you decided and did yourself, as distinct from what the team did
Stakeholder Interaction
reportedCandidates report that Stakeholder Interaction puts you in front of various stakeholders, including lead data scientists and senior leadership. That mix suggests two registers: technical depth on your models and experiments for data science peers, and business framing for leadership. Clear communication of complex ideas is listed among the role's soft skills, so be ready to describe statistical significance or a model result to someone with no statistics background. Confirm with the recruiter who attends, because the preparation differs by audience.
What to demonstrate
- The audiences reported for this stage: lead data scientists and senior leadership
- Explaining complex technical work clearly, which candidates list among the role's soft skills
How to prepare
- Prepare a short business version and a longer technical version of your best project, with the same numbers in both
- Write a plain-language explanation of statistical significance and of a confidence interval, with no jargon
- Re-read your earlier answers and list the figures you quoted so they match when repeated
- Candidates list handling constructive criticism during technical deep dives among the role's soft skills, so decide how you will respond when someone challenges a core assumption: restate it, say what evidence would change your view, and continue
PracHub editorial advice for the preparation topics above.
Answering a sudden metric drop by suggesting fixes before checking the number is real
Start with validation: check for instrumentation or logging changes, pipeline delays, and a change in the metric definition. Then slice by platform, region, acquisition channel, new versus returning users and product surface to localise the drop, and only then form causes such as a release, a campaign or seasonality. Say aloud what each check would rule out.
Picking a ranking method for the third highest salary without settling how ties count, then returning duplicates or an empty result
Ask whether tied salaries count once. If they do, use DENSE_RANK() or SELECT DISTINCT salary ... ORDER BY salary DESC LIMIT 1 OFFSET 2. If each employee counts separately, use ROW_NUMBER() (or RANK() if ties should share a place) or LIMIT/OFFSET without DISTINCT. Say what the query returns when fewer than three values exist (NULL or no row). For running totals, add a unique tiebreaker to ORDER BY (for example order_date, order_id) and then use ROWS, or keep RANGE if all rows on a date should share one cumulative value.
Choosing a one-tailed test after seeing which direction the result went, or quoting a p-value as the chance the hypothesis is true
Decide tail, significance level, primary metric and minimum detectable effect before launch. Use two-tailed unless a loss in the opposite direction truly would not change the decision. Explain a p-value as the probability of data at least this extreme if there were no effect, and give a confidence interval alongside it.
Recommending a Transformer for an NLP task by default, with no baseline or cost reasoning
Start from the constraints: labelled data volume, latency, serving cost, interpretability and how much the task depends on word order and context. Name a TF-IDF plus linear model baseline, compare it on a held-out set, and move to a pretrained Transformer only when the measured gain justifies the cost.
Telling project and behavioral stories in terms of what the team did, with no personal decision and no quantified result
Use STAR with 'I' for your own actions. Put one number in the result, know where it came from and what it excludes, and make sure the same figures appear each time you retell the story.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
What are the trade-offs between using traditional ML models versus Tra…
What are the trade-offs between using traditional ML models versus Transformer-based architectures for NLP tasks?
Approach
- Start with the traditional baseline: TF-IDF features over word or character n-grams fed into logistic regression, a linear SVM or Naive Bayes. They train in minutes on CPU, serve with very low latency, have inspectable coefficients, and are often strong on keyword- or topic-driven tasks with limited labels.
- Transformer models (fine-tuned BERT-family encoders, or prompting a large language model) use contextual representations. They handle word order, negation, synonyms and longer context much better, and pretraining transfers knowledge, so they often reach higher accuracy with fewer labels on semantic tasks.
- Name the costs: training usually needs GPUs, and self-attention cost grows quadratically with sequence length, which raises inference latency and serving cost. Encoders such as BERT have a fixed maximum input length (512 tokens), so long documents need truncation or chunking. Predictions are harder to explain, and models are larger to deploy and retrain.
- Decide on concrete criteria: label volume, how semantic the task is, the latency and throughput budget, the explainability requirement (for example, if a decision has to be justified to a customer), and the maintenance cost. Middle-ground options include frozen sentence embeddings with a linear head, or a distilled smaller transformer.
- Settle the choice empirically: compare both on the same held-out split using the metric that matters for the business, do error analysis on the cases where the transformer wins, and check whether that gain justifies the extra serving and maintenance cost.
Follow-up
- You have only 500 labelled examples. What do you try first, and why?
- How would you serve a transformer model under a tight latency budget?
- How would you explain an individual transformer prediction to a non-technical stakeholder?
Describe a situation where you had to clean a messy dataset to get it …
Describe a situation where you had to clean a messy dataset to get it model-ready.
Approach
- Pick one real project and name the concrete defects, not just 'messy data': missing values by column, duplicate records, inconsistent category spellings, mixed types or units, timezone mismatches, impossible values. Say how you found them: profiling in pandas with null rates, value counts, range checks and key-uniqueness checks.
- Explain each decision and its reason. Deduplicate on a defined business key. Treat missing values according to why they are missing: add a missing-indicator feature when missingness may carry signal, and impute with statistics computed on training data only. Separate outliers that are data errors (fix or drop) from real extreme values (keep, cap or transform).
- Show leakage awareness: fit imputers, scalers and encoders inside a scikit-learn Pipeline or ColumnTransformer so cross-validation never sees statistics from held-out folds, and drop fields recorded or updated after the prediction moment.
- Make the cleaning reproducible and reusable at prediction time: a scripted, versioned step with assertions on schema and value ranges, applied identically in training and serving, rather than one-off notebook edits.
- Quantify the outcome: how many rows were dropped or repaired, how the class balance changed, and the measured effect on the validation metric compared with training on the uncleaned data. State the business consequence in one sentence.
Follow-up
- How did you decide between dropping rows and imputing them, and how did you check that dropping did not bias the sample?
- How did you make sure the same cleaning ran on new data at prediction time?
- What would you do differently if the dataset were ten times larger?
Rebuild per-visitor ordering without groupby convenience methods
You have a DataFrame of 2 million fct_event rows with visitor_id, occurred_at_utc and event_id, unsorted and containing duplicate timestamps within a visitor. Produce three new columns: event_rank, the 1-based position of the event within its visitor ordered by occurred_at_utc; seconds_since_prev, the gap to that visitor's previous event, NULL for the first; and is_first_for_visitor. You may use sort_values, shift, cumsum, numpy and boolean masking. You may not use groupby.transform, groupby.apply, groupby.cumcount, groupby.rank or merge_asof. Break timestamp ties on event_id.
Approach
- Sort once by ['visitor_id', 'occurred_at_utc', 'event_id'] and reset the index. The whole exercise reduces to row arithmetic on a sorted frame, and the tiebreak on event_id is what makes the result reproducible across runs.
- Mark visitor boundaries with is_first = df['visitor_id'].ne(df['visitor_id'].shift()). This is the single fact every other column derives from.
- Compute seconds_since_prev as the diff of the timestamp column, then overwrite it with NaT/NaN wherever is_first is True. The shift crosses the boundary between visitors and will otherwise hand the first row of each visitor the last event of the previous one.
- Build event_rank from a running counter that resets at boundaries: take a global cumulative position (np.arange(len(df))) and subtract, per row, the global position at which that visitor started. Get the start position by forward-filling the positions where is_first is True, which is a cumsum-free reset and is O(n).
- Verify against the forbidden method once, as a test rather than as the implementation, and confirm the two agree on every row.
Worked solution 20 min
- Sort on the three-key tuple and reset_index(drop=True).
- Compute is_first via .ne(.shift()), which is True for row 0 because the shifted value is NaN.
- pos = np.arange(len(df)); start = pd.Series(np.where(is_first, pos, np.nan)).ffill(); event_rank = (pos - start + 1).astype(int).
- gap = df['occurred_at_utc'].diff().dt.total_seconds(); gap[is_first] = np.nan.
- Assert event_rank equals df.groupby('visitor_id').cumcount() + 1 on the sorted frame.
Follow-up
- The frame does not fit in memory. How does your approach change if you can only process one visitor-partitioned chunk at a time?
- occurred_at_utc is client-supplied and sometimes runs backwards within a visitor. Does your seconds_since_prev go negative, and should it?
- How would you extend this to reset the counter at every change of surface as well as visitor?
How do you use SQL window functions to perform cohort analysis or calc…
How do you use SQL window functions to perform cohort analysis or calculate running totals?
Approach
- For a per-row running total, use SUM(amount) OVER (PARTITION BY user_id ORDER BY order_date, order_id ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW). With ORDER BY and no frame, the default is RANGE, which gives all rows tied on a date one shared total. ROWS on a non-unique key makes tied rows' totals depend on an arbitrary order, so add a unique tiebreaker, or choose RANGE on purpose when a date should share one value.
- For a cohort analysis, first reduce the data to one row per user per period, then assign each user's cohort with MIN(DATE_TRUNC('month', activity_date)) OVER (PARTITION BY user_id) and compute the months elapsed since the cohort month. Count distinct users by cohort and period number.
- Get cohort size as the period-zero count: FIRST_VALUE(active_users) OVER (PARTITION BY cohort_month ORDER BY period_number) is deterministic because of the ORDER BY. MAX(active_users) OVER (PARTITION BY cohort_month) also works, since every cohort member is active in period zero by definition. Compute retention as retained users divided by cohort size, and never average per-period rates across cohorts of different sizes.
- Explain the key difference: window functions keep every row, while GROUP BY collapses rows. You cannot filter on a window result in the same SELECT's WHERE clause, so wrap it in a CTE or subquery.
- Edge cases: periods with no activity produce no rows, so join to a calendar or period spine with a LEFT JOIN when the grid must show zeros. Duplicate events per user per period inflate counts unless you reduce to the user-period grain first. The window sorts each partition, so cost is roughly O(n log n).
Follow-up
- How would you compute a 7-day rolling sum when some days have no rows at all?
- How would you use LAG to find users who churned and later reactivated?
- How does the query change if the cohort is defined by signup date stored in a separate users table?
Given a table of employee salaries, how would you find the third highe…
Given a table of employee salaries, how would you find the third highest salary using advanced SQL?
Approach
- Clarify the requirement first: the third highest distinct salary value, or the employee in third place? How should ties count? What should the query return when fewer than three distinct salaries exist (NULL or no rows)?
- Main answer: in a CTE, compute DENSE_RANK() OVER (ORDER BY salary DESC) AS rnk, then filter rnk = 3 in the outer query. DENSE_RANK gives tied values the same rank with no gaps; RANK leaves gaps after ties; ROW_NUMBER breaks ties arbitrarily. Name which one matches the stated tie rule.
- Alternative: SELECT (SELECT DISTINCT salary FROM employees ORDER BY salary DESC LIMIT 1 OFFSET 2) AS third_highest. Wrapping it in a scalar subquery returns NULL instead of an empty result when there is no third value.
- Handle NULLs explicitly: in PostgreSQL, NULLs sort first under DESC by default, so add WHERE salary IS NOT NULL or NULLS LAST. For the per-department version, add PARTITION BY department_id.
- Cost: the window approach sorts the table in O(n log n). An index on salary lets LIMIT/OFFSET read only the top rows. A correlated subquery that counts distinct higher salaries per row is O(n^2) and worth naming only to reject it.
Follow-up
- Return every employee who earns the third highest salary within each department.
- What does your query return if the table holds only two distinct salaries, and is that what the caller wants?
- How would you generalise this to the Nth highest salary with N passed as a bind parameter?
Find reactivation gaps in account paid-period history
fct_subscription_period holds account_id, subscription_id, period_start_utc, period_end_utc, period_status and change_reason. A mid-period plan or seat change closes one row and opens another, so a single continuous paid tenure is often many rows, and an account may hold two overlapping subscriptions. Collapse rows with period_status in ('active','past_due') into continuous tenures per account, treating gaps of three days or less as continuous. Return account_id, tenure_start, tenure_end, and for every tenure after the first, the gap in days that preceded it.
Approach
- Filter to paid rows only: period_status IN ('active','past_due'). Trialing periods are not tenure, and including them turns every trial that never converted into a one-period tenure followed by a fake churn.
- Order by period_start_utc and take a running maximum of all prior ends: MAX(period_end_utc) OVER (PARTITION BY account_id ORDER BY period_start_utc, period_end_utc, subscription_id ROWS BETWEEN UNBOUNDED PRECEDING AND 1 PRECEDING). Those three columns are the only stable ordering this schema exposes, so check first that they are unique within an account; if rows tie on all three, the island numbering is order-dependent between runs and you need a real row key before the result is reproducible.
- LAG on its own is wrong here because with overlapping or nested periods the immediately preceding row by start date is not the one that ends latest, so the running maximum is the part that cannot be shortcut.
- Flag a new island when prior_max_end IS NULL OR period_start_utc > prior_max_end + interval '3 days', then number islands with a running SUM of the flag over the same ordering and an explicit ROWS frame.
- Group to (account_id, island) taking MIN(period_start_utc) and MAX(period_end_utc), then LAG(tenure_end) OVER (PARTITION BY account_id ORDER BY tenure_start) to compute the preceding gap in days for every tenure after the first.
- Sanity-check with change_reason, which is the only lineage this schema carries: list its distinct values first, then confirm that rows recording a plan or seat change sit inside a tenure rather than opening one, and that every tenure after the first opens on a row whose reason records a restart rather than an ordinary renewal. Do not reconcile against a churn timestamp on dim_account, which this schema does not define; and where such a column does exist, a cancellation timestamp records when the request was made and routinely sits weeks before the period it ends.
Worked solution 35 min
- Find an account with a known mid-period upgrade and dump its period rows to use as the trace case.
- Check that (period_start_utc, period_end_utc, subscription_id) is unique per account, since the whole ordering rests on it.
- Write the paid-rows CTE and the running MAX with the explicit frame.
- Add the island flag and the running SUM, then verify the trace account yields one island.
- Group to tenures and add the LAG-based gap in days.
- List the distinct change_reason values, then count accounts with more than one tenure and compare against the count of accounts carrying a restart-flavoured reason anywhere in their history.
Follow-up
- Why three days of grace? What do 0 and 30 days each do to the count of accounts classed as reactivated?
- An account runs two concurrent subscriptions for different teams. One tenure or two, and what does the revenue reader expect?
- How would you turn these tenures into a monthly gross logo churn series without double-counting an account that churned and returned in the same month?
If a core product metric drops suddenly, walk me through your diagnost…
If a core product metric drops suddenly, walk me through your diagnostic process.
Approach
- Confirm the drop is real before explaining it: check whether the metric definition, the dashboard query or event logging changed, whether the pipeline is late or the latest day is partial, and compare against the same weekday in prior weeks so normal weekly or seasonal dips are not mistaken for a break.
- Pin down the shape and timing: a step change at a specific hour points to a release, an experiment ramp, an outage or a tracking change, while a gradual slide points to mix shift, competition or seasonality. Line the start time up against the deploy, experiment and campaign calendars.
- Decompose the metric into numerator and denominator or into funnel stages, then cut by platform, app version, region, acquisition channel and new versus returning users. Check whether one segment explains the drop, or whether segment rates are flat and the mix moved (a composition effect). If every segment improved yet the total fell, that is Simpson's paradox.
- Read neighbouring metrics to localise the cause: if upstream traffic fell, a rate can look broken while each stage is healthy; if only one funnel step dropped, the problem sits at that step. Rank hypotheses by likelihood and by how cheap they are to check.
- Close with evidence and action: confirm the cause (a rollback, a holdout or a before/after on the affected segment), size the impact in absolute terms, and name who fixes it and what alert would catch it sooner next time.
Follow-up
- Every segment rate you cut is flat, yet the total dropped. What explains that, and how would you show it?
- How would you tell a logging or instrumentation bug apart from a genuine change in user behaviour?
- What monitoring would you set up so a drop like this is caught within hours rather than days?
How would you design a metric to measure the success of a new insuranc…
How would you design a metric to measure the success of a new insurance product feature?
Approach
- Start from the feature's goal and the user behaviour it is meant to change, for example a feature that shortens the path from quote to purchase, and state the business decision the metric will inform (ship, iterate or roll back).
- Define one primary metric precisely: unit (user or policyholder), numerator, denominator, time window and exclusions, such as the share of exposed users who complete a purchase within a fixed number of days. Separate adoption (who used the feature) from impact (did outcomes improve).
- Add guardrails that reflect insurance-specific costs: early cancellations or lapses, support contacts, complaint rates, and where relevant claim or loss outcomes. Many of these mature slowly, so name leading indicators for launch and a later readout for the lagging ones.
- Avoid comparing adopters with non-adopters, because users who opt in differ from those who do not. Measure impact through a randomised rollout or a holdout group so the difference is causal.
- Check that the metric is sensitive enough to move detectably within a test, hard to game (a pushier flow can lift purchases while raising cancellations), and decomposable so a change can be traced to the step that caused it.
Follow-up
- The primary metric rises but early policy cancellations rise as well. Do you ship, and how do you weigh the two?
- Claim outcomes take months to mature. What do you measure at launch, and when do you revisit the decision?
- How would you set a success threshold for this metric before launch?
How do you determine the required sample size for a new experiment?
How do you determine the required sample size for a new experiment?
Approach
- List the inputs before computing anything: the primary metric with its baseline rate or variance, the minimum detectable effect worth shipping, the significance level (commonly 0.05, two-sided), the power (commonly 80%), and the allocation ratio between arms.
- For two proportions with equal arms: n per arm = (z_(1-alpha/2) + z_(1-beta))^2 x [p1(1-p1) + p2(1-p2)] / (p2 - p1)^2. With alpha 0.05 two-sided and 80% power, the z term squared is about 7.85, which gives the shortcut n per arm of about 16 x variance / delta^2.
- Work an example aloud: with a 10% baseline and a 1 percentage point absolute lift, n comes to roughly 14,000-15,000 users per arm. Because n scales with 1/delta^2, halving the detectable effect roughly quadruples the sample.
- Convert the sample into a duration: divide by eligible daily traffic per arm and round up to whole weeks so weekday and weekend behaviour are both covered. Fix the duration and the analysis plan before launch rather than stopping when p first dips below 0.05.
- Check the randomisation unit against the analysis unit: if you randomise by user but measure per session, the observations are correlated, so use the delta method or cluster-robust variance. Variance reduction such as CUPED cuts the required n by roughly a factor of (1 - rho^2). Correct alpha when testing several variants.
Follow-up
- Traffic supports only half the sample you need. What are your options?
- How does the calculation change for a heavy-tailed metric such as revenue per user?
- Why is it a problem to stop the test the first time the p-value falls below 0.05?
What is the difference between a one-tailed and two-tailed test in the…
What is the difference between a one-tailed and two-tailed test in the context of product launches?
Approach
- Define both precisely. A two-tailed test has H1: effect is not zero, with alpha split across both tails (cutoff z of 1.96 at alpha 0.05). A one-tailed test has H1: effect is greater than zero (or less), with all of alpha in one tail (cutoff z of 1.645).
- State the trade-off. The one-tailed test has more power in the direction you chose. At 80% power it needs about 21% fewer users ((1.645+0.84)^2 / (1.96+0.84)^2 is about 0.79). It cannot flag an effect in the other direction, so a launch that hurts the metric shows up only as not significant.
- Tie it to launch decisions: use a two-sided test for the primary metric unless the decision is truly asymmetric. For guardrails, a one-sided non-inferiority test with a stated margin (for example, not worse than -0.5 percentage points) is a legitimate and common use.
- The direction must be fixed before seeing data. When the effect falls in the predicted direction, the one-tailed p-value is half the two-tailed one, so switching after the fact doubles the effective false-positive rate.
- Mention the confidence-interval view: a two-sided 95% interval matches the two-tailed test at 0.05. Reporting the interval lets stakeholders see both the direction and the size of the plausible effect.
Follow-up
- Your two-sided p-value is 0.08 and the product manager asks you to call it one-tailed. What do you say?
- When would a one-sided test be the right choice for a product launch?
- How would you choose the non-inferiority margin for a guardrail metric?
How do you balance user experience with business-driven data collectio…
How do you balance user experience with business-driven data collection goals?
Approach
- Start from the decision each data point supports: for every field or event, name the model, metric or business process that consumes it, and drop fields with no consumer. Collecting less also reduces privacy and security exposure.
- Measure both sides rather than arguing in the abstract. The user cost is the drop-off or completion rate at the step where the field is asked, which can be A/B tested as required versus optional. The business value is what the field adds, for example an ablation showing how much a pricing or risk model improves with and without it.
- Use design alternatives that lower friction: ask progressively at the moment the data is needed rather than up front, derive values from data already held, pre-fill sensible defaults, or rely on passive event logging instead of asking the user. Respect consent and applicable privacy requirements, especially for sensitive personal data.
- Account for the analytical side effect: optional fields are rarely missing at random, so the users who answer differ from those who do not. Model the gap with missing indicators and avoid treating the responders as representative.
- Frame the recommendation as an explicit trade-off: the expected value of the information against the measured conversion loss, with completion rate and complaints as guardrails whenever a new collection step ships.
Follow-up
- Product wants five new fields added to the signup form. How do you decide which to keep?
- A field improves the model but lowers form completion by two points. How do you decide?
- What changes in your approach when the data is sensitive, such as health information?
Randomise a shared workspace feature without contaminating control
A feature changes a collaborative surface inside a workspace: when one member uses it, other members of the same account see the result in their own view. You have dim_user (user_id, account_id, is_internal), dim_account (account_id, seats_assigned, lifecycle_status) and fct_event. Among active accounts the mean seats_assigned is 6, the coefficient of variation of that count is 1.5, and the intraclass correlation of the weekly core-action rate within an account is 0.10. Choose the randomisation unit, quantify what that choice costs in sample, and specify how you would compute inference.
Approach
- State the interference before choosing anything: a treated user changes what an untreated colleague sees, so user-level randomisation puts both arms inside one account and biases the contrast toward zero. Randomise on account_id.
- Price the clustering properly. With equal clusters the design effect is 1 + (m - 1) rho = 1 + 5(0.10) = 1.5. Sizes here are far from equal, so use 1 + ((CV^2 + 1) m - 1) rho = 1 + (3.25 x 6 - 1)(0.10) = 2.85. The equal-size shortcut understates the cost by nearly half.
- Decide the estimand before the estimator. An account-weighted mean gives every workspace one vote; a user-weighted mean lets the largest workspaces dominate. With this size skew the two can move in opposite directions, so pick the one the decision needs and write it down.
- Compute standard errors on the account, not the user: cluster-robust on account_id, or collapse each account to a single number and test those. Below roughly 40 clusters per arm, cluster-robust errors are biased downward, so use a wild cluster bootstrap or randomisation inference over the assignment.
- Buy back variance where you can. Stratify assignment by seat band and lifecycle_status before randomising, and decide in advance how the handful of very large accounts are handled, since one enterprise workspace can carry more users than a hundred single-seat ones.
Worked solution 30 min
- Write the interference down: the outcome for user i depends on the treatment of other users in account(i), so the no-interference assumption fails at the user level and holds at the account level.
- Compute both design effects, 1.5 equal-size and 2.85 unequal-size, and use 2.85.
- Take the user-level sample requirement from the proportion shortcut, multiply by 2.85, then divide by the mean of 6 users per account to express it in accounts per arm.
- Specify the analysis: collapse to one row per account, regress the account-level outcome on variant with stratum fixed effects, and use a wild cluster bootstrap for inference.
- State the stopping rule up front: if the required account count exceeds the eligible population, the test is not runnable, and the alternatives are a longer window, a larger target effect, or a non-experimental read.
Follow-up
- Suppose the feature is not workspace-scoped but changes a globally shared ranking model, so no clean cluster exists. What design gets you a causal read, and what does it cost you?
- You have 900 eligible active accounts in total. Given the design effect, what absolute lift can this test detect, and is the honest answer 'do not run it'?
- The intraclass correlation is an estimate from last quarter. What happens to your sizing if the true value is 0.25?
Pooled signup conversion fell while every segment rose
Weekly visit-to-signup conversion, counted on distinct fct_session.visitor_id with is_bot_flagged = TRUE and consent_state = 'denied' sessions excluded, fell from 4.4% to 3.9% week over week. Split by device_type and referrer_channel, all twelve cells are flat or up. A paid_social campaign launched on Monday. Using fct_session and fct_event, quantify how much of the 0.5-point fall is mix and how much is within-segment rate, then state what you would tell the growth lead.
Approach
- Write the pooled rate explicitly as the sum over segments of weight times segment rate, and materialise both weeks' weights and rates into one table. Until that table exists there is nothing to decompose, only opinions.
- Compute three quantities and report all three: the rate effect holding the prior week's weights fixed, the mix effect holding the prior week's rates fixed, and the interaction residual. Reporting only the first two hides a term that can be material when both weights and rates move a lot.
- Rank segments by their individual mix contribution, computed as the change in that segment's weight multiplied by its prior-period rate. This is what lets you say one cell caused the move rather than gesturing at the campaign.
- Verify the new traffic is human and countable before accepting the mix story: check is_bot_flagged coverage on the new channel, the distribution of duration_seconds and event_count for its sessions, and whether its consent_state profile differs from the rest.
- Deliver the conclusion as a definition change rather than a diagnosis: a pooled rate over a mix that moves is not comparable week over week, so the recurring report should carry per-channel rates plus absolute signups, with the pooled figure demoted or dropped.
Follow-up
- Paid social converts at roughly a quarter of organic but absolute signups rose. Is the campaign working, and what would you need to answer that properly?
- Would you reach the same conclusion if the campaign had moved the mix by two points instead of sixteen? Where is your threshold and why?
- How would you present this to someone who has been watching the pooled number in a weekly meeting for a year?
The plan follows the four reported rounds and the reported question types: product metrics first, then SQL, then experimentation and statistics, then machine learning, then data-collection trade-offs, then behavioral stories, ending with a full rehearsal. Each day produces something you can show.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Map the loop and write two cold product answers
- List the four reported rounds and the skills named for the role (Python, SQL, scikit-learn, TensorFlow or PyTorch, A/B testing), and mark each as strong, rusty or new
- Write a cold outline for the question about a sudden drop in a core product metric, as a sequence of checks
- Write a cold outline for designing a metric to measure a new insurance product feature, naming a primary metric, a guardrail and the decision it supports
- Write three project impact lines, each with one number and a note on where it came from
Deliverable: A one-page skills map, two cold outlines for the product questions, and three quantified project lines.
Practice prompt ↗Practice prompt ↗Worked solution ↗02SQL window functions
- Write the third-highest-salary query three ways (DENSE_RANK, a DISTINCT subquery, OFFSET without DISTINCT) and test each with tied salaries and with fewer than three distinct values
- Write a running total with SUM() OVER (PARTITION BY ... ORDER BY order_date, order_id ROWS ...), then compare it with RANGE on order_date alone and note how tied dates differ
- Write a cohort retention query that assigns each user a cohort from their first activity date, then counts active users per cohort and month
- Write the top 3 users by spend per month using a ranking function partitioned by month, and complete the reactivation-gap practice problem
Deliverable: Three third-highest-salary variants tested on ties and short tables, a running total with a ROWS versus RANGE comparison, a cohort retention query, a top-3-per-month query, and the reactivation-gap practice solution.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Experimentation, statistics and logic puzzles
- Work out the sample size per arm for a conversion experiment from a baseline rate and a minimum detectable effect, first by formula and then with a library call, and check they agree
- Write when a one-tailed test is defensible for a launch decision and why a two-tailed test is the safer default, then write a plain-language explanation of statistical significance for a non-technical stakeholder
- List the pitfalls you have met in your own projects (peeking, sample-ratio mismatch, novelty effects, multiple comparisons, seasonality, outliers), write how you detected or handled each, then complete the shared-workspace randomisation practice problem
- Solve two logic puzzles aloud, naming each deduction before you make it, and do not get pulled into the puzzle's mechanics
Deliverable: A sample-size worksheet, a one-tail versus two-tail note with a plain-language significance explanation, a pitfalls list with a detection method for each plus the randomisation practice solution, and notes from two spoken puzzle solutions.
Practice prompt ↗Practice prompt ↗Practice prompt ↗04Machine learning end to end
- Build a multi-class classifier in scikit-learn on a public dataset: baseline, preprocessing pipeline, cross-validation, macro-averaged metric and confusion matrix
- Using pandas, clean a messy public dataset (null rates, duplicates, value counts, a groupby, a merge, per-user ordering), write what you repaired, the leakage risks and what you checked afterwards, then complete the per-visitor ordering practice problem
- Fine-tune a small pretrained text model in PyTorch or TensorFlow on a public text-classification dataset and compare it with a TF-IDF plus logistic regression baseline on accuracy, training time, inference latency and interpretability
- Prepare a two-sentence account of how you used Open CV in a past project, or of the closest computer vision work you have done
Deliverable: A working classifier notebook, a pandas cleaning notebook with a written account, the per-visitor ordering practice solution, a fine-tuned text model with a comparison table against the TF-IDF baseline, and the Open CV answer.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗05Metric diagnosis and data-collection trade-offs
- Rehearse the drop answer for a core metric that falls 10% overnight: list the first three checks and what each would rule out
- Write your answer to balancing user experience with business-driven data collection, covering what you would collect, what you would drop and who decides
- Work the practice problem where a pooled conversion rate falls while every segment rises, and write the mix-shift explanation in two sentences
- Record yourself answering one product case aloud and note each assertion that has no stated basis
Deliverable: A three-check metric-drop script, a data-collection trade-off answer, the mix-shift explanation and a recording with annotated assumptions.
Practice prompt ↗Practice prompt ↗Practice prompt ↗06Behavioral stories
- Write one STAR story each for defending your methodology to a skeptical stakeholder, pivoting after an unexpected data finding, and mentoring or growing a team
- Write the disagreement story about interpreting model results: both positions, the agreed check, what it showed and what you changed
- Practise the impact-ownership and flat-experiment practice problems so you can state your own contribution without claiming the topline
- Say each story aloud, shorten the setup until it takes two sentences, and list every figure you quote so you repeat them consistently
Deliverable: Four STAR stories with a result number and its source, written answers to the two practice problems, and a list of the figures you quote.
Practice prompt ↗Practice prompt ↗Practice prompt ↗07Full mock and taper
- Run a mixed mock: talk through one logic puzzle aloud, write one SQL query, answer one experiment question and one machine learning question
- Finish with a stakeholder-style question: explain one result to a listener with no statistics background
- Write a single page holding your case structure (objective, metrics, edge cases), your project numbers and the questions you will ask the recruiter
- Read only your own notes from the week and open no new material
Deliverable: A mock log with three weak moments and a fix for each, plus a one-page review sheet.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Candidates describe the Behavioral Assessments as the stage covering soft skills and cultural fit, and reports advise being ready to talk about feedback, disagreement and contributing to a team. The reported prompts cover defending a methodology, changing direction when the data surprises you, disagreeing over what model results mean, and mentoring. For a data scientist, the strongest stories show your reasoning and your own decisions, end in a measurable result, and show you can explain technical choices to a non-technical person.
How do you handle disagreements with team members regarding the interp…
How do you handle disagreements with team members regarding the interpretation of model results?
Approach
- Pick a disagreement where you owned part of the outcome and the subject was how to read results, for example whether a lift was real, which metric should decide, or whether a model's errors were acceptable for a segment. Avoid a story where you were simply right and the colleague was uninformed.
- Describe both positions in their strongest form in a sentence each, so the listener sees you understood the other reading. For instance: I thought the offline AUC gain would not survive deployment, and my colleague thought it would.
- Move quickly to how the disagreement was settled with evidence. Name the specific check you both agreed to, such as a holdout comparison, a calibration plot, an error breakdown by segment, a sensitivity analysis or a pre-agreed decision rule, and say what it showed.
- Say what you did that was personal: how you raised it, what you changed in your own analysis, and whether your position moved. Include a case where the other person was right, or partly right, because a story where you changed your mind shows you separate your ego from the result.
- Close with the outcome in business terms, any quantified effect, and what you do now to avoid the same disagreement, such as agreeing on the evaluation metric and decision threshold before training starts.
Follow-up
- What would you have done if the data could not settle the disagreement?
- How did you keep the working relationship intact afterwards?
- Would you handle it differently if the other person were a senior stakeholder rather than a peer?
Quantify your own impact without claiming the topline you touched
You are writing the impact section of your own review. Over the year you ran four experiments, one of which shipped and three of which were flat; you corrected the definition of gross monthly revenue churn so that cancellation is recognised at period_end_utc; and you built a self-serve funnel dashboard. Weekly active accounts rose 14% over the same period. Your reviewer knows the data well. Write the three impact claims you would defend, stating for each what you contributed, what evidence supports it, and what portion of the outcome you are not claiming.
Approach
- Recognise what is being probed: whether you apply to your own work the causal standard you would apply to somebody else's roadmap claim. Nearly everyone who would reject 'accounts that do Y retain better' will write 'I drove a 14% increase' without noticing it is the same error with a friendlier subject.
- Sort the work by the kind of evidence it can carry. The shipped experiment is the only item with a randomised estimate, so it is the only one where an effect size is defensible, and you claim the interval rather than the point estimate.
- Claim the three flat experiments as decisions prevented and price them. Features not built, or built differently, on evidence, with the engineering weeks reallocated as the number somebody else can verify. A defensible null is a delivered decision and should be written as one.
- Claim the definition fix as correctness, not as improvement. The old figure was overstated by a specific percentage and appeared in a specific set of recurring documents; the impact is the change it produced in the forecast built on top of it, not a change in churn itself.
- Claim the dashboard on usage and displacement: distinct weekly users of it, and the ad-hoc request count for six months before against six months after. If the request log does not exist, record the claim as unverified rather than estimating it upward.
- Disclaim the 14% explicitly and once. State that it cannot be separated from seasonality, other teams' launches and a pricing change, and bound your own contribution from above using the shipped experiment's interval converted into headline units.
Follow-up
- Your shipped experiment's interval was +0.2pp to +1.4pp on activation. How much of the 14% can that account for, and how do you say so without undercutting yourself?
- A peer in the same cycle claims the full 14%. What, if anything, do you do about it?
- If you could only keep two of your three claims, which do you drop, and why that one?
Defend a flat experiment readout against a post-hoc segment
A feature you evaluated is flat on seven-day activation: +0.05pp with a 95% interval of [-0.47pp, +0.57pp], from 61,000 exposed users per arm in fct_experiment_exposure joined to dim_user and fct_event. Baseline activation is 32%. The launch team asks you to drop every surface except mobile_web, where the point estimate is +1.1pp, and re-run. You have ten minutes in their planning meeting. Deliver a spoken position: what you will and will not do, and the decision you recommend.
Approach
- Recognise what is being probed: whether you hold a statistical position under social pressure without becoming either rigid or apologetic. A generic answer says the segment is not significant; a strong one separates the request into a question that is answerable (is the mobile_web number real?) and one that is not (can we ship on it?), and answers both.
- Price the multiplicity out loud. The slice was chosen after seeing the results, so its estimate is selected on favourable noise and is biased away from zero. With k independent looks at a nominal 5% level, the chance of at least one false positive is 1 - 0.95^k: 26% at six segments, 64% at twenty. Quote the k you actually inspected, not the k you reported.
- Use the arithmetic already in front of you. On the point estimates, a +1.1pp mobile_web effect combined with a pooled +0.05pp implies the remaining surfaces average negative in proportion to mobile_web's share of exposures. State that as a testable implication of their story rather than as a rebuttal of it.
- Ask the one question that settles the category: was mobile_web named in the analysis plan before launch? If it was, it is a planned comparison and gets a corrected reading. If it was not, it is a hypothesis, and the honest move is to size the test that would confirm it.
- Convert the refusal into a cost. Size a mobile_web-only confirmatory test at the claimed effect, state the weeks of mobile_web traffic it needs, and close with the recommendation: do not ship this as a lift, and note that the interval already rules out anything at or above +0.6pp, which is itself a useful input to the roadmap.
Follow-up
- The confirmatory test you sized needs nine weeks of mobile_web traffic and the team has three. What do you recommend instead?
- Suppose mobile_web was pre-registered. How does your reading change, and what correction do you apply?
- Your interval excludes +0.6pp. Is that the same as saying the feature does nothing?
- 01
Tell me about a time you had to defend your technical methodology to a skeptical stakeholder.
- 02
Describe a challenging project where you had to pivot your approach due to unexpected data findings.
- 03
How do you handle disagreements with team members regarding the interpretation of model results?
- 04
Tell me about a time you mentored a colleague or contributed to a team's growth.
- 05
Tell me about a project where you could show the business impact of your work in numbers, and explain where the numbers came from.
- 06
Describe a time you received tough feedback on an analysis and what you changed afterwards.
Is this an official 1 digit technology Pvt interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at 1 digit technology Pvt. Rounds and questions reflect what candidates have reported, not a process 1 digit technology Pvt has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How many rounds are reported, and what are they?
Candidates report four rounds over roughly 3-5 weeks: Initial Screening, Technical Assessments, Behavioral Assessments and Stakeholder Interaction. Reports say some stages may be combined or sped up depending on the team, and mention both remote assessments and, in many cases, in-office meetings.
PracHub Data Scientist practice ↗What SQL should I practise?
Window functions and complex joins are the reported focus. Practise ranking with DENSE_RANK and ROW_NUMBER (the third highest salary), running totals with an explicit frame and a unique tiebreaker, cohort analysis based on a first-activity date, and top-N per group, such as the top 3 users by spend per month. Test each query with ties, missing rows and empty groups.
PracHub Data Scientist practice ↗Which tools and topics should I know?
Candidates list Python, advanced SQL, scikit-learn and TensorFlow or PyTorch, along with A/B testing. Reported topic areas are SQL, statistical knowledge, logical reasoning, general machine learning and deep learning, and the questions touch NLP, Transformers and Open CV. Be ready to give a concrete example from your own work for each one you claim.
PracHub Data Scientist practice ↗Do I need insurance domain knowledge?
A reported tip suggests researching the company's product line, especially in insurance, since it helps with product-sense questions. At minimum, think through what a success metric for a new insurance feature could be, and what could go wrong with it, such as a metric that improves because of who chooses to use the feature.
PracHub Data Scientist practice ↗How should freshers and experienced candidates present themselves differently?
Candidates report that freshers should emphasise logical reasoning and mathematical foundations, while experienced hires should highlight how they influenced product outcomes. In both cases, expect to state the business impact of resume projects in quantified terms, so prepare numbers you can explain and source.
PracHub Data Scientist practice ↗How should I approach a case study question?
Candidates are told to open with what the business wants from the analysis, then agree which numbers will show it and which unusual cases could distort them, before choosing a method. They are also told to expect their assumptions to be challenged, so state your assumptions explicitly and say which evidence would change your answer.
PracHub Data Scientist practice ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22