2nd Order Solutions · Data Scientist
Updated · 2026-10-02

2nd Order Solutions Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

Candidate write-ups do not say what 2nd Order Solutions sells or builds, so this guide does not guess. What they do describe is the role itself. A Data Scientist analyses data to inform product development and strategic decisions, builds predictive models aimed at user engagement and operational efficiency, maintains dashboards and reports that show key performance indicators to stakeholders, and works with engineering, product management and marketing teams. The write-ups also mention clients, and they describe the work as business-oriented, so link each analysis to a decision.

This guide is for candidates interviewing for the Data Scientist role at 2nd Order Solutions. It covers the four reported stages and gives worked approaches for ten of the questions candidates report (product and prioritisation, SQL and coding, experiment design, data quality, missing data); the seven-day plan covers the rest, including churn, the decision tree, overfitting and feature selection. It also covers the stack the requirements list: Python or R, SQL, and Tableau or Power BI, with cloud platforms, Hadoop or Spark, deep learning and NLP as extras. The plan points to PracHub practice drills on analytics problems; those are PracHub's own exercises, not reported company questions.

Candidates report four stages over roughly three to five weeks: a modeling case study first, then technical interviews, case studies and behavioral evaluations. These are reported stages, not a published process, and the later ones vary by team.

Cluster inference at the account, not the engagementSeparate bookings, recognised revenue and collected cashReconstruct pipeline stage history from mutable rows

47 min read

Practice 17 Data Scientist prompts
17Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

The role as candidates describe it mixes three kinds of work: analysis that informs product and strategy decisions, predictive modelling for engagement and efficiency, and reporting through dashboards. The topics reported for the loop line up with that: machine learning fundamentals, model building, a data science case interview, a modeling case study and model evaluation metrics. Preparation therefore leans on modelling judgement (what to predict, how to validate, which metric fits the cost of errors) more than on puzzle-style algorithm drills.

The reported questions come in several shapes. Some are conceptual and short: supervised versus unsupervised learning, overfitting, feature selection, handling missing data. Some are small coding or SQL tasks, such as a function for the mean and median of a list or a query for the top three products by sales volume. Some are open case prompts, such as customer churn, predicting sales for a new product, cleaning a messy dataset or designing an A/B test for a new feature. A smaller group is about data systems: data quality at scale, real-time streaming, and relational versus NoSQL storage. Prepare each shape differently: a crisp definition with an example for the first, working code with edge cases for the second, and a structured walk to a recommendation for the third.

Candidates describe case analysis in three moves: define the problem, explore the data to find the key variables, and turn the findings into an action. Communication is also reported as an evaluation area, including explaining a complex model to a non-technical audience and saying which metrics show a model is succeeding. Pick one language (Python or R) and SQL as the tools you can write live, know what Tableau or Power BI would show for the KPIs you propose, and rehearse every answer out loud, because you will need to explain your reasoning as you work.

01

Modeling Case Study

reported

Candidates report this as the opening stage: a structured modeling case study where you show analytical skill. Model building, machine learning fundamentals and model evaluation metrics are among the topics reported for the loop, so prepare to carry a case from framing through to evaluation. Open by stating the target, the unit you predict for, the moment the prediction is made and the action the score would drive. Then cover features that exist at prediction time, a simple baseline, one or two candidate models, the metric, the validation scheme and what you would ship.

What to demonstrate

  • Whether you frame the target, prediction time and decision before naming an algorithm
  • Whether the metric and validation scheme match the cost of errors and the data's structure, such as a time-based split for forecasting
  • Whether the case moves through a structure you announce and finish

How to prepare

  • Write a reusable skeleton: target, unit, prediction time, features, baseline, model, metric, validation, failure modes, monitoring. Apply it to the new-product sales prediction and to customer churn.
  • For regression and classification, list the metrics you would report and when each misleads: MAPE with zero-sales weeks, accuracy on imbalanced churn labels, AUC without calibration.
  • On a public tabular dataset, fit a baseline and one stronger model in Python or R with a time-ordered split, and summarise the comparison in a single table.
PracHub interview research ↗
02

Technical Interviews

reported

Candidates describe technical interviews, plural, that assess data science knowledge and skills. The reported areas are statistical analysis (hypothesis testing, p-values, confidence intervals), machine learning (regression, classification, clustering and their use cases) and data manipulation in Python, R or SQL. Questions of the conceptual, small-coding and SQL types in this guide are what you rehearse for this stage. Say your plan before typing: which table, what grain, which filter defines the population, then write the code.

What to demonstrate

  • Statistical reasoning, including what a p-value and a confidence interval do and do not claim
  • Working knowledge of regression, classification and clustering and when each applies
  • Working code in Python, R or SQL for retrieving, cleaning and summarising data

How to prepare

  • Write the mean-and-median function by hand: odd and even lengths, an empty list, unsorted input. State the cost of sorting versus a selection approach.
  • Rehearse one-minute explanations of p-value, confidence interval, bias versus variance, regularisation and supervised versus unsupervised learning, each with a concrete example.
  • Write the top-three-products-by-sales query three ways, with ROW_NUMBER, RANK and DENSE_RANK, and say how each treats ties. Then add a per-category version.
PracHub interview research ↗
03

Case Studies

reported

Candidates say additional rounds may include case studies that test problem-solving. The reported shape of case analysis is problem definition, data exploration to find the key variables, and insight generation that ends in an actionable recommendation. Case-style prompts reported for this loop include customer churn, a messy dataset, deriving insights from a dataset, predicting sales for a new product and an A/B test for a feature. Turn the open prompt into a decision first, name the quantity that would settle it, state assumptions when you rely on them, and finish with a recommendation.

What to demonstrate

  • A clear problem definition that is narrower than the prompt and still worth answering
  • An exploration plan that finds the key variables and checks data quality before analysis
  • A recommendation someone could act on, with the result that would reverse it

How to prepare

  • Outline a churn case on one page: how churn is defined, the cohort view, candidate drivers, whether you would model it or run an experiment, and the intervention.
  • Write a messy-data checklist (duplicate keys, units, impossible values, missingness pattern, outliers, time zones) and narrate it on a real public dataset.
  • End every practice case with one closing sentence: the action, the result that supports it and the result that would change it.
PracHub interview research ↗
04

Behavioral Evaluations

reported

Candidates describe the final stage as discussions about your experience and fit with the team. Reported evaluation areas include leadership, meaning influencing decisions and communicating with colleagues and stakeholders, plus culture fit. Reported scenarios in the loop include explaining a complex machine learning model to a non-technical audience and a time your communication cleared up a misunderstanding within a team. Prepare stories that carry numbers, and be ready to state the baseline, the period and the comparison behind each impact claim.

What to demonstrate

  • Influence: how you persuaded a stakeholder or changed a decision with evidence
  • How you handle team conflict and competing deadlines
  • Whether you can explain technical work and its success metrics in plain language

How to prepare

  • Write five stories in situation, action, result form: persuading a stakeholder, a difficult team dynamic, competing deadlines, a modelling project with real challenges, and a time your communication cleared up a misunderstanding.
  • For each story with numbers, write the impact line as metric, value before, value after, period and how you attributed the change to your work.
  • Prepare a three-sentence plain-language explanation of one model you built, then a version that adds the metric used to judge it.
PracHub interview research ↗

PracHub editorial advice for the preparation topics above.

01

Naming an algorithm in the modeling case study before stating the target, prediction time and baseline

Open every modelling case with the label, the unit, the moment the prediction is made and the action the score drives. Then state a simple baseline (last period, group average or logistic regression) so any complex model has something to beat.

02

Judging a model with a default metric, such as accuracy on imbalanced churn labels or R-squared alone on a sales forecast

Tie the metric to the cost of each error type. For churn, discuss precision, recall and a threshold chosen with the retention team. For forecasts, report MAE or RMSE on a time-ordered holdout and say why MAPE fails when actuals can be zero.

03

Finishing a case study with a list of observations or "it depends"

Convert the prompt into a decision at the start, state assumptions at the point you use them, and close with the action, the number that supports it and the result that would reverse it.

04

Answering foundational questions with a single tactic: "fill with the mean" for missing data, "add regularisation" for overfitting

Give the diagnosis before the fix. For missing data, say which mechanism you suspect and how you would check it. For overfitting, show the train-versus-validation gap, then list remedies (more data, simpler model, regularisation, early stopping, feature pruning) with their costs.

05

Telling behavioral stories with impact figures that have no baseline, period or comparison, or explaining a model in jargon

Write each impact line as metric, before, after, period and attribution method, and concede where the link was correlational. Practise the plain-language model explanation until it needs no unexplained term.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

14 technical prompts3 include a worked solution

What is the time complexity of your favorite sorting algorithm?

medium
machine learning and modelling

What is the time complexity of your favorite sorting algorithm?

Approach
  1. Pick one and give the full profile. Merge sort is O(n log n) in the best, average and worst case, needs O(n) extra space and is stable. Quicksort is O(n log n) on average but O(n^2) in the worst case with bad pivots (randomised or median-of-three pivots make that unlikely), uses O(log n) expected stack space and is not stable.
  2. Mention what you actually call: Python's sorted and list.sort use Timsort, which is stable, O(n log n) in the worst case, O(n) on input that is already sorted, and uses up to O(n) extra memory. Heapsort is O(n log n) worst case with O(1) extra space but is not stable and is cache-unfriendly.
  3. Explain the lower bound. Any comparison sort needs Omega(n log n) comparisons in the worst case, because a decision tree must distinguish n! orderings and log2(n!) grows as n log n.
  4. Show when to beat it. Counting sort runs in O(n + k) for integer keys in a range of size k, and radix sort in O(d(n + k)), so neither is limited by the comparison bound. Insertion sort is O(n^2) in general but O(n + inversions), so it is fast on small or nearly sorted input, which is why Timsort uses it on short runs.
  5. Connect it to data work. For the top three products by sales you do not need a full sort: a heap of size k gives O(n log k), and in SQL ORDER BY with LIMIT lets the engine do a top-k sort.
Follow-up
  • Why can no comparison-based sort beat n log n in the worst case?
  • How would you sort a file far larger than memory?
  • When does stability matter in practice? Give an example from sorting tabular data on multiple keys.

Measure the timesheet backfill curve and pick a reporting cutoff

easy
late-arriving datadata qualitycohort curves

time_entries has work_date (date), entered_at (timezone-aware UTC timestamp), hours and status. Given a snapshot_date, restrict to work_date in [snapshot_date - 180 days, snapshot_date - 60 days] so every cohort is fully observed. For k = 0..45, compute F(k): the share of a work_date cohort's final hours that already existed as of work_date + k days, pooled across cohorts. Return the 46-point curve and the smallest k with F(k) >= 0.99. Some rows are entered before the work date; those lags are real, not errors.

Approach
  1. Compute lag = (entered_at converted to the reporting timezone and taken as a date) - work_date in whole days, then clip negative lags to 0 instead of dropping them; leave and planned time are routinely entered ahead of the work date and dropping them deflates the early curve.
  2. Take cohort totals as groupby(work_date).hours.sum() over the restricted window. These are final only because the window stops 60 days short of the snapshot, which is why the restriction is in the prompt.
  3. Build the numerator by summing hours per (work_date, lag), sorting by lag, taking a per-cohort cumsum, then reindexing each cohort onto the full 0..45 lag grid and forward-filling, so a cohort with no entries at a given lag holds its previous level rather than disappearing.
  4. Pool as sum(numerators) / sum(denominators) at each k, not as the mean of per-cohort shares. Holiday weeks are tiny cohorts and would otherwise carry the same weight as a full week.
  5. Read k* off the pooled curve and report F(45) with it: if F(45) is below about 0.995 the tail runs past the grid and k* is a lower bound, not the answer.
Follow-up
  • The dashboard refreshes daily. Would you hold the window back past k*, or publish an as-of-entered_at series instead, and what does each choice cost the reader?
  • One practice area has a tail twice as long as the rest. Does that change the firm-wide cutoff, or does it change what you publish per practice area?

Simulate the chance a capped engagement crosses its cap

mediumWorked solution
monte carlobootstrap resamplingpricing caps

A capped time-and-materials engagement has not_to_exceed_usd = 400000, has billed 250000 to date, and bills at a blended bill_rate_usd of 250. Fourteen delivery weeks remain before planned_end_date. You have that engagement's last twenty weekly totals of approved billable hours as a pandas Series. Estimate the probability that cumulative billable value crosses the cap before the planned end, and the expected unbillable hours if it does. Resample the observed weeks; do not assume normality. Report a Monte Carlo standard error and justify the number of draws.

Approach
  1. Reduce the deterministic part first: the cap still affords (400000 - 250000) / 250 = 600 hours, so the whole question is the distribution of the sum of fourteen future weekly hour totals against a fixed threshold of 600.
  2. Before resampling, check the twenty observed weeks for trend and lag-1 autocorrelation. Independent resampling is defensible only if the weeks are exchangeable; on a ramping engagement use a moving-block bootstrap or model the ramp, because i.i.d. draws understate the upper tail, which is exactly the tail the question is about.
  3. Draw one (B, 14) array with np.random.default_rng().choice(observed, size=(B,14), replace=True), sum along axis 1, and compute p_hat = mean(total > 600) in one vectorised pass.
  4. Report expected unbillable hours two ways: unconditional mean(maximum(total - 600, 0)) for expected loss, and the mean conditional on a breach for how bad a breach is when it happens. The second is the number that drives a change-order conversation.
  5. Attach se(p_hat) = sqrt(p_hat(1 - p_hat)/B) and pick B from the precision you need: plus or minus one point at 95% confidence needs about 1.96^2 * 0.25 / 0.01^2, roughly 9600 draws at the worst case p = 0.5.
  6. State the assumptions that would flip the answer: constant blended rate, no scope change, no holiday weeks inside the fourteen, and a cap that applies to fees rather than to fees plus expenses.
Worked solution 30 min
  1. affordable_hours = (400000 - 250000) / 250 = 600.0; print it, because a wrong threshold makes every later number wrong in a way no simulation reveals.
  2. Plot or regress the twenty weekly totals against week index and compute lag-1 autocorrelation; record the result as the stated justification for i.i.d. versus block resampling.
  3. rng = np.random.default_rng(7); draws = rng.choice(weekly.values, size=(10000, 14), replace=True); totals = draws.sum(axis=1).
  4. p_hat = (totals > 600).mean(); se = np.sqrt(p_hat * (1 - p_hat) / 10000); overrun = np.maximum(totals - 600, 0); report overrun.mean() and overrun[totals > 600].mean().
  5. Re-run with a second seed and confirm p_hat moves by less than two standard errors before reporting.
EXPECTED RESULTA probability with its Monte Carlo standard error (about 0.005 at B = 10000 near p = 0.5), expected unbillable hours both unconditionally and conditional on a breach, and an explicit statement of the exchangeability assumption behind the resampling.
Follow-up
  • Hours past the cap still cost money. Restate the result as expected gross margin rather than a probability.
  • Two weeks of time are entered but not yet approved. How do you fold them in without double-counting?
  • A change order that raises the cap is judged 60% likely. How does that change the number you present, and to whom?

The week follows the reported stages: a modelling case first, then ML fundamentals, coding and SQL, statistics and A/B tests, case studies, data systems, and finally behavioral stories. Each day ends with an artifact you can reread before the loop.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Frame a modeling case from target to metric
  • Write a reusable skeleton: target, unit, prediction time, features, baseline, model, metric, validation, monitoring.
  • Apply it to predicting sales for a new product. List the factor groups you would consider (price, promotion, seasonality, channel, comparable past launches, cannibalisation of existing products) and say how you would forecast with no history, such as by borrowing from analog products.
  • Choose a metric and a time-ordered validation scheme for that forecast and justify both in two sentences.

Deliverable: A one-page skeleton filled in for the new-product sales prediction, ending with the chosen metric and validation split.

Practice prompt ↗Worked solution ↗
02ML fundamentals and evaluation metrics
  • Write a comparison of supervised and unsupervised learning with one use case each, and list regression, classification and clustering methods with when each applies.
  • Build a metrics sheet: MAE, RMSE, R-squared and residual plots for regression; precision, recall, F1, ROC AUC, PR AUC and calibration for classification, each with a situation where it misleads.
  • Implement the split search of a decision tree from scratch in Python (Gini impurity, best threshold per feature, recursion with a depth limit) and compare it with a library tree on a small dataset.
  • Write the overfitting answer: how to see it (train versus validation gap) and five remedies with their costs, plus filter, wrapper and embedded feature selection.

Deliverable: A notebook with the from-scratch tree, plus a written supervised-versus-unsupervised comparison, an overfitting and feature-selection write-up, and a one-page sheet of metrics and when each fails.

03Coding and SQL under narration
  • Write the mean-and-median function with tests for odd length, even length, one element, empty input and unsorted input. State the complexity of the sorting version and of a selection-based alternative.
  • Write a short table comparing merge sort, quicksort, heapsort and insertion sort on worst case, memory, stability and nearly-sorted input, so you can answer the favourite-sorting-algorithm question for any choice.
  • Write the top-three-products-by-sales query with ROW_NUMBER, RANK and DENSE_RANK, then a per-category version. Then work the PracHub drill on avoiding fan-out across two fact tables, narrating the plan before you type.

Deliverable: A tested function, a one-page sorting comparison and a SQL file with three ranking variants plus the drill solution.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
04Statistics, A/B tests and messy data
  • Write explanations of the p-value and the confidence interval in plain language, including one common misreading of each.
  • Design an A/B test for a new product feature: randomisation unit, primary metric, guardrail metrics, the sample-size formula for a binary metric, the decision rule and what you would do if the result is flat.
  • Work through a messy dataset in Python or R: check key uniqueness, types and units, impossible values and the missingness pattern, then write what you did about each and why.
  • Answer the missing-data question out loud using mechanism, method and a sensitivity check.

Deliverable: A one-page test design with a sample-size calculation, a cleaning log of each issue and decision, a half-page of p-value and confidence-interval explanations, and a short written outline of your missing-data answer.

Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗
05Case studies that end in a recommendation
  • Outline a customer churn case: churn definition, cohort retention view, candidate drivers, model versus experiment, and the intervention you would recommend.
  • Practise the metric-diagnosis case in the PracHub drill where days sales outstanding improved while collections got worse. Name the decomposition you would run first.
  • Work the PracHub drill on a staggered rollout across practice areas, and write the identifying assumption and the strongest threat to it.
  • For each case, write the closing sentence: action, supporting number, reversing result.

Deliverable: A churn case outline, a written diagnosis for the metric drill, and a rollout memo, each ending in a recommendation sentence.

Practice prompt ↗Practice prompt ↗
06Data systems and reporting
  • Write a data-quality plan for a large processing system: schema checks, freshness, row-count and null-rate monitors, uniqueness and referential checks, distribution drift, and where each check runs.
  • Prepare the relational versus NoSQL answer for an analytical workload: joins, schema flexibility, consistency, query patterns and scale.
  • Sketch a KPI dashboard in Tableau or Power BI for a product you know, defining each metric's numerator, denominator and refresh logic.
  • Outline two designs candidates report. Streaming: source, broker such as Kafka, stream processor with watermarks, sink and late-data handling. Model deployment: batch versus online serving, feature parity between training and serving, model registry, canary rollout and drift monitoring.

Deliverable: A data-quality checklist, a relational-versus-NoSQL comparison table, a dashboard sketch with metric definitions, and a one-page streaming and model-deployment outline.

Practice prompt ↗Practice prompt ↗
07Behavioral stories and a full mock
  • Finalise five stories: persuading a stakeholder, a difficult team dynamic, competing deadlines, a modelling project, and a time your communication cleared up a misunderstanding. Write an impact line (metric, before, after, period, attribution) for each that has numbers, and test the persuasion story against the PracHub drill on defending an impact claim without randomisation.
  • Write short answers to the prioritisation, motivation and staying-updated questions, each with specifics: how you rank competing deadlines, a named source you follow and something you tried because of it.
  • Deliver the plain-language model explanation to a friend outside data and ask them to repeat it back.
  • Run a mock of a case followed by two technical questions, speaking throughout.

Deliverable: Five written stories with impact lines, three short answers, the plain-language model explanation as tested on a friend, and a mock recording with the places you went silent marked.

Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

Candidates describe the behavioral stage as a discussion of experience and team fit, with leadership defined as influencing decisions and communicating with colleagues and stakeholders. For a data role, that means stories where an analysis or model changed a decision, with the measurement attached, and an explanation of technical work that a non-technical listener can follow. Rehearse the reported prompts first.

How do you handle missing data in a dataset?

medium
behavioural and stakeholder questions

How do you handle missing data in a dataset?

Approach
  1. Start by measuring the gap: missing share per column, per segment and over time, and whether a NULL means unknown, not applicable or zero events logged. A structural NULL (a field that does not exist for that user type) should become its own category or be left alone, not imputed.
  2. Classify the mechanism: missing completely at random, at random given observed columns, or not at random (high earners skipping an income field). Test it by comparing other variables between rows with and without the value, or by fitting a quick model that predicts the missing flag. MNAR cannot be proven from the data, so say so and run a sensitivity check.
  3. Choose by purpose. For a small MCAR share, dropping rows is unbiased but costs power. For prediction, impute with the median or mode plus a missing-indicator column, KNN or iterative imputation, or use a tree-based model that accepts NaN natively. For inference, use multiple imputation so standard errors carry the imputation uncertainty; mean imputation shrinks variance and weakens correlations.
  4. Keep the imputer inside the pipeline: fit it on training folds only (a scikit-learn Pipeline inside cross-validation, or the R equivalent) and reuse the fitted values at serving time. Fitting on the full dataset leaks information, and a column rarely missing in training but often missing in production needs its own handling.
  5. Close with your own example: the column, how much was missing, the mechanism you found, the method you chose, and a comparison of results under deletion versus imputation showing the conclusion held.
Follow-up
  • How would you check whether values are missing at random, and what could you not rule out?
  • Why is mean imputation a problem for a regression coefficient or a variance estimate?
  • The column is rarely missing in training data but often missing in production. What do you change?

Defending your own impact claim without randomisation or clean units

hard
self-selectionclustered inferenceimpact measurement

At your review you plan to claim that the realisation dashboard you built recovered 1.4 million dollars. The evidence is that engagement-month realisation rose four points over two quarters among engagements whose leads used it. Adoption was voluntary. There are about sixty client accounts and the top five carry most fees. A new rate card shipped in the same quarter. Write the claim you can defend, the estimate you would actually produce, and what you say when asked for a causal number you cannot get.

Approach
  1. Name the probe: whether you can separate the number you want from the number the data supports, under review pressure, without either inflating it or retreating to saying nothing can be known.
  2. State both identification problems concretely. Voluntary adoption means adopting leads are plausibly the ones who already manage realisation, so the comparison is confounded at the person level. The rate card changes bill_rate_usd, which sits in the realisation denominator, so part of the four-point move is arithmetic rather than behavioural.
  3. Neutralise what you can. Recompute realisation with bill rates snapshotted on work_date, or hold the denominator at the old rate card, so the rate-card change cannot move the metric by construction. Then rerun the comparison.
  4. Get the inference right for the unit count. Cluster at client_id, not engagement, because engagements in one account share a partner, a rate card and a team. With sixty accounts and five carrying most fees, the effective cluster count is far below sixty, so report a wild cluster bootstrap interval rather than plain cluster-robust standard errors, which are biased downward in that regime.
  5. Report both weightings and explain the divergence: an account-weighted estimate describes the typical account, a value-weighted one describes the revenue, and if they disagree a small number of accounts is carrying the result. Then give the decision-relevant sentence: the defensible range, whether its lower bound still clears the build cost, and what a proper staggered rollout would have bought.
Follow-up
  • The pre-period trends for adopters and non-adopters are not parallel. What do you report then?
  • You get to design the next rollout. What do you change so the same question is answerable, without randomising individual accounts?

Disagreeing with a proposed utilisation target using realisation evidence

medium
realisationguardrail metricsobservational evidence

A delivery lead proposes raising the billable utilisation target for analyst through senior_consultant from 72 to 85 percent. You have fct_time_entry, including is_billable, bill_rate_usd and written_off_hours, and fct_invoice_line. You believe the target will raise reported utilisation and lower fees. Prepare the disagreement: the evidence you pull, the mechanism you name, the metric pair you propose instead, and the condition under which you would concede that the target is correct.

Approach
  1. Name the probe: whether you disagree with a mechanism and a measurement, or with an opinion about a metric being bad.
  2. State the substitution precisely. Utilisation counts approved hours with is_billable = TRUE. An hour that is charged to the client and later written off stays in that numerator, so utilisation is unaffected while realisation, fees divided by hours times bill_rate_usd, falls and margin falls with it. That is the exact channel by which a higher target can raise the reported number and lower revenue.
  3. Pull the evidence at consultant-month grain: plot realisation and the write-off share, written_off_hours over billable hours, against utilisation decile. If the current top decile already shows lower realisation, the proposed target moves a large share of the staff into that regime.
  4. Stratify before concluding. Fixed_fee teams can show high utilisation and high realisation for reasons that have nothing to do with the proposal, so run the comparison within pricing_model and report the mix.
  5. Propose the pair rather than the veto: utilisation published with realisation and write-off rate as standing guardrails, with the threshold at which the combination is net positive stated in advance. Then name your concession condition: if the top utilisation decile shows no realisation penalty and bench hours are the binding constraint, the target is right and you will say so.
Follow-up
  • Utilisation and realisation are computed from overlapping hours. Does that make the relationship you found mechanical rather than behavioural?
  • How many consultant-months would you need to detect a three-point realisation move, and does the firm have them?
  • 01

    Tell me about a time when you had to persuade a stakeholder to adopt your recommendation.

  • 02

    How do you prioritize your work when facing multiple deadlines?

  • 03

    Describe a challenging team dynamic you encountered and how you resolved it.

  • 04

    What motivates you to work as a Data Scientist?

  • 05

    How do you stay updated with the latest trends in data science?

  • 06

    How would you explain a complex machine learning model to a non-technical audience?

PracHub interview preparation framework ↗
Is this an official 2nd Order Solutions interview guide?

No. It is PracHub's own preparation material for the Data Scientist role at 2nd Order Solutions. The rounds and questions reflect what candidates have reported, not a process the company has published, and they can change between teams and over time. Confirm the current format with your recruiter.

PracHub interview research ↗
How difficult are the interviews for the Data Scientist position?

Candidates describe the technical interviews and the case studies as the hardest parts. The practical response is to rehearse out loud: a modelling case from target to metric, a statistics explanation, a SQL query with narration, and a case that ends in a recommendation.

PracHub interview research ↗
How long does the process take?

Candidates report roughly three to five weeks across four rounds, and say they hear back within a few weeks of the final interview. This varies by team, so ask your recruiter for the schedule and plan your preparation to cover all four stages rather than only the first.

PracHub interview research ↗
Which tools should I be ready to use?

The requirements candidates describe list Python or R, SQL, and a visualisation tool such as Tableau or Power BI, with statistics and machine learning as core skills. Cloud platforms such as AWS or Google Cloud, big-data tools such as Hadoop or Spark, and deep learning or NLP are listed as extras. Be able to write SQL and one of Python or R live; treat the extras as a way to talk about past projects.

PracHub Data Scientist practice ↗
How should I structure a case study answer?

Candidates describe three moves: define the problem, explore the data to find the key variables, and generate insights as recommendations. In practice: restate the prompt as a decision, ask clarifying questions, name the quantity that would settle it, check data quality before analysis, state assumptions as you use them, and close with the action, supporting number and reversing result.

PracHub Data Scientist practice ↗
Will there be coding and system design questions?

Candidates report coding and algorithm questions (a mean and median function, a decision tree from scratch, sorting complexity, a SQL ranking query) and data system design topics such as streaming, model deployment, analytical databases and data quality. These appear as question types that may apply depending on the team, so prepare small, correct code and a clear checklist-style design answer.

PracHub Data Scientist practice ↗
Are remote or hybrid arrangements available?

Candidate reports do not confirm a specific policy; they say arrangements may vary by role. Ask your recruiter early, along with questions about the team, the data you would work with and how analyses reach decisions.

PracHub Data Scientist practice ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.