A Data Scientist at Gartner plays a pivotal role in translating complex, high-volume data into actionable insights that empower executive leaders worldwide to make mission-critical decisions. Unlike traditional technology firms where data science might focus solely on optimizing ad click-through rates or app engagement, Gartner leverages data science to fuel its industry-leading research, advisory services, and proprietary benchmarking platforms. You will work at the intersection of advanced analytics, machine learning, and business strategy to build models that extract deep meaning from vast repositories of structured and unstructured data.
In this role, your models and insights will directly influence products and tools used by Fortune 500 executives. Whether you are developing natural language processing (NLP) pipelines to analyze thousands of proprietary research documents, building predictive engines to forecast technology trends, or designing recommendation algorithms for client portals, your work will have a high-leverage impact. The data environment at Gartner is rich and intellectually stimulating, requiring a balance of rigorous scientific methodology and practical business application.
To succeed as a Data Scientist here, you must possess not only technical excellence but also a strong consultative mindset. You will collaborate closely with product managers, software engineers, and research analysts. The hiring team looks for candidates who can look past the math to understand the "why" behind a business problem, designing scalable machine learning solutions that directly address real-world client challenges.
Recruiter Phone Screen
reportedA screening call is a matching exercise run by someone who will not evaluate your statistics. They are checking that the work described on your resume is work you personally did, and that its scope matches the level the role is written for. Logistics get settled in the same half hour so nobody spends an interviewer's afternoon on a mismatch. The answer that fails is the one narrated in the plural. If every sentence is 'we built' and 'the team decided', there is nothing specific to write down about you. Name the piece that was yours, the decision you made inside it, and what changed after.
What to demonstrate
- Whether the ownership implied by your resume survives one round of follow-up about who actually did which part
- Whether your described scope (data size, stakeholders, what shipped) matches the seniority the role is written at
- Whether timeline, location and compensation expectations make the rest of the loop worth scheduling
How to prepare
- Rewrite your top three resume bullets in the first person singular, each with the decision you made and what moved afterwards, then say them out loud once so the 'we' does not return under pressure
- Attach one number to each project: the baseline, the change, and the window it was measured over. Where impact was never measured, say that plainly rather than inventing a figure
- Settle your compensation range before the call and give it as a range with a reason behind it, such as current total comp or a competing timeline, instead of deflecting the question twice
HR Technical & AI Screening
reportedBefore anything else, this round is a reading test. You are given a small schema and a question phrased in business language, and most of the difficulty sits in the gap between them. Who counts as an active user, does a refunded order still count as an order, is that date column an event time or a load time. Weak answers start typing immediately and compute something precise about the wrong population. Strong ones pin the definition in one sentence, name the column that encodes it, then write the query. On a timed assessment with nobody to tell, write the definition in a comment anyway.
What to demonstrate
- Whether an ambiguous term becomes a specific column and filter before any computation happens
- Whether you read the schema for keys and cardinality rather than only for column names
- Whether the result answers the question at the grain it was asked at, per user or per session or per day
How to prepare
- Take three metrics you already use and write down the exact filter and exact grain behind each, then practise stating one of them in a single sentence out loud
- On a schema you have never seen, spend the first minute writing what one row of each table means and which key it is unique on, then predict which joins can duplicate rows
- Rehearse a version where the definition changes halfway through, and edit the query you have instead of starting over
Data Science Director Interview
reportedBecause the format is not fixed, prepare the reasoning rather than the ritual. Nearly every version of this round draws on the same underlying material: a design you can defend, a metric you can define exactly, an analysis whose assumptions you can state out loud. Only the wrapper changes, whether that is a take-home, a live case, a deep dive on past work, or a rough estimate on a whiteboard. Answers rehearsed to fit one shape stall the moment the shape differs. Practise naming the assumption behind a number, then saying how much the conclusion moves if that assumption is wrong.
What to demonstrate
- Whether your justification for a method survives the question 'why not the simpler thing', including when the simpler thing would have worked
- Precision under pressure: what exactly counts as an active user, a conversion or a success, over what window, with what exclusions
- Whether you carry an argument through to a recommendation instead of stopping at a list of tradeoffs
How to prepare
- For each project you plan to mention, write the metric definition in one sentence: numerator, denominator, time window, exclusions. Say it out loud once, because vagueness shows up in speech before it shows up on paper.
- Rehearse the same project at three lengths: two minutes, ten minutes, and a deep dive on one technical decision. Cutting live is harder than it sounds.
- For your headline result, write down what would have had to be true for it to be wrong, and how you ruled that out.
Live Coding & Case Study Round
reportedA case round is decided by whether you leave the interviewer with a recommendation, not by how much analysis you narrate on the way there. The prompt is open on purpose, so the first job is to convert it into a decision someone could act on: ask what would be done differently depending on the answer. From there name the quantity that would settle it, state the assumptions you need, and commit. Candidates who cover more ground than anyone expected and still end on "it depends" score below candidates who scoped narrowly and said what they would do.
What to demonstrate
- Whether the version of the question you choose to answer is genuinely narrower than the prompt and still worth answering
- Whether the recommendation arrives as an action with a number attached, rather than as a summary of what you looked at
- Whether assumptions are stated at the moment you rely on them, instead of collected into a disclaimer at the end
- Whether you notice when a branch you are exploring would not change the decision either way
How to prepare
- Take six open prompts and write only the scoping move for each: the one-sentence question you would actually answer and the decision it feeds. Give yourself three minutes per prompt and stop there.
- Put a five-minute warning into every practice case and force a closing statement that names the action, the result that would justify it, and the result that would reverse it.
- Record one case and count how long you talked before naming a measurable quantity. Past roughly five minutes, what you are calling scoping is narration.
PracHub editorial advice for the preparation topics above.
Computing days-to-pay or proposal cycle time over completed records only
At any snapshot date, invoices that have already been paid are disproportionately the fast ones, and proposals that already have a decision are disproportionately the quick ones. Averaging over the completed set alone biases both numbers downward, and the bias grows exactly when the business is deteriorating, because the slow cases are the ones still open. Unpaid and undecided records are right-censored; use Kaplan-Meier or a restricted mean up to a fixed horizon, and never fill paid_at with a placeholder.
Trending utilisation or revenue on work_date without accounting for timesheet backfill
Time entries are created days to weeks after the work happens, and the backfill tail often runs two to six weeks. A dashboard keyed on work_date therefore shows the most recent weeks as a decline that reverses on every refresh. The fix is either to hold the reporting window back past the observed backfill tail (measure the tail with the timesheet submission lag metric rather than guessing) or to report an as-of-entered_at snapshot so the series is internally consistent, and to state which one you used.
Dropping rows with missing values without naming the mechanism
Say whether the values are missing at random, missing by a known process, or missing in a way that depends on the outcome, and handle them accordingly. Deleting incomplete rows silently redefines the population whenever missingness correlates with what you are measuring.
Optimising accuracy on a heavily imbalanced target
State the base rate first, then choose the metric from the relative cost of a false positive against a false negative: precision and recall at the operating threshold, PR-AUC, or expected cost. At a 1 percent positive rate, predicting the majority class for everyone scores 99 percent accuracy and is worthless.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
How do you handle highly imbalanced datasets when training a classific…
How do you handle highly imbalanced datasets when training a classification model?
Approach
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Set a baseline first, so any model has something honest to beat.
- Say how the offline result would be validated online before it is trusted.
Follow-up
- How would you choose the decision threshold, and who owns that choice?
- Where could label leakage enter this setup?
How would you collaborate with a Data Engineer to transition a prototy…
How would you collaborate with a Data Engineer to transition a prototype model from a local notebook into a production-ready pipeline?
Approach
- Set a baseline first, so any model has something honest to beat.
- Say how the offline result would be validated online before it is trusted.
- Check what information would not exist at prediction time, and exclude it.
Follow-up
- What would you monitor after launch to know the model is still valid?
- How would you choose the decision threshold, and who owns that choice?
Measure the timesheet backfill curve and pick a reporting cutoff
time_entries has work_date (date), entered_at (timezone-aware UTC timestamp), hours and status. Given a snapshot_date, restrict to work_date in [snapshot_date - 180 days, snapshot_date - 60 days] so every cohort is fully observed. For k = 0..45, compute F(k): the share of a work_date cohort's final hours that already existed as of work_date + k days, pooled across cohorts. Return the 46-point curve and the smallest k with F(k) >= 0.99. Some rows are entered before the work date; those lags are real, not errors.
Approach
- Compute lag = (entered_at converted to the reporting timezone and taken as a date) - work_date in whole days, then clip negative lags to 0 instead of dropping them; leave and planned time are routinely entered ahead of the work date and dropping them deflates the early curve.
- Take cohort totals as groupby(work_date).hours.sum() over the restricted window. These are final only because the window stops 60 days short of the snapshot, which is why the restriction is in the prompt.
- Build the numerator by summing hours per (work_date, lag), sorting by lag, taking a per-cohort cumsum, then reindexing each cohort onto the full 0..45 lag grid and forward-filling, so a cohort with no entries at a given lag holds its previous level rather than disappearing.
- Pool as sum(numerators) / sum(denominators) at each k, not as the mean of per-cohort shares. Holiday weeks are tiny cohorts and would otherwise carry the same weight as a full week.
- Read k* off the pooled curve and report F(45) with it: if F(45) is below about 0.995 the tail runs past the grid and k* is a lower bound, not the answer.
Worked solution 25 min
- Restrict rows to the [snapshot - 180d, snapshot - 60d] window and compute lag_days = (entered_at.dt.tz_convert(tz).dt.normalize().dt.date - work_date).dt.days, then lag_days = lag_days.clip(lower=0).
- cohort_total = df.groupby('work_date').hours.sum(); by_lag = df.groupby(['work_date','lag_days']).hours.sum().
- Reindex by_lag onto MultiIndex.from_product([cohorts, range(0,46)]), fill 0, cumsum within work_date to get hours_by_k.
- F = hours_by_k.groupby(level='lag_days').sum() / cohort_total[cohorts_in_grid].sum(); assert F is non-decreasing.
- k_star = int(F[F >= 0.99].index.min()) if any, else report 'not reached within 45 days' along with F(45).
Follow-up
- The dashboard refreshes daily. Would you hold the window back past k*, or publish an as-of-entered_at series instead, and what does each choice cost the reader?
- One practice area has a tail twice as long as the rest. Does that change the firm-wide cutoff, or does it change what you publish per practice area?
How would you optimize a slow-running SQL query that joins multiple la…
How would you optimize a slow-running SQL query that joins multiple large tables containing research document metadata?
Approach
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Say which table is the grain you start from, and join outward from it.
Follow-up
- How does the query change if the join becomes one-to-many?
- What breaks if events arrive late or out of order?
Write a Python function using Pandas to clean, merge, and aggregate tw…
Write a Python function using Pandas to clean, merge, and aggregate two disparate client datasets with mismatched timestamps.
Approach
- State the window function and its partition and ordering out loud before writing it.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
Follow-up
- How does the query change if the join becomes one-to-many?
- How would you verify this result without re-running the same query?
Rebuild proposal stage timeline from a mutable event stream
fct_proposal holds proposal_id, client_id, expected_value_usd, stage, is_competitive, created_at and decided_at, and its stage column is overwritten in place. fct_proposal_stage_event holds stage_event_id, proposal_id, from_stage, to_stage and occurred_at. For proposals created in the trailing year, rebuild the ordered stage timeline with days spent in each stage, flag durations still open at the snapshot as censored, and build a funnel of distinct proposals and expected value ever reaching qualifying, scoping, submitted and decided. Reconcile each timeline's terminal stage against fct_proposal.stage and report disagreements.
Approach
- Order events per proposal and derive each stage's exit with LEAD(occurred_at) OVER (PARTITION BY proposal_id ORDER BY occurred_at, stage_event_id). The tiebreak on the event key matters because two transitions can share a second. LEAD returns NULL on the last event of every proposal, decided or not, so a NULL lead marks the end of the stream and nothing more — it is not by itself a censoring indicator.
- Close that final interval by state rather than by the NULL. Where decided_at IS NOT NULL, close it at decided_at with is_censored = FALSE; where decided_at IS NULL, close it at snapshot_date with is_censored = TRUE. Only the second group is censored, and marking it lets a consumer choose Kaplan-Meier or a restricted mean. A median computed over completed transitions alone is biased downward, because at any snapshot the still-open records are disproportionately the slow ones, and the bias grows exactly when the pipeline is slowing. A decided_at earlier than the last event's occurred_at yields a negative final interval: report those proposals, do not clamp them.
- Handle re-entry. A proposal can go scoping to submitted and back to scoping, so per-stage entries outnumber proposals. Build the funnel on COUNT(DISTINCT proposal_id) for 'ever reached stage X', and keep the per-entry rows separate for duration work.
- Build the funnel from the reconstructed timeline, never from fct_proposal.stage. That column records only where a proposal came to rest, so every proposal that passed through scoping and moved on is invisible in it and the funnel's middle collapses.
- Reconcile the last to_stage per proposal against fct_proposal.stage. Disagreements mean events are missing or a row was edited outside the event path, and the size of that set is the ceiling on how far anything else here can be trusted.
- If asked for win rate by value, exclude stage IN ('withdrawn','no_decision') from both numerator and denominator and filter is_competitive = TRUE. Those outcomes are not losses, they are not missing at random, and sole-sourced follow-on work inflates the rate.
Worked solution 45 min
- Count events per proposal and check for proposals with zero events; those exist only in the mutable table and must be reported, not dropped.
- Add the LEAD-based exit timestamp with the event-key tiebreak, then close each final interval with COALESCE(lead_occurred_at, decided_at, snapshot_date) and set is_censored only where decided_at IS NULL, so days_in_stage is defined on every row.
- Aggregate to the funnel with COUNT(DISTINCT proposal_id) and SUM of expected_value_usd per 'ever reached' stage.
- Compute the median days in stage over completed transitions only, and report the censored count beside it so the reader sees what the median excludes.
- Take the last to_stage per proposal with a window function and diff it against fct_proposal.stage, returning the mismatched proposal_ids.
Follow-up
- Estimate median proposal cycle time with the censored proposals included. What does the estimator need from your output?
- Two events for one proposal share occurred_at to the second. How does your ordering resolve it, and how would you detect that it happened?
- Which fields on a won proposal are populated only after the decision, and why does that rule them out of a win-rate model?
How would you design a recommendation engine for the Gartner client po…
How would you design a recommendation engine for the Gartner client portal to suggest relevant research papers based on a user's reading history and job profile?
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- What would you do if the primary metric and the guardrail moved in opposite directions?
Explain the difference between Precision, Recall, and AUC-ROC. Under w…
Explain the difference between Precision, Recall, and AUC-ROC. Under what specific business scenarios would you prioritize one metric over the others?
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- State what result would change your recommendation, so the answer is falsifiable.
- Fix the population and the time window before naming any metric.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- Which segment would you cut first, and what would that rule out?
Describe how you would build a system to automatically categorize and …
Describe how you would build a system to automatically categorize and extract key technology trends from thousands of unstructured research articles.
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- State what result would change your recommendation, so the answer is falsifiable.
- Decompose the metric into the rates that drive it, and say which one you would check first.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- How would you detect that the metric is being gamed rather than genuinely improving?
Given an array of integers representing historical client engagement s…
Given an array of integers representing historical client engagement scores, write an efficient algorithm to find the longest contiguous subarray of increasing engagement.
Approach
- State what result would change your recommendation, so the answer is falsifiable.
- Fix the population and the time window before naming any metric.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- Which segment would you cut first, and what would that rule out?
What are the core differences between the responsibilities of a Data S…
What are the core differences between the responsibilities of a Data Scientist, a Data Engineer, and a Business/Data Analyst?
Approach
- Clarify what is being asked and what a complete answer would contain.
- State your assumptions explicitly before working the problem.
- Say what you would check first and why it is the highest-information step.
Follow-up
- What assumption would you test first?
- How would you know your answer was wrong?
Size a rate-card pilot when accounts are the cluster
The firm wants to pilot a revised rate card. Assignment happens at the client account: one rate card per account, every engagement on it inherits that card. The outcome is realisation rate per engagement, fees billed over the standard-rate value of approved billable hours. You have 120 eligible accounts averaging three engagements each inside the pilot window, a per-engagement standard deviation of 0.12, and an intraclass correlation of 0.35 within account. Leadership will change the card only if realisation moves by at least 2 percentage points. Two-sided alpha 0.05, 80% power. Say whether this pilot can answer the question.
Approach
- Write the minimum detectable effect for a two-sample difference in means with equal allocation: MDE = (z_0.975 + z_0.80) * sigma * sqrt(4/N), coefficient 2.802, and be explicit that N counts engagements while randomisation happens on accounts.
- Convert engagements into independent information with the design effect DEFF = 1 + (m - 1) * rho. At m = 3 and rho = 0.35 that is 1.70, so 360 engagements carry the information of roughly 212.
- Compute the MDE on the effective sample, compare it to the 2-point decision threshold, and express the shortfall as a sample-size ratio rather than a vague 'underpowered'.
- Test whether a longer window rescues the design. Effective observations per cluster are capped at 1/rho, so total effective sample asymptotes at 120/0.35 = 343 no matter how long the pilot runs. Duration cannot fix an ICC problem; only more clusters can.
- Price the real options: more accounts, variance reduction from a pre-period covariate, a larger accepted effect size, or a design that uses within-account variation instead of between-account randomisation.
Worked solution 20 min
- DEFF = 1 + (3 - 1) * 0.35 = 1.70; effective n = 360 / 1.70 = 211.8.
- MDE constant = z_0.975 + z_0.80 = 1.960 + 0.8416 = 2.802.
- MDE = 2.802 * 0.12 * sqrt(4 / 211.8) = 2.802 * 0.12 * 0.1374 = 0.046.
- Invert for the 2-point target: required effective n = 4 * (2.802 * 0.12 / 0.02)^2 = 1131, so accounts needed = 1131 * 1.70 / 3 = 641.
- Take m to infinity for the duration ceiling: effective n approaches 120 / 0.35 = 343, giving an MDE floor of 2.802 * 0.12 * sqrt(4/343) = 0.036.
Follow-up
- Revenue is concentrated in a handful of accounts. How does a value-weighted estimate change the power calculation, and which estimand does leadership actually want?
- The ICC of 0.35 came from last year's data on the same accounts. What would make that estimate optimistic, and how would you stress it?
- What breaks if you randomise at the engagement instead of the account?
Firm margin rose while every pricing model's margin fell
Firm-wide engagement gross margin rose from 31% to 34% quarter over quarter. Split by fct_engagement.pricing_model, margin fell inside time_and_materials, fixed_fee, retainer and outcome_based. Sources are fct_invoice_line (engagement_id, line_type, amount_usd, period_start_date, period_end_date) and fct_time_entry (engagement_id, hours, cost_rate_usd, status). Explain the aggregate move, quantify how much of the 3-point change is mix and how much is within-model, and state which number you put in front of practice leadership and why.
Approach
- Write the aggregate as a weighted mean: M = sum over pricing models p of w_p * m_p, where m_p is margin within p and w_p is p's share of the fee base. The weights must be the fee base, because margin is a fees-weighted ratio; weighting by engagement count answers a different question.
- Apply the exact three-term decomposition: dM = sum (w_p1 - w_p0) * m_p0 + sum w_p0 * (m_p1 - m_p0) + sum (dw_p)(dm_p). Verify the three components sum to the reported +3.0 points before interpreting any of them.
- Identify which weights moved and in which direction; the aggregate can only rise this way if fee share shifted toward the structurally higher-margin models, typically fixed_fee delivered under budget or retainer.
- Guard against composition inside each stratum: repeat the same decomposition within pricing_model by practice_area and account_tier, so a within-model decline is not itself an unexamined mix effect.
- Confirm both sides of the ratio are on the same clock and the same scope: numerator includes line_type IN ('fees','milestone','credit_note') with credits signed negative; the cost side includes all approved delivery hours, billable and non-billable alike.
- Report within-model margin as the headline delivery signal and mix as a separate, named line with its own owner, since mix is a sales decision and margin is a delivery one.
Follow-up
- Next quarter the mix reverts. What happens to the aggregate, and what will you have already told leadership so this is not a surprise?
- Which of the four pricing models should not be pooled with the others even for a mix-adjusted number, and why?
- How would you present this if the mix shift were deliberate strategy rather than accident?
For a candidate whose interviews will centre on A/B testing, metric movement and causal claims. Design comes before arithmetic, arithmetic before analysis, and the week ends by rehearsing the readout rather than the derivation.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Design one test end to end on paper
- Take a single feature change and write the full design: randomization unit, the exact point of exposure, the primary metric with its grain, guardrails, allocation, planned duration, and the decision rule committed before any data exists.
- Write why the randomization unit must sit at or above the level where treatment can spill over, and give one case where user-level randomization is still contaminated (shared accounts or devices, or two participants in the same marketplace).
- State in advance what you will do if the primary metric is flat while a secondary metric is significant.
Deliverable: A one-page test design with a decision rule written before launch.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Power arithmetic until it is automatic
- Compute required sample size per arm for a binary metric with the normal approximation, n is approximately 2 times (z for alpha/2 plus z for power) squared times p(1 minus p) divided by delta squared, for baselines of 2, 10 and 40 percent at a 5 percent relative lift, and note that for a fixed relative lift the requirement falls as the baseline rises because delta grows proportionally with p.
- Redo the calculation for a continuous metric using variance in place of p(1 minus p), and show why a heavy-tailed quantity such as revenue per user needs either far more traffic or a capped version with a stated cap.
- Convert one of the results into weeks given a weekly eligible traffic figure, then list the two honest ways to shorten it (accept a larger detectable effect, or reduce variance) and write why quietly lowering the power target is a decision to miss more real wins, not a speedup.
Deliverable: A small script or sheet that maps baseline, minimum detectable effect, alpha and power to sample size and weeks, cross-checked against a published calculator.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Variance and the unit-of-analysis problem
- Take a ratio metric whose denominator is not the randomization unit (clicks per session, randomized by user) and compute the standard error twice, once naively at session level and once by the delta method or a user-level bootstrap, then record how much the naive version understates it.
- Implement CUPED on simulated data: choose a pre-period covariate X measured before assignment, estimate theta as Cov(Y, X) divided by Var(X), and analyse Y minus theta times (X minus its mean) in place of Y. Confirm the variance of the adjusted outcome equals the raw variance multiplied by one minus the squared correlation between Y and X, so a correlation of 0.45 removes about 20 percent of the variance and not 80.
- Now run that simulation a few hundred times and confirm the adjusted effect estimate is unbiased for the same effect rather than numerically identical to the raw one. Within any single run the two differ, sometimes by a large fraction of the true effect, because the two arms' pre-period covariate means never coincide exactly in a finite sample; they agree in expectation, which is the property that matters and the one to state out loud.
Deliverable: A notebook showing the adjusted estimator with a measurably smaller variance than the raw one, plus a repeated-simulation table showing the two estimators agreeing on average while differing run by run.
Practice prompt ↗Practice prompt ↗04Validity threats you can actually test for
- Run a sample ratio mismatch check as a chi-square goodness-of-fit test against the intended allocation, and write the three causes you would chase first (assignment logged before exposure, an arm-specific redirect or load failure, bot filtering applied asymmetrically).
- Simulate peeking: generate A/A data, test daily at alpha 0.05 across 14 looks, record the inflated false positive rate, then apply an alpha-spending boundary or commit to a fixed horizon and confirm the rate returns to nominal.
- Write how you would separate a novelty effect from a durable lift using the treatment effect plotted against days since first exposure, and what shape would change your recommendation.
Deliverable: One table showing the peeking false positive rate before and after correction, plus a written SRM triage list.
Practice prompt ↗Practice prompt ↗Worked solution ↗05When randomization is not available
- Write the identifying assumption for difference-in-differences (parallel trends in the absence of treatment), then plot pre-period trends for two candidate control groups and justify rejecting one of them.
- Design a switchback test for a change where user-level randomization would leak across participants, choosing a time-block length against the carryover you expect and saying how you would detect carryover in the data.
- List what an interrupted time series or a synthetic control buys you and the one thing neither can rule out: an unobserved shock that coincides with the launch.
Deliverable: A one-page memo recommending a single quasi-experimental design and naming its weakest assumption explicitly.
Practice prompt ↗Practice prompt ↗06The readout query
- Write the assignment-to-exposure join that returns exactly one row per unit per experiment, and handle units appearing in both arms by excluding and counting them rather than silently keeping one.
- Compute the per-arm metric, its variance and the relative lift with a confidence interval in SQL, then reproduce the identical numbers in a notebook as a cross-check.
- Add a segment breakdown and write the sentence that keeps it from being p-hacking: segments declared in advance, everything else reported as exploratory and corrected for multiplicity.
Deliverable: A single query that outputs the full readout table, matched to a notebook recomputation.
Practice prompt ↗Practice prompt ↗07Present it to someone who will not read the appendix
- Give a 10-minute readout of a real or simulated experiment in the order decision, number, uncertainty, caveat.
- Have your listener ask "can we ship it" in the case where the primary is flat and a guardrail moved, and answer with a recommendation rather than a request for more data.
- Rewrite your opening line so the recommendation lands before any methodology.
Deliverable: A one-page readout whose first line is the recommendation.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Most of the questions in this section reduce to one thing: can you be handed a vague request and come back with something useful? Prepare an example where the ask was underspecified, you chose an interpretation, and you said out loud which interpretation you chose. Describing how you narrowed the question matters more than the technique you eventually used.
If a business stakeholder asks for a descriptive dashboard versus a pr…
If a business stakeholder asks for a descriptive dashboard versus a predictive model, how do you determine which solution is more appropriate for their needs?
Approach
- State the situation in two sentences and spend the rest on your reasoning.
- Quantify the outcome, including what you would not claim credit for.
- Pick a story where you drove the decision, not one where you observed it.
Follow-up
- What did you decide not to do, and why?
- How did you know the outcome was caused by your change?
Disagreeing with a proposed utilisation target using realisation evidence
A delivery lead proposes raising the billable utilisation target for analyst through senior_consultant from 72 to 85 percent. You have fct_time_entry, including is_billable, bill_rate_usd and written_off_hours, and fct_invoice_line. You believe the target will raise reported utilisation and lower fees. Prepare the disagreement: the evidence you pull, the mechanism you name, the metric pair you propose instead, and the condition under which you would concede that the target is correct.
Approach
- Name the probe: whether you disagree with a mechanism and a measurement, or with an opinion about a metric being bad.
- State the substitution precisely. Utilisation counts approved hours with is_billable = TRUE. An hour that is charged to the client and later written off stays in that numerator, so utilisation is unaffected while realisation, fees divided by hours times bill_rate_usd, falls and margin falls with it. That is the exact channel by which a higher target can raise the reported number and lower revenue.
- Pull the evidence at consultant-month grain: plot realisation and the write-off share, written_off_hours over billable hours, against utilisation decile. If the current top decile already shows lower realisation, the proposed target moves a large share of the staff into that regime.
- Stratify before concluding. Fixed_fee teams can show high utilisation and high realisation for reasons that have nothing to do with the proposal, so run the comparison within pricing_model and report the mix.
- Propose the pair rather than the veto: utilisation published with realisation and write-off rate as standing guardrails, with the threshold at which the combination is net positive stated in advance. Then name your concession condition: if the top utilisation decile shows no realisation penalty and bench hours are the binding constraint, the target is right and you will say so.
Follow-up
- Utilisation and realisation are computed from overlapping hours. Does that make the relationship you found mechanical rather than behavioural?
- How many consultant-months would you need to detect a three-point realisation move, and does the firm have them?
Three requests, one week, and the one you defer
Three requests arrive in one week. Finance needs days sales outstanding recomputed for a board meeting in four days. A partner wants a proposal win-rate model for a pursuit review in three weeks. Delivery wants a staffing forecast, with no date attached. You have one week of your own capacity and no analyst. Write the prioritisation you send back, the request you defer, and the message you send to the person whose request you defer.
Approach
- Name the probe: whether you prioritise on decision dates and reversibility, or on who asked most forcefully.
- Score each request on four things: what decision it changes, the date that decision is made, the cost of being late, and the cost of being wrong. A board figure has a hard date and a high cost of being wrong; a forecast with no date has neither.
- Reduce scope rather than dropping work. The win-rate model is the largest piece and the easiest to get wrong, because features written after the decision, such as engagement_id and revised pricing, leak the label, and because withdrawn and no_decision proposals are not missing at random. A two-day descriptive win-rate cut by is_competitive and loss_reason answers most of what a pursuit review needs, with the model scoped separately.
- Defer explicitly, with a date and a smaller substitute, rather than leaving a request to decay quietly. Silence is read as agreement and then as failure.
- Put the trade-off in writing so it can be overturned by someone with more context than you have, and say what you would drop if the deferred request becomes urgent.
Follow-up
- The partner escalates to the practice lead. What do you change, and what do you refuse to change?
- What evidence would make you drop the finance work instead?
- 01
If a business stakeholder asks for a descriptive dashboard versus a predictive model, how do you determine which solution is more appropriate for their needs?
- 02
A delivery lead proposes raising the billable utilisation target for analyst through senior_consultant from 72 to 85 percent. You have fct_time_entry, including is_billable, bill_rate_usd and written_off_hours, and fct_invoice_line. You believe the target will raise reported utilisation and lower fees. Prepare the disagreement: the evidence you pull, the mechanism you name, the metric pair you propose instead, and the condition under which you would concede that the target is correct.
- 03
Three requests arrive in one week. Finance needs days sales outstanding recomputed for a board meeting in four days. A partner wants a proposal win-rate model for a pursuit review in three weeks. Delivery wants a staffing forecast, with no date attached. You have one week of your own capacity and no analyst. Write the prioritisation you send back, the request you defer, and the message you send to the person whose request you defer.
Is this an official Gartner interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Gartner. Rounds and questions reflect what candidates have reported, not a process Gartner has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult is the Data Scientist interview process at Gartner?
The interview process is generally rated as average to difficult. While the theoretical machine learning and fundamental coding questions are standard, the process becomes challenging due to the high emphasis on structured problem-solving, role boundaries, and your ability to explain complex technical concepts to senior directors.
PracHub interview research ↗What is the typical timeline from the initial application to an offer?
The entire process typically takes between 3 to 6 weeks. The HR team is highly communicative and structured, ensuring that candidates are kept informed of their status and next steps after each round.
PracHub interview research ↗How deeply should I study coding algorithms versus machine learning theory?
You need a balanced preparation. You must be able to write clean, efficient code for data manipulation (using Pandas/SQL) and basic algorithmic challenges, but you will also face detailed conceptual questions on machine learning metrics, model validation, and system boundaries.
PracHub interview research ↗Does Gartner support remote or hybrid working arrangements for Data Scientists?
Gartner typically operates under a hybrid model, combining remote work flexibility with structured in-office collaboration days. Specific arrangements depend heavily on the team, office location, and regional guidelines.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22