As a Data Scientist at BlackRock, you are stepping into a pivotal role at the intersection of advanced technology and global finance. BlackRock is the world’s largest asset manager, and data is the lifeblood of its investment strategies, risk management protocols, and client solutions. In this role, you will be leveraging massive, complex datasets to uncover alpha, optimize portfolios, and build predictive models that directly influence billions of dollars in assets.
Your impact will extend across various high-stakes products and teams, most notably within the Aladdin ecosystem—BlackRock’s industry-leading investment and risk management platform. You will build machine learning models to forecast market trends, natural language processing pipelines to parse financial reports, and optimization algorithms to balance risk and reward. The work you do scales globally, empowering portfolio managers, quantitative analysts, and institutional clients to make data-driven decisions with confidence.
Expect a highly collaborative, fast-paced environment where technical rigor meets deep financial intuition. You will not just be writing code; you will be solving some of the most complex, ambiguous problems in the financial sector. This role requires a unique blend of mathematical excellence, engineering proficiency, and the strategic foresight to understand how macroeconomic factors translate into actionable data insights.
Initial Recruiter Screen
reportedMost candidates lose this call inside the first two minutes, during the walkthrough of their own background. The account runs chronologically, sits at the level of tools and titles, and never arrives at a decision anyone could have disagreed with. Anchor on a problem instead of a timeline: what the team could not answer, what you did about it, what happened next. Ninety seconds is enough, and stopping on time leaves room for the half of the call that belongs to you. What you ask about how work gets prioritised signals your level more reliably than the walkthrough does.
What to demonstrate
- Whether your background summary has a shape (problem, decision, consequence) or is a chronological list of tools and employers
- Whether you can account for gaps, short stints and the reason you are looking, unprompted and without hedging
- The substance of the questions you ask back, which an experienced screener reads as a level signal
How to prepare
- Time your opening walkthrough against a clock. If it runs past two minutes, compress the earliest role into a single clause and spend the recovered time on the most recent one
- Write one honest sentence for every gap or short stint visible on your resume and offer it before being asked about it
- Prepare questions about how work arrives and gets prioritised: who writes the request, how often priorities change, and what happens to an analysis after it is delivered
Technical Screening Round
reportedMuch of what gets scored here happens out loud while you type. Nobody can see your reasoning inside a half-written query, so five silent minutes read as being stuck even when they are not. State the plan in plain language first: which tables, what grain you are aggregating to, and the one filter that defines the population. Then write it. The narration doubles as insurance, because a wrong plan gets caught early and cheaply while a wrong query gets caught at the end with no time left to redo it. A timed statistics section, where one exists, is a separate test with its own clock.
What to demonstrate
- Whether the query you write matches the plan you just described
- What you do with a hint, meaning whether the correction gets absorbed or the first approach gets defended
- Whether you can debug your own wrong output by reading the result set and naming which part of the query produced the anomaly
How to prepare
- Solve three problems while screen-sharing into a recording, then watch it back and mark every stretch longer than thirty seconds where you said nothing
- Practise compressing the plan into one sentence before typing, then check afterwards whether the finished query actually matched it
- Time yourself on statistics questions that carry a business reading, such as what a confidence interval does and does not claim, rather than re-reading notes without a clock
Onsite/Virtual Final Rounds
reportedWhere a loop ends with a senior leader, that conversation is rarely another skills test. The technical signal already exists by then, so the questions tend to open up: what you would look at first, where a metric you have heard about could mislead, what you would push back on. The decision being made is scope, which in practice means level and how much you would be trusted to own unsupervised. Treating it as a formality is the usual mistake. An open question late in the day is still being scored, and a vague answer reads as someone who has not run anything themselves.
What to demonstrate
- Whether your view of the business has anything specific behind it, given that you are working only from what is public and are expected to say so
- Whether the scope of work you describe owning matches the scope of the role, instead of sitting a level below it
- Whether you can disagree with something concrete and stay useful about it, rather than agreeing with everything said in the room
- Whether your questions are ones only this person could answer, as opposed to ones the recruiter already covered
How to prepare
- Build one view you could defend for two minutes using only public information: what the funnel probably looks like, which metric likely drives decisions, and where that metric could mislead. Being wrong for a stated reason survives this round; having no view does not
- Write down the largest piece of work you have owned from question to decision, who else touched it, and what you decided alone, then check that it reads at the level you are interviewing for
- Prepare one thing you would want changed if you joined and phrase it as a question rather than a verdict, so it opens a conversation instead of closing one
PracHub editorial advice for the preparation topics above.
Computing a t-statistic on daily observations of an h-day forward return as if the observations were independent.
Sampling an h-day forward return every day means consecutive observations share h-1 days of the same return, which induces strong positive autocorrelation. The naive standard error is too small by a factor on the order of sqrt(h), so a 5-day-horizon signal with a genuine t of 1.3 can present as 2.9. Either use non-overlapping samples, which costs power, or use a Newey-West or Hansen-Hodrick covariance with at least h-1 lags, and state which one was used.
Filtering on as_of_date rather than knowledge_ts, so restated fundamentals, revised index constituents and retroactively applied split and dividend adjustments enter the backtest before they were knowable.
Vendors overwrite history in place. A quarterly figure filed 45 days after period end is stored against period end, an index addition announced five business days before it takes effect is stored against the effective date, and a split applied tonight rewrites every prior close in the adjusted series. Each of those gives the strategy information it could not have had, and the resulting lift is concentrated in the highest-turnover, highest-apparent-alpha names. The signal_score table separates the two timestamps precisely so this filter can be written correctly.
Ending an analysis without a recommendation or next step
Close with what you would do and what would change your mind, stated as a condition you can check later. If the evidence is genuinely inconclusive, recommend the specific next measurement and say what it costs in time or exposure.
Reading an observational correlation as a causal effect
Name the confounder you are most worried about and the design that would remove it: an experiment, a difference-in-differences with a checked pre-period trend, an instrument, or a regression discontinuity. When none is available, state which direction the bias likely runs and bound the claim accordingly.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Permutation test for mean IC with overlapping forward returns
panel has as_of_date, instrument_id, zscore_xs and fwd_ret_5d_excess, the five-trading-day forward return in excess of the universe's cap-weighted mean, for 252 dates and roughly 1,500 instruments a date. Compute the daily Spearman rank IC and its mean. Then build a permutation null from scratch and report a two-sided p-value. Compare it with a naive t-test on the 252 daily ICs and with a Newey-West t using four lags. State which null hypothesis your permutation scheme actually tests.
Approach
- Compute the daily IC by ranking both columns within as_of_date with method='average' and taking the Pearson correlation of the rank vectors. Decide explicitly how to treat dates where the universe shrinks, because an unweighted mean gives a 40-name date the same weight as a 1,500-name date.
- The obvious permutation shuffles zscore_xs within each as_of_date. That is the right exchangeability for the null 'the signal carries no cross-sectional ordering on any given day', and it preserves the daily sample size and the cross-sectional return structure, factor effects included.
- But it makes the permuted daily ICs independent across dates, and the real ones are not. A five-day forward return sampled every day shares four days with its neighbour, and the signal itself is persistent, so the observed IC series is strongly autocorrelated. Under an i.i.d. daily return null the variance of the mean of overlapping h-period observations inflates by a factor of h, so the within-day null is too narrow by about sqrt(5) and its p-value is anti-conservative.
- Fix it by permuting at the date level: circularly shift the whole signal panel's as_of_date labels by a random offset, keeping each day's cross-section and the signal's serial persistence intact while breaking its alignment with the returns. The 252 distinct shifts give an exact randomization test with a p-value floor of 1/252.
- Report all three numbers. The naive t and the Newey-West t with 4 lags (h-1 for h = 5) should differ by a factor near sqrt(5), and the shift-permutation p should agree with the Newey-West figure rather than the naive one. If it does not, one of the two is implemented wrongly and you have a cheap way to find out which.
Follow-up
- With 252 circular shifts and the observed statistic the largest among them, what is the smallest p-value you can report, and is it small enough for the decision being made?
- How would you redo this on non-overlapping five-day samples, and how much power do you surrender?
- The signal has a 12-day decay half-life. Does that change your shift block length or your Newey-West lag choice?
How large a Sharpe does pure noise produce over many trials
A research team tested 250 strategy variants on the same three years of daily returns (756 observations). The best variant has an annualized Sharpe of 1.4. Under the null that every variant has zero expected return, estimate by simulation the probability that the maximum of 250 Sharpe estimates is at least 1.4. Then repeat with the variants' daily returns equicorrelated at rho = 0.7, which is closer to the truth when variants share a universe. Report both probabilities with Monte Carlo error and say which one belongs in the memo.
Approach
- Simulate at the frequency the statistic is estimated at: 756 daily draws per variant, not three annual ones. The Sharpe is scale-free, so standard normal draws suffice and the choice of sigma cannot change the answer.
- Build the correlated case as X = sqrt(rho)*Z0 + sqrt(1-rho)*E, with Z0 one common daily draw shared by all variants and E independent per variant. That is exact equicorrelation for one extra column, rather than a 250x250 Cholesky per replication.
- Per replication compute all 250 annualized Sharpes as mean/std(ddof=1)*sqrt(252) along the time axis, take the maximum, and count how often it clears 1.4. Use at least 20,000 replications so the Monte Carlo standard error on a probability near 0.85 is about 0.0025.
- Check the independent case analytically before trusting the simulation: 1 - Phi(z)^250 with z = 1.4/SE and SE = sqrt(252/756) = 0.577 gives z = 2.43 and p close to 0.85. The simulation should land inside two Monte Carlo standard errors of that.
- Get the direction of the correlation effect right. Correlated variants behave like fewer independent trials, so the null maximum is smaller and an observed 1.4 becomes less likely under the null, not more. Present the correlated p-value as the smaller, more favourable number and state plainly that it depends on an assumed rho you did not measure.
Follow-up
- The team says it only ran six configurations because it discarded the rest early. How do you count trials that were abandoned after somebody looked at the result?
- What Sharpe would the best of 250 have to reach for you to call it significant at 5%, and is that number attainable at this strategy's turnover?
- How would you carve out a holdout the search has genuinely not touched, given the team has already seen the full sample?
Block bootstrap confidence interval for an annualized Sharpe
daily_pnl has business_date, strategy_id and net_return: five strategies, 1,260 daily observations each, returns net of all costs. Build a 95% confidence interval for each strategy's annualized Sharpe using a moving-block bootstrap you write yourself, with no library resampler. Choose the block length from the data and justify it. Report each interval beside the i.i.d. normal approximation, sqrt((1 + SR_daily^2/2)/T) scaled by sqrt(252), and say which strategies the two methods disagree about and why.
Approach
- Measure the dependence before resampling: compute lag-1 through lag-20 autocorrelation of net_return per strategy. The block bootstrap only earns its cost where that is non-zero, and on a strategy where it is flat the two intervals should agree, which is your implementation check.
- Form the n - L + 1 overlapping blocks as a strided view of the return array, draw ceil(n/L) block starts with replacement, concatenate and truncate back to n. Recompute the annualized Sharpe on each resample; 5,000 resamples is enough for a 95% percentile interval.
- Start at L near n^(1/3), about 11 for n = 1,260, then tabulate interval width against L over 5 to 30. A width still climbing at L = 30 means the dependence outruns the block and the interval is still too narrow, which is information about the strategy, not a bug.
- Take either the percentile interval or the basic (reverse-percentile) interval 2*theta_hat minus the quantiles, and say which. The block bootstrap distribution centres on the sample statistic, so the two differ whenever the resample distribution is skewed, and for a ratio it is.
- Cross-check against the closed form. For an AR(1) with coefficient rho, the long-run variance of the mean inflates by (1+rho)/(1-rho), so the interval should widen by about the square root of that. It is one line of arithmetic that tells you whether the block machinery is doing what the autocorrelation says it should.
Worked solution 35 min
- r = df.loc[df.strategy_id == s, 'net_return'].to_numpy(); n = r.size; sr = r.mean()/r.std(ddof=1)*sqrt(252).
- blocks = np.lib.stride_tricks.sliding_window_view(r, L), shape (n-L+1, L).
- Per resample: idx = rng.integers(0, n-L+1, size=ceil(n/L)); x = blocks[idx].ravel()[:n]; store its annualized Sharpe.
- ci = np.percentile(boot_sr, [2.5, 97.5]); repeat for L in {5, 11, 21, 30} and tabulate the widths.
- Put the i.i.d. approximation, the block interval and the lag-1 autocorrelation in one table per strategy.
Follow-up
- One strategy holds corporate bonds marked with mark_source = 'vendor_eval'. What does mark smoothing do to the lag-1 autocorrelation, to the annualized Sharpe itself, and to which of your two intervals you believe?
- Would you use the same block length to bootstrap maximum drawdown? What breaks?
- How does the answer change if you bootstrap 60 non-overlapping 21-day blocks instead of overlapping daily blocks?
Deduplicate a restated signal table into a leak-free panel
signal_score may hold several rows per (signal_id, instrument_id, as_of_date) because reruns rewrite history; each row carries model_version, knowledge_ts, computed_at, is_backfilled, zscore_xs, decile_rank, coverage_flag and universe_id. Build the panel a backtest may legitimately use on trading date D: for every instrument in universe_id = 'liquid_us_1500', the single most recent score whose knowledge_ts is at or before D 21:00 UTC, preferring a live row over a backfilled one. Return instrument_id, as_of_date, zscore_xs, decile_rank.
Approach
- Filter on
knowledge_ts, never onas_of_date.as_of_datesays what the score describes;knowledge_tssays the earliest instant every input was observable. Put it in the WHERE so it prunes before the window is evaluated. - Rank within the key:
ROW_NUMBER() OVER (PARTITION BY signal_id, instrument_id ORDER BY as_of_date DESC, is_backfilled ASC, knowledge_ts DESC, computed_at DESC)and keep rank = 1. In Postgres FALSE sorts before TRUE, sois_backfilled ASCis the live-row preference, written down rather than assumed. - Prefer ROW_NUMBER to a
MAX(computed_at)group-then-rejoin: the rejoin duplicates rows whenever two reruns share acomputed_at, which is exactly what a batch job produces. - Decide what
coverage_flag IN ('stale','imputed')means for this panel and encode it. Dropping those rows silently changes the universe size day to day, which surfaces later as an unexplained jump in measured IC rather than as a missing-data problem. - Express the cutoff as a parameter and the whole thing as a CTE so the backtest loop reuses one query text per date instead of a hand-edited copy.
Follow-up
- Write the query that proves the panel is leak-free. What would a violation look like in measured IC, and roughly how large would you expect the inflation to be?
- Produce every trading date in one pass instead of one query per date. What does that cost in plan shape?
- Two
model_versionvalues are live at once during a migration. How does the tie-break change, and who decides?
Sessionise child fills into execution bursts across venues
execution_fill holds fill_id, order_id, exec_ts TIMESTAMPTZ(6) in UTC, venue_mic, fill_qty, fill_px and liquidity_flag. Within each order_id, a burst is a maximal run of consecutive fills separated by under 90 seconds. Return one row per burst: order_id, burst sequence, first and last exec_ts, fill count, total quantity, quantity-weighted price, and the running share of parent_order.order_qty completed at the burst's end. Attach the exchange-local session date derived from instrument_master.primary_exchange_mic, not the UTC date.
Approach
- Order fills inside each parent and take
LAG(exec_ts) OVER (PARTITION BY order_id ORDER BY exec_ts, fill_id). Thefill_idtie-break is load-bearing: an algo slicing into a venue produces genuine microsecond ties, and without a deterministic second key the burst boundaries move between runs. - Flag a boundary where the lag is NULL or the gap exceeds 90 seconds, then convert flags to a dense burst id with
SUM(is_new_burst) OVER (PARTITION BY order_id ORDER BY exec_ts, fill_id ROWS UNBOUNDED PRECEDING). One pass, no self-join, and it degrades gracefully on a single-fill order. - Aggregate to burst grain with
SUM(fill_qty * fill_px) / SUM(fill_qty), notAVG(fill_px); averaging prices weights a 100-share child the same as a 50,000-share block and misstates the burst by the spread. - Add the completion curve as a second window over the aggregated bursts:
SUM(burst_qty) OVER (PARTITION BY order_id ORDER BY burst_seq ROWS UNBOUNDED PRECEDING) / order_qty. Doing it after aggregation keeps it at the grain you are reporting. - Resolve the session date through a MIC-to-IANA-zone lookup and
(exec_ts AT TIME ZONE zone)::date. The UTC date happens to coincide with the local session date for US cash equities and most Asian cash sessions, but an Australian morning sits on the previous UTC day, and a futures trade date that rolls at 17:00 US Central does not align with either.
Worked solution 35 min
- Write the LAG and boundary flag, then the running SUM that produces burst_seq.
- Aggregate to (order_id, burst_seq) with quantity-weighted price and first/last exec_ts.
- Join
parent_orderfororder_qtyand add the cumulative completion share window. - Join
instrument_masterforprimary_exchange_mic, map to a zone, and compute the local session date. - Verify the islands: every gap between one burst's end and the next burst's start must exceed 90 seconds.
Follow-up
- Two venues report the same execution microsecond with different
received_ts. Which timestamp do you sort by for sessionisation, and which for latency analysis? - Derive the 90-second threshold from the data instead of asserting it. What distribution would you look at, and what would tell you the threshold is wrong?
- How does the burst structure change under a POV algo versus a close-auction order, and what would you expect the completion curve to look like in each?
Design a firm-level client-outcome metric that survives audit
The firm wants one number reported monthly to the board for whether clients are getting what they were sold. The proposal is: AUM held in accounts whose trailing 36-month net-of-fee return beats their own benchmark_id, divided by AUM in accounts with at least 36 months of history. Using account_mandate (benchmark_id, inception_date, funded_date, close_date, status, mgmt_fee_bps, perf_fee_rate, high_water_mark_base, effective_from, effective_to) and position_daily, critique it, name three ways it rises without any client being better off, and specify the guardrails and the exact denominator you would publish.
Approach
- Attack the denominator first, because that is where this metric is won or lost. Requiring 36 months of history excludes accounts that closed during the window, and accounts close disproportionately after bad performance. Fix it by including accounts whose close_date falls inside the window, measured to their close, and carrying them for 36 months afterwards, so the series cannot be improved by attrition.
- Name the three inflation paths concretely. First, survivorship through the close_date exclusion. Second, benchmark choice: benchmark_id is per account and amendable, and account_mandate is SCD2, so an amendment silently restates history unless the metric evaluates each date against the benchmark in force via effective_from and effective_to, with a count of amendments published beside the ratio. Third, AUM-weight concentration: one large mandate can carry the number, so publish the effective number of accounts as the inverse Herfindahl of AUM weights next to it.
- Set the guardrails from the spine: cross-sectional dispersion of trailing 12-month net return within each strategy composite, and trailing 12-month dollar redemption rate. Dispersion catches the case where the headline is met by a favoured subset, and the redemption rate catches the attrition path from the client's side of the relationship.
- Fix the numerator's arithmetic. Net of fee means management fee accrued plus performance fee crystallised, and under a high-water mark with a hurdle two accounts in one strategy legitimately pay different fees, so compute per account from its own terms rather than deducting a composite fee. Then decide the return convention deliberately: time-weighted answers whether the manager delivered, money-weighted answers whether the client made money, and publishing only one answers only one question.
- State the residual honestly. Even a clean version measures relative return, not whether the product matched the client's purpose. Pair it with a small set of hard flags from the same tables, such as breaches of max_tracking_error_bps or max_single_name_weight_pct and recon_status values other than 'matched', rather than pretending one ratio covers the question.
Worked solution 45 min
- Build the account-month panel from account_mandate with SCD2 ranges resolved, tagging each account-month with the benchmark_id in force that month rather than the current one.
- Compute per-account trailing 36-month net-of-fee time-weighted return and the matching benchmark return over the same dates, including accounts closed inside the window up to their close_date.
- Compute the headline ratio under three denominators, surviving accounts only, surviving plus closed-to-date, and surviving plus closed with a 36-month tail, and tabulate the gaps.
- Compute the guardrails: within-composite standard deviation of trailing 12-month net return, trailing 12-month dollar redemption rate, and the effective number of accounts as one over the sum of squared AUM weights.
- Recompute the headline as of a month six months in the past and compare it to what was published then; any drift is a restatement and needs a named cause.
Follow-up
- An account funded 20 months ago has beaten its benchmark throughout and is excluded by the 36-month rule. Is that the right treatment, and what would including it cost you?
- Two accounts in the same composite differ by 180 basis points over 12 months. List the legitimate causes before calling it an error.
- How does the published ratio behave in a month when the firm funds one very large new mandate, and what should the board see alongside it?
Cut short-side financing cost without cutting the short alpha
Management wants short-side financing cost down. From position_daily you have financing_cost_base_day, borrow_rate_bps, borrow_status ('general_collateral', 'hard_to_borrow', 'unavailable', 'n_a'), quantity and market_value_base per instrument per business_date; from signal_score you have decile_rank per as_of_date. Design the metric set for this project: a primary, a guardrail, and the decision rule when they disagree. The complication is that the most negative decile is disproportionately hard to borrow, so the cheapest achievable borrow book is the one that has stopped taking the trade the strategy exists to take.
Approach
- Refuse total financing_cost_base_day as the primary: it is minimised perfectly by closing the short book. Denominate it instead as short-notional-weighted borrow_rate_bps, so the metric measures the price paid per unit of short exposure rather than the quantity of exposure taken.
- Benchmark that rate against a broad market borrow rate. A firm-wide easing of lending supply lowers the weighted rate without anyone doing anything, and a project that books that as a win will also book the reverse as a failure.
- Set the guardrail on the quantity the primary can be improved by destroying: the short book's notional-weighted mean decile_rank, joined on instrument_id and business_date equal to as_of_date. If the rate falls while that mean drifts up toward the universe midpoint, the saving came from abandoning the trade.
- Convert the conflict into one comparison rather than two metrics fighting. Hold a short only where expected annualized alpha exceeds borrow_rate_bps plus impact amortised over the expected holding period. That rule keeps expensive bottom-decile names whose alpha covers the borrow and drops mid-decile names whose alpha does not, which is the opposite of what a raw cost target produces.
- Separate in reporting the wins that do not touch the book: rate shopping across lenders for the same instrument on the same day, early warning from borrow_status transitions into hard_to_borrow, and term borrow for names held long enough to justify it. Without that split the project will claim credit for alpha it destroyed.
Follow-up
- The weighted borrow rate falls 40 basis points and mean short-book decile_rank is unchanged. Name two explanations other than good work by this project.
- How do you price a name whose borrow_status is 'unavailable' rather than merely expensive, given it never appears in the cost series at all?
- A recall forces a buy-in on a bottom-decile short. Where does that cost land in your metric set, and should it?
Every algo improved yet blended shortfall got worse
Notional-weighted implementation shortfall by algo_name improved month over month for all five algos, the best by 3.1 bps and the worst by 0.4 bps, while the blended desk number rose from 21.0 to 25.2 bps. Tables: parent_order (order_id, order_qty, filled_qty, arrival_mid_px, avg_fill_px, adv_20d, algo_name, broker_code, order_type) and execution_fill. Produce the arithmetic that reconciles the per-algo improvements to the blended deterioration, name what actually changed, and state what you would and would not recommend to the head of trading.
Approach
- Write the decomposition rather than arguing about it. With w_i the share of month notional in algo i and x_i its shortfall, the blended change is the sum over i of w1_i times (x1_i minus x0_i), plus the sum over i of x0_i times (w1_i minus w0_i). The first term is the within-algo effect, the second is mix, and the two sum to the blended change exactly, with no residual term to argue about.
- Evaluate both terms. A negative within-effect and a larger positive mix effect is the whole answer arithmetically, and the real question becomes why the weights moved.
- Do not stop at the weights. Re-run the same decomposition inside fixed participation buckets (order_qty divided by adv_20d) and inside volatility buckets, to test whether algo_name is standing in for order difficulty. If every algo still improves within every bucket, the per-algo improvements are real and the desk simply traded a harder month.
- Separate the two stories a single mix shift can carry: the desk chose to route more notional to an expensive algo, or the incoming order flow changed (larger orders, more urgency, a newly onboarded strategy) and pushed itself into the algo built for it. Only the first is a routing decision.
- Recommend on the within-bucket numbers, and say plainly that routing flow to the algo with the lowest average would move the easy orders rather than the cost.
Follow-up
- The mix shift is driven by orders above 10 percent of adv_20d. Is the right response to change routing, to change how the strategy releases those orders to the desk, or neither?
- How would you present this so a reader who only ever sees the blended series does not conclude the desk got worse?
- What would make you suspect the per-algo improvements are themselves a selection artefact?
For someone who has spent the last year in notebooks, dashboards or modelling work and has not written raw SQL under time pressure. The first four days rebuild query fluency against a fixture you control and can verify by hand; the last three attach that fluency to the rest of the loop.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Build a fixture you can check answers against
- Create a local Postgres or SQLite database with four tables (users, sessions, events, orders) holding roughly 200 rows you generated yourself, so you know the contents well enough to predict every result.
- Deliberately seed the cases that break queries: a user with no sessions, a session with no events, two orders sharing a timestamp, a NULL in one join key, and one duplicated user row.
- Before writing any SQL, hand-compute five answers on paper (how many users placed at least one order, median orders per ordering user, and three others) and save them as the ground truth for the week.
Deliverable: A one-command seed script plus a text file of five hand-computed answers to grade every later query against.
Practice prompt ↗Practice prompt ↗Worked solution ↗02Joins, filters and NULL semantics
- Answer "which users have no orders" three ways (LEFT JOIN with IS NULL, NOT EXISTS, NOT IN) and confirm that the NOT IN version returns zero rows once the subquery contains a NULL, because the comparison is never TRUE.
- Reproduce the LEFT JOIN that silently collapses to an inner join by putting a right-table predicate in WHERE, then fix it by moving the predicate into the ON clause, and record both row counts.
- Create a fan-out bug on purpose by joining orders to order_items and summing the order total, then correct it with a pre-aggregated subquery and explain in one line which table changed the grain.
Deliverable: One annotated .sql file holding the three join traps, each with the wrong result and the corrected result side by side.
Practice prompt ↗Practice prompt ↗03Window functions and frames
- Write three window queries against the fixture: a running order total per user, the rank of each order within its user by value, and the day gap to that user's previous order, then check each against the day-one ground truth.
- Run ROW_NUMBER, RANK and DENSE_RANK over a column containing ties, print all three side by side, and write one sentence on when each is the correct choice.
- Switch one query from the default frame (RANGE BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, which is what you get when ORDER BY is present and no frame is written) to ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, and explain why the output differs only when the ORDER BY column has duplicates.
Deliverable: Three verified window queries plus a short note explaining the RANGE versus ROWS difference in your own words.
Practice prompt ↗Practice prompt ↗04The four analytical query patterns
- Write a monthly retention grid: first order month per user, then months-since-first as the column, and verify that month zero equals the cohort size exactly.
- Sessionize the events table under a 30-minute inactivity rule using LAG plus a cumulative sum over a new-session flag.
- Build a four-step funnel that counts distinct users rather than events at each step, and state the rule you applied to a user who reaches step three without ever logging step two.
Deliverable: One file with the retention, sessionization and funnel patterns, each carrying a one-line note on the assumption it bakes in.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Write SQL the way you will have to write it live
- Set a 12-minute timer and solve three medium prompts in a plain editor with no execution and no autocomplete, then run them and tally syntax errors separately from logic errors.
- Narrate one solution aloud while writing it, stating the grain of each intermediate result (one row per user, one row per user-day) before you type its body.
- Rewrite your slowest solution as a CTE chain where every CTE name states its grain, and time yourself re-solving it from blank.
Deliverable: A recording of one narrated solution plus an error tally that separates syntax from logic.
Practice prompt ↗06One day for everything that is not SQL
- Write the preconditions of the two-sample t-test from memory, then check them: independent observations, and a difference in means whose sampling distribution is approximately normal, which at large sample sizes follows from the central limit theorem rather than from normality of the raw values.
- Write the difference between an odds ratio from logistic regression and a relative risk, and state the condition under which the two are close (low outcome prevalence).
- Prepare a 90-second answer to "how would you know this model is any good" that names the metric, the baseline you would beat, and the cost of the errors you care about.
Deliverable: One page of notes covering test preconditions, the odds-ratio caveat and the model-quality answer.
Practice prompt ↗07Full loop rehearsal
- Run a 45-minute mock with someone willing to interrupt: 20 minutes of SQL, 15 minutes defining a metric, 10 minutes on a past project.
- Re-solve from blank the two queries you were slowest on this week and compare the times against day five.
- Write a five-line answer to "walk me through a project" that puts a number in the first sentence and names the decision the work changed.
Deliverable: Mock feedback notes plus a timed project narrative you can deliver without reading it.
Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Data people depend on systems owned by other teams, and much of the job is negotiating for instrumentation, access, or a fix to a broken pipeline. Prepare an example of getting something changed upstream that you did not control. Describe what you asked for, what you traded, and how you worked while you waited.
Explaining why two accounts in one strategy diverged 210 bps
Two separately managed accounts run the identical strategy. Over the trailing twelve months one returned 8.4 percent net and the other 6.3 percent, a gap of 210 bps. The client who owns the lower one has asked in writing why. You have position_daily, account_mandate (SCD2) and the fill history for both accounts. Produce two things: a reconciliation that accounts for the gap down to a residual you state, and a reply of at most 200 words that a non-specialist can act on, without jargon and without blaming the client.
Approach
- What is probed: whether you reconcile to the total before you explain anything. An explanation that does not sum to the observed gap is a guess with numbers attached, and the client's next analyst will find the difference.
- Reconcile mechanically and in this order, because each step has a clean source: fee difference from mgmt_fee_bps plus any performance fee crystallised above the high-water mark; average cash weight, as one minus the summed weight_pct_nav; restricted names, as is_restricted days multiplied by what those names returned; single-name cap differences across account_mandate SCD2 versions; and the ramp between funded_date and the date gross exposure first reached 90 percent of target.
- Carry a residual and state it plainly. If five components explain 170 of 210 bps, the reply says 40 bps unexplained rather than largely explained by. The residual is usually trade timing across the two accounts, and naming it is cheaper than having it discovered.
- Order the reply by what the client can do rather than by size of component: what is structural and will persist, such as the fee schedule and their own restricted list; what was one-off and will not repeat, such as the ramp; and what they can change if they choose to.
- Remove every term that would need looking up. Your restricted list kept the account out of three names that contributed 62 bps is actionable. Negative selection effect from compliance constraints is not, and the decision to relax a restriction belongs to the client rather than in a recommendation you push.
Follow-up
- The residual is 120 bps rather than 40. What goes in the letter, and what do you do before sending it?
- The client asks whether their account was deliberately disadvantaged. How do you answer that specific question?
- Does the other client need to be told anything, and who decides?
Answering whether a six-week-old signal is working yet
A new signal has been live six weeks: 30 trading days of realized cross-sectional IC against a 5-day forward return, mean 0.030, standard deviation across days 0.12. An executive with no statistics background asks in a Monday meeting whether it is working and wants a yes or a no. You have the daily IC series and nothing else. Give an answer in three sentences plus one number the executive can hold onto, and say when the question becomes answerable.
Approach
- What is probed: whether you can be honest about statistical power without hiding behind the word significant and without giving a yes that gets quoted back at you in three months.
- Compute the interval before you speak. The standard error of the mean daily IC is 0.12 divided by the square root of 30, which is 0.022, so a mean of 0.030 sits about 1.4 standard errors from zero. That is the optimistic bound and it is already not a yes.
- Adjust for overlap and say that you did. A 5-day forward return sampled every day shares four of five days with its neighbour, so the honest standard error uses a Newey-West estimator with at least 4 lags and lands materially above 0.022. Presenting the naive figure without that caveat is the same error as the signal's own author would make.
- Convert power into a date rather than a verdict. Detecting a true mean IC of 0.03 at two standard errors needs roughly (2 x 0.12 / 0.03)^2 = 64 independent days, and with the overlap inflation of a 5-day horizon that is on the order of 300 trading days, so the question becomes answerable around fifteen months in, not six weeks.
- Give one number and one decision, because wait is useless on its own. Offer a tripwire that makes waiting active: a pre-committed stop if the trailing 60-day mean IC turns negative, and a named review date.
Follow-up
- Another desk called their signal working after four weeks. What do you say when the executive raises that?
- What single observation before the review date would make you stop the signal early?
- The six-week mean is minus 0.03 instead. Does your answer change in substance or only in sign?
Writing an impact statement that survives a hostile reading
Write your own annual impact statement. Your work: a market-impact recalibration the desk adopted in March; a signal you researched that a portfolio manager sized and traded; and a rewrite of the nightly position reconciliation that reduced unreconciled rows in position_daily. Realized implementation shortfall fell from 21 bps to 14 bps after March. Market volatility also fell over the same period. State what you added, in basis points where the attribution supports it, and be explicit about where it does not.
Approach
- What is probed: whether you can separate correlation from contribution when the correlation favours you, which is the one place almost everyone's standards slip.
- Build a counterfactual for the shortfall claim instead of a before-and-after. Shortfall scales with volatility, so a pre and post comparison across a regime change partly measures the market. Use orders that kept the old routing or the old parameters as a control over the same window, matched on participation bucket (order_qty over adv_20d) and side, and report the difference-in-differences rather than the raw 7 bps.
- State the part you cannot claim before anyone asks. The signal was sized by the portfolio manager, so its P and L is a joint product. Claim the research decision itself: what you tested, what you rejected, the number of configurations tried, and the standard error you attached. Volunteering the boundary is what makes the claims inside it credible.
- Give the reconciliation work a number that is not basis points. Report unreconciled and break rows in position_daily before and after, plus the downstream consequence: marks that fell back to stale_prior_day, and client reports restated. Inventing a basis-point figure for operational work costs you the basis-point figures that are real.
- Write a falsifier next to each claim, naming the evidence that would show you added nothing. A reviewer who watches you name your own weakest claim stops auditing the strong ones.
Follow-up
- Your control group is 8 percent of order flow. Is the difference-in-differences credible at that size, and what would you need to make it so?
- The signal lost money this year. Does it appear in the statement, and in what form?
- What did you get wrong this year, and what did it cost?
- 01
Two separately managed accounts run the identical strategy. Over the trailing twelve months one returned 8.4 percent net and the other 6.3 percent, a gap of 210 bps. The client who owns the lower one has asked in writing why. You have position_daily, account_mandate (SCD2) and the fill history for both accounts. Produce two things: a reconciliation that accounts for the gap down to a residual you state, and a reply of at most 200 words that a non-specialist can act on, without jargon and without blaming the client.
- 02
A new signal has been live six weeks: 30 trading days of realized cross-sectional IC against a 5-day forward return, mean 0.030, standard deviation across days 0.12. An executive with no statistics background asks in a Monday meeting whether it is working and wants a yes or a no. You have the daily IC series and nothing else. Give an answer in three sentences plus one number the executive can hold onto, and say when the question becomes answerable.
- 03
Write your own annual impact statement. Your work: a market-impact recalibration the desk adopted in March; a signal you researched that a portfolio manager sized and traded; and a rewrite of the nightly position reconciliation that reduced unreconciled rows in position_daily. Realized implementation shortfall fell from 21 bps to 14 bps after March. Market volatility also fell over the same period. State what you added, in basis points where the attribution supports it, and be explicit about where it does not.
Is this an official BlackRock interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at BlackRock. Rounds and questions reflect what candidates have reported, not a process BlackRock has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult are the interviews, and how much should I prepare?
The interviews are highly rigorous and considered difficult. You should expect in-depth technical grilling alongside domain-specific questions. Plan for several weeks of focused preparation, dedicating time to coding practice, reviewing ML fundamentals, and brushing up on financial terminology.
PracHub interview research ↗Do I need a background in finance to get hired?
While a formal background in finance is not strictly required, a strong demonstrated interest and understanding of basic financial concepts are expected. You will be asked about finance terms, so spending time learning market fundamentals will significantly improve your chances.
PracHub interview research ↗What differentiates a successful candidate from an average one?
Successful candidates seamlessly blend deep technical expertise with strong communication skills. They do not just write code; they can clearly articulate the business problem, defend their mathematical choices, and explain how their models would behave in a live financial market.
PracHub interview research ↗What is the culture like for Data Scientists at BlackRock?
The culture is highly collaborative, intellectually stimulating, and fast-paced. You will work alongside incredibly smart people who value data-driven decision-making. There is a strong emphasis on continuous learning, given the ever-changing nature of global markets.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22