As a Data Scientist at Jane Street, you are positioned at the intersection of quantitative research, statistical modeling, and algorithmic development. This role is not about traditional business intelligence; it is about using massive datasets to refine Jane Street's trading strategies, test market hypotheses, and build robust models that navigate high-frequency environments. You will work closely with traders and engineers to identify signals, manage risk, and optimize the execution of trades across global financial markets.
The impact of this role is immediate and measurable. You will be expected to handle uncertainty with rigor, applying machine learning and statistical techniques to solve problems that are often ill-defined. Whether you are analyzing market microstructure or designing experiments to validate a new strategy, your work directly influences the firm's competitive edge. You will thrive here if you enjoy deep technical challenges, possess a relentless curiosity for how systems behave, and value a collaborative environment where the best idea wins.
While the title is Data Scientist, recognize that the interview process is heavily weighted toward quant-style problem-solving. You are not just being hired for your coding ability; you are being hired for your mathematical intuition.
Virtual Screen
reportedAn extra round usually exists because something is still open after the standard loop: a skill the earlier interviews did not sample, a level decision, or two interviewers who disagreed. It is rarely a rerun of what you already did well. Ask the recruiter who you are meeting, what function they sit in, and how long the session runs. That is an ordinary scheduling question, and the answer changes what you should prepare. What separates a strong candidate here is treating the round as a fresh evaluation with its own bar, rather than assuming earlier performance carries you through or sinks you.
What to demonstrate
- Whether you can answer well on ground the earlier rounds did not cover, without leaning on what you already said to someone else
- Consistency of the facts in your stories: the same sample size, timeframe, team size and scope of your own role as in earlier conversations
- How you handle an unfamiliar format live, including whether you ask what kind of answer is wanted before producing one
How to prepare
- Ask the recruiter for the interviewer's function, the length, and whether to expect a coding surface, a discussion, or a presentation. Preparing for a 30 minute conversation with a partner team is not the same work as preparing for a 60 minute technical block.
- Write out what each earlier round actually covered, then list the two or three areas nobody probed. That gap is the most likely subject of the extra round.
- Re-read the numbers in the project stories you have already told, so a second telling does not quietly contradict the first.
Technical Rounds
reportedMuch of what gets scored here happens out loud while you type. Nobody can see your reasoning inside a half-written query, so five silent minutes read as being stuck even when they are not. State the plan in plain language first: which tables, what grain you are aggregating to, and the one filter that defines the population. Then write it. The narration doubles as insurance, because a wrong plan gets caught early and cheaply while a wrong query gets caught at the end with no time left to redo it. A timed statistics section, where one exists, is a separate test with its own clock.
What to demonstrate
- Whether the query you write matches the plan you just described
- What you do with a hint, meaning whether the correction gets absorbed or the first approach gets defended
- Whether you can debug your own wrong output by reading the result set and naming which part of the query produced the anomaly
How to prepare
- Solve three problems while screen-sharing into a recording, then watch it back and mark every stretch longer than thirty seconds where you said nothing
- Practise compressing the plan into one sentence before typing, then check afterwards whether the finished query actually matched it
- Time yourself on statistics questions that carry a business reading, such as what a confidence interval does and does not claim, rather than re-reading notes without a clock
Onsite Superday
reportedA loop is not scored one interview at a time. The people you meet compare notes afterwards, usually in a meeting you are not in, and the outcome turns on what each of them can say about you when asked. That rewards something other than survival: every room needs one specific thing worth repeating, and none of them can contradict another. The common way to lose is to tell the same project four times with different numbers in it, or to be uniformly fine in a way that leaves nobody with anything to argue for.
What to demonstrate
- Whether your account of a project survives being told twice, with the same scale, the same metric definition and the same numbers each time
- Whether each interviewer leaves with one concrete claim they could make on your behalf later, rather than an absence of complaints
- Whether a question you already answered in an earlier room gets the same answer at the same depth, without visible impatience
How to prepare
- Write a one-page fact sheet for your two or three main projects that fixes the numbers you will quote: rows of data, the metric as a single sentence, the effect you measured and how long the work took. Say them aloud from the sheet until they come out identical every time
- For each kind of room you expect, decide the one sentence you want that interviewer repeating in a debrief, then check during the mock that you said it outright instead of implying it
- Rehearse answering the same project question twice in one sitting, the second time as though you had not just answered it, because the thing that needs fixing is the flatness that creeps into a repeated story
22 candidate reports. Individual accounts describe a particular role and hiring cycle.
Jane Street Intern Software Engineer Interview Experience — Exchange Arbitrage and Quote Throughput
It was a low-level-design problem. Core problem: An Exchange Arbitrage System for stocks. Assume there are multiple exchanges, each with a market feed continuously publishing bid and ask prices for different instruments or stocks. Design and implement an arbitrage strategy. The example gave AAPL on exchange A a bid of \$1.00 and an ask of \$1.10, while exchange B had a bid of \$0.90 and an ask of…
Read full experienceJane Street Quantitative Analyst interview experience: probability progression
I went through math-focused interviews that started with easier questions and became harder. The rounds centered on probability. Early questions established the fundamentals, then the difficulty increased. Each round also began with a behavioral check-in, which set expectations before the technical work. In a similar run, I had three phone interviews along the same probability track. The first co…
Read full experienceJane Street Quantitative Analyst interview: probability games before superday
I had several Zoom interviews, with three technical rounds before a superday. The questions were built around games presented by the interviewers and relied heavily on probability and game thinking. The early games were more straightforward and closer to deterministic setups. Later rounds brought more randomness and more game theory. In the more open-ended game rounds, I needed to optimize outcom…
Read full experienceJane Street Quantitative Analyst interview: probability rounds and nerves
After applying online, I had a resume screen and a first technical round. It was my first quant interview, and I was nervous enough that I did not answer as smoothly as I wanted. I did not move past that round. I later entered through a program that sent me directly to interviews. There were two 45-minute Zoom technical rounds. The first was clearly easier than the second, and the interviewers we…
Read full experienceJane Street Software Engineer interview: final onsite mattered most
I started with a virtual round and then moved to an onsite. The questions blended coding with system-design-style work, but they emphasized real-world systems more than academic exercises. By the onsite, it was clear that this was the most important stage and that the bar was highest there. I did not perform well enough in that final stage to move forward. The process was intense, and my main tak…
Read full experiencePracHub editorial advice for the preparation topics above.
Joining research panels to the current instrument master instead of its effective-dated version, so delisted, merged and bankrupt names silently disappear from the historical universe.
The names that leave a universe leave disproportionately after bad returns, so removing them raises backtested return and lowers backtested volatility at the same time. The bias is largest exactly where the strategy claims to add value, in the tails, and it is invisible in the output: the query succeeds, the row count looks plausible, and the equity curve simply looks better than it should. The fix is to resolve universe membership with a predicate on effective_from and effective_to, never on status = 'active'.
Judging execution quality against interval VWAP and treating a favourable number as proof of good trading.
Interval VWAP is a benchmark the trader partly determines: trading in line with volume tracks VWAP almost by construction, and stretching an order over a longer interval makes the benchmark easier while exposing the position to price drift that the benchmark never charges. Arrival price is the benchmark aligned with the decision, because it charges both the spread and the drift between decision and completion, including the unfilled remainder. Reporting both, and reporting the opportunity cost of unfilled quantity, is what separates a real TCA from a flattering one.
Over-explaining the method and under-explaining the implication
Lead with the answer and what you would do about it, then give the approach when asked. Roughly one sentence of method per three of implication is the right ratio for a stakeholder-facing answer; the interviewer already knows what a regression is.
Building features from data that postdates the prediction time
Check every feature against the timestamp at which the model would actually score, and drop anything computed from a window that includes or follows the label event. For a forecasting use case, split train and test by time rather than at random, and split by entity when the same entity recurs.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
If you have two coins, one fair and one biased, how do you determine w…
If you have two coins, one fair and one biased, how do you determine which is which with minimal flips?
Approach
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
- Translate the result into the decision it informs, in one plain sentence.
- Sanity-check the answer against a simple bound or a simulated case.
Follow-up
- What sample size would you need to detect an effect half this size?
- How would you explain this result to someone who does not know statistics?
Calculate the expected number of rolls to get a specific sequence of o…
Calculate the expected number of rolls to get a specific sequence of outcomes on a die.
Approach
- Write down the assumption the method needs before you use the method.
- Sanity-check the answer against a simple bound or a simulated case.
- Translate the result into the decision it informs, in one plain sentence.
Follow-up
- What sample size would you need to detect an effect half this size?
- How would you explain this result to someone who does not know statistics?
Sessionise a corrected fill stream into trading bursts
One trading day of execution_fill: fill_id, order_id, instrument_id, exec_ts (microsecond UTC venue time), received_ts, fill_qty, fill_px, venue_mic, liquidity_flag, is_correction, corrects_fill_id. About 9.4M rows, delivered in received_ts order, with roughly 0.3% corrections including chains and zero-quantity busts. Produce one row per (order_id, burst), where a burst is a maximal run of surviving fills whose consecutive exec_ts gap is at most 90 seconds, carrying start_ts, end_ts, n_fills, qty, quantity-weighted price and the added-liquidity share. No per-order Python loop.
Approach
- Resolve corrections with set arithmetic rather than iteration: any fill_id appearing in corrects_fill_id has been superseded, so drop those rows, then drop rows with fill_qty = 0 to remove busts, including a bust that supersedes a real fill. Chains need no special case, because every intermediate row is itself somebody's target.
- Sort by (order_id, exec_ts, fill_id). The file arrives in received_ts order and venue timestamps are not monotone in arrival, so sorting is load-bearing rather than tidy; tie-break on fill_id because exec_ts collides at microsecond resolution on active names.
- gap = df.groupby('order_id')['exec_ts'].diff(); new_burst = gap.isna() | (gap > Timedelta('90s')); burst_seq = new_burst.groupby(df.order_id).cumsum(). Two passes over an already-sorted frame, no Python-level iteration.
- Aggregate in one groupby([order_id, burst_seq]): min and max of exec_ts, size, sum of fill_qty, sum of fill_qty*fill_px, and sum of fill_qty where liquidity_flag == 'added'. Derive the weighted price after aggregation as notional over quantity, never as a mean of fill_px.
- Then the concentration flag: join each burst's quantity against the order's surviving total and the order's working span (last exec_ts minus first), and mark bursts holding over 40% of quantity in under 5% of the span. Those are the auction prints and the blocks, and they are the orders whose shortfall is driven by one decision rather than by the algo.
Worked solution 40 min
- superseded = set(f.corrects_fill_id.dropna().astype('int64')); f = f[~f.fill_id.isin(superseded)]; f = f[f.fill_qty > 0].
- f = f.sort_values(['order_id','exec_ts','fill_id'], kind='mergesort').
- Build gap, new_burst and burst_seq as above, then assert burst_seq is 1 on each order's first surviving fill (gap.isna() makes new_burst True there, so the grouped cumsum is 1-based) and increases by exactly 1 at every boundary, with no gaps in the sequence.
- g = f.groupby(['order_id','burst_seq'], sort=False); build the aggregate frame, then wavg_px = notional_sum/qty_sum.
- Reconcile against parent_order.filled_qty and hand-check the three orders with the most bursts.
Follow-up
- Two bursts on one order are separated by 91 seconds. What does a 90-second threshold do to the distribution of bursts per order, and how would you pick the threshold from the data instead of by hand?
- How would you sessionise across orders instead, over all fills in one instrument in one account, and what breaks when two strategies trade the same name in opposite directions?
- One venue reports exec_ts in local time rather than UTC. What would that look like in the burst output, and which check catches it?
How would you identify and mitigate survivorship bias in a provided da…
How would you identify and mitigate survivorship bias in a provided dataset?
Approach
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- Say which table is the grain you start from, and join outward from it.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
Follow-up
- How does the query change if the join becomes one-to-many?
- How would you verify this result without re-running the same query?
Solve a coding challenge involving complex data transformations within…
Solve a coding challenge involving complex data transformations within a time limit.
Approach
- State the window function and its partition and ordering out loud before writing it.
- Say which table is the grain you start from, and join outward from it.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
Follow-up
- How would you verify this result without re-running the same query?
- What breaks if events arrive late or out of order?
Write a function to simulate a stochastic process and aggregate the re…
Write a function to simulate a stochastic process and aggregate the results for analysis.
Approach
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- State the window function and its partition and ordering out loud before writing it.
- Say which table is the grain you start from, and join outward from it.
Follow-up
- What breaks if events arrive late or out of order?
- How does the query change if the join becomes one-to-many?
Implement a class structure to manage a stream of incoming market data…
Implement a class structure to manage a stream of incoming market data.
Approach
- Say which table is the grain you start from, and join outward from it.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- State the window function and its partition and ordering out loud before writing it.
Follow-up
- How would you verify this result without re-running the same query?
- How does the query change if the join becomes one-to-many?
Sessionise child fills into execution bursts across venues
execution_fill holds fill_id, order_id, exec_ts TIMESTAMPTZ(6) in UTC, venue_mic, fill_qty, fill_px and liquidity_flag. Within each order_id, a burst is a maximal run of consecutive fills separated by under 90 seconds. Return one row per burst: order_id, burst sequence, first and last exec_ts, fill count, total quantity, quantity-weighted price, and the running share of parent_order.order_qty completed at the burst's end. Attach the exchange-local session date derived from instrument_master.primary_exchange_mic, not the UTC date.
Approach
- Order fills inside each parent and take
LAG(exec_ts) OVER (PARTITION BY order_id ORDER BY exec_ts, fill_id). Thefill_idtie-break is load-bearing: an algo slicing into a venue produces genuine microsecond ties, and without a deterministic second key the burst boundaries move between runs. - Flag a boundary where the lag is NULL or the gap exceeds 90 seconds, then convert flags to a dense burst id with
SUM(is_new_burst) OVER (PARTITION BY order_id ORDER BY exec_ts, fill_id ROWS UNBOUNDED PRECEDING). One pass, no self-join, and it degrades gracefully on a single-fill order. - Aggregate to burst grain with
SUM(fill_qty * fill_px) / SUM(fill_qty), notAVG(fill_px); averaging prices weights a 100-share child the same as a 50,000-share block and misstates the burst by the spread. - Add the completion curve as a second window over the aggregated bursts:
SUM(burst_qty) OVER (PARTITION BY order_id ORDER BY burst_seq ROWS UNBOUNDED PRECEDING) / order_qty. Doing it after aggregation keeps it at the grain you are reporting. - Resolve the session date through a MIC-to-IANA-zone lookup and
(exec_ts AT TIME ZONE zone)::date. The UTC date happens to coincide with the local session date for US cash equities and most Asian cash sessions, but an Australian morning sits on the previous UTC day, and a futures trade date that rolls at 17:00 US Central does not align with either.
Worked solution 35 min
- Write the LAG and boundary flag, then the running SUM that produces burst_seq.
- Aggregate to (order_id, burst_seq) with quantity-weighted price and first/last exec_ts.
- Join
parent_orderfororder_qtyand add the cumulative completion share window. - Join
instrument_masterforprimary_exchange_mic, map to a zone, and compute the local session date. - Verify the islands: every gap between one burst's end and the next burst's start must exceed 90 seconds.
Follow-up
- Two venues report the same execution microsecond with different
received_ts. Which timestamp do you sort by for sessionisation, and which for latency analysis? - Derive the 90-second threshold from the data instead of asserting it. What distribution would you look at, and what would tell you the threshold is wrong?
- How does the burst structure change under a POV algo versus a close-auction order, and what would you expect the completion curve to look like in each?
How do you approach a betting game where you have incomplete informati…
How do you approach a betting game where you have incomplete information about the opponent’s hand?
Approach
- Say what you would check first and why it is the highest-information step.
- Clarify what is being asked and what a complete answer would contain.
- State your assumptions explicitly before working the problem.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
Solve a brainteaser involving sequential decision-making under resourc…
Solve a brainteaser involving sequential decision-making under resource constraints.
Approach
- State your assumptions explicitly before working the problem.
- Clarify what is being asked and what a complete answer would contain.
- Work from the decision backwards to the evidence you would need.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
Measure crowding in a signal you cannot observe others trading
Research suspects the edge is eroding because other participants hold the same positions. Nobody outside the firm reports their book, so the quantity of interest is unobservable. Build the measurement anyway. Propose two or three proxies from data the firm already has in position_daily (borrow_rate_bps, borrow_status, financing_cost_base_day, market_value_base), execution_fill (fill_px, prevailing_bid_px, prevailing_ask_px, liquidity_flag) and parent_order (arrival_mid_px, adv_20d, filled_qty). State the bias in each, name the single number you would put on a monthly risk report, and say what would retire it.
Approach
- Write the estimand down before proposing anything: the notional held by others in positions correlated with ours, or operationally, the extra cost and extra drawdown we incur when those holders unwind at the same time we do. Without a stated estimand there is nothing for a bias to be a bias relative to.
- Proxy A, borrow scarcity on the short leg: short-notional-weighted borrow_rate_bps and the share of short notional in borrow_status = 'hard_to_borrow'. Bias: it is confounded with lending supply, which moves with index-fund lending programmes, corporate actions and seasonal recalls, so use the spread against a broad borrow benchmark rather than the level.
- Proxy B, impact residual: realized implementation shortfall minus the impact model's prediction at matched participation (filled_qty over adv_20d) and matched volatility, since crowded names should cost more than the curve says. Bias: the residual absorbs every misspecification of the curve, including volatility regime and schedule changes, and it is mechanically correlated with the desk's own growth, so it cannot separate others arriving from us getting bigger.
- Proxy C, co-movement: correlation between the strategy's daily long-short return and a generic version of the same factor, public or internally reconstructed. Rising correlation alongside falling own-IC is what commoditisation looks like. Bias: the reference series carries its own survivorship and construction choices, and the correlation also rises for benign reasons when factor volatility rises.
- Pick one headline and justify the choice. A composite z-score invites false precision by averaging three different biases into one clean-looking number, so prefer a single directional series such as the benchmarked borrow spread, with the other two printed as context. State the retirement condition: a direct holdings source becomes available, or the proxy moves for six months with no matching move in the consequences it claims to predict, namely cost and drawdown.
- Write on the report itself what it does not claim. None of these identifies crowding; each is consistent with several explanations. The output is a prompt to review position sizing, not a measurement of a quantity.
Worked solution 40 min
- Build a 36-month monthly panel of the three proxies, applying the benchmark differencing to the borrow series before anything else.
- Refit the impact curve on a period-matched sample and hold order size fixed when scoring the residual, so the desk's own growth in participation does not enter as crowding.
- Regress each proxy on the two most plausible confounders, broad volatility and the desk's own average participation, and keep the residualized series.
- Test all three against a known crowded-unwind episode in the sample; every proxy claiming to measure the estimand must move there.
- Write the one-page report: one headline series, two context series, an explicit bias sentence under each, and the retirement condition.
Follow-up
- The borrow spread widens 60 basis points with no move in the impact residual or in co-movement. What goes in the report?
- How would you falsify the crowding hypothesis using an event already in your data, such as a large single-day factor reversal?
- If you had one month of a direct holdings feed, what would you check first to calibrate or discard these proxies?
Turnover doubled in a week with unchanged positions
Annualized one-way turnover for a strategy reads 368 percent this month against 180 percent last month. Position_daily shows end-of-day quantities and market_value_base following their usual pattern, and neither the signal nor portfolio construction was touched. Tables: execution_fill (fill_id, order_id, side, fill_qty, fill_px, is_correction, corrects_fill_id, exec_ts, currency) and position_daily (business_date, account_id, instrument_id, quantity, market_value_base). The catalogue definition is: monthly sum of min(total buy notional, total sell notional), over average end-of-day gross market value, times 12. Find the cause and prove it.
Approach
- Split the ratio and chart the numerator and the denominator as separate monthly series before interpreting either. A ratio that moved tells you nothing about which side moved, and the two sides have unrelated failure modes.
- Read the magnitude as evidence. A jump close to a factor of two, in a book whose buy and sell notional are nearly balanced, is the signature of min(buy, sell) having become buy plus sell. No strategy change produces a suspiciously round multiple in a week.
- Reconstruct the numerator yourself from the catalogue definition and compare it with the reported figure for both months. If your recomputation matches last month and not this month, the expression changed, and the scheduled job's query or its version history says when.
- Rule out the competing mechanism before closing. Correction rows arrive as new rows referencing corrects_fill_id, so a naive SUM(fill_qty) counts the correction and never removes the original. Size the corrected notional separately instead of assuming it is small.
- Reconcile to the book of record. Day-over-day change in position_daily.quantity per instrument must equal that day's net signed fills, which bounds true traded quantity independently of whichever expression the report used.
Follow-up
- Correction rows turn out to be 0.8 percent of notional. Write the numerator expression that handles them correctly in one pass.
- Your reconciliation against position_daily leaves a residual on two instruments. Which legitimate causes would you rule out before calling it a data bug?
- What would you add to the pipeline so a metric definition cannot change without the change being visible in the report itself?
For someone who can already write the query and train the model but stalls when asked what to measure or whether a change is worth making. Metric definition and case structure come first; the technical work is kept as maintenance rather than the centre of the week.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Metric anatomy
- For three products you use daily, write one primary metric, two input metrics that plausibly move it, and one guardrail that would catch a cheap way of moving the primary at the cost of the product.
- For one of them, specify the metric precisely enough that two analysts would return the same number: numerator, denominator, unit of observation, time window, and how returning and deleted accounts are treated.
- Pick a ratio metric and write what happens to it when the denominator shrinks for reasons unrelated to the numerator, with a concrete example of that happening.
Deliverable: A one-page metric tree for one product, with the primary metric written as an unambiguous spec.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Diagnosing a drop without guessing
- Take the prompt "weekly active users fell 8 percent week over week" and write the segmentation plan before proposing any cause: platform, region, tenure cohort, acquisition channel, and whether the movement sits in the numerator or in a changed denominator.
- List the instrumentation failures that manufacture fake drops (a client release that stopped firing an event, a bot filter change, a shifted date boundary or timezone) and write the query that rules out each one.
- Rehearse stating the boring explanations first, seasonality and day-of-week composition, before reaching for a product cause.
Deliverable: A drop-diagnosis checklist short enough to recite from memory in under a minute.
Practice prompt ↗Practice prompt ↗03Should we build it
- Take a feature idea and write it as a bet: what you believe is true, what would have to be true for it to pay off, the metric that would confirm it, and the effect size that would justify the engineering cost.
- Size the opportunity top-down and bottom-up, then reconcile the two numbers in writing instead of quoting whichever is friendlier.
- Write the counter-metric that would make you kill the feature even if it wins on the primary metric.
Deliverable: A one-page product memo ending in a decision rather than a list of considerations.
Practice prompt ↗Practice prompt ↗04The places aggregate numbers lie
- Construct a Simpson's paradox numerically: two segments where the treatment wins within each segment yet loses overall, and identify the shift in segment weights that causes it.
- Take a heavy right-tailed quantity such as revenue per user and write why the mean is the wrong summary, which percentile you would report instead, and what a moving mean with a stable median tells you.
- Write your definition of a session for the product from day one, then name two real behaviours it misclassifies.
Deliverable: One page holding a worked Simpson's paradox table and a session definition with its two known failure cases.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Technical maintenance, aimed at metrics
- Solve four timed SQL prompts that all end in a ratio metric, so the question of grain stays live in every answer.
- Compute a 95 percent confidence interval for a proportion on a small sample, and state why the normal approximation is unreliable when either np or n(1 minus p) falls below roughly 10, along with which interval you would use instead.
- Take one metric from your day-one tree, write the query that computes it correctly, then write the query that computes it wrong in the most plausible way and explain how you would notice.
Deliverable: Four solved prompts plus a matched correct and plausible-wrong query for one metric.
Practice prompt ↗Practice prompt ↗06Turning engineering work into data science stories
- Write three project stories as situation, decision, trade-off, outcome, each carrying one number and one thing you got wrong.
- For the story you will lead with, prepare an answer to "what would you do differently" that names a decision you made, not a constraint you were handed.
- Practise the sentence that reframes a systems project as a question project: the question the work answered, ahead of the pipeline it shipped.
Deliverable: Three written stories with the lead story delivered aloud and timed under four minutes.
Practice prompt ↗Practice prompt ↗07Mock case and gap list
- Run a 40-minute mock case with someone playing a product manager who pushes back on your metric choice, and record it.
- Listen back and mark every moment you proposed a solution before the success metric existed.
- Rewrite those moments as the question you should have asked, and rehearse the first 90 seconds of the case until scoping comes before solving.
Deliverable: A recorded case plus a rewritten opening 90 seconds.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Have two ready. In one, the data was on your side and you had to move someone who outranked you. In the other, the pushback was correct and you changed position. The second is the harder story and it lands better, because it shows you separate being right from being attached to an answer. Name the person's actual objection.
Writing an impact statement that survives a hostile reading
Write your own annual impact statement. Your work: a market-impact recalibration the desk adopted in March; a signal you researched that a portfolio manager sized and traded; and a rewrite of the nightly position reconciliation that reduced unreconciled rows in position_daily. Realized implementation shortfall fell from 21 bps to 14 bps after March. Market volatility also fell over the same period. State what you added, in basis points where the attribution supports it, and be explicit about where it does not.
Approach
- What is probed: whether you can separate correlation from contribution when the correlation favours you, which is the one place almost everyone's standards slip.
- Build a counterfactual for the shortfall claim instead of a before-and-after. Shortfall scales with volatility, so a pre and post comparison across a regime change partly measures the market. Use orders that kept the old routing or the old parameters as a control over the same window, matched on participation bucket (order_qty over adv_20d) and side, and report the difference-in-differences rather than the raw 7 bps.
- State the part you cannot claim before anyone asks. The signal was sized by the portfolio manager, so its P and L is a joint product. Claim the research decision itself: what you tested, what you rejected, the number of configurations tried, and the standard error you attached. Volunteering the boundary is what makes the claims inside it credible.
- Give the reconciliation work a number that is not basis points. Report unreconciled and break rows in position_daily before and after, plus the downstream consequence: marks that fell back to stale_prior_day, and client reports restated. Inventing a basis-point figure for operational work costs you the basis-point figures that are real.
- Write a falsifier next to each claim, naming the evidence that would show you added nothing. A reviewer who watches you name your own weakest claim stops auditing the strong ones.
Follow-up
- Your control group is 8 percent of order flow. Is the difference-in-differences credible at that size, and what would you need to make it so?
- The signal lost money this year. Does it appear in the statement, and in what form?
- What did you get wrong this year, and what did it cost?
Explaining why two accounts in one strategy diverged 210 bps
Two separately managed accounts run the identical strategy. Over the trailing twelve months one returned 8.4 percent net and the other 6.3 percent, a gap of 210 bps. The client who owns the lower one has asked in writing why. You have position_daily, account_mandate (SCD2) and the fill history for both accounts. Produce two things: a reconciliation that accounts for the gap down to a residual you state, and a reply of at most 200 words that a non-specialist can act on, without jargon and without blaming the client.
Approach
- What is probed: whether you reconcile to the total before you explain anything. An explanation that does not sum to the observed gap is a guess with numbers attached, and the client's next analyst will find the difference.
- Reconcile mechanically and in this order, because each step has a clean source: fee difference from mgmt_fee_bps plus any performance fee crystallised above the high-water mark; average cash weight, as one minus the summed weight_pct_nav; restricted names, as is_restricted days multiplied by what those names returned; single-name cap differences across account_mandate SCD2 versions; and the ramp between funded_date and the date gross exposure first reached 90 percent of target.
- Carry a residual and state it plainly. If five components explain 170 of 210 bps, the reply says 40 bps unexplained rather than largely explained by. The residual is usually trade timing across the two accounts, and naming it is cheaper than having it discovered.
- Order the reply by what the client can do rather than by size of component: what is structural and will persist, such as the fee schedule and their own restricted list; what was one-off and will not repeat, such as the ramp; and what they can change if they choose to.
- Remove every term that would need looking up. Your restricted list kept the account out of three names that contributed 62 bps is actionable. Negative selection effect from compliance constraints is not, and the decision to relax a restriction belongs to the client rather than in a recommendation you push.
Follow-up
- The residual is 120 bps rather than 40. What goes in the letter, and what do you do before sending it?
- The client asks whether their account was deliberately disadvantaged. How do you answer that specific question?
- Does the other client need to be told anything, and who decides?
Scoping is our execution any good into one answerable question
A chief operating officer stops you in the hallway and asks whether our execution is any good. You have parent_order and execution_fill for eighteen months, roughly 400,000 parent orders across four algos and six brokers. Nobody has defined good, no deadline is set, and two people have already produced conflicting answers. Before writing any SQL, produce a one-page scoping memo: the single question you will answer, the metric defined at field level, the slices, the exclusions, and what you are explicitly not answering.
Approach
- What is probed: whether you convert a request into a decision. The deliverable of scoping is not a work plan, it is the sentence describing what the requester will be able to decide once you are done.
- Choose the benchmark explicitly and state what each one charges. Arrival mid charges spread, impact and the price drift between the desk receiving the order and completing it, including the unfilled remainder. Interval VWAP charges almost none of that and is partly determined by the trader's own volume participation. Answer against arrival and report interval VWAP alongside so nobody believes you hid the flattering number.
- Write the metric to the field level so two analysts cannot compute it differently: side_sign times (avg_fill_px minus arrival_mid_px) times filled_qty, plus commission_amt plus exchange_fee_amt minus rebate_amt, plus opportunity cost on order_qty minus filled_qty at the terminal mid, divided by order_qty times arrival_mid_px, times 10000, aggregated by weighting each order by order_qty times arrival_mid_px.
- Fix the slices in advance, because choosing them after seeing results is how a scoping memo becomes a fishing expedition: algo_name, broker_code, side, and a participation bucket defined as order_qty divided by adv_20d. Set a minimum cell size so a 12 bps difference on 40 orders never reaches a slide.
- State the exclusions and the known data problems in the same memo: corrections net against corrects_fill_id rather than being counted twice, orders with a null arrival_mid_px are reported as a coverage percentage rather than dropped quietly, and clock skew between exec_ts and received_ts is a data-quality item and not an execution-quality finding.
- Close with the questions you are not answering, each with a rough cost: whether the strategy should trade less, whether broker relationships are priced correctly, and whether the algos' internal logic is sound.
Follow-up
- The COO says 12 bps sounds fine. What do you compare it against, and where does that comparison come from?
- The two existing answers disagree. How do you determine whether they used different benchmarks or different populations?
- What changes if this has to become a monthly production number rather than a one-off study?
- 01
Write your own annual impact statement. Your work: a market-impact recalibration the desk adopted in March; a signal you researched that a portfolio manager sized and traded; and a rewrite of the nightly position reconciliation that reduced unreconciled rows in position_daily. Realized implementation shortfall fell from 21 bps to 14 bps after March. Market volatility also fell over the same period. State what you added, in basis points where the attribution supports it, and be explicit about where it does not.
- 02
Two separately managed accounts run the identical strategy. Over the trailing twelve months one returned 8.4 percent net and the other 6.3 percent, a gap of 210 bps. The client who owns the lower one has asked in writing why. You have position_daily, account_mandate (SCD2) and the fill history for both accounts. Produce two things: a reconciliation that accounts for the gap down to a residual you state, and a reply of at most 200 words that a non-specialist can act on, without jargon and without blaming the client.
- 03
A chief operating officer stops you in the hallway and asks whether our execution is any good. You have parent_order and execution_fill for eighteen months, roughly 400,000 parent orders across four algos and six brokers. Nobody has defined good, no deadline is set, and two people have already produced conflicting answers. Before writing any SQL, produce a one-page scoping memo: the single question you will answer, the metric defined at field level, the slices, the exclusions, and what you are explicitly not answering.
Is this an official Jane Street interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Jane Street. Rounds and questions reflect what candidates have reported, not a process Jane Street has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How much time should I spend preparing?
Most successful candidates spend several weeks doing focused practice on probability puzzles and coding challenges. Consistency is more important than cramming; focus on building your "mathematical intuition" rather than memorizing answers.
PracHub interview research ↗Are the interviews focused on finance knowledge?
Not specifically. While the context is trading, the interviews are designed to test your core mathematical and logical abilities. You do not need to be an expert in financial products to succeed, but you must be able to apply your skills to financial scenarios.
PracHub interview research ↗What is the culture like at Jane Street?
The culture is intellectually intense and highly collaborative. Candidates report that it suits people who are curious, enjoy solving hard problems, and are comfortable receiving direct, constructive feedback.
PracHub interview research ↗What if I don't know the answer to a question?
Don't panic. The goal is to see how you think. Articulate your assumptions, ask clarifying questions, and show the interviewer your problem-solving process. Often, the path you take to a solution is more important than the final result.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22