At Schonfeld, a Data Scientist plays a pivotal role in bridging the gap between raw, complex financial data and actionable trading intelligence. As a premier multi-manager platform, Schonfeld relies on quantitative, systematic, and fundamental strategies to navigate global markets. Data is the fundamental fuel for these strategies, and the data science team is responsible for transforming massive, noisy, and unstructured datasets into clean, predictive signals that portfolio managers and quantitative researchers can deploy.
Your work in this role directly impacts the firm’s bottom line. Whether you are working with alternative datasets—such as transactional data, web-scraped sentiment, or satellite imagery—or processing high-frequency market tick data, your primary objective is to extract alpha. This requires not only exceptional programming and data engineering capabilities but also a deep understanding of market microstructure, statistical modeling, and machine learning.
The environment at Schonfeld is highly collaborative yet entrepreneurial. You will work alongside elite quantitative researchers, software engineers, and portfolio managers. Successful candidates are those who possess the technical rigor to write production-grade code, the mathematical sophistication to derive complex statistical models, and the financial intuition to understand why a particular data pattern translates into a viable trading strategy.
Recruiter Conversation
reportedMost candidates lose this call inside the first two minutes, during the walkthrough of their own background. The account runs chronologically, sits at the level of tools and titles, and never arrives at a decision anyone could have disagreed with. Anchor on a problem instead of a timeline: what the team could not answer, what you did about it, what happened next. Ninety seconds is enough, and stopping on time leaves room for the half of the call that belongs to you. What you ask about how work gets prioritised signals your level more reliably than the walkthrough does.
What to demonstrate
- Whether your background summary has a shape (problem, decision, consequence) or is a chronological list of tools and employers
- Whether you can account for gaps, short stints and the reason you are looking, unprompted and without hedging
- The substance of the questions you ask back, which an experienced screener reads as a level signal
How to prepare
- Time your opening walkthrough against a clock. If it runs past two minutes, compress the earliest role into a single clause and spend the recovered time on the most recent one
- Write one honest sentence for every gap or short stint visible on your resume and offer it before being asked about it
- Prepare questions about how work arrives and gets prioritised: who writes the request, how often priorities change, and what happens to an analysis after it is delivered
Technical Assessments
reportedBefore anything else, this round is a reading test. You are given a small schema and a question phrased in business language, and most of the difficulty sits in the gap between them. Who counts as an active user, does a refunded order still count as an order, is that date column an event time or a load time. Weak answers start typing immediately and compute something precise about the wrong population. Strong ones pin the definition in one sentence, name the column that encodes it, then write the query. On a timed assessment with nobody to tell, write the definition in a comment anyway.
What to demonstrate
- Whether an ambiguous term becomes a specific column and filter before any computation happens
- Whether you read the schema for keys and cardinality rather than only for column names
- Whether the result answers the question at the grain it was asked at, per user or per session or per day
How to prepare
- Take three metrics you already use and write down the exact filter and exact grain behind each, then practise stating one of them in a single sentence out loud
- On a schema you have never seen, spend the first minute writing what one row of each table means and which key it is unique on, then predict which joins can duplicate rows
- Rehearse a version where the definition changes halfway through, and edit the query you have instead of starting over
Team Interviews
reportedBecause the format is not fixed, prepare the reasoning rather than the ritual. Nearly every version of this round draws on the same underlying material: a design you can defend, a metric you can define exactly, an analysis whose assumptions you can state out loud. Only the wrapper changes, whether that is a take-home, a live case, a deep dive on past work, or a rough estimate on a whiteboard. Answers rehearsed to fit one shape stall the moment the shape differs. Practise naming the assumption behind a number, then saying how much the conclusion moves if that assumption is wrong.
What to demonstrate
- Whether your justification for a method survives the question 'why not the simpler thing', including when the simpler thing would have worked
- Precision under pressure: what exactly counts as an active user, a conversion or a success, over what window, with what exclusions
- Whether you carry an argument through to a recommendation instead of stopping at a list of tradeoffs
How to prepare
- For each project you plan to mention, write the metric definition in one sentence: numerator, denominator, time window, exclusions. Say it out loud once, because vagueness shows up in speech before it shows up on paper.
- Rehearse the same project at three lengths: two minutes, ten minutes, and a deep dive on one technical decision. Cutting live is harder than it sounds.
- For your headline result, write down what would have had to be true for it to be wrong, and how you ruled that out.
Senior Discussions
reportedRounds outside the standard loop often open with something deliberately under-specified: a loose business problem, an open question about a product area, a dataset described in one sentence. The common failure is surveying, listing six plausible approaches and committing to none of them. The thing that separates a strong answer is scoping out loud. State what you are treating as the goal, name the metric you would move, say what you are choosing not to do and why, then take one path through to an actual answer. An interviewer can follow you down a narrow path. Nobody can grade a menu.
What to demonstrate
- Whether you turn an ambiguous prompt into a stated question with a measurable outcome before doing any work
- The judgement visible in what you cut, and whether you say why you cut it rather than silently dropping it
- Whether you land on a concrete recommendation with its caveat attached, rather than an unranked set of options
How to prepare
- Take three vague prompts, such as 'is this feature working', 'why did retention drop', and 'should we expand into a new segment'. For each, write one sentence of goal, one primary metric with its window, and two things you are explicitly not doing.
- Practise giving the recommendation first and the reasoning second, in five minutes. Loosely defined rounds are usually time-boxed, and an answer that arrives last often does not arrive.
- Keep a running assumption list as you talk, on paper or in the shared doc, so the interviewer can challenge one assumption instead of your whole answer.
PracHub editorial advice for the preparation topics above.
Modelling transaction cost as a constant number of basis points, independent of order size and volatility.
Temporary market impact scales approximately with volatility times the square root of participation, that is, of order quantity divided by average daily volume, so cost per share rises as size rises rather than staying flat. A constant-bps assumption is roughly right for the small orders used to calibrate it and badly wrong for the size the strategy would actually run, which is how a book that backtests well at modest notional loses money at ten times the size. It also makes capacity unmeasurable, because capacity is exactly the notional at which marginal impact equals marginal alpha.
Filtering on as_of_date rather than knowledge_ts, so restated fundamentals, revised index constituents and retroactively applied split and dividend adjustments enter the backtest before they were knowable.
Vendors overwrite history in place. A quarterly figure filed 45 days after period end is stored against period end, an index addition announced five business days before it takes effect is stored against the effective date, and a split applied tonight rewrites every prior close in the adjusted series. Each of those gives the strategy information it could not have had, and the resulting lift is concentrated in the highest-turnover, highest-apparent-alpha names. The signal_score table separates the two timestamps precisely so this filter can be written correctly.
Naming a model class before naming the deployment constraints
Set out the latency budget, the label delay, the retraining cadence, the interpretability requirement and the number of labelled examples, then pick the model that fits them. A boosted-tree answer to a problem where each decision must be explained to the affected user is a well-executed answer to the wrong question.
Dropping rows with missing values without naming the mechanism
Say whether the values are missing at random, missing by a known process, or missing in a way that depends on the outcome, and handle them accordingly. Deleting incomplete rows silently redefines the population whenever missingness correlates with what you are measuring.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Derive the mathematical formula for the Ordinary Least Squares (OLS) e…
Derive the mathematical formula for the Ordinary Least Squares (OLS) estimator. What are its core assumptions?
Approach
- Write down the assumption the method needs before you use the method.
- Say what the estimate is of, and over what population it generalises.
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
Follow-up
- Which assumption here is most likely to be violated in practice?
- What sample size would you need to detect an effect half this size?
Explain the mathematical difference between L1 (Lasso) and L2 (Ridge) …
Explain the mathematical difference between L1 (Lasso) and L2 (Ridge) regularization and when you would choose one over the other for feature selection.
Approach
- Write down the assumption the method needs before you use the method.
- Say what the estimate is of, and over what population it generalises.
- Sanity-check the answer against a simple bound or a simulated case.
Follow-up
- Which assumption here is most likely to be violated in practice?
- How would you explain this result to someone who does not know statistics?
If you observe a decay in the performance of a production model, what …
If you observe a decay in the performance of a production model, what steps do you take to diagnose and remediate the issue?
Approach
- Say how the offline result would be validated online before it is trusted.
- Set a baseline first, so any model has something honest to beat.
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
Follow-up
- What would you monitor after launch to know the model is still valid?
- How would you choose the decision threshold, and who owns that choice?
How do you address multicollinearity in a multiple linear regression m…
How do you address multicollinearity in a multiple linear regression model, and how does it affect your coefficient estimates?
Approach
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
- Check what information would not exist at prediction time, and exclude it.
- Say how the offline result would be validated online before it is trusted.
Follow-up
- What would you monitor after launch to know the model is still valid?
- Where could label leakage enter this setup?
How large a Sharpe does pure noise produce over many trials
A research team tested 250 strategy variants on the same three years of daily returns (756 observations). The best variant has an annualized Sharpe of 1.4. Under the null that every variant has zero expected return, estimate by simulation the probability that the maximum of 250 Sharpe estimates is at least 1.4. Then repeat with the variants' daily returns equicorrelated at rho = 0.7, which is closer to the truth when variants share a universe. Report both probabilities with Monte Carlo error and say which one belongs in the memo.
Approach
- Simulate at the frequency the statistic is estimated at: 756 daily draws per variant, not three annual ones. The Sharpe is scale-free, so standard normal draws suffice and the choice of sigma cannot change the answer.
- Build the correlated case as X = sqrt(rho)*Z0 + sqrt(1-rho)*E, with Z0 one common daily draw shared by all variants and E independent per variant. That is exact equicorrelation for one extra column, rather than a 250x250 Cholesky per replication.
- Per replication compute all 250 annualized Sharpes as mean/std(ddof=1)*sqrt(252) along the time axis, take the maximum, and count how often it clears 1.4. Use at least 20,000 replications so the Monte Carlo standard error on a probability near 0.85 is about 0.0025.
- Check the independent case analytically before trusting the simulation: 1 - Phi(z)^250 with z = 1.4/SE and SE = sqrt(252/756) = 0.577 gives z = 2.43 and p close to 0.85. The simulation should land inside two Monte Carlo standard errors of that.
- Get the direction of the correlation effect right. Correlated variants behave like fewer independent trials, so the null maximum is smaller and an observed 1.4 becomes less likely under the null, not more. Present the correlated p-value as the smaller, more favourable number and state plainly that it depends on an assumed rho you did not measure.
Worked solution 25 min
- rng = np.random.default_rng(0); per replication draw X with shape (756, 250).
- sr = X.mean(axis=0)/X.std(axis=0, ddof=1)*np.sqrt(252); store sr.max().
- Repeat with X = sqrt(0.7)*Z0[:, None] + sqrt(0.3)*E where Z0 has shape (756,).
- p = (max_sr >= 1.4).mean(); mc_se = sqrt(p*(1-p)/n_sims).
- Plot both null distributions of the maximum with 1.4 marked on each.
Follow-up
- The team says it only ran six configurations because it discarded the rest early. How do you count trials that were abandoned after somebody looked at the result?
- What Sharpe would the best of 250 have to reach for you to call it significant at 5%, and is that number attainable at this strategy's turnover?
- How would you carve out a holdout the search has genuinely not touched, given the team has already seen the full sample?
Explain how you would structure a data pipeline to ingest and normaliz…
Explain how you would structure a data pipeline to ingest and normalize daily alternative data from multiple unstructured sources.
Approach
- State the window function and its partition and ordering out loud before writing it.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
Follow-up
- How would you verify this result without re-running the same query?
- How does the query change if the join becomes one-to-many?
Implement a rolling volume-weighted average price (VWAP) calculation i…
Implement a rolling volume-weighted average price (VWAP) calculation in Python, optimizing for memory efficiency.
Approach
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Say which table is the grain you start from, and join outward from it.
Follow-up
- How would you verify this result without re-running the same query?
- How does the query change if the join becomes one-to-many?
How would you handle missing values or outliers in a high-frequency li…
How would you handle missing values or outliers in a high-frequency limit order book dataset without introducing lookahead bias?
Approach
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
Follow-up
- How would you verify this result without re-running the same query?
- What breaks if events arrive late or out of order?
Write a Python script using Pandas to align two asynchronous tick-data…
Write a Python script using Pandas to align two asynchronous tick-data streams by their timestamps.
Approach
- State the window function and its partition and ordering out loud before writing it.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- Say which table is the grain you start from, and join outward from it.
Follow-up
- What breaks if events arrive late or out of order?
- How does the query change if the join becomes one-to-many?
Deduplicate a restated signal table into a leak-free panel
signal_score may hold several rows per (signal_id, instrument_id, as_of_date) because reruns rewrite history; each row carries model_version, knowledge_ts, computed_at, is_backfilled, zscore_xs, decile_rank, coverage_flag and universe_id. Build the panel a backtest may legitimately use on trading date D: for every instrument in universe_id = 'liquid_us_1500', the single most recent score whose knowledge_ts is at or before D 21:00 UTC, preferring a live row over a backfilled one. Return instrument_id, as_of_date, zscore_xs, decile_rank.
Approach
- Filter on
knowledge_ts, never onas_of_date.as_of_datesays what the score describes;knowledge_tssays the earliest instant every input was observable. Put it in the WHERE so it prunes before the window is evaluated. - Rank within the key:
ROW_NUMBER() OVER (PARTITION BY signal_id, instrument_id ORDER BY as_of_date DESC, is_backfilled ASC, knowledge_ts DESC, computed_at DESC)and keep rank = 1. In Postgres FALSE sorts before TRUE, sois_backfilled ASCis the live-row preference, written down rather than assumed. - Prefer ROW_NUMBER to a
MAX(computed_at)group-then-rejoin: the rejoin duplicates rows whenever two reruns share acomputed_at, which is exactly what a batch job produces. - Decide what
coverage_flag IN ('stale','imputed')means for this panel and encode it. Dropping those rows silently changes the universe size day to day, which surfaces later as an unexplained jump in measured IC rather than as a missing-data problem. - Express the cutoff as a parameter and the whole thing as a CTE so the backtest loop reuses one query text per date instead of a hand-edited copy.
Worked solution 25 min
- Write the filtered CTE:
universe_id = 'liquid_us_1500'andknowledge_ts <= :cutoff. - Add the ROW_NUMBER window with the four-key ordering and select rank = 1.
- Assert one row per instrument with a COUNT(*) versus COUNT(DISTINCT instrument_id) probe.
- Run the same query with
as_of_date <= Dsubstituted for the knowledge_ts filter and count how many rows differ. - Inspect the differing rows: they should be predominantly
is_backfilled = TRUEwithcomputed_atafter D.
Follow-up
- Write the query that proves the panel is leak-free. What would a violation look like in measured IC, and roughly how large would you expect the inflation to be?
- Produce every trading date in one pass instead of one query per date. What does that cost in plan shape?
- Two
model_versionvalues are live at once during a migration. How does the tie-break change, and who decides?
Describe your process for backtesting a trading signal. How do you acc…
Describe your process for backtesting a trading signal. How do you account for transaction costs, market impact, and slippage?
Approach
- Clarify what is being asked and what a complete answer would contain.
- State your assumptions explicitly before working the problem.
- Work from the decision backwards to the evidence you would need.
Follow-up
- What assumption would you test first?
- How would you know your answer was wrong?
Fix a testing platform that peeks daily across fourteen variants
A research platform runs fourteen execution-parameter variants against one control, each judged on implementation shortfall. Portfolio managers read a dashboard that recomputes a two-sided t-test every morning for thirty trading days, and anything crossing p < 0.05 is shipped that day. Over the past year, nine of eleven shipped variants failed to reproduce in a follow-up test. Quantify both error sources, then specify the replacement procedure: the stopping rule, the family definition, and what the dashboard is allowed to display.
Approach
- Quantify the peeking cost first, because it is the larger of the two. Repeatedly applying a fixed nominal 0.05 threshold as data accrues drives the probability of ever crossing far above 0.05: on the classical repeated-significance figures for a normal mean it is roughly 0.19 by ten equally spaced looks and about 0.28 by thirty, and with an unbounded horizon and a true null it tends to 1.
- Quantify the multiplicity cost. Fourteen independent comparisons each at 0.05 give a 1 - 0.95^14 = 0.51 chance of at least one false positive at a single look. Compounded with daily peeking, a dashboard hit at p < 0.05 carries almost no evidential weight, which is exactly what the nine-of-eleven reproduction failure is reporting.
- Choose a stopping rule that is valid at the stopping time actually used. If the looks can be scheduled in advance, a group-sequential design with an O'Brien-Fleming spending function preserves the overall level with very little power loss at the planned maximum sample size. If managers will look whenever they like, use an always-valid procedure such as a mixture sequential probability ratio test or a confidence sequence, and accept the larger sample it demands as the price of unrestricted looking.
- Define the family before the wave starts. Fourteen variants against one control in one wave is one family. Control the false discovery rate with Benjamini-Hochberg at q = 0.10 rather than the family-wise rate, because a false positive that later fails to reproduce is cheaper here than a missed improvement. State the precondition: the Benjamini-Hochberg guarantee holds under independence or positive regression dependence, and these variants share order flow and are positively dependent, which normally satisfies it. Under arbitrary dependence the Benjamini-Yekutieli correction divides by the sum of 1/i for i = 1..14, about 3.25.
- Make the platform enforce the design rather than document it. Display the sequential boundary and the adjusted q-value instead of a raw p-value, withhold the effect estimate until the boundary is crossed or the horizon is reached, and log every look so the number of looks and the membership of the family are recorded facts rather than recollections.
Worked solution 35 min
- Simulate the current procedure under a true null: fourteen arms, thirty daily looks, ship on any p < 0.05. Record the fraction of waves that ship at least one variant.
- Reconcile the simulation against its analytic pieces: 1 - 0.95^14 for multiplicity alone at a single look, and the repeated-significance inflation for one variant across thirty looks.
- Re-run under the replacement: an O'Brien-Fleming boundary with six scheduled looks plus Benjamini-Hochberg at q = 0.10. Record the null ship rate and the power against a pre-registered 1.5 bps effect.
- Back-test the replacement on the eleven variants shipped over the past year and count how many would have crossed its boundary.
- Write the stopping rule, the family definition and the look schedule into the platform configuration, so the next wave cannot deviate without an explicit change.
Follow-up
- An effect estimated at a sequential stop is biased. In which direction, by roughly how much, and how would you report a corrected one?
- A variant that appeared in last month's wave and again in this month's wave: one hypothesis or two, and what does your answer imply about how waves are composed?
- Would you rather run fourteen variants for thirty days or four variants for one hundred and five days, given the same total order flow?
TCA traded notional fell while book turnover held
The daily TCA report's traded notional dropped about 12 percent starting on a Tuesday, while turnover computed from position_daily is unchanged. No orders were rejected and parent_order.status shows the usual filled share. You have parent_order (order_id, decision_ts, arrival_ts, terminal_ts, order_qty, filled_qty, broker_code, algo_name) and execution_fill (fill_id, order_id, exec_ts TIMESTAMPTZ(6), received_ts, venue_mic, fill_qty, fill_px, is_correction). Find the missing notional and name the mechanism. Give the ordered checklist, including the cut you would make first and the cut you would deliberately not make first.
Approach
- Reconcile at the order level first. parent_order.filled_qty is maintained by the order management system independently of the fill table, so orders where filled_qty exceeds SUM(fill_qty) over their non-correction fills enumerate exactly which fills the report lost. That turns a 12 percent aggregate into a concrete row set you can cut.
- Cut those rows by exec_ts hour in UTC and by venue_mic before cutting by broker_code. The hour cut tests a time-boundary mechanism. The broker cut will also light up, because whichever broker routes the affected region's flow looks responsible, and that correlation is what sends people down the wrong path.
- Read the report's date predicate. exec_ts is TIMESTAMPTZ(6) in UTC; if the job buckets fills by a local calendar date while bounding parent orders by a different zone's date, every fill that crosses midnight in the mismatched zone falls outside the join window and vanishes without raising an error.
- Establish whether rows were dropped or misfiled. Sum the report over two days instead of one. If the two-day total reconciles, the rows landed in the adjacent bucket; if it does not, they are gone. These look identical in a daily total and demand different fixes.
- Date the change against deploys and configuration rather than against market events, and re-run the prior week's data through the current job to confirm the code moved and the data did not.
Follow-up
- The rows turn out to be misfiled rather than dropped. Does that change your severity assessment, and what does it do to the month-to-date figure?
- Write the predicate you would use so the same report is correct for venues in every timezone, including one whose session spans the UTC date boundary.
- received_ts and exec_ts differ by up to several seconds. Which belongs in the bucketing predicate and which belongs in latency monitoring?
For someone who has spent the last year in notebooks, dashboards or modelling work and has not written raw SQL under time pressure. The first four days rebuild query fluency against a fixture you control and can verify by hand; the last three attach that fluency to the rest of the loop.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Build a fixture you can check answers against
- Create a local Postgres or SQLite database with four tables (users, sessions, events, orders) holding roughly 200 rows you generated yourself, so you know the contents well enough to predict every result.
- Deliberately seed the cases that break queries: a user with no sessions, a session with no events, two orders sharing a timestamp, a NULL in one join key, and one duplicated user row.
- Before writing any SQL, hand-compute five answers on paper (how many users placed at least one order, median orders per ordering user, and three others) and save them as the ground truth for the week.
Deliverable: A one-command seed script plus a text file of five hand-computed answers to grade every later query against.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Joins, filters and NULL semantics
- Answer "which users have no orders" three ways (LEFT JOIN with IS NULL, NOT EXISTS, NOT IN) and confirm that the NOT IN version returns zero rows once the subquery contains a NULL, because the comparison is never TRUE.
- Reproduce the LEFT JOIN that silently collapses to an inner join by putting a right-table predicate in WHERE, then fix it by moving the predicate into the ON clause, and record both row counts.
- Create a fan-out bug on purpose by joining orders to order_items and summing the order total, then correct it with a pre-aggregated subquery and explain in one line which table changed the grain.
Deliverable: One annotated .sql file holding the three join traps, each with the wrong result and the corrected result side by side.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Window functions and frames
- Write three window queries against the fixture: a running order total per user, the rank of each order within its user by value, and the day gap to that user's previous order, then check each against the day-one ground truth.
- Run ROW_NUMBER, RANK and DENSE_RANK over a column containing ties, print all three side by side, and write one sentence on when each is the correct choice.
- Switch one query from the default frame (RANGE BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, which is what you get when ORDER BY is present and no frame is written) to ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, and explain why the output differs only when the ORDER BY column has duplicates.
Deliverable: Three verified window queries plus a short note explaining the RANGE versus ROWS difference in your own words.
Practice prompt ↗Practice prompt ↗04The four analytical query patterns
- Write a monthly retention grid: first order month per user, then months-since-first as the column, and verify that month zero equals the cohort size exactly.
- Sessionize the events table under a 30-minute inactivity rule using LAG plus a cumulative sum over a new-session flag.
- Build a four-step funnel that counts distinct users rather than events at each step, and state the rule you applied to a user who reaches step three without ever logging step two.
Deliverable: One file with the retention, sessionization and funnel patterns, each carrying a one-line note on the assumption it bakes in.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Write SQL the way you will have to write it live
- Set a 12-minute timer and solve three medium prompts in a plain editor with no execution and no autocomplete, then run them and tally syntax errors separately from logic errors.
- Narrate one solution aloud while writing it, stating the grain of each intermediate result (one row per user, one row per user-day) before you type its body.
- Rewrite your slowest solution as a CTE chain where every CTE name states its grain, and time yourself re-solving it from blank.
Deliverable: A recording of one narrated solution plus an error tally that separates syntax from logic.
Practice prompt ↗Practice prompt ↗06One day for everything that is not SQL
- Write the preconditions of the two-sample t-test from memory, then check them: independent observations, and a difference in means whose sampling distribution is approximately normal, which at large sample sizes follows from the central limit theorem rather than from normality of the raw values.
- Write the difference between an odds ratio from logistic regression and a relative risk, and state the condition under which the two are close (low outcome prevalence).
- Prepare a 90-second answer to "how would you know this model is any good" that names the metric, the baseline you would beat, and the cost of the errors you care about.
Deliverable: One page of notes covering test preconditions, the odds-ratio caveat and the model-quality answer.
Practice prompt ↗Practice prompt ↗07Full loop rehearsal
- Run a 45-minute mock with someone willing to interrupt: 20 minutes of SQL, 15 minutes defining a metric, 10 minutes on a past project.
- Re-solve from blank the two queries you were slowest on this week and compare the times against day five.
- Write a five-line answer to "walk me through a project" that puts a number in the first sentence and names the decision the work changed.
Deliverable: Mock feedback notes plus a timed project narrative you can deliver without reading it.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
A number you shipped turned out to be wrong, and someone had already acted on it. That is one of the most useful stories a data person can carry. What is being scored is how fast you noticed, who you told first, and what you changed in the process so the same class of error could not repeat quietly.
Writing an impact statement that survives a hostile reading
Write your own annual impact statement. Your work: a market-impact recalibration the desk adopted in March; a signal you researched that a portfolio manager sized and traded; and a rewrite of the nightly position reconciliation that reduced unreconciled rows in position_daily. Realized implementation shortfall fell from 21 bps to 14 bps after March. Market volatility also fell over the same period. State what you added, in basis points where the attribution supports it, and be explicit about where it does not.
Approach
- What is probed: whether you can separate correlation from contribution when the correlation favours you, which is the one place almost everyone's standards slip.
- Build a counterfactual for the shortfall claim instead of a before-and-after. Shortfall scales with volatility, so a pre and post comparison across a regime change partly measures the market. Use orders that kept the old routing or the old parameters as a control over the same window, matched on participation bucket (order_qty over adv_20d) and side, and report the difference-in-differences rather than the raw 7 bps.
- State the part you cannot claim before anyone asks. The signal was sized by the portfolio manager, so its P and L is a joint product. Claim the research decision itself: what you tested, what you rejected, the number of configurations tried, and the standard error you attached. Volunteering the boundary is what makes the claims inside it credible.
- Give the reconciliation work a number that is not basis points. Report unreconciled and break rows in position_daily before and after, plus the downstream consequence: marks that fell back to stale_prior_day, and client reports restated. Inventing a basis-point figure for operational work costs you the basis-point figures that are real.
- Write a falsifier next to each claim, naming the evidence that would show you added nothing. A reviewer who watches you name your own weakest claim stops auditing the strong ones.
Follow-up
- Your control group is 8 percent of order flow. Is the difference-in-differences credible at that size, and what would you need to make it so?
- The signal lost money this year. Does it appear in the statement, and in what form?
- What did you get wrong this year, and what did it cost?
Defending a backtest correction that removes an allocated strategy
A colleague's cross-sectional equity signal received capital last month on a backtest showing annualized Sharpe 1.6 over three years. Reproducing it, you find the panel filters signal_score on as_of_date rather than knowledge_ts. Rerunning with the knowledge_ts predicate, holding universe_id, date range, cost model and rebalance schedule fixed, gives Sharpe 0.5 and moves mean daily IC from 0.041 to 0.012. The colleague is senior and presented the original result. You have one meeting with the research lead and the portfolio manager. Deliver a recommendation on whether the allocation stands, and the evidence behind it.
Approach
- What is probed: whether a quantitative objection survives social cost. Lead with the defect and the single-variable reproduction, never with a judgement about the colleague. The claim under test is one predicate, not a person's competence.
- Rerun both versions from one code path with one line changed, and say so explicitly. Holding universe_id, date range, cost model and rebalance schedule fixed leaves the 1.1 Sharpe gap exactly one candidate cause, which is what makes the result arguable on its merits instead of on whose code is trusted.
- Pair the comparison before quoting any error bar. For modest Sharpe ratios the annualized standard error of a Sharpe estimate is approximately 1 over the square root of the number of years, so three years of daily data gives roughly 0.58 and two marginal estimates 1.1 apart are only about two standard errors apart. But the two runs are the same returns except where the leak bites, so test the daily difference series directly: its standard error is far smaller, and the pairing is what turns a marginal result into a decisive one.
- Exhibit the mechanism, not just the size. Rank instrument-days by their contribution to the return difference between the two runs and show that the top contributors carry is_backfilled TRUE or coverage_flag 'stale', with knowledge_ts postdating as_of_date by the vendor's restatement lag. A named mechanism is falsifiable; a Sharpe delta alone becomes an argument about your code.
- Bring a decision rather than only a finding: the position size the corrected Sharpe supports, and an untouched out-of-sample window that would settle it either way. Being right with no path forward is how a correct objection gets overruled.
Follow-up
- Your colleague reruns it and gets 0.9 rather than 0.5. What do you do with the discrepancy before the meeting?
- The strategy is up since funding. Does live P and L change your recommendation, and how much of it would?
- What would have caught this before the allocation, and why did the existing review not catch it?
Answering whether a six-week-old signal is working yet
A new signal has been live six weeks: 30 trading days of realized cross-sectional IC against a 5-day forward return, mean 0.030, standard deviation across days 0.12. An executive with no statistics background asks in a Monday meeting whether it is working and wants a yes or a no. You have the daily IC series and nothing else. Give an answer in three sentences plus one number the executive can hold onto, and say when the question becomes answerable.
Approach
- What is probed: whether you can be honest about statistical power without hiding behind the word significant and without giving a yes that gets quoted back at you in three months.
- Compute the interval before you speak. The standard error of the mean daily IC is 0.12 divided by the square root of 30, which is 0.022, so a mean of 0.030 sits about 1.4 standard errors from zero. That is the optimistic bound and it is already not a yes.
- Adjust for overlap and say that you did. A 5-day forward return sampled every day shares four of five days with its neighbour, so the honest standard error uses a Newey-West estimator with at least 4 lags and lands materially above 0.022. Presenting the naive figure without that caveat is the same error as the signal's own author would make.
- Convert power into a date rather than a verdict. Detecting a true mean IC of 0.03 at two standard errors needs roughly (2 x 0.12 / 0.03)^2 = 64 independent days, and with the overlap inflation of a 5-day horizon that is on the order of 300 trading days, so the question becomes answerable around fifteen months in, not six weeks.
- Give one number and one decision, because wait is useless on its own. Offer a tripwire that makes waiting active: a pre-committed stop if the trailing 60-day mean IC turns negative, and a named review date.
Follow-up
- Another desk called their signal working after four weeks. What do you say when the executive raises that?
- What single observation before the review date would make you stop the signal early?
- The six-week mean is minus 0.03 instead. Does your answer change in substance or only in sign?
- 01
Write your own annual impact statement. Your work: a market-impact recalibration the desk adopted in March; a signal you researched that a portfolio manager sized and traded; and a rewrite of the nightly position reconciliation that reduced unreconciled rows in position_daily. Realized implementation shortfall fell from 21 bps to 14 bps after March. Market volatility also fell over the same period. State what you added, in basis points where the attribution supports it, and be explicit about where it does not.
- 02
A colleague's cross-sectional equity signal received capital last month on a backtest showing annualized Sharpe 1.6 over three years. Reproducing it, you find the panel filters signal_score on as_of_date rather than knowledge_ts. Rerunning with the knowledge_ts predicate, holding universe_id, date range, cost model and rebalance schedule fixed, gives Sharpe 0.5 and moves mean daily IC from 0.041 to 0.012. The colleague is senior and presented the original result. You have one meeting with the research lead and the portfolio manager. Deliver a recommendation on whether the allocation stands, and the evidence behind it.
- 03
A new signal has been live six weeks: 30 trading days of realized cross-sectional IC against a 5-day forward return, mean 0.030, standard deviation across days 0.12. An executive with no statistics background asks in a Monday meeting whether it is working and wants a yes or a no. You have the daily IC series and nothing else. Give an answer in three sentences plus one number the executive can hold onto, and say when the question becomes answerable.
Is this an official Schonfeld interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Schonfeld. Rounds and questions reflect what candidates have reported, not a process Schonfeld has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult is the coding portion of the Schonfeld interview?
The coding portion is highly practical and focuses on data manipulation rather than abstract algorithmic puzzles. You should expect an intensive, live-coding session centered around Pandas and time-series data processing. Speed, accuracy, and clean code organization are highly valued.
PracHub interview research ↗What is the balance between statistics and machine learning in the evaluation?
Schonfeld prioritizes strong statistical foundations over complex, black-box machine learning. You must be able to derive and explain classical statistical models, such as linear regression and OLS, before discussing advanced machine learning architectures.
PracHub interview research ↗How should I handle questions about my current firm's proprietary strategies?
Always protect your current employer's confidential information. If asked about specific alpha strategies or proprietary systems, politely state that the details are confidential. Offer to explain the general methodology, mathematical frameworks, or open-source equivalents instead. Interviewers respect candidates who demonstrate strong professional ethics.
PracHub interview research ↗Does Schonfeld require prior hedge fund or trading experience for this role?
While prior experience in quantitative finance or with market data is highly advantageous, it is not always a strict requirement. Candidates with exceptional backgrounds in mathematics, statistics, or computer science who demonstrate strong financial curiosity and fast learning agility are frequently successful.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22