At Akuna Capital, a Data Scientist operates at the highly competitive intersection of technology, mathematical modeling, and financial markets. Unlike traditional technology firms where data science might focus on user growth or product analytics, data science at Akuna Capital directly drives trading decisions, market predictions, and strategy execution. You will be responsible for extracting signal from noisy market data, designing predictive models, and building high-performance systems that can react to market events in microseconds.
The team specializes in pricing complex financial derivatives and trading on prediction markets, including sports, politics, and macroeconomics. As a Data Scientist, your models will directly influence capital allocation and risk management. This means the feedback loop on your work is incredibly fast: a model improvement can be deployed and start generating trading revenue almost immediately, making the role exceptionally high-impact and intellectually rewarding.
To succeed in this position, you must possess a rare combination of elite mathematical intuition, strong software engineering foundations, and a deep curiosity about market dynamics. prides itself on its flat structure and collaborative environment, meaning you will work closely with traders, quantitative researchers, and software engineers to turn abstract mathematical theories into highly profitable trading algorithms.
Online Coding Assessment
reportedMuch of what gets scored here happens out loud while you type. Nobody can see your reasoning inside a half-written query, so five silent minutes read as being stuck even when they are not. State the plan in plain language first: which tables, what grain you are aggregating to, and the one filter that defines the population. Then write it. The narration doubles as insurance, because a wrong plan gets caught early and cheaply while a wrong query gets caught at the end with no time left to redo it. A timed statistics section, where one exists, is a separate test with its own clock.
What to demonstrate
- Whether the query you write matches the plan you just described
- What you do with a hint, meaning whether the correction gets absorbed or the first approach gets defended
- Whether you can debug your own wrong output by reading the result set and naming which part of the query produced the anomaly
How to prepare
- Solve three problems while screen-sharing into a recording, then watch it back and mark every stretch longer than thirty seconds where you said nothing
- Practise compressing the plan into one sentence before typing, then check afterwards whether the finished query actually matched it
- Time yourself on statistics questions that carry a business reading, such as what a confidence interval does and does not claim, rather than re-reading notes without a clock
Video Math Assessment
reportedBefore anything else, this round is a reading test. You are given a small schema and a question phrased in business language, and most of the difficulty sits in the gap between them. Who counts as an active user, does a refunded order still count as an order, is that date column an event time or a load time. Weak answers start typing immediately and compute something precise about the wrong population. Strong ones pin the definition in one sentence, name the column that encodes it, then write the query. On a timed assessment with nobody to tell, write the definition in a comment anyway.
What to demonstrate
- Whether an ambiguous term becomes a specific column and filter before any computation happens
- Whether you read the schema for keys and cardinality rather than only for column names
- Whether the result answers the question at the grain it was asked at, per user or per session or per day
How to prepare
- Take three metrics you already use and write down the exact filter and exact grain behind each, then practise stating one of them in a single sentence out loud
- On a schema you have never seen, spend the first minute writing what one row of each table means and which key it is unique on, then predict which joins can duplicate rows
- Rehearse a version where the definition changes halfway through, and edit the query you have instead of starting over
Technical Phone Screens
reportedData Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.
What to demonstrate
- Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
- Whether you name what you have not done instead of stretching to cover every line of the posting
- Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage
How to prepare
- Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
- Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
- Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
Superday
reportedA loop is not scored one interview at a time. The people you meet compare notes afterwards, usually in a meeting you are not in, and the outcome turns on what each of them can say about you when asked. That rewards something other than survival: every room needs one specific thing worth repeating, and none of them can contradict another. The common way to lose is to tell the same project four times with different numbers in it, or to be uniformly fine in a way that leaves nobody with anything to argue for.
What to demonstrate
- Whether your account of a project survives being told twice, with the same scale, the same metric definition and the same numbers each time
- Whether each interviewer leaves with one concrete claim they could make on your behalf later, rather than an absence of complaints
- Whether a question you already answered in an earlier room gets the same answer at the same depth, without visible impatience
How to prepare
- Write a one-page fact sheet for your two or three main projects that fixes the numbers you will quote: rows of data, the metric as a single sentence, the effect you measured and how long the work took. Say them aloud from the sheet until they come out identical every time
- For each kind of room you expect, decide the one sentence you want that interviewer repeating in a debrief, then check during the mock that you said it outright instead of implying it
- Rehearse answering the same project question twice in one sitting, the second time as though you had not just answered it, because the thing that needs fixing is the flatness that creeps into a repeated story
13 candidate reports. Individual accounts describe a particular role and hiring cycle.
AKUNA CAPITAL Quantitative Analyst interview with quant-style math testing
My process started with the usual screening and assessments, then moved into online testing that leaned heavily on quant style math. I reached the later stages, including a final technical evaluation. The last question was the one that tripped me up. I couldn't land the final step cleanly, and I couldn't tell whether that meant I was done entirely or whether there was any behavioral wrap-up after…
Read full experienceAKUNA CAPITAL Quantitative Analyst interview with two hour-long probability rounds
I went through an assessment-heavy process in which the later rounds shifted from coding toward deeper probability. After passing the earlier online steps, I reached a technical screen focused on probability and statistics. It was genuinely hard and reminded me of the tougher Greenbook-style questions people mention for quant research. As I progressed, the final interviews focused almost entirely…
Read full experienceAKUNA CAPITAL Software Engineer interview with four C++ technical rounds
My process started with a HackerRank online assessment and quickly moved into several technical interviews. The first programming question didn’t feel like a typical LeetCode problem. It was more like a medium-level simulation, and it pushed me to think about how the system behaved instead of simply deriving a clean algorithm. After the assessment, I went through four technical interview rounds.…
Read full experienceAKUNA CAPITAL Software Engineer HackerRank OA and live coding screen
After applying, I was immediately routed into a HackerRank OA that I had to complete in about an hour. Once I finished it, I received an invitation to a first-round technical discussion. It took place over Zoom, and before the call I was sent a coding link where I would write the solution. The flow was fairly standard: an OA followed by a real-time coding screen focused on how I approached the pr…
Read full experienceAKUNA CAPITAL Software Engineer HackerRank OA, coding rounds, and final day
I went through an OA and then a series of technical rounds that ended with a demanding final day. The initial HackerRank assessment was followed by an exchange-engine-style coding focus in a later interview, where I had to implement a simplified version and explain what I was doing as I went. By the final stage, the interviews were intense and back-to-back. There were four in a row, including tec…
Read full experiencePracHub editorial advice for the preparation topics above.
Reporting the best backtest out of many trials as if it were a single pre-registered test.
The maximum of N noisy Sharpe estimates grows roughly like the standard error times sqrt(2 ln N) even when every underlying strategy has zero edge, so with a few hundred variants an in-sample Sharpe near 1 is the expected result of pure noise. Worse, the search is rarely counted honestly: parameter sweeps, universe changes, date-range choices and feature variants all count as trials. Quote the number of configurations tried, deflate the Sharpe for it, and keep a genuinely untouched holdout period. Note also that the asymptotic standard error of a Sharpe estimate is approximately sqrt((1 + SR^2/2)/T) for i.i.d. normal returns, which for three years of daily data is roughly 0.33, so two strategies differing by 0.3 in Sharpe are not distinguishable.
Judging execution quality against interval VWAP and treating a favourable number as proof of good trading.
Interval VWAP is a benchmark the trader partly determines: trading in line with volume tracks VWAP almost by construction, and stretching an order over a longer interval makes the benchmark easier while exposing the position to price drift that the benchmark never charges. Arrival price is the benchmark aligned with the decision, because it charges both the spread and the drift between decision and completion, including the unfilled remainder. Reporting both, and reporting the opportunity cost of unfilled quantity, is what separates a real TCA from a flattering one.
Never asking what decision the analysis will inform
Open with who makes the decision, what the options are, and by when. The answer determines the precision you need, the segments worth cutting, and whether an observational read suffices or an experiment is required.
Ignoring interference between units in a marketplace experiment
Ask whether one unit's treatment can change another unit's outcome through shared inventory, a matching pool, a social graph or a common budget. Where it can, randomise at a level that contains the spillover, such as region or time slice, and say explicitly what that costs you in statistical power.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Walk through a Bayesian probability scenario to update the likelihood …
Walk through a Bayesian probability scenario to update the likelihood of a market event given a sequence of noisy signals.
Approach
- Translate the result into the decision it informs, in one plain sentence.
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
- Write down the assumption the method needs before you use the method.
Follow-up
- What sample size would you need to detect an effect half this size?
- How would you explain this result to someone who does not know statistics?
Explain how you would set up a robust cross-validation scheme for time…
Explain how you would set up a robust cross-validation scheme for time-series market data to prevent lookahead bias.
Approach
- Say what the estimate is of, and over what population it generalises.
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
- Write down the assumption the method needs before you use the method.
Follow-up
- Which assumption here is most likely to be violated in practice?
- How would you explain this result to someone who does not know statistics?
Solve a counting problem: If you have a bowl of noodles and randomly t…
Solve a counting problem: If you have a bowl of noodles and randomly tie ends together, what is the expected number of closed loops formed?
Approach
- Translate the result into the decision it informs, in one plain sentence.
- Say what the estimate is of, and over what population it generalises.
- Sanity-check the answer against a simple bound or a simulated case.
Follow-up
- How would you explain this result to someone who does not know statistics?
- Which assumption here is most likely to be violated in practice?
Write an algorithm to remove duplicate characters from a string in-pla…
Write an algorithm to remove duplicate characters from a string in-place while minimizing memory overhead.
Approach
- Check what information would not exist at prediction time, and exclude it.
- Set a baseline first, so any model has something honest to beat.
- Pick an evaluation metric that matches the cost of each error type, not a default.
Follow-up
- What would you monitor after launch to know the model is still valid?
- Where could label leakage enter this setup?
Walk through the assumptions of ordinary least squares (OLS) linear re…
Walk through the assumptions of ordinary least squares (OLS) linear regression and explain how you diagnose and correct for heteroscedasticity.
Approach
- Say how the offline result would be validated online before it is trusted.
- Set a baseline first, so any model has something honest to beat.
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
Follow-up
- Where could label leakage enter this setup?
- What would you monitor after launch to know the model is still valid?
Annualized information ratio with an honest standard error
You are given active, one row per (account_id, business_date) with active_return: the daily arithmetic portfolio-minus-benchmark return, net of all costs, as a decimal. Three accounts, roughly 750 business days each. Rows are absent on exchange holidays and on days before an account's funded_date. Per account, compute the annualized information ratio (mean times 252, divided by standard deviation times sqrt(252)) and an approximate standard error for that ratio, then report a 95% interval. Do not reindex onto a calendar. State what T you used and why.
Approach
- Drop null active_return per account rather than filling it. T must be the number of days the account actually traded: holidays and pre-funding days are not evidence, and treating them as such changes both the point estimate and the interval.
- Compute the daily mean m and the daily standard deviation s with ddof=1. The annualized IR is (m/s)*sqrt(252), because the 252 in the numerator and the sqrt(252) in the denominator collapse to a single sqrt(252) factor. Say that out loud instead of writing two separate annualizations that can drift apart.
- Take the standard error at the frequency the statistic is estimated at: SE(SR_daily) is approximately sqrt((1 + SR_daily^2/2)/T) for i.i.d. normal returns, then scale it by sqrt(252), the same factor as the point estimate. At daily frequency SR_daily^2/2 is of order 1e-3, so the SE is effectively sqrt(252/T) and depends only on elapsed years.
- Report IR plus or minus 1.96*SE per account, and check the lag-1 autocorrelation of active_return. The i.i.d. assumption behind that SE is the same assumption that justifies the sqrt(252) scaling, so if the series is autocorrelated both numbers need widening and you should say by how much.
Worked solution 20 min
- groupby('account_id'), dropna on active_return, and record n per account before anything else.
- Per account compute m = mean, s = std(ddof=1), ir = m/s*sqrt(252).
- sr_d = m/s; se_ann = sqrt((1 + sr_d**2/2)/n)sqrt(252); interval = ir +/- 1.96se_ann.
- Compute lag-1 autocorrelation of active_return per account and print it in the same table as the interval.
Follow-up
- The shortest account has 14 months of history. How much of the spread between the best and worst account IR is explainable by sampling noise alone?
- One account carries mark_source = 'vendor_eval' on 30% of days. What does that do to the denominator, and to the independence assumption behind the standard error?
- How many years of daily data would you need to distinguish an IR of 0.8 from an IR of 1.1 at 95% confidence?
Design a custom data structure that supports `get_max`, `get_mean`, an…
Design a custom data structure that supports get_max, get_mean, and get_mode operations in O(1) time complexity.
Approach
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
Follow-up
- How would you verify this result without re-running the same query?
- What breaks if events arrive late or out of order?
Write a simulation in Python or C++ to numerically calculate the area …
Write a simulation in Python or C++ to numerically calculate the area of a circle using a Monte Carlo approach.
Approach
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Say which table is the grain you start from, and join outward from it.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
Follow-up
- How does the query change if the join becomes one-to-many?
- What breaks if events arrive late or out of order?
Solve a dynamic programming problem to find the optimal sequence of tr…
Solve a dynamic programming problem to find the optimal sequence of trades given a historical price series.
Approach
- Say which table is the grain you start from, and join outward from it.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
Follow-up
- How does the query change if the join becomes one-to-many?
- What breaks if events arrive late or out of order?
Deduplicate a restated signal table into a leak-free panel
signal_score may hold several rows per (signal_id, instrument_id, as_of_date) because reruns rewrite history; each row carries model_version, knowledge_ts, computed_at, is_backfilled, zscore_xs, decile_rank, coverage_flag and universe_id. Build the panel a backtest may legitimately use on trading date D: for every instrument in universe_id = 'liquid_us_1500', the single most recent score whose knowledge_ts is at or before D 21:00 UTC, preferring a live row over a backfilled one. Return instrument_id, as_of_date, zscore_xs, decile_rank.
Approach
- Filter on
knowledge_ts, never onas_of_date.as_of_datesays what the score describes;knowledge_tssays the earliest instant every input was observable. Put it in the WHERE so it prunes before the window is evaluated. - Rank within the key:
ROW_NUMBER() OVER (PARTITION BY signal_id, instrument_id ORDER BY as_of_date DESC, is_backfilled ASC, knowledge_ts DESC, computed_at DESC)and keep rank = 1. In Postgres FALSE sorts before TRUE, sois_backfilled ASCis the live-row preference, written down rather than assumed. - Prefer ROW_NUMBER to a
MAX(computed_at)group-then-rejoin: the rejoin duplicates rows whenever two reruns share acomputed_at, which is exactly what a batch job produces. - Decide what
coverage_flag IN ('stale','imputed')means for this panel and encode it. Dropping those rows silently changes the universe size day to day, which surfaces later as an unexplained jump in measured IC rather than as a missing-data problem. - Express the cutoff as a parameter and the whole thing as a CTE so the backtest loop reuses one query text per date instead of a hand-edited copy.
Worked solution 25 min
- Write the filtered CTE:
universe_id = 'liquid_us_1500'andknowledge_ts <= :cutoff. - Add the ROW_NUMBER window with the four-key ordering and select rank = 1.
- Assert one row per instrument with a COUNT(*) versus COUNT(DISTINCT instrument_id) probe.
- Run the same query with
as_of_date <= Dsubstituted for the knowledge_ts filter and count how many rows differ. - Inspect the differing rows: they should be predominantly
is_backfilled = TRUEwithcomputed_atafter D.
Follow-up
- Write the query that proves the panel is leak-free. What would a violation look like in measured IC, and roughly how large would you expect the inflation to be?
- Produce every trading date in one pass instead of one query per date. What does that cost in plan shape?
- Two
model_versionvalues are live at once during a migration. How does the tie-break change, and who decides?
Ramp new accounts fast or cheap, and defend the exchange rate
Newly funded accounts must reach 90 percent of target_gross_exposure_pct within 20 trading days of funded_date; that is today's primary metric. Raising it means ramping faster, which raises implementation shortfall because temporary impact scales roughly with volatility times the square root of participation. You have account_mandate (funded_date, target_gross_exposure_pct, benchmark_id, effective_from, effective_to), position_daily (market_value_base, weight_pct_nav) and parent_order. Pick the primary, pick the guardrail, and state the exchange rate at which you would trade one against the other, in basis points rather than in words.
Approach
- Name the objective both metrics are proxying: an unramped day costs expected active return plus unwanted benchmark-relative risk, and a faster ramp costs one-off impact. Both convert to basis points of account NAV, so the conflict is resolvable by arithmetic rather than by argument.
- Price the slow side. Daily expected active return is annualized expected alpha over 252, multiplied by the exposure shortfall as a fraction of target. A linear ramp averages a 50 percent shortfall over its window, so at 300 basis points of expected annual alpha a 20-day ramp forgoes roughly 12 basis points. Flag the precondition: this is an expectation with a wide error bar, which makes it the uncertain side of the trade.
- Price the fast side correctly, because this is where most answers break. Total ramp notional is fixed, so halving the ramp window doubles participation, raises per-unit impact by about the square root of two, and raises total ramp cost by about 41 percent rather than 100 percent. The window scaling is proportional to one over the square root of the number of days.
- Choose the primary as ramp completion and the guardrail as ramp shortfall in basis points of ramped notional with a hard cap, on the stated reasoning that the impact cost is realized and permanent while the forgone alpha is an estimate. Set the cap where marginal impact per day saved exceeds daily expected active return.
- Do not paper over the conflict with a blended score. Report both, and if a single number is demanded, define total ramp cost as impact paid plus expected alpha forgone, and attach the alpha assumption to the same line so the reader can see whose estimate is carrying it.
Worked solution 30 min
- From position_daily compute each account's daily gross exposure as the sum of absolute market_value_base over NAV, and find the first business_date reaching 0.9 times target_gross_exposure_pct, joining account_mandate on business_date between effective_from and effective_to.
- From parent_order and execution_fill compute arrival shortfall in basis points restricted to orders with decision_ts inside each account's ramp window.
- Fit the impact coefficient by regressing shortfall basis points on the square root of filled_qty over adv_20d interacted with trailing realized volatility, and keep the slope.
- Project total ramp impact for windows of 10, 15, 20 and 30 trading days from that curve, and project alpha forgone as 0.5 times the daily expected active return times the window length.
- Add the two columns and pick the minimum, then state the alpha level at which the ranking flips.
Follow-up
- Two accounts funded the same week ramp at different speeds because one carries a restricted list. Should the metric penalise the desk, and how would you separate that in the data?
- How does the exchange rate move for a strategy whose expected alpha is 80 basis points a year rather than 300?
- The impact curve was fitted on the desk's ordinary daily orders. What is wrong with using it for a ramp, and which direction does the error run?
Measure forced trimming at a compliance threshold you cannot randomise
Compliance forces a trim whenever a position's weight_pct_nav in position_daily closes above the account's max_single_name_weight_pct from account_mandate, a hard 5.00 percent cap evaluated at the close and trimmed the next morning. Portfolio managers claim forced trimming destroys alpha, and compliance cannot be randomised. Over three years about 4,100 account-instrument-days sit within one percentage point of the cap. Identify the effect of forced trimming on the position's contribution to active return over the following twenty business days, and state what your estimate is and is not.
Approach
- Set it up as a sharp regression discontinuity in weight_pct_nav centred on the cap. Treatment is a deterministic function of the running variable, so the identifying assumption is only that potential outcomes are continuous at 5.00 percent. Nothing has to be randomised, which is the whole reason the design applies here.
- Test for manipulation before estimating anything. A density test on the running variable around the cap asks whether managers shave positions to 4.95 percent to stay under it. If the density has a hole just above the threshold, the design is dead and the correct deliverable is that the effect is not identified, not a number with a caveat.
- Estimate with local linear regression on each side, a triangular kernel and a data-driven mean-squared-error-optimal bandwidth rather than one chosen by eye, using robust bias-corrected confidence intervals. Report the estimate at that bandwidth and at half and double it to show the result is not a bandwidth artefact.
- State the estimand precisely. This is a local average treatment effect at a 5.00 percent weight, for positions large enough to reach the cap, in accounts whose mandate sets the cap there. It says nothing about trimming a 2 percent position and nothing about accounts with a 3 percent cap, unless you pool across caps by re-centring each observation on its own threshold, which is legitimate and enlarges the sample at the cost of assuming the discontinuity is the same at every cap level.
- Handle the two schema traps explicitly. account_mandate is SCD2, so an account whose cap was amended mid-history has two thresholds and each observation must be compared to the cap in force on its own business_date. And weight_pct_nav is derived from mark_px, so rows with mark_source of 'stale_prior_day' or 'model' carry a running variable measured with error, which in a discontinuity design attenuates the estimate and smears the jump. Exclude those rows and report how many were dropped.
- Say what would make you switch designs. If the cap is soft in practice, with a tolerance band, a grace day or discretionary waivers, the design is fuzzy rather than sharp: the indicator for being above the cap instruments for actually being trimmed, the estimate scales by the compliance rate, and the first-stage strength becomes something you have to report.
Follow-up
- How many observations actually fall inside the optimal bandwidth, and what effect size can that number detect?
- If the density test fails, what fallback design would you accept, and what would it now assume?
- Positions that reach the cap are the highest-conviction names in the book. Does that break the design, and if not, precisely why not?
Return dispersion widened inside one strategy composite
Trailing-12-month net returns for accounts running one strategy used to sit within about 40 bps of each other. This quarter the cross-sectional standard deviation is 210 bps and the AUM-weighted mandate success rate fell. The investment team insists the model is identical across accounts. You have position_daily (business_date, account_id, strategy_id, instrument_id, market_value_base, weight_pct_nav, is_restricted, fx_rate_to_base) and account_mandate (account_id, funded_date, inception_date, status, target_gross_exposure_pct, max_single_name_weight_pct, base_currency, mgmt_fee_bps, effective_from, effective_to). Explain the dispersion with an ordered checklist and say which part is a problem.
Approach
- Rank accounts by trailing-12-month net return and inspect the tail before averaging anything. A standard deviation is a distribution statistic and two accounts out of forty can produce it; the number alone does not say whether the cause is systemic or concentrated.
- Cut by months since funded_date and by status. An account in 'ramping' holds cash and sits below target_gross_exposure_pct, so it earns a diluted version of the strategy's return. In a rising market that reads as underperformance and in a falling one as outperformance, and both are mechanical rather than informative.
- Compute realized gross exposure per account-day from position_daily, as the sum of absolute market_value_base over NAV, and compare it with target_gross_exposure_pct from the account_mandate version valid on that business_date. Join on the effective_from and effective_to range, not on the current row, because mandate terms change mid-life and the current row would apply today's target to last year's book.
- Sort the remaining explanations into expected and not expected. Names excluded by is_restricted, base_currency differences moving through fx_rate_to_base, and mgmt_fee_bps differences on a net-of-fee metric are all expected. A weight breaching max_single_name_weight_pct, or an account holding instruments the others do not, is not.
- Report dispersion decomposed by cause with dollars attached. The client question is never why sigma is 210 bps; it is why this account returned less than the one on the last page.
Follow-up
- Two accounts funded the same day diverge by 90 bps and both reached target exposure inside 20 trading days. Where do you look next?
- The AUM-weighted success rate only counts accounts with 36 months of history. How does that interact with what you just found, and does it flatter the number or penalise it?
- How would you present cash drag to a client so it reads as a fact about funding rather than an excuse?
For a candidate whose interviews will centre on A/B testing, metric movement and causal claims. Design comes before arithmetic, arithmetic before analysis, and the week ends by rehearsing the readout rather than the derivation.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Design one test end to end on paper
- Take a single feature change and write the full design: randomization unit, the exact point of exposure, the primary metric with its grain, guardrails, allocation, planned duration, and the decision rule committed before any data exists.
- Write why the randomization unit must sit at or above the level where treatment can spill over, and give one case where user-level randomization is still contaminated (shared accounts or devices, or two participants in the same marketplace).
- State in advance what you will do if the primary metric is flat while a secondary metric is significant.
Deliverable: A one-page test design with a decision rule written before launch.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Power arithmetic until it is automatic
- Compute required sample size per arm for a binary metric with the normal approximation, n is approximately 2 times (z for alpha/2 plus z for power) squared times p(1 minus p) divided by delta squared, for baselines of 2, 10 and 40 percent at a 5 percent relative lift, and note that for a fixed relative lift the requirement falls as the baseline rises because delta grows proportionally with p.
- Redo the calculation for a continuous metric using variance in place of p(1 minus p), and show why a heavy-tailed quantity such as revenue per user needs either far more traffic or a capped version with a stated cap.
- Convert one of the results into weeks given a weekly eligible traffic figure, then list the two honest ways to shorten it (accept a larger detectable effect, or reduce variance) and write why quietly lowering the power target is a decision to miss more real wins, not a speedup.
Deliverable: A small script or sheet that maps baseline, minimum detectable effect, alpha and power to sample size and weeks, cross-checked against a published calculator.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Variance and the unit-of-analysis problem
- Take a ratio metric whose denominator is not the randomization unit (clicks per session, randomized by user) and compute the standard error twice, once naively at session level and once by the delta method or a user-level bootstrap, then record how much the naive version understates it.
- Implement CUPED on simulated data: choose a pre-period covariate X measured before assignment, estimate theta as Cov(Y, X) divided by Var(X), and analyse Y minus theta times (X minus its mean) in place of Y. Confirm the variance of the adjusted outcome equals the raw variance multiplied by one minus the squared correlation between Y and X, so a correlation of 0.45 removes about 20 percent of the variance and not 80.
- Now run that simulation a few hundred times and confirm the adjusted effect estimate is unbiased for the same effect rather than numerically identical to the raw one. Within any single run the two differ, sometimes by a large fraction of the true effect, because the two arms' pre-period covariate means never coincide exactly in a finite sample; they agree in expectation, which is the property that matters and the one to state out loud.
Deliverable: A notebook showing the adjusted estimator with a measurably smaller variance than the raw one, plus a repeated-simulation table showing the two estimators agreeing on average while differing run by run.
Practice prompt ↗Practice prompt ↗04Validity threats you can actually test for
- Run a sample ratio mismatch check as a chi-square goodness-of-fit test against the intended allocation, and write the three causes you would chase first (assignment logged before exposure, an arm-specific redirect or load failure, bot filtering applied asymmetrically).
- Simulate peeking: generate A/A data, test daily at alpha 0.05 across 14 looks, record the inflated false positive rate, then apply an alpha-spending boundary or commit to a fixed horizon and confirm the rate returns to nominal.
- Write how you would separate a novelty effect from a durable lift using the treatment effect plotted against days since first exposure, and what shape would change your recommendation.
Deliverable: One table showing the peeking false positive rate before and after correction, plus a written SRM triage list.
Practice prompt ↗Practice prompt ↗Worked solution ↗05When randomization is not available
- Write the identifying assumption for difference-in-differences (parallel trends in the absence of treatment), then plot pre-period trends for two candidate control groups and justify rejecting one of them.
- Design a switchback test for a change where user-level randomization would leak across participants, choosing a time-block length against the carryover you expect and saying how you would detect carryover in the data.
- List what an interrupted time series or a synthetic control buys you and the one thing neither can rule out: an unobserved shock that coincides with the launch.
Deliverable: A one-page memo recommending a single quasi-experimental design and naming its weakest assumption explicitly.
Practice prompt ↗Practice prompt ↗06The readout query
- Write the assignment-to-exposure join that returns exactly one row per unit per experiment, and handle units appearing in both arms by excluding and counting them rather than silently keeping one.
- Compute the per-arm metric, its variance and the relative lift with a confidence interval in SQL, then reproduce the identical numbers in a notebook as a cross-check.
- Add a segment breakdown and write the sentence that keeps it from being p-hacking: segments declared in advance, everything else reported as exploratory and corrected for multiplicity.
Deliverable: A single query that outputs the full readout table, matched to a notebook recomputation.
Practice prompt ↗Practice prompt ↗07Present it to someone who will not read the appendix
- Give a 10-minute readout of a real or simulated experiment in the order decision, number, uncertainty, caveat.
- Have your listener ask "can we ship it" in the case where the primary is flat and a guardrail moved, and answer with a recommendation rather than a request for more data.
- Rewrite your opening line so the recommendation lands before any methodology.
Deliverable: A one-page readout whose first line is the recommendation.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
An answer without a quantity is hard to interrogate, so interviewers keep probing until they find one. Come with the baseline, the change, the window it was measured over, and how confident you were. If the effect never got measured, say so and say what you would have measured. Fabricated precision is worse than an honest gap.
Disagreeing with a product manager over an account leaderboard
A product manager wants to ship a client portal widget ranking every separately managed account against its peers in the same strategy composite, by trailing 12-month net return. You believe the ranking will mostly order accounts by mandate mechanics rather than by anything a client can act on. You have position_daily, account_mandate (SCD2) and benchmark returns for 140 accounts in the composite. Build the case and bring a counter-proposal you would ship. The product manager has a launch date and a client asking for exactly this.
Approach
- What is probed: whether you can lose the feature and keep the working relationship, meaning your disagreement arrives with a shippable alternative rather than as a veto.
- Measure the dispersion before arguing about it. Compute the cross-sectional standard deviation of trailing 12-month net return across the 140 accounts. If it is small, the product manager is right and you are not, and you want to know that before the meeting rather than during it.
- Decompose the dispersion into causes you can name from the tables: time spent ramping between funded_date and the date gross exposure reached 90 percent of target_gross_exposure_pct, average cash weight over the period, restricted names via position_daily.is_restricted, single-name cap differences across account_mandate SCD2 versions, and fee schedule including whether a performance fee crystallised above the high-water mark. Report the share each explains and the residual.
- Convert the finding into the client's decision, because that is what moves a product manager. If most dispersion is mandate mechanics, the leaderboard tells a client to change managers when the honest action is to relax a constraint or fund fully. A wrong action is an argument; a noisy statistic is a preference.
- Bring the alternative that keeps the launch date: the same widget, showing the account's return against its own benchmark and its own constraint set, with a named driver line such as your restricted list cost 34 bps, instead of a rank. It answers what the client actually asked and it survives a phone call.
- Pre-commit to being wrong. If the residual dominates the decomposition, the leaderboard is measuring something real, and saying so in the same memo is what makes the rest of it credible next time.
Follow-up
- Dispersion is 180 bps and mandate mechanics explain 40 percent of it. What do you ship?
- The client asked for a rank by name. Do they get one, and what do you put next to it?
- How do you keep this from becoming a standing veto on anything this product manager proposes?
Correcting a published TCA note that dropped cancelled orders
Three months ago you published a monthly transaction cost note concluding that algo A beat algo B by 12 bps of implementation shortfall, and the desk moved flow to A. You have since found that your query filtered parent_order to status = 'filled', so cancelled and expired orders never entered the numerator or the denominator. A cancels more often than B. Write the correction: what you tell the desk, in what order, and what changes so this class of error is caught rather than found.
Approach
- What is probed: whether you report your own error at the speed and specificity you would demand from someone else. The clock starts when you know, not when the rework is finished.
- Size the error before describing it. Recompute both algos with cancelled and expired parent orders included and opportunity cost charged on order_qty minus filled_qty at the terminal mid, as the shortfall definition requires. The gap may shrink, vanish or reverse, and I do not yet know the sign is a legitimate first message only if it arrives within hours.
- Name the mechanism that makes the bias directional rather than noisy. Orders get cancelled disproportionately when price runs away from the decision, so excluding them removes the worst outcomes, and it removes more of them from the algo that cancels more. That is why the filter flattered A specifically.
- Tell the desk head in person before the corrected note circulates, and lead with the operational consequence rather than the methodology: flow moved on a wrong number, and here is what to do with it today.
- Make the fix structural rather than a promise to be careful. Add an assertion to the TCA job that the count and notional of parent orders in the report equal the count and notional in parent_order over the window, grouped by status, so a silent population drop fails the job instead of shipping.
Follow-up
- The corrected numbers still favour A, by 3 bps. Does the desk need to hear from you at all, and why?
- Who else built on that note, and how do you find out rather than guess?
- Your review process passed a status filter. What specifically in it was supposed to catch a population change?
Defending a backtest correction that removes an allocated strategy
A colleague's cross-sectional equity signal received capital last month on a backtest showing annualized Sharpe 1.6 over three years. Reproducing it, you find the panel filters signal_score on as_of_date rather than knowledge_ts. Rerunning with the knowledge_ts predicate, holding universe_id, date range, cost model and rebalance schedule fixed, gives Sharpe 0.5 and moves mean daily IC from 0.041 to 0.012. The colleague is senior and presented the original result. You have one meeting with the research lead and the portfolio manager. Deliver a recommendation on whether the allocation stands, and the evidence behind it.
Approach
- What is probed: whether a quantitative objection survives social cost. Lead with the defect and the single-variable reproduction, never with a judgement about the colleague. The claim under test is one predicate, not a person's competence.
- Rerun both versions from one code path with one line changed, and say so explicitly. Holding universe_id, date range, cost model and rebalance schedule fixed leaves the 1.1 Sharpe gap exactly one candidate cause, which is what makes the result arguable on its merits instead of on whose code is trusted.
- Pair the comparison before quoting any error bar. For modest Sharpe ratios the annualized standard error of a Sharpe estimate is approximately 1 over the square root of the number of years, so three years of daily data gives roughly 0.58 and two marginal estimates 1.1 apart are only about two standard errors apart. But the two runs are the same returns except where the leak bites, so test the daily difference series directly: its standard error is far smaller, and the pairing is what turns a marginal result into a decisive one.
- Exhibit the mechanism, not just the size. Rank instrument-days by their contribution to the return difference between the two runs and show that the top contributors carry is_backfilled TRUE or coverage_flag 'stale', with knowledge_ts postdating as_of_date by the vendor's restatement lag. A named mechanism is falsifiable; a Sharpe delta alone becomes an argument about your code.
- Bring a decision rather than only a finding: the position size the corrected Sharpe supports, and an untouched out-of-sample window that would settle it either way. Being right with no path forward is how a correct objection gets overruled.
Follow-up
- Your colleague reruns it and gets 0.9 rather than 0.5. What do you do with the discrepancy before the meeting?
- The strategy is up since funding. Does live P and L change your recommendation, and how much of it would?
- What would have caught this before the allocation, and why did the existing review not catch it?
- 01
A product manager wants to ship a client portal widget ranking every separately managed account against its peers in the same strategy composite, by trailing 12-month net return. You believe the ranking will mostly order accounts by mandate mechanics rather than by anything a client can act on. You have position_daily, account_mandate (SCD2) and benchmark returns for 140 accounts in the composite. Build the case and bring a counter-proposal you would ship. The product manager has a launch date and a client asking for exactly this.
- 02
Three months ago you published a monthly transaction cost note concluding that algo A beat algo B by 12 bps of implementation shortfall, and the desk moved flow to A. You have since found that your query filtered parent_order to status = 'filled', so cancelled and expired orders never entered the numerator or the denominator. A cancels more often than B. Write the correction: what you tell the desk, in what order, and what changes so this class of error is caught rather than found.
- 03
A colleague's cross-sectional equity signal received capital last month on a backtest showing annualized Sharpe 1.6 over three years. Reproducing it, you find the panel filters signal_score on as_of_date rather than knowledge_ts. Rerunning with the knowledge_ts predicate, holding universe_id, date range, cost model and rebalance schedule fixed, gives Sharpe 0.5 and moves mean daily IC from 0.041 to 0.012. The colleague is senior and presented the original result. You have one meeting with the research lead and the portfolio manager. Deliver a recommendation on whether the allocation stands, and the evidence behind it.
Is this an official Akuna Capital interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Akuna Capital. Rounds and questions reflect what candidates have reported, not a process Akuna Capital has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult is the Akuna Capital Data Scientist interview process?
The process is highly difficult and intensely quantitative. Unlike general tech companies, Akuna Capital places a massive emphasis on speed-based coding assessments and rigorous, college-level probability and mathematics questions.
PracHub interview research ↗What is the typical timeline from the initial application to an offer?
The process typically takes between 4 to 8 weeks. Akuna Capital is known for moving quickly between rounds, but scheduling the multi-round final Superday can sometimes introduce minor delays.
PracHub interview research ↗Do I need to have a background in finance or trading to apply?
No, prior finance experience is not strictly required. Akuna Capital values raw mathematical talent, strong coding foundations, and structured problem-solving, and they will teach you the necessary financial concepts on the job.
PracHub interview research ↗How should I prepare for the one-sided video math assessment?
Practice solving probability and linear algebra problems while talking out loud. The graders are not just looking for the correct numerical answer; they are evaluating the clarity, structure, and logic of your verbal reasoning under a strict 5-minute per question limit.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22