This guide covers what a Data Scientist at SoFi is expected to do and how to prepare for the interview.
Recruiter Screen
reportedWhoever runs this call is usually not a practitioner. They take notes, and a hiring manager skims those notes later, so the real question is whether your work survives being written down by someone outside the field. Test every project sentence against that: could a non-specialist repeat it correctly without knowing what a propensity score is? Carry a plain-language version of each project and one reason you want this particular role that you could not copy onto another application. Vagueness at this stage reads as inexperience, even when the underlying work was genuinely deep.
What to demonstrate
- Whether a non-specialist can restate your projects accurately, since their paraphrase is what reaches the hiring manager
- Whether your reason for wanting the role points at the work itself rather than the company's reputation
- Whether your language signals the level being screened for: what you decided yourself versus what you were handed
How to prepare
- Write a two-sentence, jargon-free version of each major project: the question nobody could answer, and the decision your work changed. Read it to someone outside data and have them repeat it back
- Point your 'why this role' answer at something concrete in the job description or the product surface you would be working on, and keep it to two sentences
- Have two questions ready about measurement: which metric the team is held to, and who acts on an analysis once it lands
Hiring Manager Screen
reportedUnderneath the questions about your past work sits a resourcing question. Given four things worth doing and one of you, which gets done and what happens to the rest? Managers ask because that is the daily texture of the job, and because the answer shows whether you rank work by effort or by what it changes. The weak version sorts by personal interest or by whoever asked most insistently. The strong version ties each candidate piece of work to a decision somebody downstream is waiting on, and then names the one you would drop and who you would tell.
What to demonstrate
- Whether you rank work by the decision it unblocks or by how interesting the method is
- How you describe a request you declined, and whether you can say who you said it to
- Whether your sense of how long something takes survives one follow-up question about the messy part
- How you decide something is good enough to hand over unfinished
How to prepare
- Write out your current queue and, next to each item, the decision that stays stalled until it lands. Anything with no waiting decision becomes your example of work you would cut
- Rehearse turning down a plausible stakeholder request out loud, including the smaller alternative you offered instead
- Have one case where you shipped a rough answer early and one where you refused to, with the reason that separated them
Technical Assessment
reportedA handful of shapes account for most of what gets asked in this format: a ranking or deduplication inside groups, a running or rolling total, a period-over-period comparison, and a cohort tracked forward over time. Recognising the shape quickly is most of the speed here; deriving it from scratch while a clock runs is where the time goes. Know that a window function keeps every row while a GROUP BY collapses them, and know which one the question needs. If the exercise is in Python instead of SQL, the same shapes arrive as groupby with transform, shift and merge, and the same grain mistakes are available.
What to demonstrate
- Whether you reach the right construct without a detour, such as ROW_NUMBER over a partition to deduplicate instead of a self-join against a MAX subquery
- Whether you know what your window frame actually is, since adding ORDER BY inside OVER changes the default frame and silently changes a running total
- Whether the thing runs. A near-miss that throws an error scores below a plainer query that returns the right rows.
How to prepare
- Write each of the four shapes once from memory against a small schema and keep the working version somewhere you will reread it: dedupe with ROW_NUMBER, a running total, a month-over-month change with LAG, and a retention table
- Compute one running total twice on data with tied timestamps, once on the default frame and once with ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, and look at where the two disagree
- If Python is on the table, rebuild the dedupe and the running total with groupby and cumsum, then assert the two implementations return identical rows
Virtual Onsite Loop
reportedA day of back-to-back interviews samples your floor, not your ceiling. Four hours in, the habits that carry a good answer are the first to go: restating the question before solving it, asking what the data would have to look like, checking a number before quoting it. What the day decides is whether the tired version of you is still someone to leave alone with an ambiguous problem. The round that sinks a candidate is usually not the hardest one. It is the one immediately after the round that went badly.
What to demonstrate
- Whether the late rounds get the same clarifying questions as the first one, or whether you start answering immediately to save effort
- Whether a weak answer stays in the room it happened in, instead of following you into the next conversation as apology or distraction
- Whether the quality of your questions holds up, since fatigue removes curiosity about the problem before it removes knowledge of the method
How to prepare
- Rehearse the length, not just the content: book four mock interviews of different types in one afternoon with short gaps, because the one you need to observe is the fourth
- Put the two or three questions you ask at the start of any problem on a card in front of you, so that under fatigue it is a habit you run rather than a decision you make
- Decide in advance what the gap between rooms is for: water, one line of notes on anything you promised to follow up, and an explicit close on the round that just ended so it does not travel
- Prepare a different closing question for each interviewer, so the end of a long day does not produce the same one four times
6 candidate reports. Individual accounts describe a particular role and hiring cycle.
SoFi Full Stack Engineer panel interview experience
I started with a standard recruiter phone screen that moved quickly. A third-party technical screen focused on JavaScript, and I entered the main interviews expecting coding and role-relevant evaluation. The next step was a packed panel. It included Algorithms and Data Structures, a React frontend discussion, and system design. I also met with a hiring manager to discuss how I worked and approach…
Read full experienceSoFi Software Engineer interview experience: graph traversal and backtracking
The process felt smooth from the start. The coding screen included a graph-traversal-style question, followed by technical rounds built around familiar problem-solving patterns. One involved generating permutations with backtracking; another asked me to rearrange an array for a prefix-sum-style objective. The rounds were practical and direct. I had room to explain my thinking instead of dealing w…
Read full experienceSoFi Software Engineer Interview Experience — Three-Round Loop, System Design Swapped for a Second Coding Round
View report detailsSoFi Senior Software Engineer Interview Experience — A Relaxed HackerRank Screen, Then a 4-Round Onsite Loop
View report detailsSoFi Software Engineer Interview Experience — Four Rounds, Rejected on the Second Attempt
This was my second time trying to get into SoFi. This time I interviewed with the Invest team. I thought the conversations went well, but I still got rejected. The market really is brutal right now. There were 4 rounds total: a phone screen plus three virtual onsite (VO) rounds. Phone screen: OOD principles, and implementing a multi-threaded task executor. I was given a class TaskExecutor that im…
Read full experiencePracHub editorial advice for the preparation topics above.
Assuming a model is fair because protected attributes are not among its inputs
Postcode, device, tenure, income proxies and even transaction patterns correlate with protected characteristics, so a model can produce a disparate outcome without ever reading the attribute. Credit decisions additionally carry an explainability obligation in many jurisdictions, since a denial has to be accompanied by its principal reasons, which constrains model form and feature engineering rather than being a reporting afterthought. Treating fairness testing and reason-code generation as design constraints from the first model version is far cheaper than retrofitting them to a deployed one.
Using written premium as the denominator of a loss ratio
Premium is written at inception and earned pro rata across the exposure period, so in a growing book written premium runs ahead of earned premium and the loss ratio comes out too low, with the error reversing when the book shrinks. The numerator has the mirror-image problem if it omits incurred-but-not-reported reserves, since recent accident periods then look profitable twice over. Both sides must refer to the same exposure period, which is what an accident-period view at a fixed development age enforces.
Sizing estimates built on unnamed, unrevisable assumptions
Write each assumption as a named number you can change, then show the arithmetic so the interviewer can challenge one input instead of the whole answer. Finish by saying which assumption the result is most sensitive to, which matters more than the point estimate.
Naming a model class before naming the deployment constraints
Set out the latency budget, the label delay, the retraining cadence, the interpretability requirement and the number of labelled examples, then pick the model that fits them. A boosted-tree answer to a problem where each decision must be explained to the affected user is a well-executed answer to the wrong question.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Describe the central limit theorem and its practical application in me…
Describe the central limit theorem and its practical application in metric distribution analysis.
Approach
- Quantify uncertainty explicitly rather than reporting a point estimate alone.
- Sanity-check the answer against a simple bound or a simulated case.
- Write down the assumption the method needs before you use the method.
Follow-up
- What sample size would you need to detect an effect half this size?
- How would you explain this result to someone who does not know statistics?
Clean and transform a messy pandas DataFrame to prepare it for machine…
Clean and transform a messy pandas DataFrame to prepare it for machine learning modeling.
Approach
- Set a baseline first, so any model has something honest to beat.
- Pick an evaluation metric that matches the cost of each error type, not a default.
- Say how the offline result would be validated online before it is trusted.
Follow-up
- Where could label leakage enter this setup?
- What would you monitor after launch to know the model is still valid?
Bootstrap a fraud loss rate that clusters within merchant
You have a per-transaction frame with auth_id, merchant_id, settled_amount_reporting and net_loss_reporting, both already in one reporting currency. Most rows carry zero loss, a few carry large ones, and losses cluster within merchant. Using only numpy's random generator and no resampling helper from any library, write a bootstrap that returns a 95 percent interval for net fraud loss in basis points of settled volume, resampling merchants with replacement and taking all rows belonging to each drawn merchant. Also produce the naive row-level interval and state which you would report.
Approach
- State the estimator before writing it: total net loss divided by total settled volume, times 10,000. It is a ratio of sums, so each replicate recomputes both sums. Averaging per-transaction loss rates instead would weight a five-unit transaction like a five-thousand-unit one.
- Pre-aggregate loss and volume to merchant level once. For a ratio of sums, drawing merchants and taking all their rows is arithmetically identical to drawing merchant-level (loss_sum, volume_sum) pairs, so a replicate becomes one integer draw plus two vectorised sums rather than a groupby inside the loop.
- Draw B replicates of M merchant indices with replacement, where M is the observed merchant count, compute the ratio per replicate, and take the 2.5th and 97.5th percentiles. Say explicitly that this is a percentile interval and that BCa would correct the skew-induced bias if the decision is close.
- Repeat with independent row draws for the naive interval and compare widths on the same replicate count.
- Report the clustered interval. Rows within a merchant share an acceptance profile, a category code and a fraud exposure, so they are not independent, and the row-level interval understates variance by roughly the design effect.
Worked solution 30 min
- Compute the point estimate directly on the full data and keep it for comparison.
- Aggregate to merchant-level loss and volume arrays, record M, and set B to 2,000 with a seeded numpy Generator.
- In a vectorised loop, draw integer indices of shape (B, M), index both arrays, sum along axis 1, and take the ratio times 10,000.
- Repeat for the row-level version using the per-transaction arrays and N draws.
- Take the 2.5 and 97.5 percentiles of each replicate array and report both intervals alongside the point estimate.
Follow-up
- Your clustered interval is three times wider. How do you explain that to someone who wanted a tighter number?
- One merchant accounts for 40 percent of losses. What does that do to the interval, and what would you do about it?
- How does this change if the question is whether two months differ rather than what this month's rate is?
Write a complex SQL query using window functions to find month-over-mo…
Write a complex SQL query using window functions to find month-over-month retention cohorts.
Approach
- State the window function and its partition and ordering out loud before writing it.
- Say which table is the grain you start from, and join outward from it.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
Follow-up
- How would you verify this result without re-running the same query?
- What breaks if events arrive late or out of order?
How would you optimize a slow-running SQL query that joins multiple la…
How would you optimize a slow-running SQL query that joins multiple large tables in Snowflake?
Approach
- Say which table is the grain you start from, and join outward from it.
- State the window function and its partition and ordering out loud before writing it.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
Follow-up
- How does the query change if the join becomes one-to-many?
- How would you verify this result without re-running the same query?
Count-weighted and dollar-weighted approval rates on one currency
Using fct_payment_authorization, report the trailing 7-day authorization approval rate two ways for transaction_currency = 'EUR': count-weighted, and dollar-weighted on amount_minor. Exclude is_reversal = true, exclude incremental authorizations (parent_auth_id not null), and exclude zero-amount account verifications. auth_result = 'approved' is the numerator; the four declined_* values make up the rest of the denominator. Return channel, attempts, approved_attempts, approval_rate_count and approval_rate_value. State every exclusion and its reason before you write the SELECT.
Approach
- Say the denominator out loud first: attempts on a single transaction currency, excluding reversals, incremental authorizations and zero-amount verifications, because none of those is a purchase attempt a merchant is trying to get approved.
- Filter requested_at against a half-open interval (>= start AND < end) so the boundary day is neither dropped nor double counted.
- Compute both rates in one pass with FILTER clauses: COUNT() FILTER (WHERE auth_result = 'approved') over COUNT(), and SUM(amount_minor) FILTER (WHERE auth_result = 'approved') over SUM(amount_minor).
- Cast one side of each ratio to numeric before dividing, since amount_minor and the counts are integers and integer division silently truncates to zero.
- Group by channel and sort by the value-weighted rate, then read the gap between the two rates as a statement about where the declines sit rather than as noise.
Worked solution 20 min
- Write the exclusion list as comments above the query: is_reversal = false, parent_auth_id is null, amount_minor > 0, transaction_currency = 'EUR'.
- Build a single aggregate query over fct_payment_authorization with a half-open requested_at predicate and those four filters.
- Emit attempts, approved_attempts, approval_rate_count and approval_rate_value with FILTER clauses and a numeric cast on the numerator.
- Group by channel, order by approval_rate_value ascending so the worst channel is on top.
Follow-up
- The two rates diverge by four points on the ecommerce channel but agree on card_present. What does that tell you, and what would you cut next?
- How would you extend this to all currencies without summing amount_minor across them?
- Which of the four decline reasons belong in the denominator of a rate you would put in front of a risk team, and which are really the network's problem?
Can you describe a case study you worked on?
Can you describe a case study you worked on?
Approach
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Name one primary metric, then the guardrail that stops it being gamed.
- Restate the decision this analysis has to support, and who acts on the answer.
Follow-up
- Which segment would you cut first, and what would that rule out?
- What would you do if the primary metric and the guardrail moved in opposite directions?
How do you account for seasonality and external market trends when ana…
How do you account for seasonality and external market trends when analyzing test results?
Approach
- Decompose the metric into the rates that drive it, and say which one you would check first.
- State what result would change your recommendation, so the answer is falsifiable.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- Which segment would you cut first, and what would that rule out?
- How would you detect that the metric is being gamed rather than genuinely improving?
What key performance indicators matter most for a lending platform?
What key performance indicators matter most for a lending platform?
Approach
- Decompose the metric into the rates that drive it, and say which one you would check first.
- Name one primary metric, then the guardrail that stops it being gamed.
- Fix the population and the time window before naming any metric.
Follow-up
- What would you do if the primary metric and the guardrail moved in opposite directions?
- Which segment would you cut first, and what would that rule out?
What metrics would you track for a newly launched marketing campaign o…
What metrics would you track for a newly launched marketing campaign offering?
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- State what result would change your recommendation, so the answer is falsifiable.
- Fix the population and the time window before naming any metric.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- What would you do if the primary metric and the guardrail moved in opposite directions?
Explain p-values and statistical significance to a non-technical produ…
Explain p-values and statistical significance to a non-technical product manager.
Approach
- Name the randomisation unit first; it decides the variance and what the test can detect.
- Say whether units interfere with each other, and switch design if they do.
- State the primary metric and the minimum effect worth shipping, then size the test.
Follow-up
- How would you handle interference between treated and control units?
- What would you do if you could not randomise at all?
How do you determine the required sample size and statistical power fo…
How do you determine the required sample size and statistical power for an A/B test?
Approach
- Decide the analysis before seeing data, including how long it runs and when you look.
- Name the randomisation unit first; it decides the variance and what the test can detect.
- Name the guardrails that would stop a launch even on a positive primary result.
Follow-up
- What would you conclude if the result is positive but the test is underpowered?
- How would you handle interference between treated and control units?
Estimate the marginal effect of approval at a score cutoff
Policy auto-approves applications with bureau_score at or above 660 and routes 640 to 659 to manual review, which approves about 35 percent of them. You cannot randomise approvals. Using fct_loan_application and a bureau-sourced panel that reports 12-month serious delinquency on any trade line for approved and declined applicants alike, estimate the effect of approval at the margin. State the identifying assumptions you would test rather than assert, the estimator and its tuning choices, and exactly what the estimate does and does not license.
Approach
- Set the design up as a fuzzy regression discontinuity, because crossing 660 changes approval probability sharply but not from zero to one. Let Z = 1{bureau_score >= 660} be the instrument, approval be the treatment, and 12-month serious delinquency be the outcome. The estimate is the ratio of the jump in the outcome to the jump in approval probability, which is two-stage least squares with Z as the instrument and identifies a local average treatment effect for compliers at the cutoff.
- Insist on the bureau panel as the outcome source. Funded loans exist on both sides of 660 only because manual review approves some applicants below it, and those are selected on whatever the reviewer saw; an outcome defined only on funded loans would reintroduce exactly the selection the design is meant to remove.
- Test the assumptions instead of listing them. Check density continuity of bureau_score at 660 with a McCrary-style or local-polynomial density test, since brokers and applicants can sometimes trigger a bureau refresh and a heaped density at the threshold kills the design. Check that pre-determined covariates, including declared_annual_income_minor, channel, product_code and model_version, are continuous at the cutoff. Check policy_rule_hits and the pricing table for any other rule that fires at exactly 660, because a second discontinuity at the same point is not separable from the first.
- Choose the estimator deliberately. Local linear regression with a triangular kernel and an MSE-optimal bandwidth, with bias-corrected robust confidence intervals. Do not fit a high-order global polynomial, which puts weight on observations far from the cutoff and produces artefacts at the boundary. Because bureau_score is an integer, the running variable has mass points, so cluster standard errors by score value or use an approach designed for discrete running variables.
- Respect outcome maturity. Include only application cohorts with a full 12 months of bureau observation; a recent cohort with partial observation will look cleaner and will drag the estimate.
- State the limits in the deliverable. The estimate is the effect at 660 on applicants whose approval status is determined by the threshold. It does not license moving the cutoff to 620, because both the first stage and the outcome relationship differ away from 660 and because a policy change at scale shifts the applicant mix. Report bandwidth sensitivity alongside the point estimate.
Worked solution 40 min
- Assemble the application-level panel: bureau_score, decision, funded status, and the bureau-sourced 12-month serious delinquency flag, restricted to cohorts with 12 full months of outcome observation.
- Plot the first stage, being approval probability against bureau_score in one-point bins, and confirm a visible jump of roughly 0.65 at 660 (about 1.00 above against about 0.35 below).
- Plot the reduced form, being delinquency against bureau_score in the same bins, and read the jump at 660 off a local linear fit on each side.
- Estimate the Wald ratio as reduced form over first stage, equivalently two-stage least squares with Z = 1{score >= 660}, using a triangular kernel, an MSE-optimal bandwidth, bias-corrected robust intervals, and standard errors clustered by integer score.
- Run the assumption battery: density continuity at 660, covariate continuity, placebo cutoffs at 650 and 670, and bandwidth sensitivity at 0.5x and 2x the selected bandwidth.
- Write the result as a complier average effect at 660 with its interval, its bandwidth sensitivity table, and an explicit sentence on what it does not license.
Follow-up
- Manual reviewers below the cutoff see documents the score does not. What does that do to the monotonicity assumption, and how would you look for defiers?
- You are offered a randomised approval band across 650 to 659 for one quarter. What does it buy you that the discontinuity does not, and how would you price the expected loss of running it?
- The density test shows a modest pile-up just above 660. What are the possible mechanisms, and which of them leave the design usable?
Fraud losses appear to halve in recent transaction months
A weekly chart attributes fct_card_dispute cases to the requested_at month of the linked fct_payment_authorization row. The two most recent months show the first-chargeback rate falling by half, and a risk rule shipped six weeks ago. Columns: dispute_id, auth_id, dispute_category, dispute_stage, opened_at, disputed_amount_minor, liability_shift_flag, outcome, net_loss_minor, resolved_at. Decide whether the rule worked, and produce the version of the chart you would sign your name to.
Approach
- Separate the two dates explicitly. opened_at is when a case was filed, requested_at is when the transaction happened. Attributing by transaction month is the right causal choice and is exactly what makes the newest months structurally incomplete.
- Measure the filing lag rather than assuming it: the distribution of opened_at minus requested_at over fully developed months, split by dispute_category, and the age at which around 95 percent of cases have arrived.
- Build a development triangle of transaction month by months of development on cumulative case counts, and estimate age-to-age factors from the columns that are complete.
- Develop the immature months with those factors and plot the result as an estimate with a visible band, kept visually distinct from the matured series rather than blended into it.
- State the assumption the method needs: a stable development pattern across cohorts. A change in filing behaviour, merchant mix or the dispute team's own backlog breaks it, so inspect factor stability down each column before relying on the estimate.
- Only then evaluate the rule, comparing pre-change and post-change cohorts at equal development age.
Follow-up
- What leading indicator would you accept while the cohort matures, and what is its known bias?
- How does liability_shift_flag change which disputes you should expect to see in the first place?
- If the rule also blocked good transactions, where does that cost appear, and is any of it in this chart?
Four days spend equal time on query work, statistics, modelling and product judgement at deliberately shallow depth, which produces a scored map of where you actually stand. The last three days spend everything on the two areas the role weights most, and close by re-running day one to measure movement.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Breadth pass: query fluency
- Solve six prompts spanning aggregation, joins, window functions and date arithmetic in 60 minutes total, stopping at 10 minutes each whether or not it works, and mark every prompt as solved, solved slowly, or stuck.
- For each unsolved prompt write the single blocking sentence (I lost the grain, I did not know the frame clause, I could not express the date boundary) instead of reading the solution.
- Translate one pandas transformation you know well into SQL and one SQL query into pandas, checking that both return the same row count and the same totals.
Deliverable: A scored six-row table, one line per prompt, saved for the day-seven re-run.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Breadth pass: statistics and inference
- Answer ten short questions in writing with nothing open: what a p-value is conditional on, what a 95 percent interval covers across repeated samples, when a paired test is the right one, what the bootstrap estimates, why multiple comparisons inflate false positives, how controlling the family-wise error rate differs from controlling the false discovery rate, what power depends on, what a missed real effect costs a product, the three situations where the central limit theorem does not rescue you (small n, very heavy tails, dependent observations), and what a standard error is the standard deviation of.
- Grade yourself against a reference and count only the answers that were exactly right, not the ones that were nearly right.
- Rewrite the two weakest answers the following morning from memory in full sentences.
Deliverable: Ten graded answers with an honest count of exact hits.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Breadth pass: modelling
- Take one tabular dataset end to end in 90 minutes: a leakage-safe split, a baseline that is not a model (majority class or historical mean), one regularized linear model, one gradient-boosted tree, and a single evaluation metric chosen before you look at any result.
- Write why that metric fits the cost structure: precision at a fixed recall for alerting, calibration for anything feeding a price or a threshold, ranking metrics for retrieval, and note that area under the ROC curve is insensitive to class balance in a way that can flatter a rare-positive problem.
- Name the leak you were most likely to introduce (an encoding fit on all rows before splitting, or a feature computed after the label's timestamp) and write the check that would have caught it.
Deliverable: A notebook whose first cell states the metric and the baseline, plus two lines on what beat what and by how much.
Practice prompt ↗Practice prompt ↗Practice prompt ↗04Breadth pass: product judgement
- Answer three case prompts aloud at 15 minutes each, timing how long passes before you state a success metric.
- For one case write the first segmentation you would run and the row counts you expect per segment, so that a tiny segment cannot quietly drive the conclusion.
- Take a metric definition you did not write, from a public dashboard, a textbook, or documentation you already have open, and list every place two analysts implementing it would diverge: which rows the denominator admits, whether the unit is an account or a person, what the time window is anchored to, and what happens to data that arrives late. Then write the one question that would close the largest of those gaps.
Deliverable: Three recorded case answers plus an ambiguity list for a metric someone else defined, ending in the single question you would ask about it.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Depth, first area
- Rank the four areas by how many bullet points in the role description each one covers, pick the top one, and spend the entire day inside it.
- Work the six hardest problems you can find in that area and for each write the generalizable move you should have reached for first, rather than the answer.
- Re-solve the two you failed the same evening with notes closed.
Deliverable: Six generalizable moves written as instructions to yourself, not as solutions.
Practice prompt ↗Practice prompt ↗06Depth, second area, and the seam between them
- Repeat the depth protocol on the second-ranked area with the same six-problem structure.
- Construct one problem that requires both areas at once, for example a metric redefinition whose effect you must validate with a test whose readout you then have to query.
- Solve your own combined problem end to end and note where the handoff between the two areas cost you time.
Deliverable: One combined problem, solved end to end, with the handoff failure written down.
Practice prompt ↗Practice prompt ↗07Integration and re-measurement
- Re-run the six prompts from day one under the same clock and compare both correctness and time.
- Run a 60-minute mixed mock that moves between areas without warning, since switching cost is what breadth passes do not train.
- Write the two areas you would still fail on, and the sentence you will use in the interview when you hit one of them.
Deliverable: A before-and-after score table plus a written plan for the two remaining gaps.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Most data work is done by groups, so an interviewer has to work out which piece was yours. An answer that runs on 'we' for several minutes gets interrupted with a question about what you personally did, and by then the answer sounds defensive even when it is true. Mark your own contribution as you go, and name the parts that belonged to someone else instead of leaving them ambiguous. Keep a few specifics back as well, like the name of the metric or who actually objected, so a probe can be answered with something you had not already said.
Tell me about a disagreement you had with a product manager or enginee…
Tell me about a disagreement you had with a product manager or engineer and how you resolved it.
Approach
- Quantify the outcome, including what you would not claim credit for.
- Name the disagreement or constraint, and how you resolved it with evidence.
- State the situation in two sentences and spend the rest on your reasoning.
Follow-up
- How did you know the outcome was caused by your change?
- What would you do differently if you ran that project again?
Recommend a decision whose true outcome matures a year later
An underwriting rule change must be decided in six weeks. Its real outcome, the vintage 90-plus rate at months_on_book 12 in fct_loan_performance_monthly, matures in a year. The executive wants a yes or no, not a range. Randomising the credit decision across the whole population is not available. Name the leading indicator you would accept, state its bias and the direction of that bias, define the decision rule and stopping condition before any rollout starts, and say what reading would make you recommend reversing the change.
Approach
- Fix the readout before the rollout, because a readout chosen after the data arrives is a story rather than a decision rule: indicator, window, threshold and reversal condition all go in writing first.
- Choose the leading indicator on its measured relationship to the matured outcome in historical vintages rather than on availability. Early delinquency, typically the share reaching dpd_1_29 or missing a first scheduled payment by months_on_book 3, is the usual candidate, and you quantify how well it predicted the 12-month rate across past cohorts.
- State the bias and its direction plainly: early delinquency under-represents default that emerges later and is contaminated by servicing and payment-date effects, so treat it as a floor on risk rather than an estimate of it.
- Buy identification where full randomisation is unavailable: a narrow randomised approval band around the cutoff, or a staged rollout by channel or region read as a difference-in-differences, with the parallel-trends assumption stated and checked in the pre-period rather than assumed.
- Give the executive the binary they asked for with the trigger attached in the same sentence: yes, conditional on the month-3 indicator staying inside a stated band, with an automatic hold if it breaches.
Follow-up
- How would you validate that the month-3 indicator predicts the 12-month outcome, and what evidence would invalidate it mid-rollout?
- Compliance refuses a randomised band. What is your next-best identification strategy, and what precision do you lose by taking it?
Retract a published number after finding a currency bug
Two weeks ago you published an interchange and fraud analysis that summed amount_minor across fct_payment_authorization without converting currencies. Minor units are not two decimals everywhere: some currencies carry none and some carry three, so the sum has no interpretation. A pricing decision is already in flight on the back of it. You now have corrected figures. Produce the retraction: what you send, to whom, in what order, and what you change in the process so this class of error is caught next time rather than trusted next time.
Approach
- Size the error before announcing it, because saying the number is wrong without a magnitude and a direction forces every reader to assume the worst case.
- Check whether the conclusion actually flips: if the ranking that drove the pricing decision is unchanged, that belongs in the first sentence beside the correction rather than buried at the end.
- Tell the person acting on it first and directly, then the wider distribution, using the same text, so nobody learns about it secondhand.
- Write the correction as four parts: the old number, the cause in one clause, the effect on the pending decision, and the new number. Leave out self-flagellation, which makes the reader do emotional work instead of acting.
- Fix the class rather than the instance: a rule that a sum over amount_minor either groups by transaction_currency or passes through both conversion steps, exponent scaling and then a dated rate into one named reporting currency, plus a standing reconciliation of the settled subset to the settlement ledger inside each settlement_currency.
Follow-up
- The corrected figures do not change the decision. Do you still send the correction, and what does that choice signal?
- What automated check would have caught this, where would it live, and what would it cost in false alarms?
- 01
Tell me about a disagreement you had with a product manager or engineer and how you resolved it.
- 02
An underwriting rule change must be decided in six weeks. Its real outcome, the vintage 90-plus rate at months_on_book 12 in fct_loan_performance_monthly, matures in a year. The executive wants a yes or no, not a range. Randomising the credit decision across the whole population is not available. Name the leading indicator you would accept, state its bias and the direction of that bias, define the decision rule and stopping condition before any rollout starts, and say what reading would make you recommend reversing the change.
- 03
Two weeks ago you published an interchange and fraud analysis that summed amount_minor across fct_payment_authorization without converting currencies. Minor units are not two decimals everywhere: some currencies carry none and some carry three, so the sum has no interpretation. A pricing decision is already in flight on the back of it. You now have corrected figures. Produce the retraction: what you send, to whom, in what order, and what you change in the process so this class of error is caught next time rather than trusted next time.
Is this an official SoFi interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at SoFi. Rounds and questions reflect what candidates have reported, not a process SoFi has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How technical are the coding rounds at SoFi?
Expect rigorous, practical coding assessments focusing primarily on SQL and Python data manipulation. You should be comfortable writing complex queries with window functions, joins, and aggregations, as well as manipulating data structures in pandas.
PracHub interview research ↗What is the typical interview timeline from initial screen to final offer?
The process usually spans 3 to 5 weeks from your initial recruiter call to the final decision. However, the exact duration can vary based on team scheduling, the specific business unit, and whether you are interviewing for senior or staff levels.
PracHub interview research ↗How can I stand out during the product sense and case study rounds?
Structure your answers clearly by starting with clarifying questions, defining the primary business objective, outlining relevant metrics, and proposing a structured analytical approach. Always tie your technical solutions back to member impact and business value.
PracHub interview research ↗Are remote work options available for Data Scientists at SoFi?
SoFi operates on a hybrid model, generally requiring employees to be in the office approximately one day per week or four days per month, depending on the specific hub location such as San Francisco.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22