As a Data Scientist within the Anti-Financial Crime (AFC) function at Deutsche Bank, you occupy a critical position at the intersection of technology, regulatory compliance, and global financial security. Your work is not merely academic; it is foundational to the bank’s ability to detect and prevent money laundering, fraud, and corruption. By leveraging advanced data strategies, you ensure that Deutsche Bank remains resilient against global financial threats while maintaining the integrity of its business operations.
In this role, you will bridge the gap between complex data infrastructure and actionable business intelligence. You will be responsible for the end-to-end lifecycle of transaction monitoring models, from defining red flags to performing deep-dive investigations into data quality. You are expected to be a strategic thinker who can translate regulatory requirements into robust technical specifications, ultimately protecting both the bank and the broader financial ecosystem.
The AFC function operates in a high-stakes, regulatory-heavy environment. Demonstrating an understanding of how your models mitigate risk is as important as your technical proficiency in Python or SQL.
Application Review
reportedMost candidates lose this call inside the first two minutes, during the walkthrough of their own background. The account runs chronologically, sits at the level of tools and titles, and never arrives at a decision anyone could have disagreed with. Anchor on a problem instead of a timeline: what the team could not answer, what you did about it, what happened next. Ninety seconds is enough, and stopping on time leaves room for the half of the call that belongs to you. What you ask about how work gets prioritised signals your level more reliably than the walkthrough does.
What to demonstrate
- Whether your background summary has a shape (problem, decision, consequence) or is a chronological list of tools and employers
- Whether you can account for gaps, short stints and the reason you are looking, unprompted and without hedging
- The substance of the questions you ask back, which an experienced screener reads as a level signal
How to prepare
- Time your opening walkthrough against a clock. If it runs past two minutes, compress the earliest role into a single clause and spend the recovered time on the most recent one
- Write one honest sentence for every gap or short stint visible on your resume and offer it before being asked about it
- Prepare questions about how work arrives and gets prioritised: who writes the request, how often priorities change, and what happens to an analysis after it is delivered
Technical Review
reportedBefore anything else, this round is a reading test. You are given a small schema and a question phrased in business language, and most of the difficulty sits in the gap between them. Who counts as an active user, does a refunded order still count as an order, is that date column an event time or a load time. Weak answers start typing immediately and compute something precise about the wrong population. Strong ones pin the definition in one sentence, name the column that encodes it, then write the query. On a timed assessment with nobody to tell, write the definition in a comment anyway.
What to demonstrate
- Whether an ambiguous term becomes a specific column and filter before any computation happens
- Whether you read the schema for keys and cardinality rather than only for column names
- Whether the result answers the question at the grain it was asked at, per user or per session or per day
How to prepare
- Take three metrics you already use and write down the exact filter and exact grain behind each, then practise stating one of them in a single sentence out loud
- On a schema you have never seen, spend the first minute writing what one row of each table means and which key it is unique on, then predict which joins can duplicate rows
- Rehearse a version where the definition changes halfway through, and edit the query you have instead of starting over
Situational Assessments
reportedMuch of what gets scored here happens out loud while you type. Nobody can see your reasoning inside a half-written query, so five silent minutes read as being stuck even when they are not. State the plan in plain language first: which tables, what grain you are aggregating to, and the one filter that defines the population. Then write it. The narration doubles as insurance, because a wrong plan gets caught early and cheaply while a wrong query gets caught at the end with no time left to redo it. A timed statistics section, where one exists, is a separate test with its own clock.
What to demonstrate
- Whether the query you write matches the plan you just described
- What you do with a hint, meaning whether the correction gets absorbed or the first approach gets defended
- Whether you can debug your own wrong output by reading the result set and naming which part of the query produced the anomaly
How to prepare
- Solve three problems while screen-sharing into a recording, then watch it back and mark every stretch longer than thirty seconds where you said nothing
- Practise compressing the plan into one sentence before typing, then check afterwards whether the finished query actually matched it
- Time yourself on statistics questions that carry a business reading, such as what a confidence interval does and does not claim, rather than re-reading notes without a clock
Case Studies
reportedA case round is decided by whether you leave the interviewer with a recommendation, not by how much analysis you narrate on the way there. The prompt is open on purpose, so the first job is to convert it into a decision someone could act on: ask what would be done differently depending on the answer. From there name the quantity that would settle it, state the assumptions you need, and commit. Candidates who cover more ground than anyone expected and still end on "it depends" score below candidates who scoped narrowly and said what they would do.
What to demonstrate
- Whether the version of the question you choose to answer is genuinely narrower than the prompt and still worth answering
- Whether the recommendation arrives as an action with a number attached, rather than as a summary of what you looked at
- Whether assumptions are stated at the moment you rely on them, instead of collected into a disclaimer at the end
- Whether you notice when a branch you are exploring would not change the decision either way
How to prepare
- Take six open prompts and write only the scoping move for each: the one-sentence question you would actually answer and the decision it feeds. Give yourself three minutes per prompt and stop there.
- Put a five-minute warning into every practice case and force a closing statement that names the action, the result that would justify it, and the result that would reverse it.
- Record one case and count how long you talked before naming a measurable quantity. Past roughly five minutes, what you are calling scoping is narration.
Behavioral Interviews
reportedRounds of this kind usually include one question about work that did not go well, and it is the part that carries the most information. Anyone can narrate a shipped win. What the interviewer learns from a project that stalled is how you behave without a result to hide behind: whether you noticed the problem yourself, how long it took, and who you told. Answers that route the failure onto a data pipeline or a reorganisation close the topic without answering it, and the follow-up comes back to your own part.
What to demonstrate
- Whether you found the error yourself or someone else found it, and how long it sat before anyone knew
- What you changed afterwards, stated as a check you now run rather than a lesson you now believe
- Whether the mistake you choose has real cost attached, such as a quarter of misdirected roadmap or a metric that was reported upward, instead of one that flatters you
How to prepare
- Choose a failure you caught yourself and be ready to say what tipped you off. A story where someone else caught it is still usable, but you will be asked why you missed it.
- Write down the check you added afterwards and where it lives now, so the correction is a concrete artefact rather than a resolution.
- Rehearse saying the cost out loud. Candidates shrink the number by instinct once the interviewer is in the room.
Final Round Assessments
reportedA loop is not scored one interview at a time. The people you meet compare notes afterwards, usually in a meeting you are not in, and the outcome turns on what each of them can say about you when asked. That rewards something other than survival: every room needs one specific thing worth repeating, and none of them can contradict another. The common way to lose is to tell the same project four times with different numbers in it, or to be uniformly fine in a way that leaves nobody with anything to argue for.
What to demonstrate
- Whether your account of a project survives being told twice, with the same scale, the same metric definition and the same numbers each time
- Whether each interviewer leaves with one concrete claim they could make on your behalf later, rather than an absence of complaints
- Whether a question you already answered in an earlier room gets the same answer at the same depth, without visible impatience
How to prepare
- Write a one-page fact sheet for your two or three main projects that fixes the numbers you will quote: rows of data, the metric as a single sentence, the effect you measured and how long the work took. Say them aloud from the sheet until they come out identical every time
- For each kind of room you expect, decide the one sentence you want that interviewer repeating in a debrief, then check during the mock that you said it outright instead of implying it
- Rehearse answering the same project question twice in one sitting, the second time as though you had not just answered it, because the thing that needs fixing is the flatness that creeps into a repeated story
2 candidate reports. Individual accounts describe a particular role and hiring cycle.
Deutsche Bank Financial Analyst Interview Experience: defending valuation assumptions
After an opening conversation where I introduced myself, the interviews moved quickly into technical territory. I was asked direct questions about valuation and how to justify assumptions: EBITDA multiples, DCF setup and its assumptions, plus multi step accounting work. The difficulty came from going beyond memorized answers; I felt they wanted to see that I understood the concepts. The second st…
Read full experienceDeutsche Bank Software Engineer Interview Experience — Two LeetCode Easy Rounds, Then a BQ Round That Was Really a Technical Interview
This is from an interview I had two months ago. The first two technical rounds were each 1 hour, with 2 interviewers per round, and both interviewers asked questions. Each round had one question, both original LeetCode Easy problems — one was checking whether two trees are the same, the other was valid parentheses. Then you had to write test cases. Probably because I was using Java, the interview…
Read full experiencePracHub editorial advice for the preparation topics above.
Recalibrating an underwriting cutoff on approved and funded applicants only
Rejected applicants have no repayment outcome, and they were rejected because the incumbent model scored them badly, so the missingness depends directly on the outcome being modelled. Reject inference by augmentation or parcelling fills the gap using the incumbent model's own assumptions, which means it can confirm those assumptions but cannot test them. The only genuinely new information about the reject region comes from bureau performance on rejects who borrowed elsewhere, or from a deliberately randomised approval band around the cutoff.
Counting authorizations instead of weighting them, and summing amounts across currencies
Declines skew toward high-value, cross-border and card-not-present transactions, so an unweighted approval rate can sit flat while approved value falls. Merchant retry logic also turns one declined purchase into several rows, inflating the denominator by an amount that varies by merchant and by decline reason. Amounts are held in the minor unit of the transaction currency and that unit is not always two decimals, since some currencies have none and some have three, so summing amount_minor across currencies produces a figure with no interpretation at all.
Generalising beyond the population the sample actually supports
State the frame the sample was drawn from and where it diverges from the population you want to talk about: time window, platform, geography, opt-in. If a group is excluded from the frame, either weight to known margins or narrow the claim rather than quietly extending it.
Reporting a p-value with no effect size or interval
Give the estimated difference with a confidence interval in the units the business cares about, then say whether that whole interval is worth acting on. A p-value only addresses whether you can rule out exactly zero; it says nothing about magnitude.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
If you found a critical error in a model that was already in productio…
If you found a critical error in a model that was already in production, what steps would you take to remediate it?
Approach
- Set a baseline first, so any model has something honest to beat.
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
- Check what information would not exist at prediction time, and exclude it.
Follow-up
- What would you monitor after launch to know the model is still valid?
- Where could label leakage enter this setup?
Can you walk us through a model you implemented from scratch, specific…
Can you walk us through a model you implemented from scratch, specifically why you chose that architecture over a more basic approach?
Approach
- Set a baseline first, so any model has something honest to beat.
- Check what information would not exist at prediction time, and exclude it.
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
Follow-up
- Where could label leakage enter this setup?
- What would you monitor after launch to know the model is still valid?
Build a vintage delinquency table without pivot or unstack
fct_loan_performance_monthly gives loan_id, origination_month, months_on_book, days_past_due, charge_off_flag and restructured_flag. Produce a DataFrame with one row per origination_month and columns for months_on_book 0 through 12, each cell holding the share of that vintage's funded loans that had ever reached 90 or more days past due, or charge-off, by that age. You may not use pivot, pivot_table, crosstab or unstack. Cells for ages a cohort has not yet reached must be NaN rather than zero.
Approach
- Define the per-row indicator as days_past_due >= 90 or charge_off_flag, then take a cumulative maximum of it per loan ordered by months_on_book, because the metric is reached-by-age-m, not in-that-state-at-age-m.
- Deal with restructuring before the cumulative max. Restructuring resets days_past_due, so a restructured loan re-enters at current and, without the cumulative maximum carrying its pre-restructure worst state, reads as a cure.
- Fix the denominator once as the count of distinct loan_id per origination_month across the whole cohort. Prepaid and charged-off loans stop producing rows, so a denominator recomputed at each age silently shrinks exactly where losses land.
- Aggregate with groupby(['origination_month','months_on_book'])['ever_90'].sum(), then pre-build the output frame indexed by sorted origination months with integer columns 0 to 12 and assign from the grouped Series by .loc on its index.
- Mask cells beyond each cohort's maximum observed months_on_book so an immature cell reads NaN instead of an artificially low rate.
Worked solution 30 min
- Sort by loan_id and months_on_book, build the ever_90 indicator, then apply groupby('loan_id')['ever_90'].cummax().
- Compute cohort_size as the distinct loan_id count per origination_month, before any filtering on age.
- Group the cumulative indicator by origination_month and months_on_book and sum it to get the numerator per cell.
- Create the output frame with the sorted origination months as index and range(0, 13) as columns, fill it from the grouped Series, and divide each row by its cohort_size.
- Compute each cohort's maximum observed months_on_book and set every cell to the right of it to NaN.
Follow-up
- Two adjacent vintages diverge at months_on_book 6. How would you separate seasoning, mix shift and a genuine credit-quality change?
- The three most recent vintages look best on this table. What do you check before saying so?
- How does the table change if charge-off policy moved from 180 to 120 days past due partway through the series?
Attribute chargebacks to transaction month without fanning out
Report the matured first-chargeback rate. Numerator: fct_card_dispute rows whose dispute_stage is first_chargeback, representment, pre_arbitration or arbitration, attributed to the requested_at month of the linked fct_payment_authorization row rather than to opened_at. Denominator: settled authorizations in that same transaction month, meaning settled_at is not null and is_reversal is false. One authorization can carry more than one dispute case. Report only transaction months whose last day is at least 120 days old, and mark the remainder incomplete rather than showing them as low.
Approach
- Compute the denominator from fct_payment_authorization alone, grouped by the requested_at month, before any dispute table is mentioned, so the join cannot touch it.
- Compute the numerator as a separate aggregate: join disputes to their authorization only to recover requested_at, then count dispute rows per transaction month with the stage filter.
- Join the two aggregates on the month key, which keeps the grain at one row per month and makes double counting structurally impossible rather than something you have to remember.
- Apply the maturity gate on the month's last day plus 120 days against the current date, because a transaction on the final day of the month is the least mature one in the cohort.
- Emit an explicit status column of 'final' or 'incomplete' instead of filtering immature months away silently, so a reader cannot mistake absence for zero.
Follow-up
- Should the numerator count dispute cases or disputed authorizations? Which one reconciles to the loss line, and which one to the operations queue?
- Some reason codes allow filing well beyond 120 days. How would you estimate the tail on a month you have decided to call final?
- How would you present the incomplete months so that a weekly reader does not read the right-hand slope as an improvement?
Reconcile captured authorizations against the daily settlement total
fct_payment_authorization holds captured_amount_minor in transaction_currency, and settlement_amount_minor in settlement_currency with settlement_fx_rate applied at settlement rather than at authorization. The rate is quoted in major units of settlement_currency per major unit of transaction_currency, and dim_currency.minor_unit_exponent carries the ISO 4217 exponent for each code (0, 2 or 3 depending on the currency). Produce a daily reconciliation: for each settled_at date and settlement_currency, return settled_count, total settlement_amount_minor, and the sum of captured_amount_minor converted into settlement minor units. Flag any date and currency pair whose two totals differ by more than one minor unit per settled authorization. Do not sum amounts across currencies anywhere in the output.
Approach
- Restrict to rows that actually settled: settled_at is not null and settlement_amount_minor is not null, which is a smaller population than captured rows because a capture can still be in flight.
- Truncate settled_at to a date with an explicit time zone so the cut matches the ledger's cut, since settled_at is timestamptz and date_trunc on timestamptz silently uses the session time zone.
- Join dim_currency twice, once on transaction_currency and once on settlement_currency, so both exponents are on the row. Minor units are not a common scale: a bare captured_amount_minor * settlement_fx_rate is correct only when the two exponents are equal, and a zero-decimal currency settling into a two-decimal one is wrong by a factor of 100.
- Convert per row as ROUND(captured_amount_minor::numeric / POWER(10::numeric, exp_txn) * settlement_fx_rate * POWER(10::numeric, exp_settle)) — minor units to major in the transaction currency, apply the major-per-major rate, then back to minor units in the settlement currency. The collapsed form ROUND(captured_amount_minor::numeric * settlement_fx_rate * POWER(10::numeric, exp_settle - exp_txn)) is the same expression. Round per row and then sum, not SUM(...) * an average rate, because the rate varies row by row and rounding per row is what the settlement file did.
- Group by the settlement date and settlement_currency together, never by date alone, and carry the currency into every output column name or row.
- Compare the two totals with a tolerance scaled by settled_count, since per-row rounding accumulates linearly in the number of rows rather than being a fixed constant.
Worked solution 25 min
- Filter to settled rows and derive settlement_date from settled_at with an explicit time zone.
- Join dim_currency on transaction_currency and again on settlement_currency to pick up exp_txn and exp_settle; fail the run if either is null rather than defaulting to 2.
- Aggregate by settlement_date and settlement_currency: COUNT(*), SUM(settlement_amount_minor), and SUM(ROUND(captured_amount_minor::numeric / POWER(10::numeric, exp_txn) * settlement_fx_rate * POWER(10::numeric, exp_settle))).
- Add a derived difference column and a boolean flag where ABS(difference) > settled_count.
- Order by the flag first and then by settlement_date so the exceptions surface at the top.
Follow-up
- A partial capture means captured_amount_minor is less than amount_minor. Where does that show up in this reconciliation, and where does it not?
- On one currency pair the converted total is consistently about one hundredth of the settlement total, on every date, while the other pairs reconcile. Which two columns do you inspect first, and what single change fixes it?
- The rate is documented as major-per-major. If a feed started publishing it minor-per-minor instead, which pairs would still reconcile and which would break?
- How would you present a total across currencies to a finance partner who has asked for one number?
How do you measure the success of your models in a real-world, high-vo…
How do you measure the success of your models in a real-world, high-volume transactional context?
Approach
- Name one primary metric, then the guardrail that stops it being gamed.
- Restate the decision this analysis has to support, and who acts on the answer.
- Fix the population and the time window before naming any metric.
Follow-up
- Which segment would you cut first, and what would that rule out?
- How would you detect that the metric is being gamed rather than genuinely improving?
What are the specific challenges you have faced when deploying models …
What are the specific challenges you have faced when deploying models in a production environment?
Approach
- Work from the decision backwards to the evidence you would need.
- State your assumptions explicitly before working the problem.
- Say what you would check first and why it is the highest-information step.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
Explain the difference between supervised and unsupervised learning ap…
Explain the difference between supervised and unsupervised learning approaches in the context of fraud detection.
Approach
- Clarify what is being asked and what a complete answer would contain.
- Work from the decision backwards to the evidence you would need.
- State your assumptions explicitly before working the problem.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
Judge a rate increase without letting price flatter the ratio
A rating change raises premium on one product_line's renewal book. Using fct_policy_period_monthly (earned_premium_minor, written_premium_minor, exposure_units, paid_loss_minor, case_reserve_minor, ibnr_reserve_minor, loss_adjustment_expense_minor, rating_tier, policy_status, renewal_flag, policy_term_end), the accident-period loss ratio improves two points over the following year. Explain why that number alone cannot tell you the change worked, decompose the movement into the effects you would separate, and specify the primary metric and guardrails you would commit to before the next rating change.
Approach
- Start with the arithmetic. A rate increase raises earned premium per exposure unit, so the loss ratio falls even if every policyholder behaves identically and every claim is unchanged. Part of the two points is mechanical and carries no information about risk at all.
- Switch the risk measure to one price cannot move: pure premium, being incurred losses (paid_loss_minor plus case_reserve_minor plus ibnr_reserve_minor, with loss_adjustment_expense_minor included or excluded consistently and the choice stated) divided by exposure_units, evaluated by accident period at a fixed development age. Flat pure premium alongside an improved loss ratio means the improvement was entirely price.
- Decompose the loss ratio movement into three named components: the price effect at constant exposure and mix, the mix effect from which rating_tiers renewed and which lapsed, and the residual change in pure premium within tier. Only the third is evidence about risk selection, and it is usually the smallest.
- Name the adverse-selection risk directly. Price sensitivity and loss propensity are not independent, and the policyholders most able to leave after a rate rise are often the ones worth keeping. Retention by rating_tier is therefore a guardrail with teeth, and it has to be read at tier level because a flat blended retention hides tiers moving in opposite directions.
- Add prior-period reserve development as the second guardrail. The same two points are producible by setting case reserves or IBNR light, which surfaces only later as adverse development, so the reserve guardrail is what stops the primary metric being satisfiable by an accounting choice.
- Commit the primary before the next change: underwriting margin per exposure unit, being earned premium minus incurred losses minus loss adjustment expense minus allocated expense, over exposure_units, at a fixed development age, reported by accident period and by rating_tier with exposure volume printed beside it so that improving margin by shrinking the book is visible in the same table.
Worked solution 40 min
- By accident quarter at a fixed twelve-month development age, compute both the loss ratio (incurred over earned_premium_minor) and the pure premium (incurred over exposure_units), before and after the change.
- Build the three-way decomposition: move price only at constant mix and exposure, then move the tier mix to the post-change distribution at constant price, then take the remainder as the within-tier pure premium change.
- Compute renewal retention by rating_tier from policies reaching policy_term_end, excluding cancelled_midterm and terms where no renewal offer was made, and cross each tier's retention with its prior pure premium.
- Pull prior-period reserve development for the periods used and state whether the improvement survives it.
Follow-up
- Retention is flat overall but fell nine points in the lowest-loss tier. What do you expect next year's pure premium to do?
- Why not use written premium as the denominator, and where would you still legitimately see it used?
- How would you separate a genuine underwriting improvement from a year of mild weather?
Fraud losses appear to halve in recent transaction months
A weekly chart attributes fct_card_dispute cases to the requested_at month of the linked fct_payment_authorization row. The two most recent months show the first-chargeback rate falling by half, and a risk rule shipped six weeks ago. Columns: dispute_id, auth_id, dispute_category, dispute_stage, opened_at, disputed_amount_minor, liability_shift_flag, outcome, net_loss_minor, resolved_at. Decide whether the rule worked, and produce the version of the chart you would sign your name to.
Approach
- Separate the two dates explicitly. opened_at is when a case was filed, requested_at is when the transaction happened. Attributing by transaction month is the right causal choice and is exactly what makes the newest months structurally incomplete.
- Measure the filing lag rather than assuming it: the distribution of opened_at minus requested_at over fully developed months, split by dispute_category, and the age at which around 95 percent of cases have arrived.
- Build a development triangle of transaction month by months of development on cumulative case counts, and estimate age-to-age factors from the columns that are complete.
- Develop the immature months with those factors and plot the result as an estimate with a visible band, kept visually distinct from the matured series rather than blended into it.
- State the assumption the method needs: a stable development pattern across cohorts. A change in filing behaviour, merchant mix or the dispute team's own backlog breaks it, so inspect factor stability down each column before relying on the estimate.
- Only then evaluate the rule, comparing pre-change and post-change cohorts at equal development age.
Follow-up
- What leading indicator would you accept while the cohort matures, and what is its known bias?
- How does liability_shift_flag change which disputes you should expect to see in the first place?
- If the rule also blocked good transactions, where does that cost appear, and is any of it in this chart?
For someone who can already write the query and train the model but stalls when asked what to measure or whether a change is worth making. Metric definition and case structure come first; the technical work is kept as maintenance rather than the centre of the week.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Metric anatomy
- For three products you use daily, write one primary metric, two input metrics that plausibly move it, and one guardrail that would catch a cheap way of moving the primary at the cost of the product.
- For one of them, specify the metric precisely enough that two analysts would return the same number: numerator, denominator, unit of observation, time window, and how returning and deleted accounts are treated.
- Pick a ratio metric and write what happens to it when the denominator shrinks for reasons unrelated to the numerator, with a concrete example of that happening.
Deliverable: A one-page metric tree for one product, with the primary metric written as an unambiguous spec.
Practice prompt ↗Practice prompt ↗Worked solution ↗02Diagnosing a drop without guessing
- Take the prompt "weekly active users fell 8 percent week over week" and write the segmentation plan before proposing any cause: platform, region, tenure cohort, acquisition channel, and whether the movement sits in the numerator or in a changed denominator.
- List the instrumentation failures that manufacture fake drops (a client release that stopped firing an event, a bot filter change, a shifted date boundary or timezone) and write the query that rules out each one.
- Rehearse stating the boring explanations first, seasonality and day-of-week composition, before reaching for a product cause.
Deliverable: A drop-diagnosis checklist short enough to recite from memory in under a minute.
Practice prompt ↗Practice prompt ↗03Should we build it
- Take a feature idea and write it as a bet: what you believe is true, what would have to be true for it to pay off, the metric that would confirm it, and the effect size that would justify the engineering cost.
- Size the opportunity top-down and bottom-up, then reconcile the two numbers in writing instead of quoting whichever is friendlier.
- Write the counter-metric that would make you kill the feature even if it wins on the primary metric.
Deliverable: A one-page product memo ending in a decision rather than a list of considerations.
Practice prompt ↗Practice prompt ↗04The places aggregate numbers lie
- Construct a Simpson's paradox numerically: two segments where the treatment wins within each segment yet loses overall, and identify the shift in segment weights that causes it.
- Take a heavy right-tailed quantity such as revenue per user and write why the mean is the wrong summary, which percentile you would report instead, and what a moving mean with a stable median tells you.
- Write your definition of a session for the product from day one, then name two real behaviours it misclassifies.
Deliverable: One page holding a worked Simpson's paradox table and a session definition with its two known failure cases.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Technical maintenance, aimed at metrics
- Solve four timed SQL prompts that all end in a ratio metric, so the question of grain stays live in every answer.
- Compute a 95 percent confidence interval for a proportion on a small sample, and state why the normal approximation is unreliable when either np or n(1 minus p) falls below roughly 10, along with which interval you would use instead.
- Take one metric from your day-one tree, write the query that computes it correctly, then write the query that computes it wrong in the most plausible way and explain how you would notice.
Deliverable: Four solved prompts plus a matched correct and plausible-wrong query for one metric.
Practice prompt ↗Practice prompt ↗06Turning engineering work into data science stories
- Write three project stories as situation, decision, trade-off, outcome, each carrying one number and one thing you got wrong.
- For the story you will lead with, prepare an answer to "what would you do differently" that names a decision you made, not a constraint you were handed.
- Practise the sentence that reframes a systems project as a question project: the question the work answered, ahead of the pipeline it shipped.
Deliverable: Three written stories with the lead story delivered aloud and timed under four minutes.
Practice prompt ↗Practice prompt ↗07Mock case and gap list
- Run a 40-minute mock case with someone playing a product manager who pushes back on your metric choice, and record it.
- Listen back and mark every moment you proposed a solution before the success metric existed.
- Rewrite those moments as the question you should have asked, and rehearse the first 90 seconds of the case until scoping comes before solving.
Deliverable: A recorded case plus a rewritten opening 90 seconds.
Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Half of this section is about translation. Be ready to describe how you explained a result to someone who did not want the method, only the implication, and what you did when the simplified version started being repeated in a way that overstated it. Correcting your own simplification is a strong beat.
Describe a situation where you had to collaborate with a cross-functio…
Describe a situation where you had to collaborate with a cross-functional team to solve a business problem.
Approach
- State the situation in two sentences and spend the rest on your reasoning.
- Quantify the outcome, including what you would not claim credit for.
- Pick a story where you drove the decision, not one where you observed it.
Follow-up
- How did you know the outcome was caused by your change?
- What would you do differently if you ran that project again?
How do you stay motivated and focused in a goal-oriented, high-pressur…
How do you stay motivated and focused in a goal-oriented, high-pressure environment?
Approach
- Close with what you would do differently, concretely.
- Name the disagreement or constraint, and how you resolved it with evidence.
- State the situation in two sentences and spend the rest on your reasoning.
Follow-up
- How did you know the outcome was caused by your change?
- What would you do differently if you ran that project again?
Tell me about a time you had to explain a complex technical finding to…
Tell me about a time you had to explain a complex technical finding to a non-technical stakeholder.
Approach
- Quantify the outcome, including what you would not claim credit for.
- Name the disagreement or constraint, and how you resolved it with evidence.
- Close with what you would do differently, concretely.
Follow-up
- What would you do differently if you ran that project again?
- What did you decide not to do, and why?
- 01
Describe a situation where you had to collaborate with a cross-functional team to solve a business problem.
- 02
How do you stay motivated and focused in a goal-oriented, high-pressure environment?
- 03
Tell me about a time you had to explain a complex technical finding to a non-technical stakeholder.
Is this an official Deutsche Bank interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Deutsche Bank. Rounds and questions reflect what candidates have reported, not a process Deutsche Bank has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗Is there a lot of live coding in the interview?
Not necessarily. Recent candidates have reported a strong focus on technical discussions surrounding past projects rather than traditional whiteboard coding. However, always be prepared to explain the logic behind any code you have written.
PracHub interview research ↗How difficult is the interview process?
It is generally considered to be of average to high difficulty. The rigor comes from the depth of the technical questioning and the requirement to handle situational, behavioral, and case-study rounds.
PracHub interview research ↗How long does the process take?
Timelines can vary, and some candidates have noted that scheduling can take time. It is important to be patient but also to stay proactive in your communication with the recruiting team.
PracHub interview research ↗What is the most important factor for success?
Being able to articulate the business impact of your technical work. Deutsche Bank is looking for scientists who understand that their models serve a specific regulatory and security purpose.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22