As a Data Scientist at Duolingo, you play a pivotal role in harnessing data to enhance language learning experiences for millions of users worldwide. Your work directly impacts product development, user engagement, and strategic decision-making, making the position not only crucial but also deeply rewarding. You will leverage data to inform product features, optimize user experiences, and drive business outcomes, all while contributing to Duolingo’s mission of making education accessible to all.
The complexity of the challenges you will face is significant. You will work with diverse datasets, apply advanced analytical techniques, and collaborate closely with engineering, product, and design teams. This role is critical in shaping how users interact with Duolingo's offerings, from personalized learning pathways to gamification strategies that make learning enjoyable. The scale of the user base and the depth of data available present unique opportunities for impactful insights and innovations.
In short, as a Data Scientist at, you will not only analyze data but also craft the future of language learning through data-driven decisions that resonate with users globally.
Phone Screen
reportedData Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.
What to demonstrate
- Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
- Whether you name what you have not done instead of stretching to cover every line of the posting
- Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage
How to prepare
- Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
- Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
- Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
Take-Home Assessment
reportedThe clock is part of the test. Three to six hours is not enough to do everything the dataset supports, so the submission mostly reveals how you spend a fixed budget against an open question. A reviewer sees which paths you took and, by absence, which you abandoned. Work that runs out of time inside the analysis ships a thin conclusion, while work that cuts scope early protects the last hour for writing. The most reliable way to lose here is to leave the scoping decision implicit, so it reads as something you missed rather than something you chose.
What to demonstrate
- Whether the scope you settled on is presented as a decision with a reason, rather than left for the reader to infer from what is missing
- Whether the depth of the work is consistent with the stated time budget, instead of several half-finished directions left open
- Whether the closing section reads as something written on purpose rather than assembled from whichever cells survived
How to prepare
- Run a timed rehearsal on a public dataset with a hard stop, holding the final sixty minutes for writing no matter where the analysis has got to
- Before opening the data, list the questions it could plausibly answer, pick one, and keep the discarded ones as a short note on what you did not attempt and why
- Commit a one-line finding after each analysis step so the writeup is assembled from recorded results rather than from memory at midnight
Virtual Interviews
reportedRounds outside the standard loop often open with something deliberately under-specified: a loose business problem, an open question about a product area, a dataset described in one sentence. The common failure is surveying, listing six plausible approaches and committing to none of them. The thing that separates a strong answer is scoping out loud. State what you are treating as the goal, name the metric you would move, say what you are choosing not to do and why, then take one path through to an actual answer. An interviewer can follow you down a narrow path. Nobody can grade a menu.
What to demonstrate
- Whether you turn an ambiguous prompt into a stated question with a measurable outcome before doing any work
- The judgement visible in what you cut, and whether you say why you cut it rather than silently dropping it
- Whether you land on a concrete recommendation with its caveat attached, rather than an unranked set of options
How to prepare
- Take three vague prompts, such as 'is this feature working', 'why did retention drop', and 'should we expand into a new segment'. For each, write one sentence of goal, one primary metric with its window, and two things you are explicitly not doing.
- Practise giving the recommendation first and the reasoning second, in five minutes. Loosely defined rounds are usually time-boxed, and an answer that arrives last often does not arrive.
- Keep a running assumption list as you talk, on paper or in the shared doc, so the interviewer can challenge one assumption instead of your whole answer.
9 candidate reports. Individual accounts describe a particular role and hiring cycle.
Duolingo Product Manager interview: Take-home feature proposal and forty-five-minute product rounds
I started with an online application and then went through a recruiter step before meeting anyone live. After that, I completed a take-home product assignment. I had to choose a feature to add or improve in Duolingo and turn it into a short slide deck covering the problem, solution, success metrics, and MVP. It took a while, but the instructions were clear enough that I understood what they wante…
Read full experienceDuolingo Software Engineer interview with disputed coding feedback
The recruiter communication set a strange tone from the beginning. When I was eventually rejected, the message said that my code was nonfunctional even though it had passed the tests. That mismatch bothered me because the feedback felt more dismissive than a simple statement that I wasn't a match. The overall atmosphere, from the recruiter to the interviewers, was unpleasant. It felt as though pe…
Read full experienceDuolingo Product Manager interview: take-home app feature assignment
The process began with a take-home assignment, and I didn't speak with anyone before submitting my work. I had only a few days to devise or reimagine a Duolingo app feature and create a slide deck. It took a lot of effort to get the deck to a presentable level. After I submitted it, the live rounds still felt as if the main goal was to collect ideas from candidates. The most jarring part wasn't t…
Read full experienceDuolingo Software Engineer interview with a one-hour technical round
Once I reached that stage, the process had a strong final-round feel. It followed a kind of superday structure, starting with an online assessment on a Zoom call and ending with a behavioral round that included a couple of STAR-style questions about my experience. The decisions felt quick and decisive. As I progressed, the rounds looked like a shortlist process for engineers and finalists. There…
Read full experienceDuolingo Product Manager: take-home feature design task and delayed rejection
I went through an internship-style process where the first real step was a take-home design task focused on a Duolingo feature. It took a significant amount of time. It felt like something they expected me to work on from start to finish and turn into a polished presentation. After I finished, the follow-up dragged on. The process went silent for a while, and I didn't receive a rejection until mo…
Read full experiencePracHub editorial advice for the preparation topics above.
Pre/post gain studies that select on low pretest scores manufacture improvement.
Any measure with reliability below 1 produces regression to the mean, so a group chosen for scoring in the bottom quartile will score higher on retest with no intervention at all. The apparent gain scales with measurement error, which for a short quiz is large. The fix is a control group selected by the identical rule, or a design that models the pretest as a covariate rather than as a selection filter.
The academic calendar creates structural breaks that look like product effects.
Term start, exam weeks, holidays, and summer each shift usage by amounts far larger than any feature change. A launch timed to week one of a term will show a large lift that is entirely calendar, and a launch in the last week of term will show a collapse. Comparisons must be term-week aligned through dim_term, and any pre/post analysis over a term boundary needs a comparison group living on the same calendar.
Reporting a p-value with no effect size or interval
Give the estimated difference with a confidence interval in the units the business cares about, then say whether that whole interval is worth acting on. A p-value only addresses whether you can rule out exactly zero; it says nothing about magnitude.
Sizing estimates built on unnamed, unrevisable assumptions
Write each assumption as a named number you can change, then show the arithmetic so the interviewer can challenge one input instead of the whole answer. Finish by saying which assumption the result is most sensitive to, which matters more than the point estimate.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
How would you approach building a predictive model for user retention?
How would you approach building a predictive model for user retention?
Approach
- Check what information would not exist at prediction time, and exclude it.
- Frame the prediction: the label, the moment of prediction, and the action it triggers.
- Say how the offline result would be validated online before it is trusted.
Follow-up
- What would you monitor after launch to know the model is still valid?
- How would you choose the decision threshold, and who owns that choice?
Implement verified mastery events with a delayed check
From fct_assessment_response (learner_id, objective_id, content_item_id, submitted_at, attempt_no, is_correct, score_points, max_points, scoring_mode, scored_at, mastery_state_after), count verified mastery events for one calendar month. A mastery event is the first time a (learner_id, objective_id) pair reaches mastery_state_after 'mastered'. It is verified when the earliest scored response on that objective submitted at least 7 days later scores score_points/max_points >= 0.8. Return the monthly count and the verified rate, with the denominator restricted to events that have an eligible later check.
Approach
- Filter to rows with mastery_state_after = 'mastered', then take the earliest submitted_at per (learner_id, objective_id). This is not the same as the first row overall for the pair, since earlier attempts sit in 'learning'.
- Restrict those crossings to the target month in UTC. A pair that first crossed in an earlier month never enters this month's count even if it crosses again after a lapse, so the first-crossing filter must run over all history, not just the month.
- Build the candidate check set by joining the full response table back on (learner_id, objective_id) and keeping rows with submitted_at >= crossing + 7 days, score_points not null, scored_at not null, and max_points > 0.
- Take the earliest eligible candidate per pair, not the best one. Taking the maximum score turns a retention check into an ever-passed-later check, which is a much weaker claim and a materially higher number.
- Mark verified where score_points / max_points >= 0.8. The denominator is crossings with at least one eligible candidate; report the count of crossings with none alongside, because that share moves with the reporting lag and a large share means the month is not ready to publish.
Follow-up
- A pair crosses to 'mastered', lapses, and crosses again inside the same month. What does your code count, and what should it count?
- Why does this metric need a 30-day reporting lag, and what shape does the series take if you publish it at 7 days?
- The 0.8 ratio is a product choice. How would you check it actually separates learners rather than just passing almost everyone?
Attach the calibration in force without merge_asof
You have responses (response_id, content_item_id, submitted_at, is_correct) and calibration_history (content_item_id, valid_from, valid_to, irt_a, irt_b), where intervals within an item are non-overlapping and the open interval carries valid_to as NaT. Attach to each response the irt_a and irt_b in force at submitted_at. You may not use pandas.merge_asof and you may not apply row-wise. responses has about 5 million rows, calibration_history about 40 thousand. Return the input frame plus two columns, NaN where no interval covers the timestamp.
Approach
- Sort calibration_history by (content_item_id, valid_from) and factorise content_item_id across both frames into a shared integer code, so an unknown item on the response side is detectable immediately rather than joining to nothing.
- Convert both timestamps to int64 seconds and build one composite key per side, code * 2**32 + seconds. With codes well under two billion and seconds near 1.8e9, that stays inside int64 and makes a single global search possible. State the second-level resolution assumption; interval boundaries are set at second granularity, so nothing is lost.
- Call np.searchsorted(calibration_keys, response_keys, side='right') - 1 once over the whole array. That gives, for each response, the position of the latest interval starting at or before it within the same item, because the composite key orders by item first.
- Invalidate bad candidates in two places: index -1, and a candidate whose code does not match the response's code, which happens when a response precedes its item's first interval and lands on the previous item's last row.
- Invalidate a third case that is easy to miss: the candidate has a non-null valid_to and submitted_at is at or after it, meaning the response falls in a coverage gap. Without this the previous interval's parameters leak onto uncovered responses.
- Take the parameters positionally with np.take and write NaN where any invalidation fired.
Worked solution 30 min
- codes = pd.factorize on the concatenated item ids, applied to both frames so the mapping is shared.
- cal_key = cal_code.astype('int64') * (1 << 32) + valid_from_seconds; resp_key built the same way; sort cal by cal_key.
- pos = np.searchsorted(cal_key, resp_key, side='right') - 1.
- valid = (pos >= 0) & (cal_code[pos] == resp_code) & (cal_valid_to_seconds[pos].isna() | (resp_seconds < cal_valid_to_seconds[pos])).
- Assign irt_a and irt_b via np.take(pos) where valid, NaN elsewhere, and report the NaN count by reason.
Follow-up
- A backfill wrote overlapping intervals for 200 items. How do you detect that before the join rather than after, and what do you do with those responses?
- The join is correct but peak memory is unacceptable. What changes, and which part of the approach survives?
- How would you test this without a golden output to compare against?
Write a Python function to clean and preprocess a dataset.
Write a Python function to clean and preprocess a dataset.
Approach
- Say which table is the grain you start from, and join outward from it.
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
Follow-up
- What breaks if events arrive late or out of order?
- How would you verify this result without re-running the same query?
How would you optimize a slow-running SQL query?
How would you optimize a slow-running SQL query?
Approach
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
Follow-up
- What breaks if events arrive late or out of order?
- How would you verify this result without re-running the same query?
Write a SQL query to find the top 10 users with the highest activity i…
Write a SQL query to find the top 10 users with the highest activity in the last month.
Approach
- Compute rates by summing numerator and denominator separately, never by averaging rates.
- Handle the rows that do not match: a LEFT JOIN with a NULL check is usually the question.
- Say which table is the grain you start from, and join outward from it.
Follow-up
- How does the query change if the join becomes one-to-many?
- What breaks if events arrive late or out of order?
Consecutive instructional-week streaks and term-over-term return
Using fct_lesson_activity(learner_id, enrollment_id, activity_date, active_seconds) and a calendar dim_term_week(term_id, term_week_no, week_start_date, week_end_date, is_instructional), where holiday weeks carry is_instructional = FALSE, compute for each learner and term the longest run of consecutive instructional weeks containing at least one activity day. A non-instructional week must not break a run, and activity inside one neither extends nor starts a run. Then compute term-over-term return rate into term T+1 with a floor of five activity days in term T. Return learner_id, term_id, activity_days counted over the whole term including non-instructional weeks, longest_streak_weeks and the return flag.
Approach
- Collapse activity to (learner_id, term_id, term_week_no) with COUNT(DISTINCT activity_date). fct_lesson_activity is one row per item per session, so a learner with twelve rows on one afternoon is one activity day and counting rows inflates everything downstream.
- Re-index before differencing. Apply DENSE_RANK() OVER (PARTITION BY term_id ORDER BY term_week_no) across instructional weeks only, producing a gap-free sequence where a holiday week has simply been removed. Differencing raw term_week_no instead splits a run at every holiday, which is exactly the artefact the task forbids.
- Accept what dropping those weeks costs. Activity in a non-instructional week is removed from the streak calculation entirely, so a learner whose only activity lands in a holiday week has activity_days >= 1 and longest_streak_weeks = 0. That is what a streak over instructional weeks means, but it is why the two columns can disagree and why the checks must expect a zero streak alongside non-zero activity rather than assert a floor of 1.
- Apply the island trick on the re-indexed sequence: instructional_index - ROW_NUMBER() OVER (PARTITION BY learner_id, term_id ORDER BY instructional_index) is constant within a run. Group on that constant, count rows per group, take the max per learner and term.
- Build activity_days from the unfiltered term aggregate, then LEFT JOIN the streak result onto it and COALESCE longest_streak_weeks to 0. An inner join here deletes the holiday-only learners from the output and from the return-rate denominator, which is the same bug as the streak-floor assertion wearing different clothes.
- Set the activity floor on COUNT(DISTINCT activity_date) >= 5 across term T, applied to the return-rate denominator only. The floor exists to remove provisioned-but-unused accounts, so leaking it into the numerator silently redefines the metric as active in both terms. Note that a learner can clear the floor on holiday-week activity alone and enter the denominator with a zero streak.
- Define the return flag as EXISTS any activity row in term T+1 for that learner, pairing terms by an explicit term ordering rather than date arithmetic, since term lengths differ. Learners whose org has no term T+1 loaded must be excluded and counted, not treated as non-returners.
Worked solution 40 min
- Build the learner-week activity CTE with COUNT(DISTINCT activity_date) and confirm one row per (learner_id, term_id, term_week_no).
- Build the instructional re-index from dim_term_week and join activity onto it, dropping non-instructional weeks entirely.
- Apply the index-minus-row-number island grouping and take MAX(run_length) per learner and term.
- Compute activity_days per learner and term over the whole term, LEFT JOIN the streak onto it with COALESCE to 0, and count the learners who come out with activity_days > 0 and longest_streak_weeks = 0.
- Apply the five-day floor to the denominator and attach the term T+1 existence flag.
- Build a two-row fixture by hand and verify the holiday behaviour before trusting the full run.
Follow-up
- This streak rewards a learner doing one minute a week over one doing four hours in a single week. Which of those is the product claim, and what would you pair the streak with to catch the difference?
- Roster sync deactivates accounts in bulk at term end. How does that interact with the five-day floor and with the return denominator?
- Show how return rate moves as the floor goes from one day to five, and argue which floor is the honest one to publish.
Describe a complex problem you solved using data analysis and the impa…
Describe a complex problem you solved using data analysis and the impact it had.
Approach
- Fix the population and the time window before naming any metric.
- Restate the decision this analysis has to support, and who acts on the answer.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- What would you do if the primary metric and the guardrail moved in opposite directions?
Given a dataset with user interactions, how would you identify trends …
Given a dataset with user interactions, how would you identify trends in language learning progress?
Approach
- Fix the population and the time window before naming any metric.
- Restate the decision this analysis has to support, and who acts on the answer.
- Name one primary metric, then the guardrail that stops it being gamed.
Follow-up
- How would you detect that the metric is being gamed rather than genuinely improving?
- Which segment would you cut first, and what would that rule out?
What statistical methods do you find most useful for analyzing user en…
What statistical methods do you find most useful for analyzing user engagement data?
Approach
- Restate the decision this analysis has to support, and who acts on the answer.
- Decompose the metric into the rates that drive it, and say which one you would check first.
- State what result would change your recommendation, so the answer is falsifiable.
Follow-up
- Which segment would you cut first, and what would that rule out?
- How would you detect that the metric is being gamed rather than genuinely improving?
Walk us through how you would design an A/B test for a new feature.
Walk us through how you would design an A/B test for a new feature.
Approach
- Name the randomisation unit first; it decides the variance and what the test can detect.
- Name the guardrails that would stop a launch even on a positive primary result.
- State the primary metric and the minimum effect worth shipping, then size the test.
Follow-up
- What would you conclude if the result is positive but the test is underpowered?
- What would you do if you could not randomise at all?
Can you explain how you would handle missing data in a dataset?
Can you explain how you would handle missing data in a dataset?
Approach
- Say what you would check first and why it is the highest-information step.
- Work from the decision backwards to the evidence you would need.
- State your assumptions explicitly before working the problem.
Follow-up
- What assumption would you test first?
- How would you know your answer was wrong?
Hints raise completion and lower verified mastery
Free hints are added to practice items. In a section-randomised test, unit completion rose from 64% to 71% of enrollments while verified mastery rate among completers fell from 44% to 38%. Fields available: fct_assessment_response.hint_count, attempt_no, is_correct, response_seconds, mastery_state_after, plus the delayed retention check behind verified mastery. Declare a primary metric and a guardrail, state the ship rule you would have written before the test, and compute whether the change is net positive per 1,000 enrollments.
Approach
- Notice first that the guardrail as stated is conditioned on completion, which the treatment moves. Conditioning on a post-treatment variable opens a selection path, because both the hint change and learner ability affect who completes, so part of the 44 to 38 drop is composition rather than harm to any individual learner.
- Replace the conditional pair with one unconditional composite: verified mastery events per enrollment. It multiplies the two stages, does not condition on anything the treatment touched, and is the quantity the product is sold on.
- Compute it. Treatment 0.71 x 0.38 = 0.2698, control 0.64 x 0.44 = 0.2816, a difference of -11.8 verified mastery events per 1,000 enrollments, about -4.2% relative.
- State the conflict honestly rather than resolving it by metric choice: hints trade coverage for depth. Which side wins depends on whether the contract is renewed on completion reporting or on mastery evidence, and that is a decision for the person who owns the renewal, made before the test rather than after.
- Fix the inference: assignment is at section level, so standard errors must be cluster-robust or come from a mixed model with a section random effect. At section size 25 and intraclass correlation 0.15 the design effect is 1 + 24 x 0.15 = 4.6 on variance, which is enough to flip a naive verdict.
- Add a misuse detector so the mechanism is visible: share of correct first attempts with hint_count at or above cap and response_seconds below the item floor, which separates scaffolding from answer reveal.
Worked solution 25 min
- Write the composite definition: verified mastery events per enrollment, with the enrollment denominator fixed at randomisation.
- Compute both arms: 0.71 x 0.38 = 0.2698 and 0.64 x 0.44 = 0.2816, giving 269.8 versus 281.6 per 1,000 enrollments.
- Write the pre-registered ship rule with its non-inferiority margin on the composite and its positive-direction requirement on completion.
- State the inference method (cluster-robust at section level, design effect 4.6 at m=25 and rho=0.15) and note that the point estimate above is not yet an interval.
Follow-up
- Suppose the composite came back flat. What would you look at to break the tie?
- How would you detect that hints are being used as an answer reveal rather than as scaffolding, using only the response columns listed?
- What is the minimum detectable effect on the composite with 40 sections of 25 learners at a baseline of 0.28?
Verified mastery rate fell six points with usage flat
Verified mastery rate fell from 0.72 to 0.66 month over month while active learners and submissions per learner held flat. Tables: fct_assessment_response (learner_id, objective_id, content_item_id, content_version_no, attempt_no, submitted_at, is_correct, mastery_state_after, mastery_prob_after) and dim_content_item (content_item_id, version_no, irt_b, irt_a, calibration_n). The metric carries a 30-day reporting lag and requires a delayed retention check at least 7 days after the mastery crossing. Establish whether learning regressed. Deliverable: a ranked list of causes, each eliminated by a named check.
Approach
- Split the rate into its two counts and chart them separately: crossings with an eligible delayed check available, which is the denominator, and crossings that passed it, which is the numerator. A rate that moves while its denominator is also moving is a composition question before it is a learning question.
- Test censoring next, because it is cheap and it invalidates everything downstream. Recompute the prior month's value as it stood at the same maturity as the current month. If the prior month at equal maturity also read 0.66, there is no regression, only an early read against a fully matured comparison.
- Test the decision rule. mastery_prob_after is a posterior from the knowledge-tracing model, so compare its distribution at the crossing response across the two months rather than only counting crossings. A retrained model or a moved threshold changes which learners enter the numerator at all and shows up as a shifted posterior, not as a usage change.
- Test the instrument. Compute the response-weighted mean irt_b of items used in delayed checks in each month, restricted to items with calibration_n at or above 200, and compare irt_a alongside it. A harder or differently discriminating retention form lowers the pass rate with no change in what learners know.
- Segment only once the artefacts are excluded, and segment cluster-aware: cut by grade_band, locale and is_assigned, then check how concentrated the move is across section_id. Outcomes correlate within sections through a shared instructor, schedule and device fleet, so a handful of sections can carry a headline move.
- If a residual survives, size it with the section as the unit of analysis and report an interval that carries the design effect, roughly 1 + (m - 1) * rho for near-equal sections, rather than a learner-level interval that will overstate certainty.
Follow-up
- If the drop survives every check, how would you design the follow-up study to establish cause, and at what unit would you randomise?
- The delayed retention check is itself served by the adaptive engine. What does that do to the verified mastery definition, and how would you fix it?
Four days spend equal time on query work, statistics, modelling and product judgement at deliberately shallow depth, which produces a scored map of where you actually stand. The last three days spend everything on the two areas the role weights most, and close by re-running day one to measure movement.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Breadth pass: query fluency
- Solve six prompts spanning aggregation, joins, window functions and date arithmetic in 60 minutes total, stopping at 10 minutes each whether or not it works, and mark every prompt as solved, solved slowly, or stuck.
- For each unsolved prompt write the single blocking sentence (I lost the grain, I did not know the frame clause, I could not express the date boundary) instead of reading the solution.
- Translate one pandas transformation you know well into SQL and one SQL query into pandas, checking that both return the same row count and the same totals.
Deliverable: A scored six-row table, one line per prompt, saved for the day-seven re-run.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Breadth pass: statistics and inference
- Answer ten short questions in writing with nothing open: what a p-value is conditional on, what a 95 percent interval covers across repeated samples, when a paired test is the right one, what the bootstrap estimates, why multiple comparisons inflate false positives, how controlling the family-wise error rate differs from controlling the false discovery rate, what power depends on, what a missed real effect costs a product, the three situations where the central limit theorem does not rescue you (small n, very heavy tails, dependent observations), and what a standard error is the standard deviation of.
- Grade yourself against a reference and count only the answers that were exactly right, not the ones that were nearly right.
- Rewrite the two weakest answers the following morning from memory in full sentences.
Deliverable: Ten graded answers with an honest count of exact hits.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Breadth pass: modelling
- Take one tabular dataset end to end in 90 minutes: a leakage-safe split, a baseline that is not a model (majority class or historical mean), one regularized linear model, one gradient-boosted tree, and a single evaluation metric chosen before you look at any result.
- Write why that metric fits the cost structure: precision at a fixed recall for alerting, calibration for anything feeding a price or a threshold, ranking metrics for retrieval, and note that area under the ROC curve is insensitive to class balance in a way that can flatter a rare-positive problem.
- Name the leak you were most likely to introduce (an encoding fit on all rows before splitting, or a feature computed after the label's timestamp) and write the check that would have caught it.
Deliverable: A notebook whose first cell states the metric and the baseline, plus two lines on what beat what and by how much.
Practice prompt ↗Practice prompt ↗Practice prompt ↗04Breadth pass: product judgement
- Answer three case prompts aloud at 15 minutes each, timing how long passes before you state a success metric.
- For one case write the first segmentation you would run and the row counts you expect per segment, so that a tiny segment cannot quietly drive the conclusion.
- Take a metric definition you did not write, from a public dashboard, a textbook, or documentation you already have open, and list every place two analysts implementing it would diverge: which rows the denominator admits, whether the unit is an account or a person, what the time window is anchored to, and what happens to data that arrives late. Then write the one question that would close the largest of those gaps.
Deliverable: Three recorded case answers plus an ambiguity list for a metric someone else defined, ending in the single question you would ask about it.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Depth, first area
- Rank the four areas by how many bullet points in the role description each one covers, pick the top one, and spend the entire day inside it.
- Work the six hardest problems you can find in that area and for each write the generalizable move you should have reached for first, rather than the answer.
- Re-solve the two you failed the same evening with notes closed.
Deliverable: Six generalizable moves written as instructions to yourself, not as solutions.
Practice prompt ↗Practice prompt ↗06Depth, second area, and the seam between them
- Repeat the depth protocol on the second-ranked area with the same six-problem structure.
- Construct one problem that requires both areas at once, for example a metric redefinition whose effect you must validate with a test whose readout you then have to query.
- Solve your own combined problem end to end and note where the handoff between the two areas cost you time.
Deliverable: One combined problem, solved end to end, with the handoff failure written down.
Practice prompt ↗Practice prompt ↗07Integration and re-measurement
- Re-run the six prompts from day one under the same clock and compare both correctness and time.
- Run a 60-minute mixed mock that moves between areas without warning, since switching cost is what breadth passes do not train.
- Write the two areas you would still fail on, and the sentence you will use in the interview when you hit one of them.
Deliverable: A before-and-after score table plus a written plan for the two remaining gaps.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Sometimes the honest read is that the initiative did not work, and the person who commissioned the analysis was hoping otherwise. Interviewers want to know whether you softened it. Prepare the case where you delivered an unwelcome result, how you presented the uncertainty without hiding behind it, and what the team did next.
Tell us about a time you had a conflict with a teammate and how you re…
Tell us about a time you had a conflict with a teammate and how you resolved it.
Approach
- Close with what you would do differently, concretely.
- Quantify the outcome, including what you would not claim credit for.
- Pick a story where you drove the decision, not one where you observed it.
Follow-up
- What would you do differently if you ran that project again?
- How did you know the outcome was caused by your change?
Tell an executive the usage report they publish is wrong
A quarterly usage report an executive sends to district administrators counts an active learner as anyone with a fct_lesson_activity row in the quarter. You find two defects: accounts created by roster_sync that opened a single item during an automated import are counted in the numerator, and learners with age_gated = TRUE never reach the source table, so they are absent from the denominator. The report has already gone out twice. Deliverable: what you send, to whom, in what order, and what the definition should become. Probed: whether you can deliver an unwelcome data-quality finding without detonating it.
Approach
- Size the error before raising it. Recompute the metric both ways per org: 'the definition is wrong' and 'the number is overstated by 18 points on the largest account' get very different responses.
- Keep the two defects separate, because they push in opposite directions. Roster artefacts inflate the numerator while age-gated truncation shrinks the denominator, so the net error varies by org and a single average hides that.
- Go to the report owner first with the recomputation and a proposed replacement, not to their audience. The person who published it needs the chance to correct it themselves.
- Propose a definition that is defensible and computable today: distinct learners with at least one scored submission in the quarter, with enrollment_source in ('roster_sync','admin_bulk') excluded from activation cuts, and the age_gated share stated as a footnote rather than silently dropped.
- Say what should happen to the two reports already sent, with a recommendation and an owner. The restatement is the uncomfortable part, and leaving it for someone else to raise means the work is unfinished.
Follow-up
- The executive asks you to change the definition quietly from next quarter with no restatement. What do you do?
- How would you stop this class of defect reaching a published report again?
Allocate one week across three competing team requests
In one week you receive three requests. Sales wants a renewal-risk list for 40 institutional accounts whose period_end falls in 30 days. Curriculum suspects an item-quality problem on a published unit that is currently collecting responses. Growth wants a signup-flow test sized. You have capacity for roughly one and a half of them and cannot escalate for arbitration. Deliverable: your allocation, the reasoning you give each requester, and the one question you ask each before deciding. Probed: whether you prioritise on decisions and reversibility rather than on who asked loudest.
Approach
- Score each request on deadline and reversibility, not on seniority. The renewal list is worthless after period_end; a bad published item compounds with every response collected against it; a test sizing costs almost nothing to delay a week.
- Ask each requester the one question that could collapse their request: whether sales already has a workable heuristic list, whether the suspect unit can simply be set to retired today, whether the growth test has a launch date at all.
- Look for the cheap partial that still buys the deadline. A rules-based risk cut from first_activity_at, units_completed against units_total, and days to period_end ships in a day and captures most of the value of a model.
- Decide, then tell the person who is not getting the work directly and with a date. Unmanaged silence costs more trust than an explicit decline.
- Write the decision and its reasoning somewhere durable so next week's triage does not relitigate the same three requests.
Follow-up
- The growth PM escalates to your skip-level. What do you do, and what do you send ahead of that conversation?
- Two weeks later the risk list you shipped went unused. What changes in how you triage next time?
- 01
Tell us about a time you had a conflict with a teammate and how you resolved it.
- 02
A quarterly usage report an executive sends to district administrators counts an active learner as anyone with a fct_lesson_activity row in the quarter. You find two defects: accounts created by roster_sync that opened a single item during an automated import are counted in the numerator, and learners with age_gated = TRUE never reach the source table, so they are absent from the denominator. The report has already gone out twice. Deliverable: what you send, to whom, in what order, and what the definition should become. Probed: whether you can deliver an unwelcome data-quality finding without detonating it.
- 03
In one week you receive three requests. Sales wants a renewal-risk list for 40 institutional accounts whose period_end falls in 30 days. Curriculum suspects an item-quality problem on a published unit that is currently collecting responses. Growth wants a signup-flow test sized. You have capacity for roughly one and a half of them and cannot escalate for arbitration. Deliverable: your allocation, the reasoning you give each requester, and the one question you ask each before deciding. Probed: whether you prioritise on decisions and reversibility rather than on who asked loudest.
Is this an official Duolingo interview guide?
No. It is PracHub's own research and practice material for the Data Scientist role at Duolingo. Rounds and questions reflect what candidates have reported, not a process Duolingo has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗What is the typical interview difficulty for this position?
The interview difficulty is generally considered average to difficult, with a focus on both technical and behavioral aspects. Candidates should prepare thoroughly to showcase their skills and experiences.
PracHub interview research ↗What differentiates successful candidates?
Successful candidates typically demonstrate a strong grasp of technical skills, effective problem-solving abilities, and a clear alignment with Duolingo's mission and values. Additionally, strong communication skills and teamwork are key.
PracHub interview research ↗What is the culture and working style at Duolingo?
Duolingo fosters a collaborative and inclusive culture. Employees are encouraged to be curious, embrace feedback, and prioritize user-focused design. The work environment is dynamic, with a strong emphasis on continuous learning.
PracHub interview research ↗What is the typical timeline from the initial screen to the offer?
The timeline can vary, but candidates generally receive feedback quickly throughout the interview process. Expect a few weeks from the initial phone screen to a potential offer.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Data Scientist practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22