Duolingo · Data Scientist
Updated · 2026-09-22

Duolingo Data Scientist
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

As a Data Scientist at Duolingo, you play a pivotal role in harnessing data to enhance language learning experiences for millions of users worldwide. Your work directly impacts product development, user engagement, and strategic decision-making, making the position not only crucial but also deeply rewarding. You will leverage data to inform product features, optimize user experiences, and drive business outcomes, all while contributing to Duolingo’s mission of making education accessible to all.

Coding rounds for this role are usually data-manipulation shaped rather than data-structure shaped: group-bys, joins, time windows, ranking within a partition. Confirm the format before spending a week on graph traversal.

Duolingo candidates report 3 rounds · ≈ 3-5 weeks. The stages below are what candidates describe, not a published process.

Separate mastery signals from raw usage exposureCluster experiments at class or school levelAlign cohorts to term weeks, not calendar weeks

34 min read

Practice 17 Data Scientist prompts
9Candidate experiences ↗Read their reports
17Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

As a Data Scientist at Duolingo, you play a pivotal role in harnessing data to enhance language learning experiences for millions of users worldwide. Your work directly impacts product development, user engagement, and strategic decision-making, making the position not only crucial but also deeply rewarding. You will leverage data to inform product features, optimize user experiences, and drive business outcomes, all while contributing to Duolingo’s mission of making education accessible to all.

The complexity of the challenges you will face is significant. You will work with diverse datasets, apply advanced analytical techniques, and collaborate closely with engineering, product, and design teams. This role is critical in shaping how users interact with Duolingo's offerings, from personalized learning pathways to gamification strategies that make learning enjoyable. The scale of the user base and the depth of data available present unique opportunities for impactful insights and innovations.

In short, as a Data Scientist at, you will not only analyze data but also craft the future of language learning through data-driven decisions that resonate with users globally.

01

Phone Screen

reported

Data Scientist covers at least four different jobs: experimentation, product analytics, causal work on observational data, and applied modelling that ships into a system. A screening call is the cheapest place to find out which of them is being hired for, and doing that diagnosis openly reads as senior rather than fussy. Ask what the last few pieces of work on the team actually were, and roughly how a week splits between querying, modelling and stakeholder time. Then say which parts of that you have done and which you have not. Claiming the whole range is the fastest way to be caught one round later.

What to demonstrate

  • Whether you can distinguish the flavours of the role and locate your own experience inside one of them honestly
  • Whether you name what you have not done instead of stretching to cover every line of the posting
  • Whether your hard constraints (notice period, location, work authorisation, level) surface now rather than at offer stage

How to prepare

  • Map the last two years of your time into rough percentages across query writing, experiment design, modelling and stakeholder work, so a question about scope has a real answer
  • Mark every responsibility in the posting as done, adjacent or new, and prepare one sentence for each adjacent item naming the closest thing you have actually built
  • Decide which logistics are non-negotiable before the call so you can state them in one sentence rather than negotiating live
PracHub interview research ↗
02

Take-Home Assessment

reported

The clock is part of the test. Three to six hours is not enough to do everything the dataset supports, so the submission mostly reveals how you spend a fixed budget against an open question. A reviewer sees which paths you took and, by absence, which you abandoned. Work that runs out of time inside the analysis ships a thin conclusion, while work that cuts scope early protects the last hour for writing. The most reliable way to lose here is to leave the scoping decision implicit, so it reads as something you missed rather than something you chose.

What to demonstrate

  • Whether the scope you settled on is presented as a decision with a reason, rather than left for the reader to infer from what is missing
  • Whether the depth of the work is consistent with the stated time budget, instead of several half-finished directions left open
  • Whether the closing section reads as something written on purpose rather than assembled from whichever cells survived

How to prepare

  • Run a timed rehearsal on a public dataset with a hard stop, holding the final sixty minutes for writing no matter where the analysis has got to
  • Before opening the data, list the questions it could plausibly answer, pick one, and keep the discarded ones as a short note on what you did not attempt and why
  • Commit a one-line finding after each analysis step so the writeup is assembled from recorded results rather than from memory at midnight
PracHub interview research ↗
03

Virtual Interviews

reported

Rounds outside the standard loop often open with something deliberately under-specified: a loose business problem, an open question about a product area, a dataset described in one sentence. The common failure is surveying, listing six plausible approaches and committing to none of them. The thing that separates a strong answer is scoping out loud. State what you are treating as the goal, name the metric you would move, say what you are choosing not to do and why, then take one path through to an actual answer. An interviewer can follow you down a narrow path. Nobody can grade a menu.

What to demonstrate

  • Whether you turn an ambiguous prompt into a stated question with a measurable outcome before doing any work
  • The judgement visible in what you cut, and whether you say why you cut it rather than silently dropping it
  • Whether you land on a concrete recommendation with its caveat attached, rather than an unranked set of options

How to prepare

  • Take three vague prompts, such as 'is this feature working', 'why did retention drop', and 'should we expand into a new segment'. For each, write one sentence of goal, one primary metric with its window, and two things you are explicitly not doing.
  • Practise giving the recommendation first and the reasoning second, in five minutes. Loosely defined rounds are usually time-boxed, and an answer that arrives last often does not arrive.
  • Keep a running assumption list as you talk, on paper or in the shared doc, so the interviewer can challenge one assumption instead of your whole answer.
PracHub interview research ↗

9 candidate reports. Individual accounts describe a particular role and hiring cycle.

Product Manager

Duolingo Product Manager interview: Take-home feature proposal and forty-five-minute product rounds

HR Screen → Take-home Project → Other

I started with an online application and then went through a recruiter step before meeting anyone live. After that, I completed a take-home product assignment. I had to choose a feature to add or improve in Duolingo and turn it into a short slide deck covering the problem, solution, success metrics, and MVP. It took a while, but the instructions were clear enough that I understood what they wante…

Read full experience
Software Engineer

Duolingo Software Engineer interview with disputed coding feedback

OtherOutcome: rejected

The recruiter communication set a strange tone from the beginning. When I was eventually rejected, the message said that my code was nonfunctional even though it had passed the tests. That mismatch bothered me because the feedback felt more dismissive than a simple statement that I wasn't a match. The overall atmosphere, from the recruiter to the interviewers, was unpleasant. It felt as though pe…

Read full experience
Product Manager

Duolingo Product Manager interview: take-home app feature assignment

Take-home Project → OtherOutcome: rejected

The process began with a take-home assignment, and I didn't speak with anyone before submitting my work. I had only a few days to devise or reimagine a Duolingo app feature and create a slide deck. It took a lot of effort to get the deck to a presentable level. After I submitted it, the live rounds still felt as if the main goal was to collect ideas from candidates. The most jarring part wasn't t…

Read full experience
Software Engineer

Duolingo Software Engineer interview with a one-hour technical round

Online Assessment → Technical Screen → Other

Once I reached that stage, the process had a strong final-round feel. It followed a kind of superday structure, starting with an online assessment on a Zoom call and ending with a behavioral round that included a couple of STAR-style questions about my experience. The decisions felt quick and decisive. As I progressed, the rounds looked like a shortlist process for engineers and finalists. There…

Read full experience
Product Manager

Duolingo Product Manager: take-home feature design task and delayed rejection

Take-home ProjectOutcome: rejected

I went through an internship-style process where the first real step was a take-home design task focused on a Duolingo feature. It took a significant amount of time. It felt like something they expected me to work on from start to finish and turn into a polished presentation. After I finished, the follow-up dragged on. The process went silent for a while, and I didn't receive a rejection until mo…

Read full experience

PracHub editorial advice for the preparation topics above.

01

Pre/post gain studies that select on low pretest scores manufacture improvement.

Any measure with reliability below 1 produces regression to the mean, so a group chosen for scoring in the bottom quartile will score higher on retest with no intervention at all. The apparent gain scales with measurement error, which for a short quiz is large. The fix is a control group selected by the identical rule, or a design that models the pretest as a covariate rather than as a selection filter.

02

The academic calendar creates structural breaks that look like product effects.

Term start, exam weeks, holidays, and summer each shift usage by amounts far larger than any feature change. A launch timed to week one of a term will show a large lift that is entirely calendar, and a launch in the last week of term will show a collapse. Comparisons must be term-week aligned through dim_term, and any pre/post analysis over a term boundary needs a comparison group living on the same calendar.

03

Reporting a p-value with no effect size or interval

Give the estimated difference with a confidence interval in the units the business cares about, then say whether that whole interval is worth acting on. A p-value only addresses whether you can rule out exactly zero; it says nothing about magnitude.

04

Sizing estimates built on unnamed, unrevisable assumptions

Write each assumption as a named number you can change, then show the arithmetic so the interviewer can challenge one input instead of the whole answer. Finish by saying which assumption the result is most sensitive to, which matters more than the point estimate.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

14 technical prompts3 include a worked solution

How would you approach building a predictive model for user retention?

medium
machine learning and modelling

How would you approach building a predictive model for user retention?

Approach
  1. Check what information would not exist at prediction time, and exclude it.
  2. Frame the prediction: the label, the moment of prediction, and the action it triggers.
  3. Say how the offline result would be validated online before it is trusted.
Follow-up
  • What would you monitor after launch to know the model is still valid?
  • How would you choose the decision threshold, and who owns that choice?

Implement verified mastery events with a delayed check

medium
first crossingwindow logicmetric implementation

From fct_assessment_response (learner_id, objective_id, content_item_id, submitted_at, attempt_no, is_correct, score_points, max_points, scoring_mode, scored_at, mastery_state_after), count verified mastery events for one calendar month. A mastery event is the first time a (learner_id, objective_id) pair reaches mastery_state_after 'mastered'. It is verified when the earliest scored response on that objective submitted at least 7 days later scores score_points/max_points >= 0.8. Return the monthly count and the verified rate, with the denominator restricted to events that have an eligible later check.

Approach
  1. Filter to rows with mastery_state_after = 'mastered', then take the earliest submitted_at per (learner_id, objective_id). This is not the same as the first row overall for the pair, since earlier attempts sit in 'learning'.
  2. Restrict those crossings to the target month in UTC. A pair that first crossed in an earlier month never enters this month's count even if it crosses again after a lapse, so the first-crossing filter must run over all history, not just the month.
  3. Build the candidate check set by joining the full response table back on (learner_id, objective_id) and keeping rows with submitted_at >= crossing + 7 days, score_points not null, scored_at not null, and max_points > 0.
  4. Take the earliest eligible candidate per pair, not the best one. Taking the maximum score turns a retention check into an ever-passed-later check, which is a much weaker claim and a materially higher number.
  5. Mark verified where score_points / max_points >= 0.8. The denominator is crossings with at least one eligible candidate; report the count of crossings with none alongside, because that share moves with the reporting lag and a large share means the month is not ready to publish.
Follow-up
  • A pair crosses to 'mastered', lapses, and crosses again inside the same month. What does your code count, and what should it count?
  • Why does this metric need a 30-day reporting lag, and what shape does the series take if you publish it at 7 days?
  • The 0.8 ratio is a product choice. How would you check it actually separates learners rather than just passing almost everyone?

Attach the calibration in force without merge_asof

mediumWorked solution
as-of joinsearchsortedvectorisationslowly changing dimensions

You have responses (response_id, content_item_id, submitted_at, is_correct) and calibration_history (content_item_id, valid_from, valid_to, irt_a, irt_b), where intervals within an item are non-overlapping and the open interval carries valid_to as NaT. Attach to each response the irt_a and irt_b in force at submitted_at. You may not use pandas.merge_asof and you may not apply row-wise. responses has about 5 million rows, calibration_history about 40 thousand. Return the input frame plus two columns, NaN where no interval covers the timestamp.

Approach
  1. Sort calibration_history by (content_item_id, valid_from) and factorise content_item_id across both frames into a shared integer code, so an unknown item on the response side is detectable immediately rather than joining to nothing.
  2. Convert both timestamps to int64 seconds and build one composite key per side, code * 2**32 + seconds. With codes well under two billion and seconds near 1.8e9, that stays inside int64 and makes a single global search possible. State the second-level resolution assumption; interval boundaries are set at second granularity, so nothing is lost.
  3. Call np.searchsorted(calibration_keys, response_keys, side='right') - 1 once over the whole array. That gives, for each response, the position of the latest interval starting at or before it within the same item, because the composite key orders by item first.
  4. Invalidate bad candidates in two places: index -1, and a candidate whose code does not match the response's code, which happens when a response precedes its item's first interval and lands on the previous item's last row.
  5. Invalidate a third case that is easy to miss: the candidate has a non-null valid_to and submitted_at is at or after it, meaning the response falls in a coverage gap. Without this the previous interval's parameters leak onto uncovered responses.
  6. Take the parameters positionally with np.take and write NaN where any invalidation fired.
Worked solution 30 min
  1. codes = pd.factorize on the concatenated item ids, applied to both frames so the mapping is shared.
  2. cal_key = cal_code.astype('int64') * (1 << 32) + valid_from_seconds; resp_key built the same way; sort cal by cal_key.
  3. pos = np.searchsorted(cal_key, resp_key, side='right') - 1.
  4. valid = (pos >= 0) & (cal_code[pos] == resp_code) & (cal_valid_to_seconds[pos].isna() | (resp_seconds < cal_valid_to_seconds[pos])).
  5. Assign irt_a and irt_b via np.take(pos) where valid, NaN elsewhere, and report the NaN count by reason.
EXPECTED RESULTEvery response inside a covered interval carries that interval's irt_a and irt_b. Responses before the item's first valid_from, inside a coverage gap, after a closed final valid_to, or on an item absent from calibration_history carry NaN. The whole join runs in seconds on 5 million rows with no Python-level row loop.
Follow-up
  • A backfill wrote overlapping intervals for 200 items. How do you detect that before the join rather than after, and what do you do with those responses?
  • The join is correct but peak memory is unacceptable. What changes, and which part of the approach survives?
  • How would you test this without a golden output to compare against?

Four days spend equal time on query work, statistics, modelling and product judgement at deliberately shallow depth, which produces a scored map of where you actually stand. The last three days spend everything on the two areas the role weights most, and close by re-running day one to measure movement.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Breadth pass: query fluency
  • Solve six prompts spanning aggregation, joins, window functions and date arithmetic in 60 minutes total, stopping at 10 minutes each whether or not it works, and mark every prompt as solved, solved slowly, or stuck.
  • For each unsolved prompt write the single blocking sentence (I lost the grain, I did not know the frame clause, I could not express the date boundary) instead of reading the solution.
  • Translate one pandas transformation you know well into SQL and one SQL query into pandas, checking that both return the same row count and the same totals.

Deliverable: A scored six-row table, one line per prompt, saved for the day-seven re-run.

Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗
02Breadth pass: statistics and inference
  • Answer ten short questions in writing with nothing open: what a p-value is conditional on, what a 95 percent interval covers across repeated samples, when a paired test is the right one, what the bootstrap estimates, why multiple comparisons inflate false positives, how controlling the family-wise error rate differs from controlling the false discovery rate, what power depends on, what a missed real effect costs a product, the three situations where the central limit theorem does not rescue you (small n, very heavy tails, dependent observations), and what a standard error is the standard deviation of.
  • Grade yourself against a reference and count only the answers that were exactly right, not the ones that were nearly right.
  • Rewrite the two weakest answers the following morning from memory in full sentences.

Deliverable: Ten graded answers with an honest count of exact hits.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
03Breadth pass: modelling
  • Take one tabular dataset end to end in 90 minutes: a leakage-safe split, a baseline that is not a model (majority class or historical mean), one regularized linear model, one gradient-boosted tree, and a single evaluation metric chosen before you look at any result.
  • Write why that metric fits the cost structure: precision at a fixed recall for alerting, calibration for anything feeding a price or a threshold, ranking metrics for retrieval, and note that area under the ROC curve is insensitive to class balance in a way that can flatter a rare-positive problem.
  • Name the leak you were most likely to introduce (an encoding fit on all rows before splitting, or a feature computed after the label's timestamp) and write the check that would have caught it.

Deliverable: A notebook whose first cell states the metric and the baseline, plus two lines on what beat what and by how much.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
04Breadth pass: product judgement
  • Answer three case prompts aloud at 15 minutes each, timing how long passes before you state a success metric.
  • For one case write the first segmentation you would run and the row counts you expect per segment, so that a tiny segment cannot quietly drive the conclusion.
  • Take a metric definition you did not write, from a public dashboard, a textbook, or documentation you already have open, and list every place two analysts implementing it would diverge: which rows the denominator admits, whether the unit is an account or a person, what the time window is anchored to, and what happens to data that arrives late. Then write the one question that would close the largest of those gaps.

Deliverable: Three recorded case answers plus an ambiguity list for a metric someone else defined, ending in the single question you would ask about it.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05Depth, first area
  • Rank the four areas by how many bullet points in the role description each one covers, pick the top one, and spend the entire day inside it.
  • Work the six hardest problems you can find in that area and for each write the generalizable move you should have reached for first, rather than the answer.
  • Re-solve the two you failed the same evening with notes closed.

Deliverable: Six generalizable moves written as instructions to yourself, not as solutions.

Practice prompt ↗Practice prompt ↗
06Depth, second area, and the seam between them
  • Repeat the depth protocol on the second-ranked area with the same six-problem structure.
  • Construct one problem that requires both areas at once, for example a metric redefinition whose effect you must validate with a test whose readout you then have to query.
  • Solve your own combined problem end to end and note where the handoff between the two areas cost you time.

Deliverable: One combined problem, solved end to end, with the handoff failure written down.

Practice prompt ↗Practice prompt ↗
07Integration and re-measurement
  • Re-run the six prompts from day one under the same clock and compare both correctness and time.
  • Run a 60-minute mixed mock that moves between areas without warning, since switching cost is what breadth passes do not train.
  • Write the two areas you would still fail on, and the sentence you will use in the interview when you hit one of them.

Deliverable: A before-and-after score table plus a written plan for the two remaining gaps.

Practice prompt ↗Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

Sometimes the honest read is that the initiative did not work, and the person who commissioned the analysis was hoping otherwise. Interviewers want to know whether you softened it. Prepare the case where you delivered an unwelcome result, how you presented the uncertainty without hiding behind it, and what the team did next.

Tell us about a time you had a conflict with a teammate and how you re…

medium
behavioural and stakeholder questions

Tell us about a time you had a conflict with a teammate and how you resolved it.

Approach
  1. Close with what you would do differently, concretely.
  2. Quantify the outcome, including what you would not claim credit for.
  3. Pick a story where you drove the decision, not one where you observed it.
Follow-up
  • What would you do differently if you ran that project again?
  • How did you know the outcome was caused by your change?

Tell an executive the usage report they publish is wrong

medium
data qualitymetric definitionstakeholder communication

A quarterly usage report an executive sends to district administrators counts an active learner as anyone with a fct_lesson_activity row in the quarter. You find two defects: accounts created by roster_sync that opened a single item during an automated import are counted in the numerator, and learners with age_gated = TRUE never reach the source table, so they are absent from the denominator. The report has already gone out twice. Deliverable: what you send, to whom, in what order, and what the definition should become. Probed: whether you can deliver an unwelcome data-quality finding without detonating it.

Approach
  1. Size the error before raising it. Recompute the metric both ways per org: 'the definition is wrong' and 'the number is overstated by 18 points on the largest account' get very different responses.
  2. Keep the two defects separate, because they push in opposite directions. Roster artefacts inflate the numerator while age-gated truncation shrinks the denominator, so the net error varies by org and a single average hides that.
  3. Go to the report owner first with the recomputation and a proposed replacement, not to their audience. The person who published it needs the chance to correct it themselves.
  4. Propose a definition that is defensible and computable today: distinct learners with at least one scored submission in the quarter, with enrollment_source in ('roster_sync','admin_bulk') excluded from activation cuts, and the age_gated share stated as a footnote rather than silently dropped.
  5. Say what should happen to the two reports already sent, with a recommendation and an owner. The restatement is the uncomfortable part, and leaving it for someone else to raise means the work is unfinished.
Follow-up
  • The executive asks you to change the definition quietly from next quarter with no restatement. What do you do?
  • How would you stop this class of defect reaching a published report again?

Allocate one week across three competing team requests

medium
prioritisationimpact estimationstakeholder management

In one week you receive three requests. Sales wants a renewal-risk list for 40 institutional accounts whose period_end falls in 30 days. Curriculum suspects an item-quality problem on a published unit that is currently collecting responses. Growth wants a signup-flow test sized. You have capacity for roughly one and a half of them and cannot escalate for arbitration. Deliverable: your allocation, the reasoning you give each requester, and the one question you ask each before deciding. Probed: whether you prioritise on decisions and reversibility rather than on who asked loudest.

Approach
  1. Score each request on deadline and reversibility, not on seniority. The renewal list is worthless after period_end; a bad published item compounds with every response collected against it; a test sizing costs almost nothing to delay a week.
  2. Ask each requester the one question that could collapse their request: whether sales already has a workable heuristic list, whether the suspect unit can simply be set to retired today, whether the growth test has a launch date at all.
  3. Look for the cheap partial that still buys the deadline. A rules-based risk cut from first_activity_at, units_completed against units_total, and days to period_end ships in a day and captures most of the value of a model.
  4. Decide, then tell the person who is not getting the work directly and with a date. Unmanaged silence costs more trust than an explicit decline.
  5. Write the decision and its reasoning somewhere durable so next week's triage does not relitigate the same three requests.
Follow-up
  • The growth PM escalates to your skip-level. What do you do, and what do you send ahead of that conversation?
  • Two weeks later the risk list you shipped went unused. What changes in how you triage next time?
  • 01

    Tell us about a time you had a conflict with a teammate and how you resolved it.

  • 02

    A quarterly usage report an executive sends to district administrators counts an active learner as anyone with a fct_lesson_activity row in the quarter. You find two defects: accounts created by roster_sync that opened a single item during an automated import are counted in the numerator, and learners with age_gated = TRUE never reach the source table, so they are absent from the denominator. The report has already gone out twice. Deliverable: what you send, to whom, in what order, and what the definition should become. Probed: whether you can deliver an unwelcome data-quality finding without detonating it.

  • 03

    In one week you receive three requests. Sales wants a renewal-risk list for 40 institutional accounts whose period_end falls in 30 days. Curriculum suspects an item-quality problem on a published unit that is currently collecting responses. Growth wants a signup-flow test sized. You have capacity for roughly one and a half of them and cannot escalate for arbitration. Deliverable: your allocation, the reasoning you give each requester, and the one question you ask each before deciding. Probed: whether you prioritise on decisions and reversibility rather than on who asked loudest.

PracHub interview preparation framework ↗
Is this an official Duolingo interview guide?

No. It is PracHub's own research and practice material for the Data Scientist role at Duolingo. Rounds and questions reflect what candidates have reported, not a process Duolingo has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
What is the typical interview difficulty for this position?

The interview difficulty is generally considered average to difficult, with a focus on both technical and behavioral aspects. Candidates should prepare thoroughly to showcase their skills and experiences.

PracHub interview research ↗
What differentiates successful candidates?

Successful candidates typically demonstrate a strong grasp of technical skills, effective problem-solving abilities, and a clear alignment with Duolingo's mission and values. Additionally, strong communication skills and teamwork are key.

PracHub interview research ↗
What is the culture and working style at Duolingo?

Duolingo fosters a collaborative and inclusive culture. Employees are encouraged to be curious, embrace feedback, and prioritize user-focused design. The work environment is dynamic, with a strong emphasis on continuous learning.

PracHub interview research ↗
What is the typical timeline from the initial screen to the offer?

The timeline can vary, but candidates generally receive feedback quickly throughout the interview process. Expect a few weeks from the initial phone screen to a potential offer.

PracHub interview research ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.