As a Software Engineer at Duolingo, you will build systems that make education accessible, engaging, and effective for hundreds of millions of active learners worldwide. Your work directly powers the core gamification loops, adaptive learning engines, audio/video streaming infrastructures, and real-time user notification pipelines that drive daily engagement. Engineering at Duolingo sits at the intersection of high-scale backend services, mobile experience polish, and data-driven personalization.
In this role, you will work on complex technical problems where scalability, latency, and code maintainability are paramount. Whether you are optimizing core lesson session engines, building scalable push notification architectures for inactive user retargeting, or engineering AI-assisted features for natural language generation, your code directly influences how millions of users acquire new skills every day.
The engineering organization at Duolingo highly values autonomy, rapid execution, code quality, and user-centric problem-solving. You will collaborate closely with product managers, designers, site reliability engineers, and learning scientists in an environment that prioritizes rigorous code reviews, automated testing, and continuous deployment.
Recruiter Screen
reportedThe title covers product work, platform work, infrastructure, mobile and frontend, and those are different jobs with different loops behind them. A screening call is the cheapest place to find out which one the seat is, and asking reads as experienced rather than fussy. The questions that separate them: what the team is on call for, what the last three projects were, and whether any round happens inside an existing repository instead of a blank file. Then say which of that you have done and which you have not. Claiming the whole posting is the fastest way to be found out one round later.
What to demonstrate
- Whether you can locate your experience inside one flavour of the role honestly instead of claiming the entire requirements list
- Whether you name what you have not done, which an experienced screener reads as a level signal and can plan the loop around
- Whether what you want next matches what the seat is: someone who wants greenfield work landing on a team that mostly operates an existing system is a hire that leaves within the year
How to prepare
- Mark every line of the posting as done, adjacent or new, and write one sentence for each adjacent line naming the closest thing you actually built
- Split your last two years into rough percentages across feature work, operating and debugging live systems, and design or review, so a question about scope gets numbers rather than adjectives
- Bring three questions that discriminate between seats: what the team is paged for, how much of the work is changing existing code versus standing up something new, and what shipped in the last quarter
Automated Online Assessment
reportedThe same problem is scored by two different mechanisms depending on the format, and preparing for one does not cover the other. With a person watching, partial progress is visible and a hint is a correction you can absorb; silence is the expensive failure, because nobody can read a half-written function. With an automated grader there is no partial credit for what you were about to do, nobody to ask, and the worked examples in the prompt are the entire specification. Read them as a contract, down to whether an empty result should be an empty list or no output at all.
What to demonstrate
- In a live session, whether your commentary tracks what your hands are doing, and whether a hint redirects you or gets defended against
- In an automated one, whether you cover the cases the examples do not show, since the hidden cases are where the score moves
- Whether you manage the clock on purpose: abandoning an approach that is not converging while there is still time to write something simpler that finishes
How to prepare
- Have someone hand you a problem and feed you one deliberately wrong hint. Practise testing it against a concrete case instead of accepting or rejecting it on authority.
- Do one timed run a week in a plain browser editor with autocomplete, linting and your own snippets switched off, which is closer to what these environments give you
- For the automated format, write the harness before the solution: a main that feeds the worked examples plus an empty and a single-element case and prints expected against actual, so a wrong submission is caught by you first
Technical Screening
reportedWhat this round decides is narrow: whether you can produce code that runs and is correct on inputs nobody showed you. An elegant solution that does not compile scores below a plain one that does, so write a correct brute force first, say out loud that you know its cost, and improve it with the working version still on screen. What separates strong answers is who finds the broken case. Trace your own code against an empty input, a single element, and duplicate keys before you say you are finished, because being told is far more expensive than noticing.
What to demonstrate
- Whether degenerate inputs get checked without being asked for: an empty collection, one element, every element equal, and the extreme value the input type allows
- Whether the complexity you state matches the code you actually wrote, including a sort or a copy sitting inside a loop
- Whether the finished answer is verified against the worked examples before you call it done, rather than assumed correct because the code reads correctly
How to prepare
- Take five problems you have already solved and, without running anything, write down what each returns for empty input, a single element, and all-duplicates. Then run them and count how many you predicted wrong.
- Drill the brute force as its own skill: on ten problems, write only the obviously-correct slow version and time how long it takes to get it passing. If that is more than a few minutes, that is what to practise, not the optimal version.
- Add a fixed last step before you submit anything, reading only the loop bounds and the initial value of each accumulator, which is where most off-by-one errors live
Virtual Onsite
reportedWhere the day includes a partner from product, design or data, that conversation is weighted like the technical ones and prepared for least. They are deciding one thing: whether having you in the room makes their decisions cheaper. That means options with costs attached, not implementation detail and not "it depends". An estimate someone can plan against — a range, the assumption that would push it to the high end, and what you would drop to hit the low one — is worth more than a confident single number, which everyone present already knows is wrong.
What to demonstrate
- Whether an estimate comes as a range with the assumption most likely to break it, and states what a specific scope cut would actually buy
- Whether a technical constraint is handed over as a choice with consequences on their side, rather than as a verdict they have no standing to argue with
- Whether you establish what decision is on the table before proposing anything
- Whether risk is raised while it can still change the plan, with the trigger that would confirm it, instead of reported afterwards as a slip
How to prepare
- Take a project that shipped late and write the two-sentence warning you could have given three weeks earlier, naming what you would have needed decided at that point
- Rehearse one estimate out loud until it arrives in three parts: the range, the single assumption that would blow it, and the smallest thing you would cut to protect the date
- Rewrite an objection you have actually made — the "we can't do that" version — as two options with their costs, so the choice ends up with the person who owns it
9 candidate reports. Individual accounts describe a particular role and hiring cycle.
Duolingo Product Manager interview: Take-home feature proposal and forty-five-minute product rounds
I started with an online application and then went through a recruiter step before meeting anyone live. After that, I completed a take-home product assignment. I had to choose a feature to add or improve in Duolingo and turn it into a short slide deck covering the problem, solution, success metrics, and MVP. It took a while, but the instructions were clear enough that I understood what they wante…
Read full experienceDuolingo Software Engineer interview with disputed coding feedback
The recruiter communication set a strange tone from the beginning. When I was eventually rejected, the message said that my code was nonfunctional even though it had passed the tests. That mismatch bothered me because the feedback felt more dismissive than a simple statement that I wasn't a match. The overall atmosphere, from the recruiter to the interviewers, was unpleasant. It felt as though pe…
Read full experienceDuolingo Product Manager interview: take-home app feature assignment
The process began with a take-home assignment, and I didn't speak with anyone before submitting my work. I had only a few days to devise or reimagine a Duolingo app feature and create a slide deck. It took a lot of effort to get the deck to a presentable level. After I submitted it, the live rounds still felt as if the main goal was to collect ideas from candidates. The most jarring part wasn't t…
Read full experienceDuolingo Software Engineer interview with a one-hour technical round
Once I reached that stage, the process had a strong final-round feel. It followed a kind of superday structure, starting with an online assessment on a Zoom call and ending with a behavioral round that included a couple of STAR-style questions about my experience. The decisions felt quick and decisive. As I progressed, the rounds looked like a shortlist process for engineers and finalists. There…
Read full experienceDuolingo Product Manager: take-home feature design task and delayed rejection
I went through an internship-style process where the first real step was a take-home design task focused on a Duolingo feature. It took a significant amount of time. It felt like something they expected me to work on from start to finish and turn into a polished presentation. After I finished, the follow-up dragged on. The process went silent for a while, and I didn't receive a rejection until mo…
Read full experiencePracHub editorial advice for the preparation topics above.
Treating the launch subject identifier as a global user id, or merging identities on email.
The subject claim is stable only within its issuer, so the same human arriving from a second platform is a different subject and must be a second external identity bound to the same person. Email is worse than useless as a merge key here: privacy settings frequently suppress it from the launch entirely, districts recycle addresses between graduating and incoming students, and younger learners often have none. A merge on email eventually joins two children's work into one account, which is both a grade bug and a privacy incident.
Autoscaling on average utilisation against a bell-shaped step function.
Scale-up latency (scheduling, image pull, process and JIT warm-up, connection pool and cache fill) is typically tens of seconds to minutes, while the demand step is under a minute and repeats on a published schedule. By the time the metric crosses a threshold, the period has started and the queue has already built. The workable answer is a schedule-driven pre-warm derived from the tenant calendar plus a load shedder that protects writes, not a more aggressive threshold.
Quoting amortised or average cost as if it were a worst-case guarantee
Appending to a dynamic array is amortised O(1), but the append that triggers a resize copies every element, and hash lookup is constant only while the hash spreads the actual keys. Say which guarantee you are offering when the caller cares about the latency of one call rather than the total over many.
Finishing a solution without stating its complexity
Give time and space in the same breath as the code, and define n explicitly when there are two sizes, since n nodes and m edges are not interchangeable. Space is the half that gets skipped: count the auxiliary structures you allocate and the recursion stack at its deepest, not only the answer you hand back.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Write a function to calculate the minimum number of steps required for…
Write a function to calculate the minimum number of steps required for a square moving diagonally across a bounded screen to reach any corner.
Approach
- State the target complexity and say which constraint rules the naive version out.
- Name the brute-force solution and its complexity before improving on it.
- Choose the data structure from the access pattern, not from familiarity.
Follow-up
- How does this change if the input no longer fits in memory?
- Which test case would catch an off-by-one here?
Given a starting board state and target configuration, find the shorte…
Given a starting board state and target configuration, find the shortest path of valid tile movements using breadth-first search.
Approach
- Name the brute-force solution and its complexity before improving on it.
- State the target complexity and say which constraint rules the naive version out.
- Choose the data structure from the access pattern, not from familiarity.
Follow-up
- How does this change if the input no longer fits in memory?
- What is the worst case, and how likely is it on real data?
Given a 2D matrix representing heights, find the length of the longest…
Given a 2D matrix representing heights, find the length of the longest strictly decreasing path, and extend the solution to handle a boost item allowing one low-to-high jump.
Approach
- Choose the data structure from the access pattern, not from familiarity.
- Restate the input: its shape, its size, and what is guaranteed about it.
- Walk one small example through your approach before writing the whole thing.
Follow-up
- How does this change if the input no longer fits in memory?
- Which test case would catch an off-by-one here?
Validate an org hierarchy and resolve inherited entitlements
You hold a tenant's normalised org rows from one snapshot: org_id, parent_org_id (nullable), tenant_id, and a boolean saying whether that node carries an explicit content entitlement. There are up to 100,000 nodes and the feed is not trusted, so parent links may form a cycle, reference a node that does not exist, or point at another tenant's org. Reject the snapshot if the links do not form a forest inside one tenant, naming the offending nodes. Otherwise return, for every node, its nearest entitled ancestor including itself, or none.
Approach
- Validate the edge set before traversing anything. Each node has at most one parent, so the graph has at most V edges; reject any edge whose target is missing from the snapshot, and reject any edge whose target carries a different
tenant_id. The cross-tenant edge is the security-relevant one — it makes one district inherit another district's licensed content, which is the tenancy invariant failing through a data path rather than a query. - Detect cycles by reachability rather than by coloured recursion. Collect the roots (
parent_org_id IS NULL), build a child adjacency list in O(V), and run an iterative traversal from the roots. Any node left unvisited is on or below a cycle, because a finite functional graph where every node has one parent decomposes into trees hanging off cycles. To name a cycle, start from any unvisited node and follow parents into a visited set until a node repeats. - Make the traversal iterative with an explicit stack. A well-formed hierarchy is three or four levels deep, but a corrupt feed can produce a chain of 100,000 nodes, and recursive descent dies on it — CPython stops at a recursion limit of 1,000 by default, and a JVM thread's default stack overflows in the tens of thousands of frames. The failure arrives as a crash during ingest of a customer file, which is the worst place to learn this.
- Resolve entitlements on the way down, not on the way up. Push
(node, nearest_entitled_so_far)onto the stack; at each node the value is the node itself when it carries an explicit entitlement and the inherited value otherwise. One pass, O(V + E) time and O(V) space, no memo table needed. - State the alternative you rejected and why. Walking upward per node is O(V * depth), degrading to O(V^2) on the pathological chain the validator is there to catch; memoising the upward walk recovers O(V) amortised but needs a second map and still re-enters nodes. Top-down is strictly simpler here because you are already traversing for the cycle check.
- Report failures as a set, not the first one. Roster operators fix files in batches, so returning every missing parent, every cross-tenant edge and one representative node per cycle in a single pass saves a round trip per defect.
Worked solution 30 min
- Build three structures in one pass over the rows: an id set, a child adjacency map, and a root list, recording missing and cross-tenant parents as you go.
- Traverse iteratively from every root, carrying the nearest entitled ancestor down the stack and writing it into the result map on arrival.
- Compare visited count against node count; if they differ, walk parents from an unvisited node until a repeat to extract one concrete cycle.
- Run a fixture with a three-level district, one department whose parent is missing, one school whose parent belongs to another tenant, and a two-node cycle.
- Add a synthetic 100,000-node chain and confirm the traversal completes rather than overflowing.
Follow-up
- A node's parent changes between snapshots, moving a school to a different district. What must be recomputed, and what must not change about attempts already taken under it?
- Entitlements become time-bounded rather than boolean. How does the top-down pass change, and what is the new answer type?
- The snapshot is valid but 40 percent of nodes changed parent in one run. Is that a defect, and which gate decides?
Replace per-retry outbox keys with derived ones on a live table
score_publication holds 400 million rows and takes continuous writes: (publication_id, tenant_id, attempt_id, target, idempotency_key TEXT NOT NULL, score_given, score_timestamp, state, attempt_count, next_attempt_at). The key was minted as a fresh UUID per retry, so the same grade has several rows and duplicate deliveries have already happened. Migrate to a key derived from (attempt_id, score revision) with UNIQUE (tenant_id, target, idempotency_key), online. Give the ordered steps with the lock each takes, how you handle the existing duplicates, and what you do with rows currently in_flight.
Approach
- Add a plain nullable column, derived_key TEXT, with no default: a catalogue-only ALTER held for microseconds. Set lock_timeout to a second or two and retry, because the hazard is the queue rather than the statement — the ALTER waits behind one long reader and every query arriving afterwards stacks behind its pending ACCESS EXCLUSIVE. Resist the elegant version: adding a STORED generated column rewrites all 400 million rows under that same lock on current major versions.
- Deploy the dual write before the backfill so every new publication row populates derived_key from attempt_id and the score revision, while the relay still reads the old column. The backfill then chases a closed id range instead of a moving target, and you can bound it.
- Backfill in batches by publication_id range, on the order of 20,000 rows per statement, committing between batches, recording the high-water mark in its own table so a killed run resumes rather than restarts, and filtering WHERE derived_key IS NULL so a replayed batch is a no-op. Pace on replica lag and dead tuples, not CPU: each UPDATE writes a new row version and WAL volume tracks rows touched.
- Resolve the duplicates as a data decision, not a DELETE. Group by (tenant_id, target, derived_key) having count(*) > 1; within each group keep the row in state='succeeded' if one exists, since that is the delivery whose outcome is known, otherwise the newest 'pending'. Move the rest to state='dead' with last_error_code recording the migration, and keep the rows — support has to be able to explain a grade that posted twice, and a deleted row explains nothing.
- Leave in_flight rows alone and drain them. A row in_flight has an outstanding request whose outcome is unknown, so collapsing it either loses a delivery record or races the relay marking a row you just rewrote. Pause the claim loop or wait out the lease, assert zero in_flight for the affected keys, then finish the group.
- Build the constraint without a blocking scan and expect the first attempt to fail: CREATE UNIQUE INDEX CONCURRENTLY, which aborts if any duplicate remains and leaves an index with indisvalid = false that must be DROP INDEX CONCURRENTLY'd rather than reused. Once it builds clean, ALTER TABLE ... ADD CONSTRAINT ... UNIQUE USING INDEX adopts it under a brief ACCESS EXCLUSIVE without rebuilding — which is a second reason derived_key is a real column, since USING INDEX accepts neither an expression index nor a partial one.
Worked solution 45 min
- Rehearse on a 10-million-row copy under concurrent relay load, measuring per-batch duration, WAL generated, and replica lag.
- Deploy the dual write, then run the batched backfill and kill it mid-run to confirm it resumes from the high-water mark with no batch applied twice.
- Run the duplicate-group query, classify each group by whether it contains a succeeded row, and apply the supersede update with in_flight rows excluded.
- Attempt CREATE UNIQUE INDEX CONCURRENTLY, observe the failure on any remaining duplicate, drop the invalid index, resolve, and rebuild.
- Adopt the index with ADD CONSTRAINT ... UNIQUE USING INDEX and switch the relay to claim on derived_key.
Follow-up
- score_timestamp must strictly increase per line item and learner. Does this migration change what the receiver sees, and what would a stale retry do mid-migration?
- A tenant's external endpoint has been failing for a week and has 40,000 rows queued. What does your duplicate collapse do to that backlog?
- How do you prove afterwards that no grade was delivered twice as a result of the migration itself?
Decide whether the gradebook grid is derived or materialised
An instructor gradebook renders 35 learners by 60 published assignments in one request against a 300 ms p95 budget. Truth lives in attempt(tenant_id, assignment_id, person_id, attempt_no, state, is_late) and score(attempt_id, tenant_id, score_given, revision), and the policy is best attempt wins. Separately, a district report crosses 90,000 learners and 40 million attempts. Decide with arithmetic whether the grid is derived per request or materialised into gradebook_cell, and if materialised give its key, its staleness detector, and every event that invalidates a cell.
Approach
- Do the arithmetic before choosing anything. The grid is 2,100 cells. With an index on attempt (tenant_id, assignment_id, person_id, attempt_no) and one score lookup per selected attempt, that is a few thousand index tuples and a few thousand buffer accesses against a warm cache — single-digit milliseconds, an order of magnitude inside a 300 ms budget. Deriving it wins, and the thing you avoid buying is an invalidation set you would own forever.
- Identify the read that actually breaks, because it is not this one. The district report crosses 40 million attempts: that is scan-shaped, has no per-request deadline, and belongs in a periodic rollup with its own store and its own freshness SLA. Conflating the two is how a live-path cache gets built to fix an analytics problem.
- Name the condition that would change the answer, so the decision is falsifiable rather than a preference: a policy of average-across-attempts with per-item weighting, a grid that grows to 200 assignments across several terms, or a term-end hour in which 10^3 sections open at once. Each multiplies the per-request work rather than the data, which is the signal that a projection is now cheaper than the recomputation.
- If you do materialise, key gradebook_cell on (tenant_id, section_id, assignment_id, person_id) and store source_attempt_id and source_score_revision on the row. Those two columns are what turn 'is this stale' from an opinion into a comparison, and they let a repair job find divergence by joining rather than by rebuilding everything.
- Enumerate invalidation exhaustively, since the omissions are where this design fails: a newly graded attempt, a regrade that bumps score.revision, an attempt voided after an integrity review, a change to the assignment's policy or point total, an enrollment added or soft-deleted (a withdrawn learner's cell must stop rendering without being deleted, because the score still exists), and a content version repin. Each is an event the projection consumes; a cell written inline by the request path is how a projection quietly becomes the system of record.
- Write the consistency contract down: the projection is derived, rebuilt idempotently from (attempt_id, revision), never authoritative. When cell and score disagree, score wins and the cell is rebuilt, and a scheduled comparison over a sample proves it rather than assuming it. Carry tenant_id in the table and in every cache key — a projection keyed on assignment_id alone is exactly the shape that renders one district's grades inside another's page.
Follow-up
- An instructor edits one score by hand. How long until the grid shows it, and what does the instructor see in the meantime?
- As sections per tenant grow, what breaks first: the rebuild job, the invalidation fan-out, or the read you were trying to protect?
- Would you materialise the column totals and the class mean as well, and what does that do to the invalidation set?
Architect an interactive system designed to deliver adaptive language …
Architect an interactive system designed to deliver adaptive language learning modules embedded within streaming video content.
Approach
- Name the read and write paths separately; they rarely have the same bottleneck.
- Fix the scope first: who calls this, how often, and what they do when it fails.
- Name the failure you are designing for, then the recovery path.
Follow-up
- How does this behave when that dependency is down for an hour?
- What would you drop to keep the system up under load?
Design a scalable backend architecture for tracking daily streak state…
Design a scalable backend architecture for tracking daily streak states and synchronizing offline progress across millions of mobile clients.
Approach
- State the consistency you need, and where you are willing to be stale.
- Name the failure you are designing for, then the recovery path.
- Choose a partition key and say what query it makes expensive.
Follow-up
- What breaks first when traffic grows ten times?
- What would you drop to keep the system up under load?
Design a global push notification system that detects inactive users w…
Design a global push notification system that detects inactive users who have not logged in for several days and dispatches localized reminder alerts.
Approach
- State the consistency you need, and where you are willing to be stale.
- Name the failure you are designing for, then the recovery path.
- Choose a partition key and say what query it makes expensive.
Follow-up
- What breaks first when traffic grows ten times?
- What would you drop to keep the system up under load?
In a live pair-programming session, implement custom multi-attribute s…
In a live pair-programming session, implement custom multi-attribute sorting and filtering logic within an existing code framework.
Approach
- Work from the requirement backwards to the design.
- Say what you would check first and why it is the highest-information step.
- State your assumptions explicitly before working the problem.
Follow-up
- What assumption would you test first?
- How would you know your answer was wrong?
Verify sixty thousand signed launches a minute without replay
Launches arrive as a cross-site form POST carrying a signed token from a platform that publishes its public keys at a JWKS URL. Peak is 60,000 launches per minute sustained for about a minute, six to eight times a day, against an overnight floor near zero. The budget is p99 under 400 ms because the launch blocks first paint inside an iframe. Design the verification path: key caching by key id, the behaviour on an unknown key id, and a nonce burn that stays single-use under two simultaneous replays. State where the latency actually goes.
Approach
- Separate the costs before optimising anything. Verifying an RSA-2048 or P-256 signature is on the order of tens of microseconds of CPU, so at 1,000 launches per second the cryptography is a fraction of one core and is not the bottleneck. The bottleneck is every network round trip on the path: a JWKS fetch, the nonce store round trip, and the reads for enrollment and assignments. Budget the 400 ms as round trips, not as signature work.
- Cache the key set by (issuer, key id) with a long TTL and refresh it in the background so no launch ever waits on the provider. Rotation makes an unknown key id legitimate sometimes, so permit at most one refetch per (issuer, key id) per window using single flight plus a negative cache. Without that bound, a forged key id converts 1,000 launches per second into 1,000 outbound fetches per second against the platform you depend on, which is a denial of service you inflicted on yourself.
- Burn the nonce with one atomic conditional write, a set-if-absent with an expiry, and treat 'already present' as a replay. The expiry must be at least the token's remaining validity plus your clock-skew allowance, or the entry can lapse while the signature still verifies and the token becomes replayable. The burn must precede any side effect. A read followed by a write is not a burn: two concurrent replays both observe absence and both proceed.
- Verify the claims the nonce does not cover: the signature against that key id, the issuer matching a registered platform, the audience matching this tool's client id, expiry and issued-at inside the skew window, and a deployment id that maps to a tenant you know. Resolve identity on (issuer, subject, deployment), never on subject alone and never on email, which is frequently absent and sometimes recycled between learners.
- Name the residual honestly. If the nonce store replicates asynchronously, a failover can lose recently burned nonces, so the replay window is bounded by replication lag rather than being zero. Decide deliberately whether that is acceptable, given that a replay reproduces the same session for the same person, and say what would make it unacceptable.
Worked solution 30 min
- Convert the peak to launches per second and list every network round trip on the critical path with a latency you can defend, then sum them against the 400 ms budget.
- Write the key cache policy: cache key, TTL, background refresh, and the rate limit that bounds refetches on an unknown key id.
- Write the nonce burn as one atomic command and derive its expiry from the token's remaining validity plus a skew allowance.
- Trace two concurrent replays of one token through the burn and confirm exactly one proceeds.
- Order the claim checks cheapest and most-likely-to-fail first, and say which ones can be done before any network call.
Follow-up
- The nonce store is unreachable for 90 seconds during a bell. Do you refuse launches or accept them with no replay check, and what does each choice cost?
- A platform rotates keys and publishes the new one only at the moment it starts signing with it. What does the first launch after rotation do, and how many fetches does the whole fleet make?
- How do you hold p99 under 400 ms on the first launch of the day, when every cache in the path is cold?
Roster apply deadlocks against launches on term-start mornings
In the first week of term, between 07:45 and 08:15, deadlock errors (SQLSTATE 40P01) appear at 20-60 per minute, two roster runs abort with roster_sync_run.error set, some launches return 500, and read p99 doubles even in minutes with no deadlock. A 10^6-row roster apply started at 02:00 and is still running inside one transaction. The apply updates section metadata first and then that section's enrollment rows in file order; the launch path upserts an lti_launch enrollment and then bumps a counter on section. Give an ordered diagnostic checklist, the cycle, and the fix.
Approach
- Order the checklist to separate the two symptoms, because they have different causes and only one is the deadlock. First, pull the deadlock detail from the server log, which names both transactions, both statements and the locks each held and wanted; that is the cycle, printed for you. Second, explain the p99 rise in the minutes with no deadlock at all, which the cycle does not cover. Third, look at transaction duration and lock hold time. Fourth, look at scheduling.
- Read the cycle off the two orders: the apply takes section then enrollment; the launch takes enrollment then section. That is a two-resource cycle in opposite order, so any overlap can deadlock. Postgres detects it after a lock wait exceeds deadlock_timeout, default 1 s, and cancels one transaction with 40P01, which is why the errors are capped at a few dozen a minute rather than unbounded.
- Explain the second symptom separately: a single transaction over 10^6 enrollment rows holds every row lock it has taken until it commits, because row locks are never released early. Any launch touching those sections waits, and the wait is real even when no cycle forms. That is the doubled p99, and it would exist with perfect lock ordering.
- Fix the cycle by removing an edge rather than by reordering both sides. The launch path bumps a counter on section, which is derived data on a latency-critical path; move it to an async aggregate or a cache. With the launch path no longer writing section, the cycle cannot form regardless of the apply's order. Where two writers genuinely must touch both, fix a canonical acquisition order, and have the apply sort its batch by (section_id, person_id) before writing so file order cannot dictate lock order.
- Fix the contention by chunking the apply into bounded transactions, a few thousand rows each, committing between chunks. Lock hold time drops from hours to well under a second, an abort loses one chunk instead of a night, and the run stays restartable from the snapshot digest with per-chunk progress recorded. Retry 40P01 at chunk granularity with jittered backoff, which is safe precisely because the apply is idempotent on the digest. Chunked commits must not weaken the destructive-change gate: it is decided on the full diff before the first chunk commits, never per chunk, so a partially applied snapshot cannot deprovision a bounded fraction of a tenant ahead of the decision.
- Fix the schedule last, as defence rather than as the remedy: cap per-tenant apply concurrency and hold the apply out of the tenant's bell windows from the term calendar. Scheduling alone would hide the bug until a large tenant overran its window again, which is exactly what happened here.
Follow-up
- A chunked apply is no longer atomic. What can a reader observe mid-run, and which of those intermediate states are acceptable?
- Chunk 412 of 800 fails and the run restarts. Prove the restart neither double-applies nor skips rows.
- Raising deadlock_timeout would reduce the error count. Explain precisely what it would make worse.
Four days sample coding, design, fundamentals and the practical rounds at deliberately shallow depth, which is enough to surface the topics you did not know were in scope. That map, rather than a guess made on day one, decides where the last three days go.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Coding, one pass at shallow depth
- Solve one problem from each of six families, an array with two pointers, hash counting, binary search, a tree traversal, a graph traversal and one dynamic program, under a hard twenty-minute cap with no extensions, marking each finished, late, or stalled.
- For every stall, write the exact move you could not make rather than the subject, so the note reads could not turn the recurrence into a loop rather than bad at dynamic programming.
- Fix nothing today. The value of the pass is the unfixed record.
Deliverable: Six timed attempts marked finished, late or stalled, each stall carrying a named blocking move.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Design, one pass at shallow depth
- Spend twenty minutes each on three different shapes, a read-heavy feed, a write-heavy ingest path, and something needing a transaction across two entities, stopping each at requirements, interface and data model.
- After each, write the first question you could not answer, which is usually a number you could not estimate or a failure mode you had no vocabulary for.
- Mark which of the three you would be most relieved not to be asked, and treat that as data rather than as a preference.
Deliverable: Three shallow designs, each with the first unanswerable question written at the bottom.
Practice prompt ↗Practice prompt ↗03Fundamentals and the practical rounds
- Answer eight short questions in writing at four minutes each, covering the material that fills the gaps between the big rounds: what happens between a URL and a rendered page, what an index costs on write, when a process is preferable to a thread, and what conditions a deadlock requires.
- Do one thirty-minute practical task of the kind a take-home compresses: read an unfamiliar two-hundred-line file and write what it does, what you would change, and the one thing you remain unsure of.
- Score every answer fluent, correct but slow, or absent, and keep the absent ones visible.
Deliverable: Eight scored short answers and one written reading of unfamiliar code.
Practice prompt ↗Practice prompt ↗04The rounds that are about you, and the map
- Deliver three behavioural answers aloud against a timer, a conflict, a failure you owned, and a decision made without enough information, marking any that ran past three minutes or contained no number.
- Assemble the map: every marked item from days one to three on a single page, sorted by how likely it is to appear in your loop rather than by how uncomfortable it felt.
- Choose exactly two areas for the remaining three days and write down what you are deliberately abandoning.
Deliverable: A one-page scored map of the whole surface area with two areas chosen and the rest explicitly abandoned.
Practice prompt ↗Practice prompt ↗Worked solution ↗05First chosen area, to the depth you skipped
- Work the higher-ranked area in four focused blocks, choosing items one level above where you stalled rather than repeating what already works.
- After each block write the rule you extracted in one sentence with its precondition attached, since a rule carrying no precondition is exactly what fails under a variation.
- Re-attempt the day-one or day-two item that exposed this area and compare against the original timing.
Deliverable: Four worked blocks, a timed re-attempt against the original, and three one-sentence rules with preconditions.
Practice prompt ↗Practice prompt ↗06Second chosen area, where the gap is coverage rather than speed
- Treat the second area differently from the first. Day five drilled something you could already half-do; this one is usually a topic you had simply never met, so build one worked reference example end to end and keep it, rather than attempting six problems badly.
- Write down the vocabulary you were missing on day two or three, five terms at most, each with the one sentence that makes it usable in an answer rather than the textbook definition.
- Redo the shallow attempt that exposed this area and note whether you now fail later in the problem, because moving the failure point is the realistic gain from a single day and is worth more than a score that did not change.
Deliverable: One worked reference example for the newly covered area, a five-term vocabulary list, and a note on where the failure point moved.
Practice prompt ↗Practice prompt ↗07Reassemble the loop
- Sit two rounds back to back with no gap, ordering them so the area you chose second comes last, because the map was built from rested, isolated attempts and the loop will reach your weaker area when you are already spent.
- Write where the second round suffered from the first, which is normally the point at which structure collapses into narration.
- Reduce the week to one page holding only the rules you can state without reading them.
Deliverable: Mock notes on cross-round carryover plus a one-page card of rules you can recite from memory.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Keep one story where the bad call was yours rather than a dependency's or a manager's. Name the check that would have caught it, whether you added that check afterwards, and whether it has fired since. Answers that route blame outward end the conversation early; answers that end in a guardrail someone still relies on tend to open it up.
How do you handle situations where a peer raises conflicting feedback …
How do you handle situations where a peer raises conflicting feedback during a pull request review?
Approach
- Close with what you would do differently, concretely.
- Give the blast radius: what could have broken, and what you measured.
- Pick a story where you made the decision, not one where you watched it.
Follow-up
- What would you do differently if you ran that again?
- How did you know your change caused the improvement?
Describe a situation where you had to quickly acquire expertise in an …
Describe a situation where you had to quickly acquire expertise in an unfamiliar language or technological framework to deliver a critical project feature.
Approach
- Pick a story where you made the decision, not one where you watched it.
- Give the blast radius: what could have broken, and what you measured.
- State the situation in two sentences and spend the rest on the reasoning.
Follow-up
- What did you decide not to do, and why?
- How did you know your change caused the improvement?
Explain missing grades without absorbing a credential failure
A district reports that a week of grades never reached their gradebook. The publication rows sit in failed_retryable with last_error_code 401 dating from the day the district rotated its platform credentials, and attempt_count is at the cap. The internal score rows are correct. You are on a call with a district administrator and an engineer from the platform vendor. Describe handling this: what you say first, what evidence you put in front of them, what you own, and what you decline to own without turning the call adversarial.
Approach
- Open with what is true for a learner, before any attribution: the grades exist, they are correct, none were lost, and the repair is a republish. That sentence removes the panic driving the call, and every subsequent technical point lands better once the administrator is no longer worried about a week of missing work.
- Show state rather than narrative. Put the publication rows on screen with state, attempt_count, last_error_code and last_error_at, and place the first 401 next to the credential rotation date. Evidence someone can ask you to re-run in front of them is more persuasive than a summary, and it keeps the conversation on a shared artefact rather than on recollections.
- Own your actual defect, which is visibility and not delivery. A week of 401s for one tenant should have paged someone or reached that district's administrator on day one, and it did not. Saying that first, unprompted, is what buys the right to decline the rest; teams that defend everything get believed on nothing.
- Decline the remainder by describing mechanism instead of assigning blame. A 401 is the receiving platform refusing the credential, and no amount of retrying on the sending side recovers it. State what you need, which is a re-authorised deployment, and what happens next: a replay whose idempotency keys prevent duplicates and whose ordering tokens keep the newest score winning.
- Leave with one commitment per party and a defined verification. Who re-authorises, when you replay, how the district confirms the grades landed, and what you will have built before the next rotation so that this is a one-day problem rather than a one-week one.
Follow-up
- Some rows have exhausted retries and moved to a dead state. What does a safe replay do, and what makes it safe to run twice?
- What alert would have caught this on day one, and what would it have had to avoid firing on?
- The administrator asks whether any grade was ever wrong, not merely late. How do you answer with evidence rather than reassurance?
- 01
How do you handle situations where a peer raises conflicting feedback during a pull request review?
- 02
Describe a situation where you had to quickly acquire expertise in an unfamiliar language or technological framework to deliver a critical project feature.
- 03
A district reports that a week of grades never reached their gradebook. The publication rows sit in failed_retryable with last_error_code 401 dating from the day the district rotated its platform credentials, and attempt_count is at the cap. The internal score rows are correct. You are on a call with a district administrator and an engineer from the platform vendor. Describe handling this: what you say first, what evidence you put in front of them, what you own, and what you decline to own without turning the call adversarial.
Is this an official Duolingo interview guide?
No. It is PracHub's own research and practice material for the Software Engineer role at Duolingo. Rounds and questions reflect what candidates have reported, not a process Duolingo has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult are the technical interviews at Duolingo compared to other tech companies?
The technical bar is rigorous, combining standard data structure challenges with real-world practical evaluations like code reviews and pair programming. Success relies equally on clean code style and algorithmic correctness.
PracHub interview research ↗What programming languages am I allowed to use during the coding rounds?
Most technical rounds permit standard mainstream languages such as Python, Java, C++, or JavaScript/TypeScript. Certain pair-programming exercises may utilize pre-configured environments (Java or Python), but interviewers assist with language-specific syntax if needed.
PracHub interview research ↗What sets Duolingo's interview process apart from typical industry processes?
Duolingo places significant weight on practical engineering rounds, such as asking you to comment on realistic pull requests and collaborate via live pair programming, rather than relying exclusively on whiteboard algorithm puzzles.
PracHub interview research ↗How long does the hiring process take from initial screen to offer?
The typical process takes between three to six weeks depending on scheduling availability, candidate response velocity, and virtual onsite scheduling. Recruiters maintain fast communication turnarounds at most stages.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-24 - 02PracHub Software Engineer practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-24 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-24