As a Software Engineer at Zearn, you play a critical role in building and scaling an educational technology platform that impacts millions of K-8 students and teachers across the United States. Zearn is dedicated to helping all children love and learn math, so its engineering team builds accessible, interactive, and engaging digital learning experiences. Your work directly influences how students grasp complex mathematical concepts and how educators track and support student progress in real-time.
The engineering challenges at Zearn span across high-performance web applications, complex database design, and intuitive user interfaces that must run seamlessly on diverse classroom hardware. As a member of Zearn's engineering team, you will collaborate closely with product managers, curriculum designers, and visual artists to translate pedagogical frameworks into clean, performant, and reliable code. This is not just about writing software; it is about building equity in education through robust technology.
This role requires a blend of technical execution, systems thinking, and commitment to Zearn's mission. Whether you are optimizing database queries to handle peak classroom hours or refining the frontend interactive elements of a math lesson, your contributions will scale to support millions of active users. It is an environment where code quality, architectural durability, and user empathy are valued in equal measure.
Application Review
reportedThe person on this call usually cannot evaluate your code and does not need to. They write a short paragraph, and that paragraph is what a hiring manager skims when deciding who to put on your loop. So the test is not whether your work was hard, it is whether a non-engineer can repeat it correctly. Name systems by what they did rather than by their internal codename, give each project a shape (what was breaking, what you changed, what happened after), and keep the whole walkthrough near ninety seconds. Depth that cannot survive a paraphrase reads as vagueness.
What to demonstrate
- Whether a non-engineer can restate your projects without distorting them, since their paraphrase is what travels to the hiring manager, not your sentences
- Whether each project has a shape rather than a stack list: the failure or constraint, the change you made, the result and how it was measured
- Whether you can say what was yours inside a team project without either inflating it or disappearing into the plural
How to prepare
- Rewrite each headline project as two sentences with no internal system names and no acronyms outside your company, then say them to someone outside engineering and have them repeat them back. Fix whatever came back wrong
- Attach one measured number to each project: the baseline, the change, and the window it was measured over. Where nothing was ever measured, say that plainly rather than reaching for a plausible percentage
- Time the background walkthrough against a clock. If it runs past two minutes, compress the earliest role to a single clause and spend the recovered time on the most recent one
Phone Screen
reportedBefore anything technical happens, someone has to decide which rung of the ladder your loop is calibrated to, and that decision sets the bar for every round after it. It comes from how you describe scope, not from your title, because titles do not convert cleanly between companies. The weak version of the answer is team size and years. The strong version names the largest change you shipped where nobody reviewed the design, what would have broken if you had been wrong, and what you were paged for. Get the level said out loud on this call, because the range and the loop both follow from it.
What to demonstrate
- Whether the scope in your own account maps onto a level the team actually has an opening at, so a mismatch ends the process cheaply rather than after four interviewers have spent a day
- Whether your title needs re-mapping: the same word describes very different amounts of independent decision-making at a twenty-person company and a ten-thousand-person one
- Whether your compensation expectation can be filled at that level in the structure the role pays in, which is why the number gets asked for before any engineer is scheduled
How to prepare
- Write down two changes from the last two years: the largest one you designed with nobody reviewing the design, and the largest one where someone more senior did. Lead with the first when scope comes up, and be ready to say which parts of the second were yours
- Ask which level the loop is calibrated to and what changes at the level above it, then plan your weeks from that answer rather than from the posting
- Settle a total-compensation range beforehand with the split named, base against bonus against equity and its vesting period, so a question about numbers gets a number instead of the word market
Programming Exercise
reportedWhen a round has no standard shape, it is often there because something is still open: an area no earlier conversation reached, a round where the signal came out mixed, or a decision someone is not ready to make alone. Work out which by going back over what each earlier round actually covered rather than how it felt, and arrive able to give evidence on that point without being asked twice. Weak answers replay the loop's earlier material at the same depth. Strong ones go a level deeper and stay consistent with what you already said.
What to demonstrate
- Whether your account of a project matches the one you gave earlier in the loop, since what you said before may be available to whoever runs this round
- Whether you can go a level deeper on something already covered, reaching the decision and its alternatives rather than repeating the summary
- Whether you state your own uncertainty accurately, including parts of a system you did not build and decisions you inherited, instead of claiming even ownership across all of it
- Whether you can answer a question you handled poorly earlier by naming what you missed, rather than delivering a polished second version as if the first had not happened
How to prepare
- Reconstruct the loop on one page: for each round, the questions you were asked and the answer you actually gave, not the better one you thought of afterwards. The gaps on that page are your best available guess at why this round exists.
- Take the two claims you made earlier that carry the most weight and assemble the backing for each: the measurement, the date, what broke, the decision you would make differently now.
- Write down the three facts about your work that must not drift between tellings, such as team size, timeline and your own role, and check your stories against that list rather than trusting recall under pressure
Take-Home Assessment
reportedThe README is read before the code, and a follow-up conversation is usually built from it, so treat every sentence you put there as a question you have agreed to answer. It needs the command that runs the thing, the assumptions you made where the prompt was ambiguous, and the limits of what you built stated with the preconditions that make them true. Overclaiming is the expensive mistake here. Writing that something is thread-safe, or constant-time, or handles files larger than memory invites a reader to check that exact line, and a claim the code cannot support costs more than silence would have.
What to demonstrate
- Whether the run instructions work from a clean clone, naming the exact commands, the language version you tested on, and any environment variable the program expects
- Whether ambiguities in the prompt are resolved in writing, with the interpretation you picked and the reason, rather than settled silently in the code
- Whether documented limits match the implementation, so a stated input bound is one the code enforces or at least does not contradict
- Whether the trade-offs you list come with the condition that would make you choose the other way, instead of reading as a list of alternatives you happened to consider
How to prepare
- Write the README before the final hour, then read the code against it claim by claim and correct or delete every statement the implementation does not back
- For each ambiguity in the prompt, write one sentence fixing your interpretation and keep it; those sentences become the assumptions section and your answer when someone asks why you did it that way
- Give the repository to someone who has not seen the prompt and ask them to run it using only what is written down, treating every question they have to ask you as a gap in the document
Virtual On-Site Panel
reportedNobody in the room with you decides this. Interviewers typically write their rounds up separately, often before seeing anyone else's, and the outcome is settled later from those write-ups. A split panel gets resolved by whichever note carries specific evidence, so what you want out of each room is one concrete thing that person could write down: a bug you caught yourself, a trade-off you named, a decision you owned. The rest is arithmetic. The project you describe in a behavioural conversation is often the same system you sketched an hour earlier, and the two accounts have to agree.
What to demonstrate
- Whether the scale, team size and timeline you attach to a project hold steady when that project resurfaces in a different round
- Whether each interviewer leaves with a specific thing to cite rather than a general impression of competence
- Whether a trade-off you defended in one round survives a challenge in another, instead of being quietly swapped for the answer the new interviewer seemed to want
- Whether a question you have already answered earlier in the day gets the same answer at the same depth, without visible impatience
How to prepare
- Write a one-page sheet per project fixing the figures you will quote — request volume, data size, team size, elapsed time, what broke — and say them aloud from the sheet until they come out identical every time
- For each round on the schedule, decide in advance the one sentence you want in that person's notes, then check in a mock that you said it outright instead of leaving it to be inferred
- Have someone ask you the same project question twice, an hour apart, and diff the two answers for numbers that moved or a trade-off that reversed
PracHub editorial advice for the preparation topics above.
Editing a content item in place and thereby rewriting the meaning of every historical score.
Authors expect to fix a typo, change a distractor, or adjust a point value, and nothing warns them that thousands of stored results reference the record they are editing. If attempts and analytics join to the mutable item rather than to a version, a regrade or a rebuilt report scores past answers against an item that did not exist when they were given. Versioning has to be the default write path, because a convention that authors must remember will not hold.
Treating the launch subject identifier as a global user id, or merging identities on email.
The subject claim is stable only within its issuer, so the same human arriving from a second platform is a different subject and must be a second external identity bound to the same person. Email is worse than useless as a merge key here: privacy settings frequently suppress it from the launch entirely, districts recycle addresses between graduating and incoming students, and younger learners often have none. A merge on email eventually joins two children's work into one account, which is both a grade bug and a privacy incident.
Assuming the bug is in the framework
Suspect your own code first: read the stack trace top to bottom, check which versions are actually installed rather than which ones you believe are, and reproduce in isolation before blaming a library that thousands of people run daily. When the fault really is upstream, you need that minimal reproduction to say so credibly anyway.
Check-then-act on shared state
Read, decide, write is not safe under concurrency unless the decision and the write are one atomic step: a unique constraint with conflict handling, a compare-and-set, or a row lock held for the whole transaction. Two requests can both pass the existence check before either inserts, which shows up as duplicate rows under load and never in a single-threaded test.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Write a utility function to flatten a nested configuration object repr…
Write a utility function to flatten a nested configuration object representing a school curriculum hierarchy.
Approach
- State the target complexity and say which constraint rules the naive version out.
- Name the brute-force solution and its complexity before improving on it.
- Walk one small example through your approach before writing the whole thing.
Follow-up
- How does this change if the input no longer fits in memory?
- What is the worst case, and how likely is it on real data?
Write a function to determine if the parentheses, brackets, and braces…
Write a function to determine if the parentheses, brackets, and braces in a given string are closed and nested correctly.
Approach
- Choose the data structure from the access pattern, not from familiarity.
- Name the brute-force solution and its complexity before improving on it.
- Restate the input: its shape, its size, and what is guaranteed about it.
Follow-up
- What is the worst case, and how likely is it on real data?
- How does this change if the input no longer fits in memory?
Implement a function that parses a custom string representation of a m…
Implement a function that parses a custom string representation of a math problem and evaluates its basic arithmetic operations.
Approach
- Name the brute-force solution and its complexity before improving on it.
- State the target complexity and say which constraint rules the naive version out.
- Choose the data structure from the access pattern, not from familiarity.
Follow-up
- Which test case would catch an off-by-one here?
- What is the worst case, and how likely is it on real data?
Rank the longest-stuck grade publications in bounded memory
You are streaming score_publication rows to build an operator view: tenant_id, attempt_id, target, state, attempt_count, created_at, next_attempt_at, last_error_code. Up to 20,000,000 rows are in states pending and failed_retryable, spread across up to 5,000 tenants, and they arrive as an unsorted export you may not re-read. Return the 50 longest-stuck publications overall, the 5 longest-stuck per tenant, and a stuck count per tenant. Memory must be bounded independently of row count. State your bounds and the field you rank on.
Approach
- Pick the ranking field before the data structure, because this is where the answer is usually lost. Stuck duration is
now - created_at, not anything derived fromnext_attempt_at. Exponential backoff pushesnext_attempt_atfurther out with every failure, so ordering ascending by it surfaces the publications that just failed once and buries the row that has been retrying for six days — the ranking inverts exactly on the rows the operator opened the page for. - For the global top 50 — the 50 smallest
created_at— keep a max-heap of size 50 keyed oncreated_at. Push each row; when the heap exceeds 50, pop the maximum. The heap head is the youngest survivor, so a row older than the head displaces it and everything else is discarded in O(1) after one comparison. That is O(N log K) time and O(K) space, with the log K term only paid on the shrinking fraction of rows that beat the head. - For per-tenant results keep a map from
tenant_idto a size-5 max-heap plus an integer count. Memory is O(T * K') = 5,000 * 5 entries plus 5,000 counters, a few megabytes, and it is bounded by tenant cardinality rather than row count. Say that out loud: if tenants were unbounded this structure is not, and you would need a sketch or a two-pass job. - Do not sort. A full sort is O(N log N) over 20,000,000 rows and must materialise all of them; the heap never holds more than K. Note that a database
ORDER BY ... LIMIT 50reaches for the same bounded heap internally, so the interesting case is precisely the one stated — a stream you cannot re-read and cannot fit. - Carry the diagnostic fields through the heap entry rather than re-joining afterward. An operator needs
attempt_countandlast_error_codenext to the age to tell a receiver that is rate limiting from one that is rejecting the payload, and a second pass to fetch them defeats the single-pass constraint you just paid for. - Exclude
deadrows from these structures and count them separately. They are terminal by definition, so mixing them in lets a permanently dead publication occupy a slot in the top 50 forever while genuinely stuck rows rotate beneath it.
Worked solution 25 min
- Write the comparator explicitly: order by
created_atascending, tiebreak onpublication_idascending, and build a max-heap under that order so the head is the weakest survivor. - Stream rows once, skipping
dead, updating the global heap, the tenant heap and the tenant counter for each surviving row. - Drain both heap levels at the end and sort only the K and K' survivors for display, which is at most 25,050 items.
- Build a fixture where one tenant has a publication created six days ago with
attempt_count14 and anext_attempt_atfar in the future, a second tenant holds eight of the fifty globally oldest rows so the containment check exercises the k-above-5 case, and twenty further publications created minutes ago carry imminent retry times. - Assert the six-day-old row is first in both the global and the tenant list, then re-rank by
next_attempt_atand record that it lands last.
Follow-up
- Two publications share the same
created_atto the microsecond. What is your tiebreak, and why does a deterministic one matter for an operator page that refreshes? - The operator wants to replay the top 50 safely. What must be true of the idempotency key for that to be a no-op when the original attempt actually succeeded?
- Tenant cardinality rises to 5,000,000. Which part of your structure breaks first and what replaces it?
Decide whether the gradebook grid is derived or materialised
An instructor gradebook renders 35 learners by 60 published assignments in one request against a 300 ms p95 budget. Truth lives in attempt(tenant_id, assignment_id, person_id, attempt_no, state, is_late) and score(attempt_id, tenant_id, score_given, revision), and the policy is best attempt wins. Separately, a district report crosses 90,000 learners and 40 million attempts. Decide with arithmetic whether the grid is derived per request or materialised into gradebook_cell, and if materialised give its key, its staleness detector, and every event that invalidates a cell.
Approach
- Do the arithmetic before choosing anything. The grid is 2,100 cells. With an index on attempt (tenant_id, assignment_id, person_id, attempt_no) and one score lookup per selected attempt, that is a few thousand index tuples and a few thousand buffer accesses against a warm cache — single-digit milliseconds, an order of magnitude inside a 300 ms budget. Deriving it wins, and the thing you avoid buying is an invalidation set you would own forever.
- Identify the read that actually breaks, because it is not this one. The district report crosses 40 million attempts: that is scan-shaped, has no per-request deadline, and belongs in a periodic rollup with its own store and its own freshness SLA. Conflating the two is how a live-path cache gets built to fix an analytics problem.
- Name the condition that would change the answer, so the decision is falsifiable rather than a preference: a policy of average-across-attempts with per-item weighting, a grid that grows to 200 assignments across several terms, or a term-end hour in which 10^3 sections open at once. Each multiplies the per-request work rather than the data, which is the signal that a projection is now cheaper than the recomputation.
- If you do materialise, key gradebook_cell on (tenant_id, section_id, assignment_id, person_id) and store source_attempt_id and source_score_revision on the row. Those two columns are what turn 'is this stale' from an opinion into a comparison, and they let a repair job find divergence by joining rather than by rebuilding everything.
- Enumerate invalidation exhaustively, since the omissions are where this design fails: a newly graded attempt, a regrade that bumps score.revision, an attempt voided after an integrity review, a change to the assignment's policy or point total, an enrollment added or soft-deleted (a withdrawn learner's cell must stop rendering without being deleted, because the score still exists), and a content version repin. Each is an event the projection consumes; a cell written inline by the request path is how a projection quietly becomes the system of record.
- Write the consistency contract down: the projection is derived, rebuilt idempotently from (attempt_id, revision), never authoritative. When cell and score disagree, score wins and the cell is rebuilt, and a scheduled comparison over a sample proves it rather than assuming it. Carry tenant_id in the table and in every cache key — a projection keyed on assignment_id alone is exactly the shape that renders one district's grades inside another's page.
Worked solution 40 min
- Write the derived query and measure it on a seeded section with EXPLAIN (ANALYZE, BUFFERS), recording buffers and runtime at a cold and a warm cache.
- Multiply out the term-end scenario: sections opening per minute times cells per section times buffers per cell, and compare against the connection and buffer budget you have.
- Draft gradebook_cell with source_attempt_id and source_score_revision, and write the rebuild as an idempotent upsert keyed on the cell key.
- List every invalidation event and, for each, the cells it touches; find the one with the widest fan-out and state its cost.
- Write the divergence check as a query comparing a sample of cells against a freshly derived value.
Follow-up
- An instructor edits one score by hand. How long until the grid shows it, and what does the instructor see in the meantime?
- As sections per tenant grow, what breaks first: the rebuild job, the invalidation fan-out, or the read you were trying to protect?
- Would you materialise the column totals and the class mean as well, and what does that do to the invalidation set?
Reconstruct paused and active time from an append-only event table
attempt_event is append-only: (attempt_id, seq BIGINT monotonic per attempt, event_type IN ('start','pause','resume','autosave','submit'), occurred_at TIMESTAMPTZ, client_ts TIMESTAMPTZ). attempt carries started_at, server_deadline_at, accumulated_pause_seconds and submitted_at. Write one query returning, per attempt in a section, total paused seconds and active seconds between start and submit, and flagging attempts whose active time exceeds the assignment's time limit by more than 60 seconds. A pause may have no matching resume. State which ordering you use and which window frame, and why.
Approach
- Order by seq, not by a timestamp. seq is monotonic per attempt by construction; occurred_at is server-receive time and ties under a burst, and client_ts is attacker-controlled on a device the learner owns. Ordering on a column with ties silently reorders a pause and its resume.
- Close each pause with the next 'resume' or 'submit', not with the next event of any type, or an autosave landing during a pause ends the interval two seconds in. Either filter to the closing event types in a CTE and then LEAD over that reduced set, or keep every row and use min(occurred_at) FILTER (WHERE event_type IN ('resume','submit')) OVER (PARTITION BY attempt_id ORDER BY seq ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING). The second form works because FILTER is permitted on an aggregate used as a window function; it is not permitted on lead(), which is a window function and not an aggregate, so lead(...) FILTER (...) is a syntax error rather than a subtle bug.
- Define the unmatched pause explicitly instead of letting NULL decide: close it at submitted_at, and where there is no submit, at least(server_deadline_at, now()). An interval closed at now() makes the number change on every run, so pin it for any stored report and say so in the output.
- Aggregate with sum(EXTRACT(EPOCH FROM (closed_at - occurred_at))) FILTER (WHERE event_type = 'pause') GROUP BY attempt_id, then active = EXTRACT(EPOCH FROM (submitted_at - started_at)) - paused. Where you use a running window anywhere in this query, write the frame: with ORDER BY and no frame clause the default is RANGE UNBOUNDED PRECEDING AND CURRENT ROW, which includes every row tied on the ordering key — harmless on a unique seq, which is exactly why the tiebreak is the load-bearing decision.
- Reconcile the computed total against attempt.accumulated_pause_seconds and treat a disagreement as a finding, not as a rounding difference: it means either an event was lost, or a code path changed the counter without emitting an event. The event table is the auditable side, so it wins in the report and the counter gets repaired from it.
- Support the access with an index on attempt_event (attempt_id, seq) so each partition is an ordered range scan; then confirm no Sort node sits above it, since a sort here is O(E log E) over every event in the section rather than a walk.
Follow-up
- The same event is delivered twice with the same seq. What does your query do, and what constraint should have prevented it?
- Two tabs produce interleaved events for one attempt. How do you detect the interleaving from this table, and what does it do to the pause arithmetic?
- How would you use this output to prove a learner did not gain time by reloading?
Publish grades outward exactly once under retries and regrades
score_publication holds publication_id, tenant_id, attempt_id, target (lti_ags, webhook, sis_export), idempotency_key with UNIQUE (tenant_id, target, idempotency_key), score_given, score_maximum, score_timestamp, state (pending, in_flight, succeeded, failed_retryable, dead), attempt_count and next_attempt_at. The receiving endpoints are third-party: they rate-limit, they time out after having committed, and they can be down for days. Design the relay: how a row is claimed, what happens when the process dies after the HTTPS call but before the row is marked, how the idempotency key is derived, what stops a stale retry from overwriting a regrade, and which rows a claim is allowed to select. next_attempt_at is the only time column on the row; say what it means in each state.
Approach
- Write the score row and the publication rows in one transaction. That is the entire point of an outbox: a queue publish after the commit can be lost, and a queue publish before the commit can announce a transaction that then rolls back. There must be no window in which a grade exists with no recorded intent to deliver it.
- Claim with a single conditional statement rather than a read then a write, and let next_attempt_at carry two meanings: the backoff time for a row waiting to be sent, and the lease deadline for a row already out. UPDATE score_publication p SET state='in_flight', attempt_count = attempt_count + 1, next_attempt_at = now() + :lease FROM (SELECT publication_id FROM score_publication WHERE state IN ('pending','failed_retryable','in_flight') AND next_attempt_at <= now() ORDER BY next_attempt_at FOR UPDATE SKIP LOCKED LIMIT :n) c WHERE p.publication_id = c.publication_id RETURNING p.*. Both halves are load-bearing. Drop in_flight from the state list and that state becomes a one-way door: a row a process marked in_flight before dying matches no claim again, stalls permanently, and holds its tenant's backlog behind it. Drop the lease write and the row is re-claimable the instant it is claimed, because its next_attempt_at is already in the past, so every publication goes out once per polling worker. SKIP LOCKED is what lets N relay processes share one table without contending on the same head rows; the lease is what makes a crash recoverable. They are separate mechanisms and the design needs both. Index for the predicate with a partial index on next_attempt_at restricted to those three states, or the ORDER BY degrades into a scan over every succeeded row the table has ever accumulated.
- Accept that the transport is at-least-once and make the receiver's side idempotent instead of chasing exactly-once on the sender. A process that dies after the call and before the mark leaves the row in_flight with a lease deadline now in the past, and the next claim picks it up and re-sends it. Size :lease above the HTTP timeout plus a margin: a lease shorter than the call reclaims a request that is still running and sends the same grade twice concurrently. That duplicate is harmless only if the key is derived from (attempt_id, score revision) and nothing else. A UUID minted per retry, or a key containing a clock reading, makes every retry a fresh write at the receiver and turns one grade into several. Increment attempt_count at claim time rather than at completion, so a publication that kills its worker on every attempt walks its own counter to the cap and reaches dead, rather than being reclaimed forever.
- Treat ordering as a separate failure from duplication. A retry of an old score can land after a regrade's newer score, so score_timestamp must strictly increase per line item and learner; the AGS profile directs receivers to reject a non-increasing timestamp with a conflict, but a sender cannot assume every implementation honours it. Also suppress superseded rows at claim time, so a pending publication that is no longer the current revision for that attempt and target is retired rather than sent.
- Classify responses into distinct actions instead of one retry rule. A timeout is an unknown outcome and must be retried with the same key. A conflict on a stale timestamp is success-equivalent and should be marked, not retried. A rate-limit response honours a retry hint when present and otherwise backs off exponentially with jitter under a per-tenant concurrency cap. An expired credential is retryable only after refresh. A permanent rejection, such as a line item that no longer exists, is terminal and moves to dead with last_error_code recorded, because a row retrying forever looks identical to a healthy system on every dashboard.
Worked solution 35 min
- Write the transaction that inserts the score and its publication rows together, and state precisely what is lost if the publish happens outside it.
- Write the skip-locked claim, name every state its predicate selects, set the next lease deadline in that same statement, then trace a process that dies immediately after the HTTPS call through to the claim that picks the row up again.
- Derive the idempotency key from named columns and show two separate retries producing an identical key.
- Build the response classification table: timeout, stale-timestamp conflict, rate limit with and without a hint, credential expiry, line item gone, and server error, with the resulting state and next_attempt_at for each.
- Trace a regrade racing a stale retry and show which score the receiver ends up holding.
Follow-up
- A district's endpoint is down for four days and 200,000 publications accumulate. What happens when it returns, and what protects it from your entire backlog arriving at once?
- An instructor regrades one attempt three times inside a minute. How many outbound calls does the receiver see, and which score does it hold at the end?
- What does a support engineer see for a stuck tenant, and what exactly does a safe manual replay of a dead row do?
Fix an attempt-start endpoint whose timeout retry burns an attempt
POST /assignments/{id}/attempts allocates the next attempt_no for the caller, copies content_version_id onto the row, and computes server_deadline_at once from time_limit_seconds. The caller is the learner's browser on school wifi; it retries once when a request times out. With max_attempts = 2, a timed-out create followed by that retry consumes both attempts and hands the learner a fresh deadline. Redesign the endpoint so a retry is safe, a deliberate second attempt is still possible, and two tabs cannot each create one. Give the request shape, the storage, and the response to a replay.
Approach
- Say why the naive shape fails rather than patching it: POST-creates-a-child is non-idempotent by construction, and here the duplicate is worse than the original because it spends a countable resource (an attempt_no against max_attempts) and re-derives server_deadline_at, handing back time.
- Reject start-or-resume as the fix. Returning any existing in_progress attempt makes the retry safe but makes a legitimate second attempt impossible once the first is submitted, so the two cases have to be distinguished by something the client sends, not inferred from server state.
- Take a client-supplied idempotency key (an Idempotency-Key header, currently an IETF draft rather than a published RFC, or an explicit body field) scoped as (tenant_id, assignment_id, person_id, key). Store the key with a fingerprint of the semantic request and the created attempt_id in the same transaction that creates the attempt, so a unique index decides the race rather than application logic.
- Define replay precisely: same key and same fingerprint returns the original attempt_no and the original server_deadline_at, never a recomputed deadline. Same key with a different fingerprint is a 422 and creates nothing, because that is a caller bug and silently minting a new attempt hides it.
- Handle the concurrent loser explicitly. A second tab racing the first hits the unique constraint mid-flight, so the key row needs an in-flight state: the loser gets either the winner's completed result or an explicit retry-shortly response, never a 500 surfaced from a constraint violation.
- Set key retention against the client's real retry horizon plus clock skew. Expiring keys after an hour while an offline client drains its queue the next morning reintroduces the exact bug you removed.
Follow-up
- The learner submits attempt 1 and legitimately starts attempt 2. Where does the new key come from, and what stops the client reusing the old one?
- A replay must not reset the deadline. Write the test that fails if someone later moves the deadline computation into the response serializer.
- What status and body do you return while the key is still in flight, and what is the caller supposed to do with it?
Attempts marked late an hour early after the clock change
On the Monday after the autumn clock change, instructors in one zone report attempts submitted at 23:10 local flagged late against a 23:59 due time. attempt.is_late is true and frozen. assignment stores due_at_local, policy_time_zone ('America/Chicago'), and a materialised due_at_utc. The affected assignments were published in October and are due in November. Zones without a transition are unaffected. Produce an ordered diagnostic checklist, the root cause, the query that identifies every affected assignment, and the repair, including what learners and instructors see.
Approach
- Order the checklist to separate a read-time bug from stored bad data, because the repairs are completely different. First, confirm is_late is frozen on the attempt rather than recomputed in the report, which the schema already tells you. Second, compare each affected assignment's stored due_at_utc against the value the wall-clock policy implies. Third, check whether the divergence is exactly one hour and confined to one zone. Fourth, check the publish timestamp relative to the transition date. That ordering reaches the answer without reading the submit path at all.
- Name the defect precisely: due_at_utc was materialised at publish by applying the UTC offset in effect at publish time, not the offset in effect at the due instant. An assignment published in October under UTC-5 and due 23:59 on a November date under UTC-6 materialises as 04:59Z instead of 05:59Z, so everything submitted between 23:00 and 23:59 local is one hour past a deadline that is one hour early.
- Write the detection as a comparison against the correct conversion rather than as a guess about which rows are bad: SELECT assignment_id FROM assignment WHERE due_at_local IS NOT NULL AND due_at_utc IS DISTINCT FROM (due_at_local AT TIME ZONE policy_time_zone). In Postgres that expression interprets a naive timestamp in the named zone using the rules in effect at that local instant, which is exactly the computation that was skipped.
- Repair in two passes with different risk. Pass one recomputes due_at_utc for assignments still open, which changes only future decisions. Pass two revisits frozen is_late only for attempts whose assignment's due_at_utc moved and whose submitted_at falls inside the moved window, and emits a grade-change event for each so instructors see a reason rather than a number that changed on its own. Unfreezing every attempt would recompute history that was decided correctly.
- Close the hole permanently rather than fixing the arithmetic once. Keep both the wall-clock policy and the materialised instant, treat the materialised column as a cache with an invariant, and run a nightly assertion that due_at_utc equals the AT TIME ZONE expression for every open assignment. That same assertion catches the other source of drift: a tz database update that changes a zone's rules after you materialised.
- State the two edge cases the policy must handle before they become tickets. A due time inside a spring-forward gap has no instant, and a due time inside a fall-back repeated hour has two; resolution differs across databases and libraries, so reject or normalise such policy times at publish rather than depending on whichever rule the stack happens to implement.
Follow-up
- A section is moved from a school in one zone to a school in another mid-term. Which assignments recompute, and which must not?
- A government changes its zone's transition dates and the tz database ships an update. What does your nightly assertion report the next morning, and what is the approval path for the recompute?
- How would you present a corrected is_late to a learner who already saw a late penalty?
For someone fluent in a dynamic language who has shipped real work but has never had to say what the runtime is doing underneath. The week is built on measuring and deliberately breaking things, because the questions that expose this background are the ones where the interviewer asks why a second time.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Measure before reasoning
- Take a slow piece of your own code, write down in advance where you believe the time goes, then profile it and record how wrong the guess was. The cost is usually an allocation you did not notice or an accidental quadratic membership test.
- Replace one list membership test inside a loop with a set and measure at a thousand, ten thousand and a hundred thousand elements, confirming the shape of the curve rather than only that it got faster.
- Write down the three quantities you can now measure instead of assert: wall time, peak memory, and call count for the function you suspected.
Deliverable: A before-and-after profile of real code plus a written note on the size of the gap between the guess and the measurement.
Practice prompt ↗Practice prompt ↗Worked solution ↗02References, copies, and the bugs they produce
- Write the function with a mutable default argument, call it three times, and explain the accumulating result: the default is evaluated once when the function is defined, so every call shares one object.
- Build a nested structure, take a shallow copy, mutate an inner element, and show that both views changed, because a shallow copy duplicates the container and not the elements. Then fix it with a deep copy and state the cost you just accepted.
- Write two functions, one mutating its argument in place and one rebinding the local name, and predict the caller's view of each before running it. That single distinction produces most of the bugs that pass their tests.
Deliverable: Three small programs whose output you predicted correctly before running, each with a one-line statement of the rule underneath.
Practice prompt ↗Practice prompt ↗03Types, once, in a language that checks them
- Port one module you have already written, roughly a hundred lines, into a statically typed language, and record every place the compiler demanded an answer your original had left implicit: a value that can be absent, a numeric width, a case never handled.
- Write the same signature in both languages and state what the static one guarantees before the program runs and what it does not, since it will not save you from a wrong algorithm or an index out of range.
- Write the difference between an interface satisfied by declaration and one satisfied structurally, with one case each where the other approach would miss the mistake.
Deliverable: One module in two languages plus a list of the questions the type checker forced you to answer.
Practice prompt ↗Practice prompt ↗04Concurrency, starting with what actually runs at the same time
- Run the same CPU-bound function across four threads and four processes and measure both. Under the default CPython build the threaded version will not speed up, because only one thread executes bytecode at a time; the process version will. Check which build you are on first, since free-threaded builds remove that lock and change the result.
- Then run a blocking I/O workload across four threads and measure it speeding up, because the interpreter releases that lock around blocking calls, which is why treating threads as useless is wrong as a general claim.
- Build the lost update: two threads each incrementing a shared counter a hundred thousand times, and show a final value below the expected sum, because an increment is a load, an add and a store and the thread can be suspended between them. Fix it with a lock and then measure what the lock costs.
Deliverable: Three measurements, threads against processes on CPU work, threads on I/O work, and a demonstrated lost update, each with the mechanism written underneath.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Debugging as a procedure rather than an instinct
- Work one real failure as a bisection: find a revision or an input size where it is good and one where it is bad, halve repeatedly, and state the two assumptions bisection needs, that the property changes exactly once across the range and that the test is reliable.
- Minimise one failing input to the smallest version that still fails, and record how many rounds it took.
- Keep a hypothesis log for one bug in three columns, what I believe, what would disprove it, what I observed, and stop yourself the first time you are about to change two things at once.
Deliverable: One bug worked to root cause with a written hypothesis log and a minimised reproducing input.
Practice prompt ↗Practice prompt ↗06Tests that catch the bug you are about to write
- Implement an LRU cache with a capacity bound, then write the three test cases that would catch an off-by-one in eviction: insert exactly capacity items and assert nothing was evicted, insert one more and assert the least recently used key is the one gone, and read an old key just before that insert so the eviction victim changes.
- Add a property test comparing your implementation against a deliberately slow reference, an ordered list scanned linearly, over a few thousand random operation sequences, because a slow reference finds the cases you would not have thought to write.
- Write one numeric test that fails under exact equality and passes with a tolerance, and state why the tolerance has to be relative rather than absolute once the magnitudes grow.
Deliverable: An LRU implementation with three boundary tests, one property test against a slow reference, and one tolerance-based numeric test.
Practice prompt ↗07Debug something broken, out loud
- Have someone plant three defects in a two-hundred-line program, an off-by-one, a shared mutable state bug, and a wrong error-handling path, then find them while narrating, under a fixed rule: state the hypothesis before touching anything.
- Time each one and record which tool found it, reading, a printed value, a debugger, or a test, because the question asked in interviews is how you would find it rather than what it was.
- Write the sentence you will use when you do not yet know the cause, one that names the next measurement instead of offering a guess.
Deliverable: A recorded debugging session with time-to-find per defect and the method that found each.
Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Team size, service count and tickets closed say very little. Seniority shows in the decision you owned: what you chose not to build, which constraint you traded away, whose objection you had to resolve before anything could move. A large project where you executed someone else's plan is a small story.
Describe a time when you received constructive feedback on your code d…
Describe a time when you received constructive feedback on your code during a peer review. How did you handle it, and what did you learn?
Approach
- Give the blast radius: what could have broken, and what you measured.
- Pick a story where you made the decision, not one where you watched it.
- Name the disagreement and how you resolved it with evidence.
Follow-up
- How did you know your change caused the improvement?
- What did you decide not to do, and why?
Share an experience where you had to debug a critical production issue…
Share an experience where you had to debug a critical production issue under tight time constraints. How did you isolate the problem?
Approach
- Give the blast radius: what could have broken, and what you measured.
- State the situation in two sentences and spend the rest on the reasoning.
- Pick a story where you made the decision, not one where you watched it.
Follow-up
- What would you do differently if you ran that again?
- What did you decide not to do, and why?
How do you approach collaborating with non-technical stakeholders, suc…
How do you approach collaborating with non-technical stakeholders, such as product managers or curriculum designers, when defining technical requirements?
Approach
- Name the disagreement and how you resolved it with evidence.
- State the situation in two sentences and spend the rest on the reasoning.
- Pick a story where you made the decision, not one where you watched it.
Follow-up
- What did you decide not to do, and why?
- What would you do differently if you ran that again?
- 01
Describe a time when you received constructive feedback on your code during a peer review. How did you handle it, and what did you learn?
- 02
Share an experience where you had to debug a critical production issue under tight time constraints. How did you isolate the problem?
- 03
How do you approach collaborating with non-technical stakeholders, such as product managers or curriculum designers, when defining technical requirements?
Is this an official Zearn interview guide?
No. It is PracHub's own research and practice material for the Software Engineer role at Zearn. Rounds and questions reflect what candidates have reported, not a process Zearn has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult is the coding portion of the interview?
The coding questions at Zearn are practical and focus on core programming concepts, such as string manipulation, loops, and basic data structures. Candidates report that Zearn does not ask overly complex, theoretical competitive programming questions but does expect clean, working code and clear explanations of your logic.
PracHub interview research ↗What is the take-home assessment like?
The take-home assessment is a short, realistic coding exercise designed to let you solve a practical problem in your own environment. It focuses on clean code structure, error handling, and basic database or API design, and you will have the opportunity to discuss your solution with the team in the subsequent round.
PracHub interview research ↗How important is database design in this process?
Database design is highly critical. Because Zearn processes millions of student learning events daily, candidates report a strong emphasis on your ability to model data cleanly, write efficient queries, and design sound relational schemas.
PracHub interview research ↗What is the typical timeline from the initial screen to an offer?
The entire process typically takes between 2 to 4 weeks, depending on candidate availability and scheduling. Scheduling is usually the main factor in how long it takes.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-24 - 02PracHub Software Engineer practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-24 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-24