As a Software Engineer at the Law School Admission Council (LSAC), you play a foundational role in the technology ecosystem that supports the legal education journey for thousands of students. Your work directly impacts the systems used to facilitate law school admissions, assessments, and the delivery of critical educational resources. You are not just writing code; you are building the infrastructure that ensures reliability, security, and accessibility for a high-stakes user base.
The environment at LSAC demands a balance between technical precision and an understanding of the broader organizational mission. Whether you are working on core infrastructure, operational systems, or application development, you will contribute to projects that require high availability and rigorous attention to detail. Successful engineers here are those who can navigate complex technical landscapes while maintaining a focus on the stability of the platforms that support LSAC’s mission-critical services.
Recruiter Screen
reportedThe person on this call usually cannot evaluate your code and does not need to. They write a short paragraph, and that paragraph is what a hiring manager skims when deciding who to put on your loop. So the test is not whether your work was hard, it is whether a non-engineer can repeat it correctly. Name systems by what they did rather than by their internal codename, give each project a shape (what was breaking, what you changed, what happened after), and keep the whole walkthrough near ninety seconds. Depth that cannot survive a paraphrase reads as vagueness.
What to demonstrate
- Whether a non-engineer can restate your projects without distorting them, since their paraphrase is what travels to the hiring manager, not your sentences
- Whether each project has a shape rather than a stack list: the failure or constraint, the change you made, the result and how it was measured
- Whether you can say what was yours inside a team project without either inflating it or disappearing into the plural
How to prepare
- Rewrite each headline project as two sentences with no internal system names and no acronyms outside your company, then say them to someone outside engineering and have them repeat them back. Fix whatever came back wrong
- Attach one measured number to each project: the baseline, the change, and the window it was measured over. Where nothing was ever measured, say that plainly rather than reaching for a plausible percentage
- Time the background walkthrough against a clock. If it runs past two minutes, compress the earliest role to a single clause and spend the recovered time on the most recent one
Technical Evaluations
reportedThe same problem is scored by two different mechanisms depending on the format, and preparing for one does not cover the other. With a person watching, partial progress is visible and a hint is a correction you can absorb; silence is the expensive failure, because nobody can read a half-written function. With an automated grader there is no partial credit for what you were about to do, nobody to ask, and the worked examples in the prompt are the entire specification. Read them as a contract, down to whether an empty result should be an empty list or no output at all.
What to demonstrate
- In a live session, whether your commentary tracks what your hands are doing, and whether a hint redirects you or gets defended against
- In an automated one, whether you cover the cases the examples do not show, since the hidden cases are where the score moves
- Whether you manage the clock on purpose: abandoning an approach that is not converging while there is still time to write something simpler that finishes
How to prepare
- Have someone hand you a problem and feed you one deliberately wrong hint. Practise testing it against a concrete case instead of accepting or rejecting it on authority.
- Do one timed run a week in a plain browser editor with autocomplete, linting and your own snippets switched off, which is closer to what these environments give you
- For the automated format, write the harness before the solution: a main that feeds the worked examples plus an empty and a single-element case and prints expected against actual, so a wrong submission is caught by you first
Meet Stakeholders
reportedAn unlabelled round is first an information problem, and the cheapest information is free. Whoever schedules it can usually tell you how long it runs, who will be in the room and what they work on, whether you will be writing code and in what environment, and whether anything is being sent beforehand. Ask in writing so the answer is on record, then prepare for the two or three formats those answers still leave open instead of betting on one. What separates a strong candidate is not guessing right; it is having an opening that works whichever one it turns out to be.
What to demonstrate
- Whether you can start work from an ambiguous brief, since tolerating a vague scope without stalling is the same thing the job asks for
- Whether the questions you asked beforehand were ones that change your preparation, such as duration, medium and who is joining, rather than ones whose answers you could not have acted on
- Whether you adapt when the round turns out to be something other than what you were told, instead of spending the first ten minutes visibly recalibrating
How to prepare
- Send one short scheduling message asking four things: how long, who is joining and what they work on, whether you will be writing code and where, and whether to prepare anything in advance. Treat a vague reply as real information, since it means the round is loosely structured and you will be shaping it yourself.
- Write one opening that works in any of the formats still open: restate in your own words what you have been asked to do, then ask which of two directions is more useful to them. Say it aloud until it stops sounding recited.
- Set up for the two most likely formats before the call starts, with a blank editor in the language you would choose and a shared document you can type into, so a format surprise costs you nothing in the first minutes
Behavioral Assessments
reportedThis round is deciding whether a change you make without supervision can be allowed to reach production. It is scored on what you knew at the moment you decided, not on how it turned out, so a story that opens with the result and works backwards reads as luck retold as judgement. Say what the options were, what you did not know, what you did to shrink the unknown before committing, and what you accepted as the worst plausible case. The detail that separates answers is a bound: how many users, how much data, and for how long, if you had been wrong.
What to demonstrate
- Whether the reasoning you give was available at the time you decided rather than after the result came in, since a story whose deciding evidence arrived later describes an outcome and not a judgement
- Whether you can put units on the exposure (users, rows, minutes of degraded service) and whether the containment you chose actually bounded it: a canary bounds the request path it fronts, while a background job writing to a shared table reaches every user regardless of which version served their requests
- Whether the reversal path existed before you shipped or was improvised during the incident, and whether it restores state or only stops further damage
How to prepare
- For your three largest changes, write down the one thing you would have had to be wrong about for it to fail, and what your best estimate of it was on the day you shipped. If you never held an estimate, that is the gap the follow-up questions will find
- Write the undo procedure for one of those changes as it existed at the time, then mark which steps restore data and which only stop new damage. Turning a flag off or reverting a deploy ends the new writes; rows already written come back only from a copy you kept, and a dropped column comes back empty unless something outside the schema holds the values
- Rehearse one story from the decision point forward and stop before the outcome, then have someone ask what you would do next. If the story only works with the ending attached, it is an anecdote rather than a decision you can defend
PracHub editorial advice for the preparation topics above.
Editing a content item in place and thereby rewriting the meaning of every historical score.
Authors expect to fix a typo, change a distractor, or adjust a point value, and nothing warns them that thousands of stored results reference the record they are editing. If attempts and analytics join to the mutable item rather than to a version, a regrade or a rebuilt report scores past answers against an item that did not exist when they were given. Versioning has to be the default write path, because a convention that authors must remember will not hold.
Autoscaling on average utilisation against a bell-shaped step function.
Scale-up latency (scheduling, image pull, process and JIT warm-up, connection pool and cache fill) is typically tens of seconds to minutes, while the demand step is under a minute and repeats on a published schedule. By the time the metric crosses a threshold, the period has started and the queue has already built. The workable answer is a schedule-driven pre-warm derived from the tenant calendar plus a load shedder that protects writes, not a more aggressive threshold.
Comparing floating-point values for equality, or holding money in them
Binary floating point cannot represent 0.1 exactly, so repeated addition drifts and an equality check fails on values that are mathematically equal. Store currency as integer minor units or a decimal type, and compare floats against a tolerance you chose for a stated reason.
Assuming fixed-width integer arithmetic cannot overflow
In languages with fixed-width integers, including C, C++, Java, Go and Rust, computing a midpoint as (lo + hi) / 2 overflows once the sum passes the type's maximum, so write lo + (hi - lo) / 2 instead. Say which language you are in: arbitrary-precision integers, as in Python or Ruby, remove this specific hazard and none of the others.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Walk through the specific design choices made in your submitted work s…
Walk through the specific design choices made in your submitted work sample.
Approach
- Choose the data structure from the access pattern, not from familiarity.
- State the target complexity and say which constraint rules the naive version out.
- Walk one small example through your approach before writing the whole thing.
Follow-up
- What is the worst case, and how likely is it on real data?
- How does this change if the input no longer fits in memory?
Find peak concurrent attempts to size a pre-warm
You have one district's attempt rows for a single school day: started_at, and an end instant taken as submitted_at when present and server_deadline_at otherwise. There are up to 50,000,000 rows, all inside one local calendar day, and the district's org.time_zone is known. Compute the maximum number of simultaneously in-progress attempts and the second at which it occurs, so capacity can be pre-warmed against the bell schedule. A sort-based sweep is acceptable but is not the best answer here. State your bounds and what happens on a day with a daylight-saving transition.
Approach
- Note the structural fact that decides the algorithm: the output domain is one day at one-second resolution, so there are on the order of 86,400 distinct answer positions while there are 50,000,000 inputs. That inverts the usual sweep, because the coordinate space is far smaller than the data.
- Allocate a difference array over seconds from local midnight, with one sentinel slot past the end so the decrement for an attempt ending in the final bucket has somewhere to land. For each attempt add 1 at the bucket containing
started_atand subtract 1 at the bucket after the one containing the end instant, then prefix-sum once and take the argmax. That is O(N + T) time and O(T) space, which is 86,401 int32 values, about 346 KB — small enough to stay in L2 and to run per section if you want. - Say why this beats the sweep-line answer here. Sorting 2N endpoints is O(N log N) with 100,000,000 entries to materialise; the difference array touches each input twice and never sorts. The sweep only wins when the coordinate space is large or unbounded, which is exactly the condition this problem does not have.
- State the bias the bucketing introduces rather than hiding it. Every attempt is counted in full in every second it touches at all, so the bucketed maximum is greater than or equal to the true instantaneous maximum, never less, which is the safe direction for pre-warming. The overcount is not confined to sub-second attempts: two hour-long attempts that merely share a boundary second without ever overlapping, one ending at 10:00:05.2 and the next starting at 10:00:05.8, already put 2 in that bucket against a true instantaneous maximum of 1. The bound that does hold is per bucket. The attempts touching second s split into those that span it entirely, which by definition coexist at every instant of s, and those with an endpoint inside it, so the overcount at s is at most the number of attempts that begin or end within that second. Equality with an exact sweep is guaranteed when no attempt begins or ends inside the peak second — which holds, for instance, on a fixture whose every endpoint lands exactly on a second boundary.
- Handle the calendar honestly, and fix the indexing convention before sizing anything. Index buckets by elapsed seconds from the UTC instant of local midnight and size the array from the real UTC span between successive local midnights: in a zone that shifts by an hour that span is 82,800 seconds on the spring-forward day and 90,000 on the fall-back day, not 86,400. A hardcoded 86,400 therefore overruns on the fall-back day, whose last hour indexes up to 89,999, and merely leaves 3,600 dead slots at the tail on the spring-forward day, which is harmless. The opposite convention fails the opposite way: bucketing by local wall-clock seconds-since-midnight always stays in range, but maps the fall-back day's repeated hour onto buckets that already hold the first pass through it, folding two real hours of load into one and understating the peak.
- For per-section peaks, do not allocate T buckets per section — 10,000 sections is 3.5 GB. Partition the input by
section_idand reuse one array, or keep only sections whose total attempt count clears a threshold, since the rest cannot produce a peak worth pre-warming for.
Worked solution 25 min
- Compute the UTC instants of this local day's start and the next day's start, subtract to get the true bucket count, and allocate that many plus one sentinel slot.
- Fill the difference array in one pass, using a half-open convention and writing the decrement at end_bucket + 1 so an attempt is counted in the second it ends and the final bucket's decrement lands in the sentinel.
- Prefix-sum in place and track the running maximum and its index, converting that index back to a local wall-clock time for the report.
- Check against a brute-force O(N * T) reference — for each bucket, the number of attempts touching it — on a 500-row fixture spanning two bell periods, including one attempt that starts and ends inside a single second and one pair that abuts inside a second without overlapping.
- Re-run the fixture shifted onto a spring-forward date and onto a fall-back date, confirming the allocation follows the real UTC span each time — 82,800 and 90,000 buckets — and that the fall-back day's final hour, at indices 86,400 and above, lands inside the array rather than past its end.
Follow-up
- Now compute the peak across three districts in different time zones on one shared cluster. What is the coordinate space and does your answer survive?
- Some attempts have neither
submitted_atnorserver_deadline_atbecause they were abandoned. What end do you pick, and how does the choice bias the number you hand to capacity planning? - You need the top ten peak minutes rather than the single peak. What changes, and what does not?
Stream a quoted roster CSV without loading the file
You are handed a 2 GB roster export as a byte stream. It follows RFC 4180: any field may be quoted, a quoted field may contain commas and CRLF, a literal quote inside a quoted field is written as two quotes, and the file may open with a UTF-8 byte order mark. There are up to 10,000,000 records and a single field may reach 64 KB. Write a reader that yields one record at a time without buffering the whole file and that fails loudly on a malformed file rather than emitting a short record. State your bounds.
Approach
- Refuse the shape that fails first: splitting the stream on newlines and then splitting each line on commas. A quoted field may contain a CRLF, so line splitting cuts records in half, and the damage is silent because both halves parse into plausible short records.
- Write a byte-level state machine with four states — field start, unquoted field, quoted field, and quote-seen-inside-quoted. In the last state a second quote emits one literal quote and returns to quoted; a comma or newline ends the field; anything else is a malformed file. That table is the whole parser and it is the part to get right on paper before typing.
- Strip the BOM only at offset zero, comparing the first three bytes against EF BB BF. Left in place it becomes part of the first header name, so the header lookup for that column misses and the column reads as absent for every record in the file.
- Cap field length at the stated 64 KB and record length at a sane multiple of it. Without a cap, one unbalanced quote makes the parser treat the remaining 2 GB as a single field and the process dies of memory exhaustion rather than telling you the file is broken at byte 12,004.
- Treat end of stream inside a quoted field as a hard error, not an implicit close. A truncated upload is the most common malformed input here and it is byte-for-byte indistinguishable from a complete file if you close the field silently — the missing rows then present downstream as a mass withdrawal.
- Complexity: O(n) time in bytes with one pass and no backtracking, and O(longest field + longest record) space, which is bounded by the caps rather than by the file. Read in fixed blocks and keep the partial field across block boundaries instead of reading line-wise.
Follow-up
- The file arrives gzipped and you must report progress as a percentage. What can you actually report, and what does that do to your memory bound?
- Two exporters disagree on line endings and one emits a bare LF inside a quoted field. Does your state machine care, and should it?
- How do you distinguish a truncated file from a complete one when the byte stream itself gives you no signal?
Stop the gradebook rollup from counting attempts as learners
For one section, produce per published assignment: learners assigned, learners with a graded attempt, mean best score, and late count. Tables: enrollment as above (a person may hold two roles, and may have withdrawn and re-enrolled), assignment(assignment_id, tenant_id, section_id, status, max_attempts), attempt(attempt_id, tenant_id, assignment_id, person_id, attempt_no, state, is_late), score(attempt_id PRIMARY KEY, tenant_id, score_given, score_maximum, revision). Grading policy is best attempt wins. Write one query. State the grain of every branch before you join it, and name where a naive join multiplies rows.
Approach
- Say the grains out loud first: enrollment is one row per (section, person, role) and repeats per person; attempt is one row per (assignment, person, attempt_no); score is one per attempt. Join all three flat and the row count is enrollment rows multiplied by attempts, so every aggregate above it is weighted twice over.
- Collapse attempts to the policy grain before anything else: SELECT DISTINCT ON (a.assignment_id, at.person_id) a.assignment_id, at.person_id, at.is_late, s.score_given FROM assignment a JOIN attempt at ON at.tenant_id = a.tenant_id AND at.assignment_id = a.assignment_id JOIN score s ON s.tenant_id = at.tenant_id AND s.attempt_id = at.attempt_id WHERE a.tenant_id = $1 AND a.section_id = $2 AND at.state = 'graded' ORDER BY a.assignment_id, at.person_id, s.score_given DESC, at.attempt_no DESC. That ORDER BY is where 'best attempt wins' is written down, and it is the only place it should appear.
- Collapse the roster the same way: SELECT DISTINCT person_id FROM enrollment WHERE tenant_id = $1 AND section_id = $2 AND role = 'learner' AND deleted_at IS NULL AND begin_date <= $3 AND (end_date IS NULL OR end_date >= $3). Now re-enrolment and dual roles cannot multiply anything downstream.
- Join the two collapsed sets, LEFT from assignment so an assignment nobody submitted still returns a row of zeros, and aggregate with FILTER instead of one correlated subquery per column: count(b.person_id) AS learners_graded, avg(b.score_given) AS mean_score, count(*) FILTER (WHERE b.is_late) AS late_count, with learners_assigned taken from the roster cardinality.
- Reject COUNT(DISTINCT person_id) as the fix. It repairs the counts and leaves AVG computed over the multiplied rows, so the mean quietly weights learners who attempted more often, and the number stays plausible for as long as anyone looks at it.
- Keep the cost where it belongs: with an index on attempt (tenant_id, assignment_id, person_id, attempt_no) the DISTINCT ON is an ordered index scan over a bounded set — order 10^2 assignments times order 10^1 to 10^2 learners — rather than a sort of the table. Carry tenant_id into every join condition, not just the outer WHERE.
Worked solution 30 min
- Seed one section with 6 learners, one of them re-enrolled and one holding two roles, and give one learner three graded attempts on the same assignment.
- Run the flat three-way join first and record the inflated counts and the shifted mean as the baseline error.
- Add the DISTINCT ON collapse and the roster collapse, then re-run and diff against the baseline.
- Compute each learner's best score in a separate standalone query and compare the per-assignment mean by hand.
Follow-up
- A regrade changes score_given after the fact. Does the best attempt change, and what would you order by if the policy were latest attempt instead?
- The instructor opens twelve sections at once. Does this query shape hold, or does the DISTINCT ON become the problem?
- Half the attempts are state='voided' after an integrity review. Where in the query does that belong, and what happens to mean_score?
Replace per-retry outbox keys with derived ones on a live table
score_publication holds 400 million rows and takes continuous writes: (publication_id, tenant_id, attempt_id, target, idempotency_key TEXT NOT NULL, score_given, score_timestamp, state, attempt_count, next_attempt_at). The key was minted as a fresh UUID per retry, so the same grade has several rows and duplicate deliveries have already happened. Migrate to a key derived from (attempt_id, score revision) with UNIQUE (tenant_id, target, idempotency_key), online. Give the ordered steps with the lock each takes, how you handle the existing duplicates, and what you do with rows currently in_flight.
Approach
- Add a plain nullable column, derived_key TEXT, with no default: a catalogue-only ALTER held for microseconds. Set lock_timeout to a second or two and retry, because the hazard is the queue rather than the statement — the ALTER waits behind one long reader and every query arriving afterwards stacks behind its pending ACCESS EXCLUSIVE. Resist the elegant version: adding a STORED generated column rewrites all 400 million rows under that same lock on current major versions.
- Deploy the dual write before the backfill so every new publication row populates derived_key from attempt_id and the score revision, while the relay still reads the old column. The backfill then chases a closed id range instead of a moving target, and you can bound it.
- Backfill in batches by publication_id range, on the order of 20,000 rows per statement, committing between batches, recording the high-water mark in its own table so a killed run resumes rather than restarts, and filtering WHERE derived_key IS NULL so a replayed batch is a no-op. Pace on replica lag and dead tuples, not CPU: each UPDATE writes a new row version and WAL volume tracks rows touched.
- Resolve the duplicates as a data decision, not a DELETE. Group by (tenant_id, target, derived_key) having count(*) > 1; within each group keep the row in state='succeeded' if one exists, since that is the delivery whose outcome is known, otherwise the newest 'pending'. Move the rest to state='dead' with last_error_code recording the migration, and keep the rows — support has to be able to explain a grade that posted twice, and a deleted row explains nothing.
- Leave in_flight rows alone and drain them. A row in_flight has an outstanding request whose outcome is unknown, so collapsing it either loses a delivery record or races the relay marking a row you just rewrote. Pause the claim loop or wait out the lease, assert zero in_flight for the affected keys, then finish the group.
- Build the constraint without a blocking scan and expect the first attempt to fail: CREATE UNIQUE INDEX CONCURRENTLY, which aborts if any duplicate remains and leaves an index with indisvalid = false that must be DROP INDEX CONCURRENTLY'd rather than reused. Once it builds clean, ALTER TABLE ... ADD CONSTRAINT ... UNIQUE USING INDEX adopts it under a brief ACCESS EXCLUSIVE without rebuilding — which is a second reason derived_key is a real column, since USING INDEX accepts neither an expression index nor a partial one.
Follow-up
- score_timestamp must strictly increase per line item and learner. Does this migration change what the receiver sees, and what would a stale retry do mid-migration?
- A tenant's external endpoint has been failing for a week and has 40,000 rows queued. What does your duplicate collapse do to that backlog?
- How do you prove afterwards that no grade was delivered twice as a result of the migration itself?
How do you approach documenting your technical designs?
How do you approach documenting your technical designs?
Approach
- Clarify what is being asked and what a complete answer contains.
- Work from the requirement backwards to the design.
- Say what you would check first and why it is the highest-information step.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
Explain your approach to debugging complex system issues.
Explain your approach to debugging complex system issues.
Approach
- Say what you would check first and why it is the highest-information step.
- Clarify what is being asked and what a complete answer contains.
- Work from the requirement backwards to the design.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
Describe how you ensure high availability in your previous projects.
Describe how you ensure high availability in your previous projects.
Approach
- State your assumptions explicitly before working the problem.
- Work from the requirement backwards to the design.
- Clarify what is being asked and what a complete answer contains.
Follow-up
- What assumption would you test first?
- How would you know your answer was wrong?
Fix an attempt-start endpoint whose timeout retry burns an attempt
POST /assignments/{id}/attempts allocates the next attempt_no for the caller, copies content_version_id onto the row, and computes server_deadline_at once from time_limit_seconds. The caller is the learner's browser on school wifi; it retries once when a request times out. With max_attempts = 2, a timed-out create followed by that retry consumes both attempts and hands the learner a fresh deadline. Redesign the endpoint so a retry is safe, a deliberate second attempt is still possible, and two tabs cannot each create one. Give the request shape, the storage, and the response to a replay.
Approach
- Say why the naive shape fails rather than patching it: POST-creates-a-child is non-idempotent by construction, and here the duplicate is worse than the original because it spends a countable resource (an attempt_no against max_attempts) and re-derives server_deadline_at, handing back time.
- Reject start-or-resume as the fix. Returning any existing in_progress attempt makes the retry safe but makes a legitimate second attempt impossible once the first is submitted, so the two cases have to be distinguished by something the client sends, not inferred from server state.
- Take a client-supplied idempotency key (an Idempotency-Key header, currently an IETF draft rather than a published RFC, or an explicit body field) scoped as (tenant_id, assignment_id, person_id, key). Store the key with a fingerprint of the semantic request and the created attempt_id in the same transaction that creates the attempt, so a unique index decides the race rather than application logic.
- Define replay precisely: same key and same fingerprint returns the original attempt_no and the original server_deadline_at, never a recomputed deadline. Same key with a different fingerprint is a 422 and creates nothing, because that is a caller bug and silently minting a new attempt hides it.
- Handle the concurrent loser explicitly. A second tab racing the first hits the unique constraint mid-flight, so the key row needs an in-flight state: the loser gets either the winner's completed result or an explicit retry-shortly response, never a 500 surfaced from a constraint violation.
- Set key retention against the client's real retry horizon plus clock skew. Expiring keys after an hour while an offline client drains its queue the next morning reintroduces the exact bug you removed.
Worked solution 30 min
- Write the DDL for attempt_idempotency(tenant_id, assignment_id, person_id, key, request_fingerprint, attempt_id, state, created_at) with a unique index on the first four columns.
- Implement create as one transaction that inserts the key row and the attempt row together; on unique violation, read the existing row and branch on its state and fingerprint.
- Run the same request twice sequentially and assert one attempt row, with identical attempt_no and identical server_deadline_at in both responses.
- Run two requests concurrently and assert one attempt row, with the loser receiving the winner's result or an explicit in-flight status rather than a 500.
- Send the same key with a changed assignment_id and assert a 422 and zero new rows.
- Submit attempt 1, send a fresh key to create attempt 2, then send a third key and assert max_attempts refuses it.
Follow-up
- The learner submits attempt 1 and legitimately starts attempt 2. Where does the new key come from, and what stops the client reusing the old one?
- A replay must not reset the deadline. Write the test that fails if someone later moves the deadline computation into the response serializer.
- What status and body do you return while the key is still in flight, and what is the caller supposed to do with it?
Attempts marked late an hour early after the clock change
On the Monday after the autumn clock change, instructors in one zone report attempts submitted at 23:10 local flagged late against a 23:59 due time. attempt.is_late is true and frozen. assignment stores due_at_local, policy_time_zone ('America/Chicago'), and a materialised due_at_utc. The affected assignments were published in October and are due in November. Zones without a transition are unaffected. Produce an ordered diagnostic checklist, the root cause, the query that identifies every affected assignment, and the repair, including what learners and instructors see.
Approach
- Order the checklist to separate a read-time bug from stored bad data, because the repairs are completely different. First, confirm is_late is frozen on the attempt rather than recomputed in the report, which the schema already tells you. Second, compare each affected assignment's stored due_at_utc against the value the wall-clock policy implies. Third, check whether the divergence is exactly one hour and confined to one zone. Fourth, check the publish timestamp relative to the transition date. That ordering reaches the answer without reading the submit path at all.
- Name the defect precisely: due_at_utc was materialised at publish by applying the UTC offset in effect at publish time, not the offset in effect at the due instant. An assignment published in October under UTC-5 and due 23:59 on a November date under UTC-6 materialises as 04:59Z instead of 05:59Z, so everything submitted between 23:00 and 23:59 local is one hour past a deadline that is one hour early.
- Write the detection as a comparison against the correct conversion rather than as a guess about which rows are bad: SELECT assignment_id FROM assignment WHERE due_at_local IS NOT NULL AND due_at_utc IS DISTINCT FROM (due_at_local AT TIME ZONE policy_time_zone). In Postgres that expression interprets a naive timestamp in the named zone using the rules in effect at that local instant, which is exactly the computation that was skipped.
- Repair in two passes with different risk. Pass one recomputes due_at_utc for assignments still open, which changes only future decisions. Pass two revisits frozen is_late only for attempts whose assignment's due_at_utc moved and whose submitted_at falls inside the moved window, and emits a grade-change event for each so instructors see a reason rather than a number that changed on its own. Unfreezing every attempt would recompute history that was decided correctly.
- Close the hole permanently rather than fixing the arithmetic once. Keep both the wall-clock policy and the materialised instant, treat the materialised column as a cache with an invariant, and run a nightly assertion that due_at_utc equals the AT TIME ZONE expression for every open assignment. That same assertion catches the other source of drift: a tz database update that changes a zone's rules after you materialised.
- State the two edge cases the policy must handle before they become tickets. A due time inside a spring-forward gap has no instant, and a due time inside a fall-back repeated hour has two; resolution differs across databases and libraries, so reject or normalise such policy times at publish rather than depending on whichever rule the stack happens to implement.
Follow-up
- A section is moved from a school in one zone to a school in another mid-term. Which assignments recompute, and which must not?
- A government changes its zone's transition dates and the tz database ships an update. What does your nightly assertion report the next morning, and what is the approval path for the recompute?
- How would you present a corrected is_late to a learner who already saw a late penalty?
For someone fluent in a dynamic language who has shipped real work but has never had to say what the runtime is doing underneath. The week is built on measuring and deliberately breaking things, because the questions that expose this background are the ones where the interviewer asks why a second time.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Measure before reasoning
- Take a slow piece of your own code, write down in advance where you believe the time goes, then profile it and record how wrong the guess was. The cost is usually an allocation you did not notice or an accidental quadratic membership test.
- Replace one list membership test inside a loop with a set and measure at a thousand, ten thousand and a hundred thousand elements, confirming the shape of the curve rather than only that it got faster.
- Write down the three quantities you can now measure instead of assert: wall time, peak memory, and call count for the function you suspected.
Deliverable: A before-and-after profile of real code plus a written note on the size of the gap between the guess and the measurement.
Practice prompt ↗Practice prompt ↗Worked solution ↗02References, copies, and the bugs they produce
- Write the function with a mutable default argument, call it three times, and explain the accumulating result: the default is evaluated once when the function is defined, so every call shares one object.
- Build a nested structure, take a shallow copy, mutate an inner element, and show that both views changed, because a shallow copy duplicates the container and not the elements. Then fix it with a deep copy and state the cost you just accepted.
- Write two functions, one mutating its argument in place and one rebinding the local name, and predict the caller's view of each before running it. That single distinction produces most of the bugs that pass their tests.
Deliverable: Three small programs whose output you predicted correctly before running, each with a one-line statement of the rule underneath.
Practice prompt ↗Practice prompt ↗03Types, once, in a language that checks them
- Port one module you have already written, roughly a hundred lines, into a statically typed language, and record every place the compiler demanded an answer your original had left implicit: a value that can be absent, a numeric width, a case never handled.
- Write the same signature in both languages and state what the static one guarantees before the program runs and what it does not, since it will not save you from a wrong algorithm or an index out of range.
- Write the difference between an interface satisfied by declaration and one satisfied structurally, with one case each where the other approach would miss the mistake.
Deliverable: One module in two languages plus a list of the questions the type checker forced you to answer.
Practice prompt ↗Practice prompt ↗04Concurrency, starting with what actually runs at the same time
- Run the same CPU-bound function across four threads and four processes and measure both. Under the default CPython build the threaded version will not speed up, because only one thread executes bytecode at a time; the process version will. Check which build you are on first, since free-threaded builds remove that lock and change the result.
- Then run a blocking I/O workload across four threads and measure it speeding up, because the interpreter releases that lock around blocking calls, which is why treating threads as useless is wrong as a general claim.
- Build the lost update: two threads each incrementing a shared counter a hundred thousand times, and show a final value below the expected sum, because an increment is a load, an add and a store and the thread can be suspended between them. Fix it with a lock and then measure what the lock costs.
Deliverable: Three measurements, threads against processes on CPU work, threads on I/O work, and a demonstrated lost update, each with the mechanism written underneath.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Debugging as a procedure rather than an instinct
- Work one real failure as a bisection: find a revision or an input size where it is good and one where it is bad, halve repeatedly, and state the two assumptions bisection needs, that the property changes exactly once across the range and that the test is reliable.
- Minimise one failing input to the smallest version that still fails, and record how many rounds it took.
- Keep a hypothesis log for one bug in three columns, what I believe, what would disprove it, what I observed, and stop yourself the first time you are about to change two things at once.
Deliverable: One bug worked to root cause with a written hypothesis log and a minimised reproducing input.
Practice prompt ↗Practice prompt ↗06Tests that catch the bug you are about to write
- Implement an LRU cache with a capacity bound, then write the three test cases that would catch an off-by-one in eviction: insert exactly capacity items and assert nothing was evicted, insert one more and assert the least recently used key is the one gone, and read an old key just before that insert so the eviction victim changes.
- Add a property test comparing your implementation against a deliberately slow reference, an ordered list scanned linearly, over a few thousand random operation sequences, because a slow reference finds the cases you would not have thought to write.
- Write one numeric test that fails under exact equality and passes with a tolerance, and state why the tolerance has to be relative rather than absolute once the magnitudes grow.
Deliverable: An LRU implementation with three boundary tests, one property test against a slow reference, and one tolerance-based numeric test.
Practice prompt ↗Practice prompt ↗07Debug something broken, out loud
- Have someone plant three defects in a two-hundred-line program, an off-by-one, a shared mutable state bug, and a wrong error-handling path, then find them while narrating, under a fixed rule: state the hypothesis before touching anything.
- Time each one and record which tool found it, reading, a printed value, a debugger, or a test, because the question asked in interviews is how you would find it rather than what it was.
- Write the sentence you will use when you do not yet know the cause, one that names the next measurement instead of offering a guess.
Deliverable: A recorded debugging session with time-to-find per defect and the method that found each.
Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Team size, service count and tickets closed say very little. Seniority shows in the decision you owned: what you chose not to build, which constraint you traded away, whose objection you had to resolve before anything could move. A large project where you executed someone else's plan is a small story.
Describe a time you had to troubleshoot an issue within a codebase you…
Describe a time you had to troubleshoot an issue within a codebase you did not originally write.
Approach
- Give the blast radius: what could have broken, and what you measured.
- Close with what you would do differently, concretely.
- Name the disagreement and how you resolved it with evidence.
Follow-up
- How did you know your change caused the improvement?
- What did you decide not to do, and why?
Estimate a roster re-key you have never run
A district changes its roster identifier scheme, so every external_sourced_id arrives in a new format and an ordinary snapshot diff reads as roughly 100,000 persons deleted and 100,000 created, with about a million enrollment rows attached. You have never run a re-key of this kind. The apply must fit an overnight window and must not create duplicate persons. Produce an estimate with a range, the first measurement you would take, and the thing most likely to make you wrong.
Approach
- Refuse to estimate the whole thing and decompose into terms with different epistemic status: matching, which is unknown; applying, which is measurable; verifying, which is measurable; and the human queue that clears whatever does not match, which is usually the largest term and the one people leave out of the number entirely.
- Measure before estimating, and measure the right thing. Run the matcher over a one percent sample and report both the match rate and precision on a hand-checked subset, because the schedule is set by the match rate rather than by throughput. Ninety-eight percent matching on a stable attribute pair leaves a manual queue you can size in people-hours; eighty percent is a different project with a different plan.
- Size the apply from a timed batch rather than a guess: run ten thousand updates against a copy carrying the real indexes, including the partial unique index on live enrollments and the foreign keys, since those are what you are actually paying for. Extrapolate to a million and compare against the window. If it does not fit, the answer is a resumable multi-night apply keyed on snapshot_digest, not a faster loop.
- Give a range with named drivers instead of a padded single number. State the low end assuming the high match rate and no schema change, the high end assuming a manual queue of the size the sample implies at a stated minutes per case, and say explicitly which measurement collapses the range. An estimate whose endpoints you cannot attribute is just anxiety expressed in days.
- Name the most likely way you are wrong: that this is not a pure re-key. Some rows are genuinely new and some genuinely gone, so the run is simultaneously a re-key and a real diff, and any plan assuming a bijection silently mishandles exactly those rows. Bound the risk rather than the estimate: the destructive gate holds the run regardless and the apply is restartable, so being wrong costs another night instead of a deprovisioned district.
Follow-up
- Two persons under the old scheme map to one under the new. What do you do, and what happens to the attempts already attached to each?
- The district asks for a date on day one, before you have the match rate. What do you commit to?
- How do you verify a finished re-key beyond comparing row counts?
Reverse a deletion gate that operators stopped reading
You shipped a gate holding any roster run that proposes to soft-delete more than five percent of a tenant's active enrollments, recorded on roster_sync_run.gate_state as held_for_review. In the first two weeks of a term, when withdrawals are genuinely large, it fired on roughly forty percent of runs and operators began approving without opening the diff. Describe reversing a decision of this shape: the evidence that changed your mind, what you replaced it with, and how you avoided reversing into the failure the gate existed to prevent.
Approach
- Lead with the measurement that reversed you, not with the discomfort. Hold rate, and more importantly the distribution of time between a hold and its approval. A control approved in a median of a few seconds is not a control, it is a log line with a button, and that number is the argument. Say plainly that the gate was still technically doing its job and had already stopped being a defence.
- Separate the two situations the single threshold cannot tell apart. The gate exists to catch a truncated file or a changed identifier scheme; it is not meant to catch a real withdrawal wave. Both present as many rows missing, so proportion alone is the wrong discriminator. The signals that do separate them are rows_in against the tenant's trailing snapshot sizes, whether missing persons have same-shaped replacements present, and whether deletions cluster on whole orgs or spread thinly across sections.
- Replace one threshold with a small set of rules that keeps the fail-safe direction. Hold unconditionally when rows_in collapses against the trailing median or falls below an absolute floor, hold when a large share of missing sourcedIds have plausible replacements in a new format, auto-apply when the run looks like the tenant's own history. Keep the hold non-destructive: deletes are retained, the run stays restartable from snapshot_digest, and nothing is discarded by a rejection.
- Guard the reversal with a measurement rather than confidence. Run the old rule in shadow, log disagreements, and track hold precision, meaning of the runs held how many were genuinely bad, alongside hold rate. Loosening a safety control is only defensible if you can show it was loosened where it was wrong, and a shadow comparison is what turns that from an assertion into a number.
- Say what would reverse you again and when you look. Term-start weeks are the regime where this rule earns or loses its keep, so the review cadence is per term, not quarterly. Then name the original mistake without softening it: the threshold was designed and validated against the forty-six quiet weeks and met its real workload for the first time in production.
Follow-up
- A district legitimately closes a school mid-year and every enrollment under one org ends. Does your rule hold it, and is that the behaviour you want?
- What does an operator see that lets them decide in under a minute, and what would make them read it rather than approve it?
- A run was approved that should not have been. What does unwinding it touch beyond the enrollment rows?
- 01
Describe a time you had to troubleshoot an issue within a codebase you did not originally write.
- 02
A district changes its roster identifier scheme, so every external_sourced_id arrives in a new format and an ordinary snapshot diff reads as roughly 100,000 persons deleted and 100,000 created, with about a million enrollment rows attached. You have never run a re-key of this kind. The apply must fit an overnight window and must not create duplicate persons. Produce an estimate with a range, the first measurement you would take, and the thing most likely to make you wrong.
- 03
You shipped a gate holding any roster run that proposes to soft-delete more than five percent of a tenant's active enrollments, recorded on roster_sync_run.gate_state as held_for_review. In the first two weeks of a term, when withdrawals are genuinely large, it fired on roughly forty percent of runs and operators began approving without opening the diff. Describe reversing a decision of this shape: the evidence that changed your mind, what you replaced it with, and how you avoided reversing into the failure the gate existed to prevent.
Is this an official Law School Admission Council interview guide?
No. It is PracHub's own research and practice material for the Software Engineer role at Law School Admission Council. Rounds and questions reflect what candidates have reported, not a process Law School Admission Council has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How long does the interview process typically take?
The process can span several weeks, involving multiple rounds of interviews and potential technical assessments. It is best to remain patient and maintain consistent communication with your recruiter.
PracHub interview research ↗What makes a candidate stand out?
Successful candidates demonstrate a deep understanding of their own past projects, clear communication skills, and a genuine interest in the stability and reliability of the systems they build.
PracHub interview research ↗Is the technical work environment collaborative?
Yes, engineering at LSAC involves working closely with peers and management to solve complex problems, making effective communication and teamwork essential traits.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Software Engineer practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22