As a Software Engineer at Western Governors University, you play a vital role in building and scaling the digital learning platforms that empower thousands of non-traditional students nationwide. This position directly impacts user experiences across complex web applications, learning systems, cloud platforms, and automated assessment tools. By driving high-availability software development, you help bridge the gap between innovative educational technology and student success.
The work encompasses modernizing legacy infrastructure, designing robust RESTful web services, and deploying scalable applications within cloud ecosystems like AWS and Salesforce environments. Teams operate in collaborative, mission-driven spaces where technical excellence meets a commitment to individualized instruction and accessible education. You will tackle multifaceted engineering challenges ranging from pipeline automation to frontend performance optimization.
Expect a fast-paced yet supportive environment where continuous learning and adaptability are highly valued. Success in this role requires a blend of rigorous technical execution, clear communication, and a genuine enthusiasm for leveraging technology to transform higher education. You will collaborate closely with product managers, system architects, and academic leaders to deliver reliable and secure software solutions.
Recruiter Screen
reportedHalf of this call is the part candidates treat as small talk: start date, notice period, work authorisation and its timing, location and time zone, on-call, and the number. Those are what kill offers late, after several engineers have each spent a day. Surfacing a hard constraint now costs you nothing and occasionally buys you something, since a loop compressed to fit a competing deadline can usually only be arranged if it is asked for early. The common failure is deflecting the compensation question twice, then discovering at offer stage that the band never reached your number.
What to demonstrate
- Whether your hard constraints are compatible with the role before a loop gets booked: earliest start, notice period, what authorisation you hold and when it needs action, days on site, willingness to carry a pager
- Whether you give a compensation range with something behind it, such as current total compensation or a competing timeline, rather than leaving the band untested
- Whether your stated timeline is real, since a competing deadline raised now is something scheduling can sometimes work around and the same deadline raised at offer stage usually is not
How to prepare
- Write each constraint down in one line before the call and state them as facts rather than negotiating them live under a question you were not expecting
- Set your range from two or three current data points for that level and location, and name the structure you are quoting in, so the number is comparable to the one they are holding
- If another process is running, say where it stands and by when, and ask directly whether this loop can be scheduled inside that window
Hiring Manager Discussion
reportedThe design portion here is shorter and lower-stakes than a dedicated design round, which changes what it tests. There is rarely time to reach a full component diagram, so what gets read is your first two minutes: whether you pin down constraints, meaning request rate, data size, what must not be lost and how stale a read may be, before naming any technology. Opening with a stack list invites being steered back. Once the numbers are on the table, say what breaks first if they grow tenfold, and defend the plain option where the load does not justify more.
What to demonstrate
- Whether constraints come before components: peak request rate, data volume, what must survive a process dying, and the staleness the product can tolerate
- Whether you can name what saturates first when traffic grows by an order of magnitude, and whether that matches the design you just sketched
- Whether a cache is reasoned about on both paths, since a cold or recently flushed cache sends the full request rate to the origin, so capacity has to cover the miss case and not only the steady state
- Whether you distinguish what you have operated from what you have only read about, which usually shows up in the answer to why a particular component is there
How to prepare
- Take one system you worked on and write down the numbers you would need to defend it: requests per second at peak, rows in the largest table, the latency you were actually held to. Not having them is the common stall in this part.
- Practise the tenfold question on that system out loud, naming the first bottleneck you would hit, whether that is a single writer, connection limits, disk, or a queue that grows faster than it drains, and the smallest change that buys headroom.
- Prepare one decision where the plain option was correct: the cache you did not add or the queue you did not introduce, with the load figure that made that the right call. Being able to argue for less is rarer than being able to argue for more.
Technical Deep-Dive
reportedMost of the time lost in this format is not lost to thinking. It goes to a standard-library call you half-remember, an off-by-one in a loop bound, and a debugging loop that mutates code at random until something passes. When output is wrong, stop re-reading the whole function: take the smallest input that reproduces it and walk the state through by hand, printing intermediates if the environment allows. Guessing at a fix without a failing case you understand is how a five-minute bug becomes twenty, and the clock does not pause while you do it.
What to demonstrate
- Whether you reach the right structure without a detour, and can write it from memory rather than only recall that one exists
- Whether overflow is considered where the language has fixed-width integers, since a signed 32-bit value stops at 2,147,483,647 and then wraps in Java, is undefined behaviour in C++, and does not arise in Python, whose integers grow instead
- Whether recursion depth is treated as a constraint on large inputs, given that CPython's default limit is 1000 frames and a deep recursion can exhaust the stack in any language where an iterative version would not
- Whether a failing case is isolated and explained before any edit is made to the code
How to prepare
- From an empty file and with no references open, implement the pieces you lean on most: a heap push and pop, an iterative DFS with an explicit stack, and a binary search whose midpoint is written lo + (hi - lo) / 2, which avoids the overflow that (lo + hi) / 2 can hit in a fixed-width integer type
- Time yourself on the ten library calls you look up most, such as sorting with a custom comparator, splitting and joining strings, and finding the next key at or above a value in an ordered map, until the lookup is gone
- Take a solution you know is broken and, before touching it, write one sentence naming the input, the expected value and the actual value. Repeat until you do it without deciding to.
Practical Coding Exercise
reportedInput bounds are the part of the prompt most often skimmed, and they usually contain the answer. They tell you which complexity class is admissible, which narrows the search before you have thought about the problem itself. As a rough planning figure, a compiled language does on the order of 10^8 simple operations per second and an interpreted one roughly an order of magnitude less. So n up to about twenty admits enumerating subsets, a few thousand admits a quadratic pass, and a million admits neither: you need near-linear, or linear with a log factor. If the bounds are missing, ask for them.
What to demonstrate
- Whether the approach is justified by the stated input size rather than by whichever pattern you recognised first
- Whether you ask about the properties that change the algorithm: whether the input arrives sorted, whether duplicates occur, whether values are bounded integers, whether it all fits in memory
- Whether you can name the bottleneck in your own solution and what would remove it, even when you deliberately leave it in place
- Whether a claimed speedup is real, since memoising a recursion only helps when subproblems genuinely overlap and the state can be keyed cheaply
How to prepare
- For each algorithm you rely on, write down the largest n it handles in roughly a second, then check two of those figures by timing them in the language you will actually type in
- For two weeks, write one line naming your target complexity and the bound that justifies it before you write any code, then compare that line with what you ended up submitting
- Practise the conversion backwards: given a required O(n log n), list the mechanisms that get you there (sorting, a heap, an ordered map, divide and conquer) and choose by what the problem needs to query, not by what you used last
PracHub editorial advice for the preparation topics above.
Treating the launch subject identifier as a global user id, or merging identities on email.
The subject claim is stable only within its issuer, so the same human arriving from a second platform is a different subject and must be a second external identity bound to the same person. Email is worse than useless as a merge key here: privacy settings frequently suppress it from the launch entirely, districts recycle addresses between graduating and incoming students, and younger learners often have none. A merge on email eventually joins two children's work into one account, which is both a grade bug and a privacy incident.
Applying a bulk roster snapshot as authoritative, including its absences.
A truncated file, a partially completed export, or a change in the upstream identifier scheme all present as a large set of rows that vanished, and nothing in the file distinguishes that from a real mass withdrawal. Applied naively it deprovisions enrollments in a single run, and the recovery is not just an undo because in-progress work and derived caches have already moved. Every pipeline that survives has a proportional gate, a recorded diff, and a restartable apply keyed on the snapshot digest.
Writing code before the input contract is pinned down
Before the first line, state the types, the size bounds, whether duplicates, negatives or an empty input are possible, whether the input is sorted, whether you may mutate it, and what the function returns when nothing matches. Every one of those answers changes the code, and discovering one at minute twenty costs a rewrite you no longer have time for.
Choosing a schema before the access patterns are known
Write the queries first, with their filters, sort orders, cardinalities and which ones sit on the latency-critical path, then design tables and indexes to serve them. An index nothing queries still costs write throughput and storage, and a hot query with no supporting index becomes a full scan that only hurts once the table is big.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Find peak concurrent attempts to size a pre-warm
You have one district's attempt rows for a single school day: started_at, and an end instant taken as submitted_at when present and server_deadline_at otherwise. There are up to 50,000,000 rows, all inside one local calendar day, and the district's org.time_zone is known. Compute the maximum number of simultaneously in-progress attempts and the second at which it occurs, so capacity can be pre-warmed against the bell schedule. A sort-based sweep is acceptable but is not the best answer here. State your bounds and what happens on a day with a daylight-saving transition.
Approach
- Note the structural fact that decides the algorithm: the output domain is one day at one-second resolution, so there are on the order of 86,400 distinct answer positions while there are 50,000,000 inputs. That inverts the usual sweep, because the coordinate space is far smaller than the data.
- Allocate a difference array over seconds from local midnight, with one sentinel slot past the end so the decrement for an attempt ending in the final bucket has somewhere to land. For each attempt add 1 at the bucket containing
started_atand subtract 1 at the bucket after the one containing the end instant, then prefix-sum once and take the argmax. That is O(N + T) time and O(T) space, which is 86,401 int32 values, about 346 KB — small enough to stay in L2 and to run per section if you want. - Say why this beats the sweep-line answer here. Sorting 2N endpoints is O(N log N) with 100,000,000 entries to materialise; the difference array touches each input twice and never sorts. The sweep only wins when the coordinate space is large or unbounded, which is exactly the condition this problem does not have.
- State the bias the bucketing introduces rather than hiding it. Every attempt is counted in full in every second it touches at all, so the bucketed maximum is greater than or equal to the true instantaneous maximum, never less, which is the safe direction for pre-warming. The overcount is not confined to sub-second attempts: two hour-long attempts that merely share a boundary second without ever overlapping, one ending at 10:00:05.2 and the next starting at 10:00:05.8, already put 2 in that bucket against a true instantaneous maximum of 1. The bound that does hold is per bucket. The attempts touching second s split into those that span it entirely, which by definition coexist at every instant of s, and those with an endpoint inside it, so the overcount at s is at most the number of attempts that begin or end within that second. Equality with an exact sweep is guaranteed when no attempt begins or ends inside the peak second — which holds, for instance, on a fixture whose every endpoint lands exactly on a second boundary.
- Handle the calendar honestly, and fix the indexing convention before sizing anything. Index buckets by elapsed seconds from the UTC instant of local midnight and size the array from the real UTC span between successive local midnights: in a zone that shifts by an hour that span is 82,800 seconds on the spring-forward day and 90,000 on the fall-back day, not 86,400. A hardcoded 86,400 therefore overruns on the fall-back day, whose last hour indexes up to 89,999, and merely leaves 3,600 dead slots at the tail on the spring-forward day, which is harmless. The opposite convention fails the opposite way: bucketing by local wall-clock seconds-since-midnight always stays in range, but maps the fall-back day's repeated hour onto buckets that already hold the first pass through it, folding two real hours of load into one and understating the peak.
- For per-section peaks, do not allocate T buckets per section — 10,000 sections is 3.5 GB. Partition the input by
section_idand reuse one array, or keep only sections whose total attempt count clears a threshold, since the rest cannot produce a peak worth pre-warming for.
Follow-up
- Now compute the peak across three districts in different time zones on one shared cluster. What is the coordinate space and does your answer survive?
- Some attempts have neither
submitted_atnorserver_deadline_atbecause they were abandoned. What end do you pick, and how does the choice bias the number you hand to capacity planning? - You need the top ten peak minutes rather than the single peak. What changes, and what does not?
Stream a quoted roster CSV without loading the file
You are handed a 2 GB roster export as a byte stream. It follows RFC 4180: any field may be quoted, a quoted field may contain commas and CRLF, a literal quote inside a quoted field is written as two quotes, and the file may open with a UTF-8 byte order mark. There are up to 10,000,000 records and a single field may reach 64 KB. Write a reader that yields one record at a time without buffering the whole file and that fails loudly on a malformed file rather than emitting a short record. State your bounds.
Approach
- Refuse the shape that fails first: splitting the stream on newlines and then splitting each line on commas. A quoted field may contain a CRLF, so line splitting cuts records in half, and the damage is silent because both halves parse into plausible short records.
- Write a byte-level state machine with four states — field start, unquoted field, quoted field, and quote-seen-inside-quoted. In the last state a second quote emits one literal quote and returns to quoted; a comma or newline ends the field; anything else is a malformed file. That table is the whole parser and it is the part to get right on paper before typing.
- Strip the BOM only at offset zero, comparing the first three bytes against EF BB BF. Left in place it becomes part of the first header name, so the header lookup for that column misses and the column reads as absent for every record in the file.
- Cap field length at the stated 64 KB and record length at a sane multiple of it. Without a cap, one unbalanced quote makes the parser treat the remaining 2 GB as a single field and the process dies of memory exhaustion rather than telling you the file is broken at byte 12,004.
- Treat end of stream inside a quoted field as a hard error, not an implicit close. A truncated upload is the most common malformed input here and it is byte-for-byte indistinguishable from a complete file if you close the field silently — the missing rows then present downstream as a mass withdrawal.
- Complexity: O(n) time in bytes with one pass and no backtracking, and O(longest field + longest record) space, which is bounded by the caps rather than by the file. Read in fixed blocks and keep the partial field across block boundaries instead of reading line-wise.
Follow-up
- The file arrives gzipped and you must report progress as a percentage. What can you actually report, and what does that do to your memory bound?
- Two exporters disagree on line endings and one emits a bare LF inside a quoted field. Does your state machine care, and should it?
- How do you distinguish a truncated file from a complete one when the byte stream itself gives you no signal?
Backfill frozen lateness without converting a hundred million times
A backfill must set due_at_utc on 1,000,000 assignment rows where it is NULL, from due_at_local plus policy_time_zone, and then set is_late on the 100,000,000 attempt rows beneath them whose is_late is NULL, by comparing submitted_at to that instant. Converting one local time per attempt row is correct and far too slow. Give a plan that is both faster and reproducible when re-run months later, state the speed-up you expect, and define the result for a local due time that is ambiguous or nonexistent across a daylight-saving transition.
Approach
- Name why the naive plan is slow rather than asserting it. A local-to-UTC conversion is a binary search over that zone's transition table plus an object construction, on the order of one to three microseconds in a managed runtime; 100,000,000 of them is a few minutes of pure conversion and, once per-row fetch and allocation are included, hours. Nothing about it is wrong — it is the row count multiplied by a cost that did not have to be paid per row.
- Collapse the cardinality. The conversion depends only on
(policy_time_zone, due_at_local), and due dates cluster hard on end-of-day instants for school days, so the distinct pair count is on the order of 10,000 against 100,000,000 attempts. Convert the distinct pairs once into a small mapping, then resolve each attempt with a hash probe. That is roughly a 10,000-fold reduction in conversions, and the per-row cost drops from a conversion to a lookup. - Do it in two stages so the second one is a join, not a computation: materialise
due_at_utcon the 1,000,000 assignment rows from the mapping, then setis_lateon attempts by joining onassignment_id. The attempt stage then performs no time-zone work at all, which is also what makes it re-runnable without the tz database on the path. - Pin reproducibility explicitly. The local-to-UTC mapping is a function of the zone, the local time, and the time-zone database version, because releases correct historical transitions. Record the tzdata version on the backfill run, or a re-run months later can produce a different instant for the same row and a grade will move with no grade event behind it.
- Define the two degenerate cases instead of letting a library default decide. An ambiguous local time in a repeated hour maps to two instants; a nonexistent one in a skipped hour maps to none. Pick the rule — the later instant when ambiguous, the first instant after the gap when nonexistent — encode it as an explicit flag rather than relying on a fold default, and assert both in tests. An 11:59 PM due time is outside the affected hour for most zones, so this is rare, which is exactly why it will not be noticed when it is wrong.
- Respect the freeze. Update only rows where
is_late IS NULL; a row that already carries the flag was decided at submit and must not be recomputed, or the backfill becomes a retroactive regrade. Then chunk the update by primary-key range with a bounded transaction size, because one 100,000,000-row statement holds locks and generates dead tuples for the whole run and cannot be resumed after a failure.
Worked solution 40 min
- Count the distinct
(policy_time_zone, due_at_local)pairs over the 1,000,000 affected assignments and write the number down; the whole plan rests on it being small. - Convert those pairs once into an explicit mapping table, recording the tzdata version and the ambiguity rule alongside it.
- Update
due_at_utcfrom that mapping, then updateis_lateon attempts by joining onassignment_idwith aWHERE is_late IS NULLpredicate, chunked by primary-key range. - Seed a fixture with a due time of 01:30 local on a fall-back date, one of 02:30 local on a spring-forward date, and one ordinary 23:59 due time, each with attempts on both sides of the boundary.
- Run the whole backfill twice and diff the resulting columns row by row.
Follow-up
- A district is moved to a different school with a different time zone mid-term. Which rows, if any, change, and what does the freeze rule say?
- The tzdata version on the backfill host differs from the one in the application image. How would you have found that out before the run rather than after?
late_policyisblockfor some assignments, so a late submission should not exist. What do you write for an attempt that submitted after the due instant under that policy?
Explain why the grading queue index stopped serving its ORDER BY
attempt holds 2 billion rows and 99.7% are state='graded'. The grading queue runs SELECT attempt_id FROM attempt WHERE tenant_id = $1 AND state IN ('submitted','grading') ORDER BY submitted_at LIMIT 100 against index ix_attempt_queue (tenant_id, state, submitted_at). At the bell the query takes 30 seconds and EXPLAIN shows a Sort above a bitmap heap scan, even though the queue is only a few thousand rows. Explain why that index cannot serve the ordering, give the replacement definition, and say how the replacement behaves as rows leave the queue.
Approach
- Start from what the index actually orders. Its tuples are sorted by (tenant_id, state, submitted_at), so with two state values the scan yields all 'grading' rows in submitted_at order, then all 'submitted' rows in submitted_at order. The query wants one submitted_at order across both, which that physical order does not provide, so the planner must sort the whole matching set and the LIMIT cannot stop the scan early. With a single state constant the same index satisfies the ordering perfectly — the plan changed because of the IN list, not because of the data.
- Replace it with a partial index that moves the predicate out of the key: CREATE INDEX CONCURRENTLY ix_attempt_queue ON attempt (tenant_id, submitted_at) WHERE state IN ('submitted','grading'). Ordering is now satisfied directly, the LIMIT stops after 100 index tuples, and the index covers roughly 0.3% of rows so it stays resident in cache instead of competing with the table.
- State the precondition, because it is the part people skip: the planner uses a partial index only when it can prove the query's predicate implies the index predicate. A literal IN list matching the definition proves; state = ANY($1) with the list bound as a parameter does not, so the queue's state set has to be literal in the SQL, not passed in.
- Fix the estimate separately from the access path. With five states the planner already has each one in its most-common-values list, so single-column selectivity is fine; what it gets wrong is the dependence between tenant_id and state, because a bell means one tenant owns nearly the whole queue while the planner multiplies the two selectivities as if independent. CREATE STATISTICS ON tenant_id, state FROM attempt corrects it, and you ANALYZE explicitly after a bulk transition rather than waiting for autovacuum, whose analyze threshold of 50 + 0.1 x reltuples needs 200 million changed rows on this table.
- Account for the lifecycle. A row leaves the queue by UPDATE state='graded', so its new version is not in the partial index at all while the old index entry remains until vacuum reclaims it. The index therefore churns in proportion to throughput rather than to depth: size it and vacuum it against how many attempts pass through per day, not against how many sit in the queue at once.
- Keep the alternative in your pocket for when the partial index is not acceptable: two explicit scans merged, (SELECT ... WHERE state='submitted' ORDER BY submitted_at LIMIT 100) UNION ALL (SELECT ... WHERE state='grading' ORDER BY submitted_at LIMIT 100) ORDER BY submitted_at LIMIT 100, which keeps the existing index and pushes the LIMIT into both branches. Confirm whichever you choose with EXPLAIN on your major version rather than from memory, since array-key handling in b-tree scans has changed between releases.
Follow-up
- The queue drains and refills eight times a day. What is the vacuum strategy for this index, and what do you monitor to know it is losing?
- A second consumer claims rows with FOR UPDATE SKIP LOCKED. Does the plan survive, and what does skipping do to the ordering guarantee?
- Support wants oldest-first fairness across tenants rather than within one. Does the same index serve that query?
Stop the gradebook rollup from counting attempts as learners
For one section, produce per published assignment: learners assigned, learners with a graded attempt, mean best score, and late count. Tables: enrollment as above (a person may hold two roles, and may have withdrawn and re-enrolled), assignment(assignment_id, tenant_id, section_id, status, max_attempts), attempt(attempt_id, tenant_id, assignment_id, person_id, attempt_no, state, is_late), score(attempt_id PRIMARY KEY, tenant_id, score_given, score_maximum, revision). Grading policy is best attempt wins. Write one query. State the grain of every branch before you join it, and name where a naive join multiplies rows.
Approach
- Say the grains out loud first: enrollment is one row per (section, person, role) and repeats per person; attempt is one row per (assignment, person, attempt_no); score is one per attempt. Join all three flat and the row count is enrollment rows multiplied by attempts, so every aggregate above it is weighted twice over.
- Collapse attempts to the policy grain before anything else: SELECT DISTINCT ON (a.assignment_id, at.person_id) a.assignment_id, at.person_id, at.is_late, s.score_given FROM assignment a JOIN attempt at ON at.tenant_id = a.tenant_id AND at.assignment_id = a.assignment_id JOIN score s ON s.tenant_id = at.tenant_id AND s.attempt_id = at.attempt_id WHERE a.tenant_id = $1 AND a.section_id = $2 AND at.state = 'graded' ORDER BY a.assignment_id, at.person_id, s.score_given DESC, at.attempt_no DESC. That ORDER BY is where 'best attempt wins' is written down, and it is the only place it should appear.
- Collapse the roster the same way: SELECT DISTINCT person_id FROM enrollment WHERE tenant_id = $1 AND section_id = $2 AND role = 'learner' AND deleted_at IS NULL AND begin_date <= $3 AND (end_date IS NULL OR end_date >= $3). Now re-enrolment and dual roles cannot multiply anything downstream.
- Join the two collapsed sets, LEFT from assignment so an assignment nobody submitted still returns a row of zeros, and aggregate with FILTER instead of one correlated subquery per column: count(b.person_id) AS learners_graded, avg(b.score_given) AS mean_score, count(*) FILTER (WHERE b.is_late) AS late_count, with learners_assigned taken from the roster cardinality.
- Reject COUNT(DISTINCT person_id) as the fix. It repairs the counts and leaves AVG computed over the multiplied rows, so the mean quietly weights learners who attempted more often, and the number stays plausible for as long as anyone looks at it.
- Keep the cost where it belongs: with an index on attempt (tenant_id, assignment_id, person_id, attempt_no) the DISTINCT ON is an ordered index scan over a bounded set — order 10^2 assignments times order 10^1 to 10^2 learners — rather than a sort of the table. Carry tenant_id into every join condition, not just the outer WHERE.
Worked solution 30 min
- Seed one section with 6 learners, one of them re-enrolled and one holding two roles, and give one learner three graded attempts on the same assignment.
- Run the flat three-way join first and record the inflated counts and the shifted mean as the baseline error.
- Add the DISTINCT ON collapse and the roster collapse, then re-run and diff against the baseline.
- Compute each learner's best score in a separate standalone query and compare the per-assignment mean by hand.
Follow-up
- A regrade changes score_given after the fact. Does the best attempt change, and what would you order by if the policy were latest attempt instead?
- The instructor opens twelve sections at once. Does this query shape hold, or does the DISTINCT ON become the problem?
- Half the attempts are state='voided' after an integrity review. Where in the query does that belong, and what happens to mean_score?
These questions examine your familiarity with specialized stacks, depl…
These questions examine your familiarity with specialized stacks, deployment pipelines, and cloud systems.
Approach
- Name the failure you are designing for, then the recovery path.
- Fix the scope first: who calls this, how often, and what they do when it fails.
- Name the read and write paths separately; they rarely have the same bottleneck.
Follow-up
- What breaks first when traffic grows ten times?
- What would you drop to keep the system up under load?
What strategies do you employ when designing secure RESTful web servic…
What strategies do you employ when designing secure RESTful web services?
Approach
- Name the failure you are designing for, then the recovery path.
- Choose a partition key and say what query it makes expensive.
- State the consistency you need, and where you are willing to be stale.
Follow-up
- What breaks first when traffic grows ten times?
- What would you drop to keep the system up under load?
Describe your ideal pipeline and what tools you would use to build it.
Describe your ideal pipeline and what tools you would use to build it.
Approach
- Fix the scope first: who calls this, how often, and what they do when it fails.
- State the consistency you need, and where you are willing to be stale.
- Name the read and write paths separately; they rarely have the same bottleneck.
Follow-up
- What breaks first when traffic grows ten times?
- How does this behave when that dependency is down for an hour?
Interviewers use these scenarios to test your approach to scalability,…
Interviewers use these scenarios to test your approach to scalability, optimization, and system architecture.
Approach
- Choose a partition key and say what query it makes expensive.
- Name the failure you are designing for, then the recovery path.
- Name the read and write paths separately; they rarely have the same bottleneck.
Follow-up
- What would you drop to keep the system up under load?
- How does this behave when that dependency is down for an hour?
Explain composition in the context of object-oriented programming.
Explain composition in the context of object-oriented programming.
Approach
- State your assumptions explicitly before working the problem.
- Work from the requirement backwards to the design.
- Clarify what is being asked and what a complete answer contains.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
This category evaluates your core programming competencies, language f…
This category evaluates your core programming competencies, language fluency, and architectural understanding.
Approach
- Say what you would check first and why it is the highest-information step.
- State your assumptions explicitly before working the problem.
- Work from the requirement backwards to the design.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
Absorb a hundredfold submission burst without starving one section
Submissions reach the autograder in the last 90 seconds of a period: assume 2,000 sections of 30 learners in one time zone, so roughly 60,000 jobs in 90 seconds against a flat mean two orders of magnitude lower. Per-job wall clock is capped between 2 and 15 seconds depending on language. Design the queue and admission path: the partition key, how you size worker concurrency for a stated drain time, what backpressure the submit endpoint applies, and what prevents one district's 10,000-job upload from delaying every other district behind it.
Approach
- Fix the target before the topology. Derive a drain deadline from the product promise, say feedback within five minutes of submit, then size with Little's law: 60,000 jobs at a 5-second mean service time drained in 300 seconds needs 60,000 x 5 / 300 = 1,000 concurrent sandboxes. Redo it at the 15-second cap to see the worst case. A queue with no stated drain time is an unbounded backlog with extra machinery.
- Separate submit from grade. The submit write must be durable whether or not a worker is free, so the job is an outbox row or a queue message written in the same transaction as the attempt state change. Backpressure then lives on grading latency, expressed as queue depth and a visible grading state, rather than on the submit endpoint, because refusing a submit at the end of a period destroys work that cannot be recreated.
- Choose the partition key for fairness rather than throughput. Keying lanes on tenant gives each district its own ordered stream so a bulk upload cannot head-of-line block a neighbour, and a per-tenant concurrency cap bounds any one tenant's share of the pool. Keying on section is finer but shatters the pool into thousands of mostly empty lanes and makes next-job selection expensive. Tenant lanes plus a cap match the axis along which complaints actually arrive.
- Keep the dispatcher work-conserving. Workers pull across eligible lanes, by weighted round robin or a least-recently-served scan using a skip-locked claim, instead of being pinned to a lane. Without that, fairness degenerates into reserved-but-idle capacity at precisely the minute the burst lands.
- Write the degradation ladder explicitly. Pre-warm sandbox capacity against the bell schedule, place instructor-triggered regrades and re-runs below first attempts in priority, and surface a queued position rather than an indefinite spinner. The job is never dropped: the attempt is already submitted and the score is owed.
Worked solution 30 min
- Compute the arrival rate across the 90-second window and the steady-state rate, and state the ratio between them.
- Pick a drain deadline and apply Little's law at your assumed mean service time, then repeat at the 15-second cap.
- Write the partition key and the per-tenant concurrency cap, then demonstrate the head-of-line case it prevents with a two-district example.
- Describe next-job selection and show it stays work-conserving when one lane empties.
- Write the degradation ladder in priority order and mark the one thing that is never shed.
Follow-up
- One submission deterministically hits the 15-second wall clock and is retried three times. What does that do to drain time, and does a timeout retry differ from a crashed-worker retry?
- An instructor queues regrades for an entire section during a bell. Which lane do those jobs enter, and at what priority relative to first attempts?
- Workers are pre-warmed for the bell and then idle for 20 minutes. What is the cost model, and would you rather hold the capacity or accept a longer drain?
One tenant's grade publications stall behind a single row
One tenant's grade publications stopped at 09:14. Its score_publication pending count has grown to 38,000, succeeded is flat, and other tenants are unaffected. Worker logs show publication_id 8841 processed 41,000 times. That row is state pending, attempt_count 0, last_error_code NULL, and the assignment behind it has external_line_item_url NULL. The relay claims 100 rows per transaction ordered by next_attempt_at, publishes each, writes the outcome onto the row, and commits once at the end. Give an ordered diagnostic checklist, the root cause including why attempt_count never moved, and the fix.
Approach
- Order the checklist from blast radius inward. First, is the stall tenant-scoped or fleet-scoped: it is one tenant, which matches a per-tenant relay shard, so the transport and the receiving platform are not obviously at fault. Second, is the oldest pending row advancing: it is not. Third, read the counters on the stuck row. attempt_count 0 and last_error_code NULL after 41,000 log lines is the whole diagnosis, because those columns are written on every failure path.
- Name the mechanism from that contradiction. The failure bookkeeping is written inside the same transaction that the uncaught exception aborts, so the increment and the error code roll back with everything else. The row can therefore never reach the attempt_count threshold that promotes it to dead, and the relay reclaims the identical batch forever. This is a hot loop, not a slow queue.
- Explain the blast radius: the claim is a batch in one transaction ordered by next_attempt_at, so one unpublishable row at the head blocks the 37,999 behind it. The queue is not backed up by volume; it is head-of-line blocked by a single row.
- Fix the transaction boundaries first. Claim per row with SELECT ... FOR UPDATE SKIP LOCKED so a locked or failing row cannot block its neighbours, and commit each row's outcome in its own transaction. Advance next_attempt_at as part of claiming, before the outbound call, so that a crash mid-publish still leaves a scheduled row rather than an immediately reclaimable one.
- Then fix the classification, because retrying was never going to work. A NULL external_line_item_url is a configuration fault with no time-based recovery, so it is a permanent rejection: set state dead with last_error_code, not failed_retryable. Reserve backoff for timeouts, 429s and 5xx; treat an authorisation expiry as retryable only after a token refresh; treat a conflict on a stale ordering token as terminal for that revision. Throughout, the at-most-once property rests on the idempotency key staying derived from attempt_id and score revision, so a row re-published after a crash between the outbound call and the commit cannot double-apply; a per-retry UUID here would convert that crash into a duplicate grade.
- Change the alert to the thing that was actually wrong: age of the oldest pending publication per tenant, plus a dead-letter count support can see. A total pending count would have fired six hours late, and a success-rate alert would never have fired at all because the failing call never returned a status.
Follow-up
- The relay dies after the receiving platform commits but before the row is marked succeeded. Walk through what happens on restart and why no grade is duplicated.
- Support wants to replay the dead rows after the line item is configured. What does a safe manual replay do, and what must it not regenerate?
- How do you bound a tenant whose publications legitimately fail for two days, so its retries do not starve other tenants sharing the worker pool?
Four days sample coding, design, fundamentals and the practical rounds at deliberately shallow depth, which is enough to surface the topics you did not know were in scope. That map, rather than a guess made on day one, decides where the last three days go.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Coding, one pass at shallow depth
- Solve one problem from each of six families, an array with two pointers, hash counting, binary search, a tree traversal, a graph traversal and one dynamic program, under a hard twenty-minute cap with no extensions, marking each finished, late, or stalled.
- For every stall, write the exact move you could not make rather than the subject, so the note reads could not turn the recurrence into a loop rather than bad at dynamic programming.
- Fix nothing today. The value of the pass is the unfixed record.
Deliverable: Six timed attempts marked finished, late or stalled, each stall carrying a named blocking move.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Design, one pass at shallow depth
- Spend twenty minutes each on three different shapes, a read-heavy feed, a write-heavy ingest path, and something needing a transaction across two entities, stopping each at requirements, interface and data model.
- After each, write the first question you could not answer, which is usually a number you could not estimate or a failure mode you had no vocabulary for.
- Mark which of the three you would be most relieved not to be asked, and treat that as data rather than as a preference.
Deliverable: Three shallow designs, each with the first unanswerable question written at the bottom.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Fundamentals and the practical rounds
- Answer eight short questions in writing at four minutes each, covering the material that fills the gaps between the big rounds: what happens between a URL and a rendered page, what an index costs on write, when a process is preferable to a thread, and what conditions a deadlock requires.
- Do one thirty-minute practical task of the kind a take-home compresses: read an unfamiliar two-hundred-line file and write what it does, what you would change, and the one thing you remain unsure of.
- Score every answer fluent, correct but slow, or absent, and keep the absent ones visible.
Deliverable: Eight scored short answers and one written reading of unfamiliar code.
Practice prompt ↗Practice prompt ↗04The rounds that are about you, and the map
- Deliver three behavioural answers aloud against a timer, a conflict, a failure you owned, and a decision made without enough information, marking any that ran past three minutes or contained no number.
- Assemble the map: every marked item from days one to three on a single page, sorted by how likely it is to appear in your loop rather than by how uncomfortable it felt.
- Choose exactly two areas for the remaining three days and write down what you are deliberately abandoning.
Deliverable: A one-page scored map of the whole surface area with two areas chosen and the rest explicitly abandoned.
Practice prompt ↗Practice prompt ↗Worked solution ↗05First chosen area, to the depth you skipped
- Work the higher-ranked area in four focused blocks, choosing items one level above where you stalled rather than repeating what already works.
- After each block write the rule you extracted in one sentence with its precondition attached, since a rule carrying no precondition is exactly what fails under a variation.
- Re-attempt the day-one or day-two item that exposed this area and compare against the original timing.
Deliverable: Four worked blocks, a timed re-attempt against the original, and three one-sentence rules with preconditions.
Practice prompt ↗Practice prompt ↗06Second chosen area, where the gap is coverage rather than speed
- Treat the second area differently from the first. Day five drilled something you could already half-do; this one is usually a topic you had simply never met, so build one worked reference example end to end and keep it, rather than attempting six problems badly.
- Write down the vocabulary you were missing on day two or three, five terms at most, each with the one sentence that makes it usable in an answer rather than the textbook definition.
- Redo the shallow attempt that exposed this area and note whether you now fail later in the problem, because moving the failure point is the realistic gain from a single day and is worth more than a score that did not change.
Deliverable: One worked reference example for the newly covered area, a five-term vocabulary list, and a note on where the failure point moved.
Practice prompt ↗Practice prompt ↗07Reassemble the loop
- Sit two rounds back to back with no gap, ordering them so the area you chose second comes last, because the map was built from rested, isolated attempts and the loop will reach your weaker area when you are already spent.
- Write where the second round suffered from the first, which is normally the point at which structure collapses into narration.
- Reduce the week to one page holding only the rules you can state without reading them.
Deliverable: Mock notes on cross-round carryover plus a one-page card of rules you can recite from memory.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Team size, service count and tickets closed say very little. Seniority shows in the decision you owned: what you chose not to build, which constraint you traded away, whose objection you had to resolve before anything could move. A large project where you executed someone else's plan is a small story.
Tell me about a time you had to learn a completely new framework under…
Tell me about a time you had to learn a completely new framework under a tight deadline.
Approach
- State the situation in two sentences and spend the rest on the reasoning.
- Close with what you would do differently, concretely.
- Give the blast radius: what could have broken, and what you measured.
Follow-up
- What would you do differently if you ran that again?
- What did you decide not to do, and why?
Describe your previous experience and technical projects related to th…
Describe your previous experience and technical projects related to this role.
Approach
- Close with what you would do differently, concretely.
- Pick a story where you made the decision, not one where you watched it.
- Name the disagreement and how you resolved it with evidence.
Follow-up
- How did you know your change caused the improvement?
- What did you decide not to do, and why?
Reverse a deletion gate that operators stopped reading
You shipped a gate holding any roster run that proposes to soft-delete more than five percent of a tenant's active enrollments, recorded on roster_sync_run.gate_state as held_for_review. In the first two weeks of a term, when withdrawals are genuinely large, it fired on roughly forty percent of runs and operators began approving without opening the diff. Describe reversing a decision of this shape: the evidence that changed your mind, what you replaced it with, and how you avoided reversing into the failure the gate existed to prevent.
Approach
- Lead with the measurement that reversed you, not with the discomfort. Hold rate, and more importantly the distribution of time between a hold and its approval. A control approved in a median of a few seconds is not a control, it is a log line with a button, and that number is the argument. Say plainly that the gate was still technically doing its job and had already stopped being a defence.
- Separate the two situations the single threshold cannot tell apart. The gate exists to catch a truncated file or a changed identifier scheme; it is not meant to catch a real withdrawal wave. Both present as many rows missing, so proportion alone is the wrong discriminator. The signals that do separate them are rows_in against the tenant's trailing snapshot sizes, whether missing persons have same-shaped replacements present, and whether deletions cluster on whole orgs or spread thinly across sections.
- Replace one threshold with a small set of rules that keeps the fail-safe direction. Hold unconditionally when rows_in collapses against the trailing median or falls below an absolute floor, hold when a large share of missing sourcedIds have plausible replacements in a new format, auto-apply when the run looks like the tenant's own history. Keep the hold non-destructive: deletes are retained, the run stays restartable from snapshot_digest, and nothing is discarded by a rejection.
- Guard the reversal with a measurement rather than confidence. Run the old rule in shadow, log disagreements, and track hold precision, meaning of the runs held how many were genuinely bad, alongside hold rate. Loosening a safety control is only defensible if you can show it was loosened where it was wrong, and a shadow comparison is what turns that from an assertion into a number.
- Say what would reverse you again and when you look. Term-start weeks are the regime where this rule earns or loses its keep, so the review cadence is per term, not quarterly. Then name the original mistake without softening it: the threshold was designed and validated against the forty-six quiet weeks and met its real workload for the first time in production.
Follow-up
- A district legitimately closes a school mid-year and every enrollment under one org ends. Does your rule hold it, and is that the behaviour you want?
- What does an operator see that lets them decide in under a minute, and what would make them read it rather than approve it?
- A run was approved that should not have been. What does unwinding it touch beyond the enrollment rows?
- 01
Tell me about a time you had to learn a completely new framework under a tight deadline.
- 02
Describe your previous experience and technical projects related to this role.
- 03
You shipped a gate holding any roster run that proposes to soft-delete more than five percent of a tenant's active enrollments, recorded on roster_sync_run.gate_state as held_for_review. In the first two weeks of a term, when withdrawals are genuinely large, it fired on roughly forty percent of runs and operators began approving without opening the diff. Describe reversing a decision of this shape: the evidence that changed your mind, what you replaced it with, and how you avoided reversing into the failure the gate existed to prevent.
Is this an official Western Governors University interview guide?
No. It is PracHub's own research and practice material for the Software Engineer role at Western Governors University. Rounds and questions reflect what candidates have reported, not a process Western Governors University has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult are the technical interviews at Western Governors University?
The interviews are generally rated as average in difficulty, focusing heavily on practical engineering skills, problem-solving approaches, and past project experience rather than grueling algorithmic puzzles.
PracHub interview research ↗How much preparation time should I plan for?
Allocate 2 to 3 weeks of focused preparation, concentrating on your primary programming language, system design fundamentals, and reviewing your own past project architectures.
PracHub interview research ↗What differentiates successful candidates from others?
Successful candidates combine solid technical competence with a clear understanding of how software reliability directly impacts student success and institutional efficiency.
PracHub interview research ↗What is the typical interview timeline?
The process typically moves from an initial HR screen to a hiring manager conversation and concludes with a technical panel, taking roughly 3 to 5 weeks from start to offer.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Software Engineer practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22