At Drexel University, a Software Engineer plays a vital role in building, maintaining, and optimizing the digital ecosystem that supports thousands of students, faculty, and administrative staff. Working within Drexel’s central IT division or specific academic departments, you will develop applications that streamline academic operations, power research initiatives, and enhance the student portal experience. Your work ensures that the university's digital infrastructure remains secure, reliable, and highly accessible.
This role is highly unique because it operates within a world-class cooperative education (co-op) ecosystem. You won't just write code in isolation; you will build systems that facilitate experiential learning, coordinate co-op placements, and directly impact the daily lives of the university community. Whether you are working on modern frontend interfaces using ReactJS or backend services following the MVC architecture, your contributions directly support Drexel's mission of academic excellence and civic engagement.
The engineering environment at Drexel University is highly collaborative and supportive, making it an excellent place to grow professionally. Teams often work closely with student co-ops and academic advisors, meaning you will have the opportunity to mentor emerging talent while delivering production-grade software. It is a rewarding space where your technical skills directly contribute to the educational success of the next generation.
Application Review
reportedThe person on this call usually cannot evaluate your code and does not need to. They write a short paragraph, and that paragraph is what a hiring manager skims when deciding who to put on your loop. So the test is not whether your work was hard, it is whether a non-engineer can repeat it correctly. Name systems by what they did rather than by their internal codename, give each project a shape (what was breaking, what you changed, what happened after), and keep the whole walkthrough near ninety seconds. Depth that cannot survive a paraphrase reads as vagueness.
What to demonstrate
- Whether a non-engineer can restate your projects without distorting them, since their paraphrase is what travels to the hiring manager, not your sentences
- Whether each project has a shape rather than a stack list: the failure or constraint, the change you made, the result and how it was measured
- Whether you can say what was yours inside a team project without either inflating it or disappearing into the plural
How to prepare
- Rewrite each headline project as two sentences with no internal system names and no acronyms outside your company, then say them to someone outside engineering and have them repeat them back. Fix whatever came back wrong
- Attach one measured number to each project: the baseline, the change, and the window it was measured over. Where nothing was ever measured, say that plainly rather than reaching for a plausible percentage
- Time the background walkthrough against a clock. If it runs past two minutes, compress the earliest role to a single clause and spend the recovered time on the most recent one
Structured Interview
reportedYou cannot drill a format you do not know, so put the preparation into material that travels. Three pieces of your own work, each rehearsed until you can take a follow-up you did not anticipate, will carry a conversation or a code walkthrough equally well. Specificity is what separates that from filler. A number needs its definition before it means anything: a p99 is over some window and measured at some hop, and a server-side figure excludes the queueing and network time a client would see. The number you cannot qualify is the one to leave out.
What to demonstrate
- Whether your examples carry detail only someone who did the work would hold, such as what the binding constraint actually was, which alternative you rejected and why it was worse, and what you measured on each side of the change
- Whether a number survives one follow-up, meaning you can say what it was measured over and whether it moved because of your change or merely alongside it
- Whether a failure is described with the specific change that followed it, rather than a lesson stated in general terms
- Whether your part in a team effort is stated accurately, including what other people did
How to prepare
- Write a page on each of three projects covering the constraint, the option you rejected, the measurement before and after, and what went wrong. Cut any line you cannot take a follow-up on, since you are writing the parts you will be pressed on rather than a summary.
- Recover the real figures while you still have access: request volume, data size, latency with its percentile and window, team size, timeline. Note where each came from, whether a dashboard, a design document or memory, and mark the estimates so you can say which they are out loud.
- Take your weakest project story to someone who works in a different area and have them ask why four times in succession. The point where you run out of answer is the part to go and re-read before the round.
PracHub editorial advice for the preparation topics above.
Autoscaling on average utilisation against a bell-shaped step function.
Scale-up latency (scheduling, image pull, process and JIT warm-up, connection pool and cache fill) is typically tens of seconds to minutes, while the demand step is under a minute and repeats on a published schedule. By the time the metric crosses a threshold, the period has started and the queue has already built. The workable answer is a schedule-driven pre-warm derived from the tenant calendar plus a load shedder that protects writes, not a more aggressive threshold.
Editing a content item in place and thereby rewriting the meaning of every historical score.
Authors expect to fix a typo, change a distractor, or adjust a point value, and nothing warns them that thousands of stored results reference the record they are editing. If attempts and analytics join to the mutable item rather than to a version, a regrade or a rebuilt report scores past answers against an item that did not exist when they were given. Versioning has to be the default write path, because a convention that authors must remember will not hold.
Starting work without saying what you are about to spend time on
State the plan before executing it: the approach, roughly how long it will take, and what you intend to leave hand-waved. That gives the interviewer a chance to redirect you in ten seconds rather than watching you spend fifteen minutes on the wrong sub-problem.
Assuming fixed-width integer arithmetic cannot overflow
In languages with fixed-width integers, including C, C++, Java, Go and Rust, computing a midpoint as (lo + hi) / 2 overflows once the sum passes the type's maximum, so write lo + (hi - lo) / 2 instead. Say which language you are in: arbitrary-precision integers, as in Python or Ruby, remove this specific hazard and none of the others.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Detect two-device autosave interleaving in an event stream
Autosave events arrive on a stream as (attempt_id, client_instance_id, seq, received_at) at roughly 4,000 events per second, out of event-time order by at most 120 seconds. For each attempt, report the earliest 600-second window in which that attempt's autosaves alternate between client instances — A, then B, then A again — which indicates a second tab or device rather than a reload. A plain handoff, where one instance stops and another starts, must not be reported. Memory must be bounded by concurrency, not by total events. State your bounds.
Approach
- Partition by
attempt_idso every event for one attempt lands on one worker, then buffer per attempt until the watermark — the maximumreceived_atseen minus the 120-second lag bound — passes an event's timestamp. Only then is that event's event-time position final. Anything arriving after its watermark goes to a side output and is counted, not silently dropped, because a rising late rate is the signal that the lag bound is wrong. - Inside one attempt, slide two pointers over the event-time-ordered buffer, advancing the right edge one event at a time and evicting from the left while
right.t - left.t > 600. Each event enters and leaves exactly once, so the window maintenance is O(1) amortised per event. - Represent the window as a run-length-compressed deque of
(client_instance_id, count)rather than a raw list. Pushing on the right either increments the tail run or appends a new one; evicting on the left decrements the head run and pops it at zero. Runs never merge, because adjacent runs have different ids by construction, so both operations stay O(1). - Read the alternation criterion straight off that structure: three or more runs in the window means A, B, A or three distinct instances, which is interleaving. Exactly two runs is a clean handoff — a reload where one instance stops and the next begins — and must not fire. Counting distinct ids instead of runs is what produces a false positive on every browser refresh.
- Bound the memory arithmetically. At 4,000 events per second with one autosave per learner every 25 seconds, about 100,000 attempts are active, and a 600-second window holds roughly 24 events each. That is 2.4 million live events; packed at 48 bytes it is about 115 MB, and several times that with boxed objects. The run-compressed deque is smaller still, since a single-instance window collapses to one entry.
- Separate this from the write-path rule it resembles. The attempt rejects an autosave whose
seqis belowlast_autosave_seq, which protects the stored state; it does not detect anything, because two tabs sharing a monotonic counter produce strictly increasing sequences and never trip it. Detection and rejection answer different questions and both are needed.
Follow-up
- A learner's connection drops for ten minutes and the client replays its buffered autosaves on reconnect. Does your detector fire, and should it?
- You must now also report how many seconds of overlap the two instances had. What changes in the window structure?
- The lag bound is violated and events arrive 15 minutes late for one tenant. What does your side output let you do that a dropped-event counter would not?
Dispatch a bell-driven autograde burst under per-tenant caps
Thirty thousand autograde jobs enqueue over 90 seconds as a district's periods end. Each carries tenant_id, enqueued_at, a language-dependent wall_limit_seconds between 2 and 15 with a mean near 3, and a deadline equal to its section's period end. You have 400 one-shot workers, jobs cannot be preempted once started, and no tenant may hold more than 20 percent of the workers at once. Choose the dispatch order that maximises jobs returning feedback before their deadline, give the data structure and its per-dispatch cost, and say plainly whether that order is optimal.
Approach
- Do the capacity arithmetic first, because it reframes the question. Four hundred workers at a 3-second mean serve about 133 jobs per second; 30,000 arrivals over 90 seconds is 333 per second. Twelve thousand are served during the burst, leaving 18,000 queued, which drains in about 135 more seconds. The system is oversubscribed by roughly two and a half times, so the scheduler is not deciding whether anyone misses a deadline — it is deciding who.
- Choose earliest-deadline-first and then state its standing precisely. On a single machine with preemption, EDF meets every deadline whenever any schedule can. This is neither: 400 machines and non-preemptive jobs of unequal length. Minimising the number of late jobs on parallel identical machines is NP-hard — even minimising makespan on two machines reduces from PARTITION — so EDF here is a defensible heuristic, not an optimum. On one machine, Moore-Hodgson would give the exact optimum in O(n log n); that exactness does not survive the move to 400.
- Enforce the cap with a heap of queues rather than by scanning. Keep one min-heap per tenant ordered by deadline, and a global min-heap holding at most one entry per eligible tenant keyed on that tenant's head deadline. Dispatch pops the global head, pops that tenant's head job, increments its in-flight count, and re-pushes the tenant only if it is still under 80 workers. On completion, decrement and re-insert the tenant if it was parked with work waiting. Each dispatch is O(log T + log J_t) instead of O(J), and the cap costs a counter comparison rather than a rescan.
- Add admission control, which is where most of the win actually is. A job whose
now + wall_limit_secondsalready exceeds its deadline cannot return feedback in time no matter what, so running it burns a worker slot that a still-feasible job could have used. Route it to a graded-after-the-bell path immediately. This is the Moore-Hodgson instinct — discard the jobs you cannot save — applied as a filter rather than as an exact algorithm. - Break ties on
enqueued_atascending so the order is deterministic and a job cannot be starved by a stream of equal-deadline arrivals. Without a deterministic tiebreak the same burst replays differently, and an operator comparing two incidents cannot tell a scheduling change from noise. - Give the honest ceiling. No dispatch order makes a 2.5x oversubscription feasible; the levers are worker count, wall limits per language, and whether feedback is required before the bell at all. A candidate who presents EDF as the fix, rather than as the allocation of an unavoidable shortfall, has answered a different question.
Worked solution 45 min
- Compute service rate, arrival rate, peak backlog and drain time from the stated numbers, and write the oversubscription factor down before designing anything.
- Implement the two-level heap with per-tenant in-flight counters, plus the admission filter that rejects jobs already past feasibility at dispatch time.
- Simulate the 90-second burst against 400 workers with a sampled wall-limit distribution, recording on-time count, per-tenant miss rate and the tail completion time.
- Re-run the same trace three ways — FIFO, pure EDF with no cap, and EDF with the cap and admission filter — and tabulate the three outcomes.
- Re-run the skewed case where one tenant owns two thirds of the jobs, and compare the cap's effect on that tenant against its effect on everyone else.
Follow-up
- One tenant submits 20,000 of the 30,000 jobs. Show what your cap does to the other tenants' deadline-miss rate versus pure EDF.
- Workers become preemptible with a checkpoint. Does EDF's optimality argument come back, and under exactly what conditions?
- A job exceeds its wall limit and is killed. Does it re-enter the queue, and what does your answer do to the deadline guarantees of everything behind it?
Rank the longest-stuck grade publications in bounded memory
You are streaming score_publication rows to build an operator view: tenant_id, attempt_id, target, state, attempt_count, created_at, next_attempt_at, last_error_code. Up to 20,000,000 rows are in states pending and failed_retryable, spread across up to 5,000 tenants, and they arrive as an unsorted export you may not re-read. Return the 50 longest-stuck publications overall, the 5 longest-stuck per tenant, and a stuck count per tenant. Memory must be bounded independently of row count. State your bounds and the field you rank on.
Approach
- Pick the ranking field before the data structure, because this is where the answer is usually lost. Stuck duration is
now - created_at, not anything derived fromnext_attempt_at. Exponential backoff pushesnext_attempt_atfurther out with every failure, so ordering ascending by it surfaces the publications that just failed once and buries the row that has been retrying for six days — the ranking inverts exactly on the rows the operator opened the page for. - For the global top 50 — the 50 smallest
created_at— keep a max-heap of size 50 keyed oncreated_at. Push each row; when the heap exceeds 50, pop the maximum. The heap head is the youngest survivor, so a row older than the head displaces it and everything else is discarded in O(1) after one comparison. That is O(N log K) time and O(K) space, with the log K term only paid on the shrinking fraction of rows that beat the head. - For per-tenant results keep a map from
tenant_idto a size-5 max-heap plus an integer count. Memory is O(T * K') = 5,000 * 5 entries plus 5,000 counters, a few megabytes, and it is bounded by tenant cardinality rather than row count. Say that out loud: if tenants were unbounded this structure is not, and you would need a sketch or a two-pass job. - Do not sort. A full sort is O(N log N) over 20,000,000 rows and must materialise all of them; the heap never holds more than K. Note that a database
ORDER BY ... LIMIT 50reaches for the same bounded heap internally, so the interesting case is precisely the one stated — a stream you cannot re-read and cannot fit. - Carry the diagnostic fields through the heap entry rather than re-joining afterward. An operator needs
attempt_countandlast_error_codenext to the age to tell a receiver that is rate limiting from one that is rejecting the payload, and a second pass to fetch them defeats the single-pass constraint you just paid for. - Exclude
deadrows from these structures and count them separately. They are terminal by definition, so mixing them in lets a permanently dead publication occupy a slot in the top 50 forever while genuinely stuck rows rotate beneath it.
Follow-up
- Two publications share the same
created_atto the microsecond. What is your tiebreak, and why does a deterministic one matter for an operator page that refreshes? - The operator wants to replay the top 50 safely. What must be true of the idempotency key for that to be a no-op when the original attempt actually succeeded?
- Tenant cardinality rises to 5,000,000. Which part of your structure breaks first and what replaces it?
Replace per-retry outbox keys with derived ones on a live table
score_publication holds 400 million rows and takes continuous writes: (publication_id, tenant_id, attempt_id, target, idempotency_key TEXT NOT NULL, score_given, score_timestamp, state, attempt_count, next_attempt_at). The key was minted as a fresh UUID per retry, so the same grade has several rows and duplicate deliveries have already happened. Migrate to a key derived from (attempt_id, score revision) with UNIQUE (tenant_id, target, idempotency_key), online. Give the ordered steps with the lock each takes, how you handle the existing duplicates, and what you do with rows currently in_flight.
Approach
- Add a plain nullable column, derived_key TEXT, with no default: a catalogue-only ALTER held for microseconds. Set lock_timeout to a second or two and retry, because the hazard is the queue rather than the statement — the ALTER waits behind one long reader and every query arriving afterwards stacks behind its pending ACCESS EXCLUSIVE. Resist the elegant version: adding a STORED generated column rewrites all 400 million rows under that same lock on current major versions.
- Deploy the dual write before the backfill so every new publication row populates derived_key from attempt_id and the score revision, while the relay still reads the old column. The backfill then chases a closed id range instead of a moving target, and you can bound it.
- Backfill in batches by publication_id range, on the order of 20,000 rows per statement, committing between batches, recording the high-water mark in its own table so a killed run resumes rather than restarts, and filtering WHERE derived_key IS NULL so a replayed batch is a no-op. Pace on replica lag and dead tuples, not CPU: each UPDATE writes a new row version and WAL volume tracks rows touched.
- Resolve the duplicates as a data decision, not a DELETE. Group by (tenant_id, target, derived_key) having count(*) > 1; within each group keep the row in state='succeeded' if one exists, since that is the delivery whose outcome is known, otherwise the newest 'pending'. Move the rest to state='dead' with last_error_code recording the migration, and keep the rows — support has to be able to explain a grade that posted twice, and a deleted row explains nothing.
- Leave in_flight rows alone and drain them. A row in_flight has an outstanding request whose outcome is unknown, so collapsing it either loses a delivery record or races the relay marking a row you just rewrote. Pause the claim loop or wait out the lease, assert zero in_flight for the affected keys, then finish the group.
- Build the constraint without a blocking scan and expect the first attempt to fail: CREATE UNIQUE INDEX CONCURRENTLY, which aborts if any duplicate remains and leaves an index with indisvalid = false that must be DROP INDEX CONCURRENTLY'd rather than reused. Once it builds clean, ALTER TABLE ... ADD CONSTRAINT ... UNIQUE USING INDEX adopts it under a brief ACCESS EXCLUSIVE without rebuilding — which is a second reason derived_key is a real column, since USING INDEX accepts neither an expression index nor a partial one.
Follow-up
- score_timestamp must strictly increase per line item and learner. Does this migration change what the receiver sees, and what would a stale retry do mid-migration?
- A tenant's external endpoint has been failing for a week and has 40,000 rows queued. What does your duplicate collapse do to that backlog?
- How do you prove afterwards that no grade was delivered twice as a result of the migration itself?
Explain why the grading queue index stopped serving its ORDER BY
attempt holds 2 billion rows and 99.7% are state='graded'. The grading queue runs SELECT attempt_id FROM attempt WHERE tenant_id = $1 AND state IN ('submitted','grading') ORDER BY submitted_at LIMIT 100 against index ix_attempt_queue (tenant_id, state, submitted_at). At the bell the query takes 30 seconds and EXPLAIN shows a Sort above a bitmap heap scan, even though the queue is only a few thousand rows. Explain why that index cannot serve the ordering, give the replacement definition, and say how the replacement behaves as rows leave the queue.
Approach
- Start from what the index actually orders. Its tuples are sorted by (tenant_id, state, submitted_at), so with two state values the scan yields all 'grading' rows in submitted_at order, then all 'submitted' rows in submitted_at order. The query wants one submitted_at order across both, which that physical order does not provide, so the planner must sort the whole matching set and the LIMIT cannot stop the scan early. With a single state constant the same index satisfies the ordering perfectly — the plan changed because of the IN list, not because of the data.
- Replace it with a partial index that moves the predicate out of the key: CREATE INDEX CONCURRENTLY ix_attempt_queue ON attempt (tenant_id, submitted_at) WHERE state IN ('submitted','grading'). Ordering is now satisfied directly, the LIMIT stops after 100 index tuples, and the index covers roughly 0.3% of rows so it stays resident in cache instead of competing with the table.
- State the precondition, because it is the part people skip: the planner uses a partial index only when it can prove the query's predicate implies the index predicate. A literal IN list matching the definition proves; state = ANY($1) with the list bound as a parameter does not, so the queue's state set has to be literal in the SQL, not passed in.
- Fix the estimate separately from the access path. With five states the planner already has each one in its most-common-values list, so single-column selectivity is fine; what it gets wrong is the dependence between tenant_id and state, because a bell means one tenant owns nearly the whole queue while the planner multiplies the two selectivities as if independent. CREATE STATISTICS ON tenant_id, state FROM attempt corrects it, and you ANALYZE explicitly after a bulk transition rather than waiting for autovacuum, whose analyze threshold of 50 + 0.1 x reltuples needs 200 million changed rows on this table.
- Account for the lifecycle. A row leaves the queue by UPDATE state='graded', so its new version is not in the partial index at all while the old index entry remains until vacuum reclaims it. The index therefore churns in proportion to throughput rather than to depth: size it and vacuum it against how many attempts pass through per day, not against how many sit in the queue at once.
- Keep the alternative in your pocket for when the partial index is not acceptable: two explicit scans merged, (SELECT ... WHERE state='submitted' ORDER BY submitted_at LIMIT 100) UNION ALL (SELECT ... WHERE state='grading' ORDER BY submitted_at LIMIT 100) ORDER BY submitted_at LIMIT 100, which keeps the existing index and pushes the LIMIT into both branches. Confirm whichever you choose with EXPLAIN on your major version rather than from memory, since array-key handling in b-tree scans has changed between releases.
Worked solution 25 min
- Reproduce on a scaled copy: 20 million attempts, 0.3% in the queue, and confirm EXPLAIN (ANALYZE, BUFFERS) shows a Sort node above the scan.
- Re-run the identical query with state = 'submitted' alone and observe the Sort disappear, which isolates the IN list as the cause.
- Build the partial index CONCURRENTLY, re-run, and compare buffers touched and actual rows.
- Change the literal list to a bound parameter and confirm the partial index is no longer chosen.
Follow-up
- The queue drains and refills eight times a day. What is the vacuum strategy for this index, and what do you monitor to know it is losing?
- A second consumer claims rows with FOR UPDATE SKIP LOCKED. Does the plan survive, and what does skipping do to the ordering guarantee?
- Support wants oldest-first fairness across tenants rather than within one. Does the same index serve that query?
How do you utilize Bootstrap or other CSS frameworks to build responsi…
How do you utilize Bootstrap or other CSS frameworks to build responsive and accessible user interfaces?
Approach
- State your assumptions explicitly before working the problem.
- Work from the requirement backwards to the design.
- Say what you would check first and why it is the highest-information step.
Follow-up
- What assumption would you test first?
- How would you know your answer was wrong?
What is your approach to debugging a frontend application when a compo…
What is your approach to debugging a frontend application when a component fails to render correctly?
Approach
- State your assumptions explicitly before working the problem.
- Work from the requirement backwards to the design.
- Clarify what is being asked and what a complete answer contains.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
Publish grades outward exactly once under retries and regrades
score_publication holds publication_id, tenant_id, attempt_id, target (lti_ags, webhook, sis_export), idempotency_key with UNIQUE (tenant_id, target, idempotency_key), score_given, score_maximum, score_timestamp, state (pending, in_flight, succeeded, failed_retryable, dead), attempt_count and next_attempt_at. The receiving endpoints are third-party: they rate-limit, they time out after having committed, and they can be down for days. Design the relay: how a row is claimed, what happens when the process dies after the HTTPS call but before the row is marked, how the idempotency key is derived, what stops a stale retry from overwriting a regrade, and which rows a claim is allowed to select. next_attempt_at is the only time column on the row; say what it means in each state.
Approach
- Write the score row and the publication rows in one transaction. That is the entire point of an outbox: a queue publish after the commit can be lost, and a queue publish before the commit can announce a transaction that then rolls back. There must be no window in which a grade exists with no recorded intent to deliver it.
- Claim with a single conditional statement rather than a read then a write, and let next_attempt_at carry two meanings: the backoff time for a row waiting to be sent, and the lease deadline for a row already out. UPDATE score_publication p SET state='in_flight', attempt_count = attempt_count + 1, next_attempt_at = now() + :lease FROM (SELECT publication_id FROM score_publication WHERE state IN ('pending','failed_retryable','in_flight') AND next_attempt_at <= now() ORDER BY next_attempt_at FOR UPDATE SKIP LOCKED LIMIT :n) c WHERE p.publication_id = c.publication_id RETURNING p.*. Both halves are load-bearing. Drop in_flight from the state list and that state becomes a one-way door: a row a process marked in_flight before dying matches no claim again, stalls permanently, and holds its tenant's backlog behind it. Drop the lease write and the row is re-claimable the instant it is claimed, because its next_attempt_at is already in the past, so every publication goes out once per polling worker. SKIP LOCKED is what lets N relay processes share one table without contending on the same head rows; the lease is what makes a crash recoverable. They are separate mechanisms and the design needs both. Index for the predicate with a partial index on next_attempt_at restricted to those three states, or the ORDER BY degrades into a scan over every succeeded row the table has ever accumulated.
- Accept that the transport is at-least-once and make the receiver's side idempotent instead of chasing exactly-once on the sender. A process that dies after the call and before the mark leaves the row in_flight with a lease deadline now in the past, and the next claim picks it up and re-sends it. Size :lease above the HTTP timeout plus a margin: a lease shorter than the call reclaims a request that is still running and sends the same grade twice concurrently. That duplicate is harmless only if the key is derived from (attempt_id, score revision) and nothing else. A UUID minted per retry, or a key containing a clock reading, makes every retry a fresh write at the receiver and turns one grade into several. Increment attempt_count at claim time rather than at completion, so a publication that kills its worker on every attempt walks its own counter to the cap and reaches dead, rather than being reclaimed forever.
- Treat ordering as a separate failure from duplication. A retry of an old score can land after a regrade's newer score, so score_timestamp must strictly increase per line item and learner; the AGS profile directs receivers to reject a non-increasing timestamp with a conflict, but a sender cannot assume every implementation honours it. Also suppress superseded rows at claim time, so a pending publication that is no longer the current revision for that attempt and target is retired rather than sent.
- Classify responses into distinct actions instead of one retry rule. A timeout is an unknown outcome and must be retried with the same key. A conflict on a stale timestamp is success-equivalent and should be marked, not retried. A rate-limit response honours a retry hint when present and otherwise backs off exponentially with jitter under a per-tenant concurrency cap. An expired credential is retryable only after refresh. A permanent rejection, such as a line item that no longer exists, is terminal and moves to dead with last_error_code recorded, because a row retrying forever looks identical to a healthy system on every dashboard.
Worked solution 35 min
- Write the transaction that inserts the score and its publication rows together, and state precisely what is lost if the publish happens outside it.
- Write the skip-locked claim, name every state its predicate selects, set the next lease deadline in that same statement, then trace a process that dies immediately after the HTTPS call through to the claim that picks the row up again.
- Derive the idempotency key from named columns and show two separate retries producing an identical key.
- Build the response classification table: timeout, stale-timestamp conflict, rate limit with and without a hint, credential expiry, line item gone, and server error, with the resulting state and next_attempt_at for each.
- Trace a regrade racing a stale retry and show which score the receiver ends up holding.
Follow-up
- A district's endpoint is down for four days and 200,000 publications accumulate. What happens when it returns, and what protects it from your entire backlog arriving at once?
- An instructor regrades one attempt three times inside a minute. How many outbound calls does the receiver see, and which score does it hold at the end?
- What does a support engineer see for a stuck tenant, and what exactly does a safe manual replay of a dead row do?
Roster apply deadlocks against launches on term-start mornings
In the first week of term, between 07:45 and 08:15, deadlock errors (SQLSTATE 40P01) appear at 20-60 per minute, two roster runs abort with roster_sync_run.error set, some launches return 500, and read p99 doubles even in minutes with no deadlock. A 10^6-row roster apply started at 02:00 and is still running inside one transaction. The apply updates section metadata first and then that section's enrollment rows in file order; the launch path upserts an lti_launch enrollment and then bumps a counter on section. Give an ordered diagnostic checklist, the cycle, and the fix.
Approach
- Order the checklist to separate the two symptoms, because they have different causes and only one is the deadlock. First, pull the deadlock detail from the server log, which names both transactions, both statements and the locks each held and wanted; that is the cycle, printed for you. Second, explain the p99 rise in the minutes with no deadlock at all, which the cycle does not cover. Third, look at transaction duration and lock hold time. Fourth, look at scheduling.
- Read the cycle off the two orders: the apply takes section then enrollment; the launch takes enrollment then section. That is a two-resource cycle in opposite order, so any overlap can deadlock. Postgres detects it after a lock wait exceeds deadlock_timeout, default 1 s, and cancels one transaction with 40P01, which is why the errors are capped at a few dozen a minute rather than unbounded.
- Explain the second symptom separately: a single transaction over 10^6 enrollment rows holds every row lock it has taken until it commits, because row locks are never released early. Any launch touching those sections waits, and the wait is real even when no cycle forms. That is the doubled p99, and it would exist with perfect lock ordering.
- Fix the cycle by removing an edge rather than by reordering both sides. The launch path bumps a counter on section, which is derived data on a latency-critical path; move it to an async aggregate or a cache. With the launch path no longer writing section, the cycle cannot form regardless of the apply's order. Where two writers genuinely must touch both, fix a canonical acquisition order, and have the apply sort its batch by (section_id, person_id) before writing so file order cannot dictate lock order.
- Fix the contention by chunking the apply into bounded transactions, a few thousand rows each, committing between chunks. Lock hold time drops from hours to well under a second, an abort loses one chunk instead of a night, and the run stays restartable from the snapshot digest with per-chunk progress recorded. Retry 40P01 at chunk granularity with jittered backoff, which is safe precisely because the apply is idempotent on the digest. Chunked commits must not weaken the destructive-change gate: it is decided on the full diff before the first chunk commits, never per chunk, so a partially applied snapshot cannot deprovision a bounded fraction of a tenant ahead of the decision.
- Fix the schedule last, as defence rather than as the remedy: cap per-tenant apply concurrency and hold the apply out of the tenant's bell windows from the term calendar. Scheduling alone would hide the bug until a large tenant overran its window again, which is exactly what happened here.
Follow-up
- A chunked apply is no longer atomic. What can a reader observe mid-run, and which of those intermediate states are acceptable?
- Chunk 412 of 800 fails and the run restarts. Prove the restart neither double-applies nor skips rows.
- Raising deadlock_timeout would reduce the error count. Explain precisely what it would make worse.
For someone who has spent the last few years shipping features and reading other people's code, and who has not solved a timed problem from a blank file in a long time. Five days rebuild the primitives and the patterns that sit on them, working from invariants rather than remembered solutions, and the last two attach that back to the rest of the loop.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Rebuild the primitives by implementing them
- Implement a dynamic array with doubling growth and an operation counter, then change the growth rule to add a fixed sixteen slots instead, and time both for n of ten thousand, a hundred thousand and a million. The fixed-increment version resizes n/16 times at O(n) each, so its total work is quadratic; doubling is what makes append amortised constant.
- Implement a hash map with separate chaining and a load-factor resize, then insert ten thousand keys engineered to land in one bucket and record what happens to lookup time, so that average-case O(1) becomes a claim with a stated precondition rather than a reflex.
- For dynamic-array append and hash-map insert, write down which cost is amortised rather than worst-case, which single operation pays the whole bill, and what a system with a hard per-operation deadline would have to do instead.
Deliverable: Two working implementations plus a timing table showing the input at which each structure's advertised complexity stops holding.
Practice prompt ↗Practice prompt ↗Worked solution ↗02Arrays under an invariant: two pointers, sliding window, binary search
- Solve longest-subarray-with-sum-at-most-K using a sliding window, then run it on an input containing negative numbers and watch it return the wrong answer: extending the window only moves the sum monotonically when every element is non-negative, and that precondition is the whole reason the technique works.
- Write the binary search that finds the first index satisfying a predicate rather than an exact value, put the loop invariant above the loop in a comment, and verify termination on the two inputs that break careless versions: the empty range, and a range where every element satisfies the predicate.
- Compute the midpoint as lo + (hi - lo) / 2 and write one line on why the obvious (lo + hi) / 2 is a genuine defect in a fixed-width integer type and a non-issue in a language with arbitrary-precision integers.
Deliverable: Three solved problems, each with its invariant written above the loop, plus one recorded input on which the sliding window is provably wrong.
Practice prompt ↗Practice prompt ↗03Sorting, heaps, and the greedy argument that has to be proved
- Solve one top-k problem three ways, by full sort, by a size-k heap, and by quickselect, then write the values of n and k at which each becomes the right choice, along with quickselect's quadratic worst case and why a randomised pivot makes that unlikely rather than impossible.
- Implement bottom-up heapify and count sift-down steps to confirm it does linear work rather than n log n, because most nodes sit near the bottom of the tree and therefore move only a short distance.
- Take interval scheduling by earliest finishing time and write the exchange argument out in full: given any optimal schedule, swapping in the earliest-finishing interval keeps it feasible and no smaller. Then construct the weighted variant where that same greedy fails and name what has to replace it.
Deliverable: A three-way top-k comparison with measured crossover points, one written exchange argument, and one counterexample to a greedy rule that looks almost identical.
Practice prompt ↗Practice prompt ↗04Recursion, memoisation, and the step to a table
- Take one problem with overlapping subproblems, such as edit distance or coin change, instrument the plain recursion with a call counter to show the blow-up, then add memoisation and re-count.
- Convert the memoised version to a bottom-up table and state the two properties you relied on: each subproblem's result depends only on its arguments, and the dependencies form a DAG you can enumerate in order.
- Rewrite one deep recursion with an explicit stack, then find the input length at which the original hits the interpreter's frame limit, which defaults to about a thousand frames in CPython, so you know when the rewrite is required rather than decorative.
Deliverable: One problem in three forms, naive, memoised and tabulated, with call counts for each and the input length at which recursion depth becomes the binding constraint.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Graphs, where most of the work is choosing the traversal
- Implement BFS and DFS over one adjacency list, then answer for each which finds a shortest path in an unweighted graph and which you would use to detect a cycle in a directed graph, including why the in-progress versus finished distinction matters for the second.
- Implement topological sort by in-degree, feed it a graph containing a cycle, and confirm the failure signature is that fewer than V nodes come out rather than an exception, then note that the order it produces is one of several valid ones.
- Run a shortest-path search on a graph with a single negative edge weight and show the wrong answer, then write the precondition Dijkstra actually needs, non-negative weights, because it finalises a node's distance the first time that node is popped, and name the algorithm you would switch to and its own limit.
Deliverable: A small graph library with BFS, DFS and topological sort, plus two inputs that produce documented wrong answers under the wrong algorithm choice.
Practice prompt ↗Practice prompt ↗06One day for everything that is not an algorithm
- Sketch one system only to the depth a coding-heavy loop tends to reach: the endpoints, what the service stores, and the single query pattern that decides the schema. Stop at twenty-five minutes.
- Prepare the project answer for an interviewer who codes, which means rehearsing the two levels they push to: the specific thing you built, and why you chose that approach over the alternative they will name. Open with a number and be ready to say what it excludes.
- Prepare the answer to what you would do differently, choosing a real technical mistake with a specific fix rather than a complaint about process or staffing.
Deliverable: One design sketch at endpoint-and-schema depth, plus a project answer rehearsed to two levels of follow-up.
Practice prompt ↗Practice prompt ↗07Solve out loud, under time
- Do three timed problems at twenty-five minutes each in a plain editor with no autocomplete and no execution until the end, then tally separately the failures that were syntax and the ones that were approach, because those two numbers call for different fixes.
- Narrate one solution from the first sentence, stating the approach and its complexity before writing any code, and rehearse the sentence you will use when you realise mid-solution that the approach is wrong.
- Re-solve from blank the two problems you were slowest on this week and compare the times against the day they first appeared.
Deliverable: A recording of one fully narrated solution and a tally that separates syntax failures from approach failures.
Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Keep one story where the bad call was yours rather than a dependency's or a manager's. Name the check that would have caught it, whether you added that check afterwards, and whether it has fired since. Answers that route blame outward end the conversation early; answers that end in a guardrail someone still relies on tend to open it up.
What are your top three professional strengths, and how do you apply t…
What are your top three professional strengths, and how do you apply them to your development work?
Approach
- State the situation in two sentences and spend the rest on the reasoning.
- Name the disagreement and how you resolved it with evidence.
- Give the blast radius: what could have broken, and what you measured.
Follow-up
- What would you do differently if you ran that again?
- What did you decide not to do, and why?
What is your experience with modern frontend frameworks like ReactJS o…
What is your experience with modern frontend frameworks like ReactJS or Angular?
Approach
- Give the blast radius: what could have broken, and what you measured.
- Name the disagreement and how you resolved it with evidence.
- Close with what you would do differently, concretely.
Follow-up
- How did you know your change caused the improvement?
- What did you decide not to do, and why?
How do you handle a heavy workload, particularly if you need to balanc…
How do you handle a heavy workload, particularly if you need to balance professional commitments with coursework or external academic responsibilities?
Approach
- Give the blast radius: what could have broken, and what you measured.
- Name the disagreement and how you resolved it with evidence.
- Pick a story where you made the decision, not one where you watched it.
Follow-up
- What would you do differently if you ran that again?
- What did you decide not to do, and why?
Can you describe a time when you had to work closely with non-technica…
Can you describe a time when you had to work closely with non-technical stakeholders, such as academic advisors or department heads, to deliver a software solution?
Approach
- Close with what you would do differently, concretely.
- Pick a story where you made the decision, not one where you watched it.
- Give the blast radius: what could have broken, and what you measured.
Follow-up
- How did you know your change caused the improvement?
- What did you decide not to do, and why?
- 01
What are your top three professional strengths, and how do you apply them to your development work?
- 02
What is your experience with modern frontend frameworks like ReactJS or Angular?
- 03
How do you handle a heavy workload, particularly if you need to balance professional commitments with coursework or external academic responsibilities?
- 04
Can you describe a time when you had to work closely with non-technical stakeholders, such as academic advisors or department heads, to deliver a software solution?
Is this an official Drexel University interview guide?
No. It is PracHub's own research and practice material for the Software Engineer role at Drexel University. Rounds and questions reflect what candidates have reported, not a process Drexel University has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult is the Software Engineer interview process at Drexel University?
The interview process is generally rated as easy to average. The focus is much more on your practical web development skills, problem-solving approach, and work ethic rather than highly theoretical or competitive coding challenges.
PracHub interview research ↗What is the typical timeline from the first interview to an offer?
The process is known to be relatively fast and smooth. Depending on the department's urgency, candidates often progress from the initial application review to a final decision within a few weeks.
PracHub interview research ↗How does the co-op system impact the Software Engineer role?
If you are hired as a student or co-op software engineer, Drexel expects you to maintain your academic coursework while gaining hands-on, real-world experience. If you are a full-time staff engineer, you will likely work alongside and help mentor these talented co-op students.
PracHub interview research ↗What frameworks and technologies are most commonly used?
Most teams at Drexel utilize standard web technologies, focusing heavily on frontend frameworks like ReactJS and Angular, styled with Bootstrap or similar CSS libraries, backed by MVC-structured backend systems.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Software Engineer practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22