Headway works on universal access to affordable mental healthcare by removing the financial and administrative barriers that keep therapists from accepting insurance. Software Engineers are described as building the product and infrastructure behind this: digitising and automating insurance billing, claims reconciliation, credentialing and provider practice operations, across areas named Provider tools, Payer integrations and Trust Foundations.
Example work described for the role includes low-latency search that connects patients with care providers, access-control platforms (RBAC/ABAC) that protect sensitive healthcare data, and REST APIs that connect to legacy insurance carriers. The languages and frameworks named are Python, TypeScript/JavaScript, Java and React, with cloud deployments that have to meet HIPAA compliance and reliability requirements together.
The reported interview questions follow the same practical pattern. The coding questions cover appointment-slot scheduling, markdown-to-HTML conversion, fixing and extending existing code, simulation and rate-limiting middleware. The design questions cover a location- and date-range rental search, a paginated REST API with JWT auth, diagnosing a traffic-spike slowdown, capacity estimation and a walkthrough of a system you built. Prepare to write correct, readable code for everyday business logic, and to explain data models and trade-offs out loud.
Recruiter Screen
reportedCandidates describe the recruiter screen as a conversation to line up your background, career goals, location expectations, team openings and compensation expectations. Candidates also report that the order of rounds and any specialised screens can vary by team, with Trust Foundations and Provider engineering given as examples. Use this call to find out which team the loop is for and what the next technical step looks like, so you can plan your preparation around it.
What to demonstrate
- Whether your background and career goals match a team that currently has an opening
- Whether your location and compensation expectations fit the role
- Whether you can give a clear, short account of your recent work
How to prepare
- Ask which team the loop is for and whether the technical screen is in-house or on an external platform, since candidates report both
- Prepare a two-minute summary of recent work that points to APIs, data modelling, integrations, access control or scheduling-style product logic where you have it
- Settle a compensation range and your location constraints before the call so you can answer directly
Technical Screening
reportedThe technical screen is reported as a live coding challenge on practical algorithms or systems integration. It runs either in-house or on an external technical platform, and covers practical coding logic or system design fundamentals. The reported coding questions for this role involve interval and time arithmetic, text parsing, simulation, and request-limiting middleware, so practise those problem types by writing runnable code while explaining your reasoning. Do not start coding before you have clarified input constraints and edge cases with your interviewer.
What to demonstrate
- Whether you clarify inputs, constraints and edge cases before writing code
- Whether the code runs correctly on boundary cases such as the start and end of a time window, empty input and touching intervals
- Whether you explain your reasoning out loud so the interviewer can follow and work with you
How to prepare
- Solve the appointment-slot question with explicit half-open intervals, then test an empty calendar, a fully booked day and a gap exactly equal to the requested duration
- Build a rate limiter with an injectable clock so you can show window resets and rejections in a test without sleeping
- Confirm with the recruiter whether the screen is coding or design fundamentals, and which languages are allowed
Virtual Onsite Loop
reportedThe virtual onsite is reported to combine coding, system architecture and behavioral assessments. Candidates say it can run as one multi-hour block or be split across two consecutive days. Candidates who pass the screen reportedly receive preparation materials describing what the final loop expects, so read them closely and adjust your plan. Reported design questions for this role, which candidates do not tie to a particular round, include a rental or hotel search optimised by location and date range, a REST API with pagination and JWT auth, diagnosing slowdowns under a request spike, capacity estimation, and walking through a production system you built.
What to demonstrate
- Whether design answers start from the domain data model and the actual queries (date ranges, geography, pagination) before adding infrastructure
- Whether you can explain the data flows, storage trade-offs and scaling bottlenecks of a system you built
- Whether you state assumptions and confirm requirements with the interviewer instead of designing alone
- Whether your coding stays clean and correct across more than one problem
How to prepare
- Read the preparation materials sent after the screen and list the topics they name next to this guide's question categories
- Diagram one backend project from the last two years: its boundaries, schema and trade-offs, with no confidential details
- Practise the rental-search design by writing the availability and location queries first, then choosing indexes, caching and how double-booking is prevented
- Practise a capacity estimate from stated access patterns, saying each assumption and number out loud
Behavioral Interview
reportedThe behavioral interview is reported as a structured 'career chapters' round, also called the WHO interview. You lead by presenting your work history as chronological chapters, and the conversation goes into the intent behind each move, what you achieved, mistakes you made and what you learned. Prepare the chapters as a structured narrative, not a list of jobs. Other reported behavioral questions for this role, which candidates do not tie to a particular round (the onsite also includes behavioral assessments), cover resolving a conflict, a major architectural trade-off or mistake, explaining technical constraints to non-technical stakeholders, and mentoring or improving developer tooling.
What to demonstrate
- Whether each career move has a clear reason and a result you can explain
- Whether you take ownership of mistakes and can say what you changed afterwards
- Whether you can explain technical trade-offs to product managers and other non-engineering partners
How to prepare
- Write your career as three or four chapters, each with the goal, what you owned, a concrete result, a regret and why you moved on
- Choose one real architectural mistake or trade-off and be ready to cover its impact, how you found out and how you fixed it
- Prepare one mentoring or developer-tooling example and one conflict you resolved, each with your own actions named explicitly
Reference Checks
reportedCandidates report reference checks with past managers and colleagues as the last step before an offer. Pick references who worked with you on the projects you plan to discuss and who can speak to your technical abilities and work ethic. Arrange them early so this step does not hold up an offer.
What to demonstrate
- Whether past managers and colleagues can speak to your technical ability and how you work
- Whether you can provide suitable references promptly when asked
How to prepare
- Line up references who worked with you recently, ideally a manager and a peer, before the onsite
- Pick references who worked on the projects you plan to discuss, and tell each one about the role so they know what kind of work to speak to
7 candidate reports. Individual accounts describe a particular role and hiring cycle.
Headway Software Engineer interview: withdrew after a delayed process
I initially had an easier path into the process, but it became frustrating quickly. The recruiter seemed unprepared, and the timeline took too long to get moving. Eventually, I decided it wasn’t worth my time and withdrew to focus on other interviews. The process felt unprofessional to me. I didn’t like how long things dragged before anything substantial happened. Because I withdrew and didn’t co…
Read full experienceHeadway Customer Success Engineer interview with an abrupt manager round
I was contacted through LinkedIn, and someone from the recruiting team arranged a recruiter screen. That interview was fairly normal, covering my work experience along with an overview of the company and role. Afterward, I moved straight to a hiring manager interview that felt very different. The hiring manager barely introduced themselves, didn't ask much about me, and instead went straight to a…
Read full experienceHeadway Software Engineer interview with role homework and employee interviews
I went through a multi-step process that felt unusual in how quickly it ramped up. It started with a recruiter meeting where I answered basic questions. Then I was assigned homework tied to the role, which I had to complete before the next technical conversation. After that, I had a technical interview with the hiring manager. The format shifted again when I interviewed with current employees. Th…
Read full experienceHeadway Software Engineer interview: unclear onsite design expectations
My process started with a recruiter touchpoint, followed by a technical screen and then a three-round onsite. Scheduling and feedback were responsive, but the actual evaluation didn't line up with what I expected. The onsite design segment especially didn't feel like a classic system design round. The constraints made it hard to tell what direction they wanted, so I spent more time guessing than…
Read full experienceHeadway Backend Engineer interview: recruiter screen and unclear follow-up
My process started with a quick recruiter screen focused on behavioral questions. We covered the usual background topics, along with how I’d troubleshoot technical issues when something breaks and how I’d explain a technical concept to someone without a technical background. I made it into the hiring process, but it ended without an offer. In one experience, the recruiter interaction had a good t…
Read full experiencePracHub editorial advice for the preparation topics above.
Getting the edges wrong in appointment-slot code: the 8 AM start, the 5 PM end, and back-to-back bookings
Before writing code, say whether intervals are half-open [start, end). Then check three kinds of gap: from 8 AM to the first appointment, between each pair of appointments, and from the last appointment to 5 PM. A meeting ending exactly at 5 PM fits, and appointments that only touch do not overlap. Ask whether the input can contain overlapping entries even though it is sorted, and merge them if so. Test an empty calendar, a fully booked day and a gap exactly equal to the duration. Time-slot logic is a recurring theme in the reported questions, so check these edge cases explicitly and out loud.
Spending the global quota on a request that the per-endpoint limit then rejects
The reported middleware has a global cap of 100 requests per minute and a per-endpoint limit of 3 requests per second. Check both windows first, and increment the counters only if both allow the request. Otherwise rejected calls use up the global budget and later valid requests fail. Say how windows are aligned and that fixed windows allow a burst across a boundary. Inject the clock so tests can move time forward, and say whether the counters live in memory or in a shared store once there is more than one instance.
Rewriting an inherited codebase instead of finding and fixing the bug
The reported best_of_bests question involves fixing bugs in existing code, and the bank has a matching 'debug and extend obstacle-run statistics classes' question. Read the existing code and any tests first. Reproduce the failure with the smallest input you can, explain the cause in one sentence, then make the smallest fix that keeps the existing interface. Extend the code only once the fix passes, and say why each change was needed.
Drawing service boxes for a search or booking design before writing the location, date-range and conflict queries
For the rental-search and booking questions, write the read path first: filter by location, filter by availability over a date range, and paginate. Then choose the data model and indexes that serve those queries, and add caching or read replicas only where you can point to the bottleneck. For bookings, say how two concurrent requests for the same slot are prevented, for example a unique constraint or a conditional update. A check followed by a separate insert passes single-request tests and fails under concurrency.
Turning the career-chapters round into a readout of your résumé
For each chapter, give the intent behind the move, what you owned, one concrete result, one mistake or regret, and what you learned. A list of employers and technologies does not give the interviewer the reasoning this round is built around. Keep facts such as team size, timelines and your own role the same every time you tell a story, across every round.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Given a sorted list of appointments, write a function that finds an av…
Given a sorted list of appointments, write a function that finds an available slot for a meeting of a specified duration between 8 AM and 5 PM.
Approach
- Walk one small example through your approach before writing the whole thing.
- Restate the input: its shape, its size, and what is guaranteed about it.
- Name the brute-force solution and its complexity before improving on it.
Follow-up
- Which test case would catch an off-by-one here?
- How does this change if the input no longer fits in memory?
Write a function that converts raw markdown text into valid HTML, incl…
Write a function that converts raw markdown text into valid HTML, including support for nested lists, combined elements, and line breaks.
Approach
- Name the brute-force solution and its complexity before improving on it.
- State the target complexity and say which constraint rules the naive version out.
- Walk one small example through your approach before writing the whole thing.
Follow-up
- What is the worst case, and how likely is it on real data?
- Which test case would catch an off-by-one here?
Write a simulation function `chance_of_personal_best` to determine the…
Write a simulation function chance_of_personal_best to determine the statistical likelihood of an athlete achieving a personal best time.
Approach
- State the target complexity and say which constraint rules the naive version out.
- Restate the input: its shape, its size, and what is guaranteed about it.
- Choose the data structure from the access pattern, not from familiarity.
Follow-up
- How does this change if the input no longer fits in memory?
- Which test case would catch an off-by-one here?
Implement a function `best_of_bests` to process race course data and c…
Implement a function best_of_bests to process race course data and calculate the fastest times for each obstacle, handling bug fixes in an existing codebase.
Approach
- Name the brute-force solution and its complexity before improving on it.
- Walk one small example through your approach before writing the whole thing.
- State the target complexity and say which constraint rules the naive version out.
Follow-up
- How does this change if the input no longer fits in memory?
- What is the worst case, and how likely is it on real data?
Order a job dependency graph and find its critical path
A workspace defines up to 50,000 jobs with up to 200,000 dependency edges and an estimated duration_seconds per job. Given the edge list, reject the graph if it contains a cycle and name one cycle's nodes; otherwise return a valid execution order, the earliest possible completion time with unlimited workers, and the set of jobs whose slack is zero. Then say which single job to shorten in order to cut the completion time, and by exactly how much. State the complexity of each part.
Approach
- Kahn's algorithm for the order: compute indegrees, seed a queue with zero-indegree nodes, emit and decrement. O(V + E), which at 50,000 and 200,000 is milliseconds. If fewer than V nodes are emitted, the graph contains a cycle.
- Kahn detects a cycle but cannot name one. The nodes left with indegree above zero contain every cycle, so run one DFS restricted to that residual subgraph with three-colour marking and report the stack slice from the grey node the back edge points at. That is the difference between a usable error message and 'dependency cycle detected'.
- Earliest completion with unlimited workers is the longest path, which is NP-hard on a general graph and linear on a DAG. State the precondition, then relax in topological order:
earliest_finish[v] = duration[v] + max(earliest_finish[u] for u in preds(v)), taking the max over an empty predecessor set as zero. The makespan T is the maximum over all nodes. O(V + E). - Second pass in reverse topological order for
latest_finish, thenslack[v] = latest_finish[v] - earliest_finish[v]. Zero-slack nodes form the critical path, and there can be several disjoint critical paths, so return the set rather than one chain.slack[v] = 0is exactly the statement that some longest path runs through v; equivalently, the longest path through v has lengthT - slack[v]. - The speed-up bound is the point of the question, and the obvious form of it is wrong. Shortening a zero-slack job v by d, with 0 <= d <= duration[v], cuts the makespan by
min(d, T - L_avoid(v)), whereL_avoid(v)is the longest path in the graph with v deleted: the longest path that avoids v, not the second-longest path overall. The two coincide only when the runner-up path misses v. Counterexample: A of 10 s feeds both B of 5 s and C of 4 s, so T = 15 s and the second-longest path is 14 s, yet shortening A by 10 s leaves a makespan of 5 s. The realised gain is the full 10 s, because both paths ran through A and shrank together, whilemin(10, 15 - 14)predicts 1 s. The reason is structural: shortening v reduces every path through v by d and leaves every other path alone, so the new makespan ismax(T - d, L_avoid(v)). - Compute
L_avoid(v)the direct way: delete v and re-run the same forward relaxation, O(V + E) per candidate. The cheaper equivalent skips the deletion, sinceL_avoid(v)only ever matters through that max: setduration[v] := 0, recompute the makespan asT0(v) = max(T - duration[v], L_avoid(v)), and the gain ismin(d, T - T0(v)), which is identical for every d <= duration[v]. Only zero-slack jobs are candidates, because shortening a job with positive slack changes the completion time not at all. One relaxation is milliseconds at this size, so ranking a critical set in the hundreds costs O(k(V + E)) and is worth doing exactly; a critical set in the tens of thousands is not, and there you evaluate a shortlist, longest jobs first, and say that the answer is the best of that shortlist rather than the optimum.
Worked solution 30 min
- Build four fixtures. A: 12 jobs, two branches of 100 s and 95 s that share no job. B: fixture A plus one back edge. C: two disjoint paths tied at 100 s. D: the shared-prefix case, one job of 10 s feeding a 5 s job and a 4 s job, so the longest path is 15 s and the runner-up is 14 s.
- Run Kahn; on fixture B confirm it emits fewer than V nodes, then run the residual-subgraph DFS and print the actual cycle.
- Compute
earliest_finishforward andlatest_finishbackward, and list the zero-slack set for each fixture. - For each zero-slack job v, recompute the makespan with
duration[v] := 0to getT0(v), and record both the correct boundT - T0(v)and the wrong one,T - second_longest_path, side by side. - Apply the shortening for real (20 s off the critical branch of A, 10 s off the shared prefix of D) and diff the recomputed makespan against each prediction.
Follow-up
- Only m workers are available. What happens to your answer, and what can you still promise about the schedule you produce?
- Edges arrive incrementally as the customer edits the pipeline. How do you detect a cycle at insert time without re-running Kahn over 250,000 elements?
- Durations are estimates. How would you express completion time as a distribution, and what breaks about the critical path once you do?
Enforce a concurrent-run quota that survives simultaneous requests
A plan allows at most 20 concurrently running rows in job_run per tenant. The table holds run_id, tenant_id, workspace_id, status (queued, leased, running, succeeded, failed, timed_out, cancelled, lost), lease_token, leased_until, started_at and finished_at. Today the service runs select count(*) from job_run where tenant_id = $1 and status = 'running', compares the result to 20, then inserts. Under load a tenant exceeds the cap by exactly the number of concurrent requests. Name the anomaly, say which isolation levels do and do not prevent it, and give a version that holds, as SQL.
Approach
- Name it: write skew. Each transaction reads a predicate (the count of running rows), neither modifies what the other read, and both then insert rows that jointly violate an invariant no single row expresses. Read committed permits it. So does repeatable read, because snapshot isolation's first-updater-wins check fires only on conflicting row updates, and these are inserts touching disjoint rows.
- Enumerate the fixes with their real costs. SERIALIZABLE works: PostgreSQL's SSI tracks the predicate read and aborts one transaction with SQLSTATE 40001, which obliges the caller to retry and makes the abort rate rise with contention on a hot tenant. Folding the predicate into the write as
insert ... select ... where (select count(*) ...) < 20narrows the race to the statement's snapshot but does not close it under read committed. - Give the version that holds at read committed: serialise on a row both transactions must touch.
update tenant_concurrency set running = running + 1 where tenant_id = $1 and running < 20 returning runningupdates zero rows when the cap is reached, and zero rows is the rejection. This works because at read committed a blocked UPDATE re-evaluates its WHERE clause against the newly committed row; at repeatable read the same statement raises a serialisation error instead, so the isolation level changes the calling contract. - State the cost you just bought. That row is now a per-tenant serialisation point, so admission throughput for the tenant is bounded by one divided by the lock hold time; at a 2 ms hold that is roughly 500 admissions/second. Keep the critical section to the single UPDATE, with no network call or scheduling decision inside the transaction, and decrement in the same transaction that writes the terminal status.
- Close the leak the status enum implies: a run can end as
lost, so a crashed worker otherwise consumes a slot forever. Reconcile on a schedule againststatus = 'running' and leased_until < now(), and treat the counter as a fast path overjob_run, which stays the system of record.
Worked solution 25 min
- Seed a tenant with 19 running rows, then fire 8 concurrent sessions each running the select-then-insert, and count the resulting running rows.
- Repeat at REPEATABLE READ and confirm the count still exceeds 20.
- Repeat at SERIALIZABLE, count the 40001 aborts, and note that without a retry loop those requests fail rather than queue.
- Implement the atomic counter UPDATE, re-run the 8-way test, and confirm exactly 20 running rows with zero over-admissions.
- Kill a worker mid-run, let the lease expire, and check whether the slot comes back without intervention.
Follow-up
- Write the retry loop for the SERIALIZABLE version. What does the caller see when it keeps aborting, and what bounds the retries?
- Two regions each keep a counter. What is the effective cap, and what does admission do when the counter store is unreachable?
- The cap changes mid-flight on a plan upgrade. Do running jobs get killed, and what does the counter row look like during the change?
Migrate a live partitioned event table without blocking ingest
usage_event is range-partitioned daily on ingested_at, holds roughly 250M rows per day across 400 live partitions, and is written at 10-40k rows/second. Two changes are required: quantity must move from double precision to numeric(20,6), and a new environment column must become NOT NULL with a default of 'production'. Ingest cannot stop. Give the ordered plan, naming for each step the lock it takes, what that lock blocks, and roughly how long it is held. Identify the one step that cannot be rolled back cleanly once traffic depends on it.
Approach
- Classify the two changes before planning anything. Adding a column with a non-volatile default has been metadata-only since PostgreSQL 11, so it is cheap. Changing double precision to numeric is not binary-coercible, so
alter column ... typerewrites every partition under ACCESS EXCLUSIVE and rebuilds its indexes; on this volume that is hours of blocked ingest and is simply not an option, which is why the plan is expand-and-contract rather than one statement. - Expand: add
quantity_numeric numeric(20,6)andenvironmentwith its default on the parent. Both are catalogue-only but both take a brief ACCESS EXCLUSIVE that cascades to partitions, so run each withlock_timeoutset to a second or two and retry on failure. A queued ACCESS EXCLUSIVE request blocks every reader behind it, which is how a metadata-only change turns into an outage. - Dual-write: deploy producer code that populates both columns on every insert, and leave it running before anything reads the new column. This is the step that cannot be reverted cleanly. Once readers depend on quantity_numeric, reverting the writer leaves rows with a null there, and the gap is only discoverable by re-reading the old column, which the readers have stopped doing.
- Backfill older partitions in batches keyed on the primary key, oldest first, committing every few thousand rows with a pause between batches, and skipping the partition still receiving writes until it rotates. Each batch is an ordinary UPDATE taking row locks only. The cost is bloat and WAL rather than blocking, so watch dead tuples and let autovacuum keep pace instead of wrapping 400 partitions in one transaction.
- Make NOT NULL cheap with the three-step form:
add constraint ... check (environment is not null) not valid(brief ACCESS EXCLUSIVE, no scan), thenvalidate constraint(SHARE UPDATE EXCLUSIVE, scans while reads and writes continue), thenset not null, which from PostgreSQL 12 uses the validated check and skips its own full scan. Do this per partition, then on the parent. - Switch and contract: move reads to the new column behind a flag, verify over a full period that both columns agree on freshly written rows, drop the old column (metadata-only), and only then remove the dual-write. Any index on the new column goes on with CREATE INDEX CONCURRENTLY per partition, since CIC is not supported on a partitioned parent: create the parent index with ONLY, build each child concurrently, then ALTER INDEX ... ATTACH PARTITION until the parent index becomes valid.
Follow-up
- A CREATE INDEX CONCURRENTLY fails halfway through the partition list. What state is the table in, how do you detect it, and what do you run?
- The producer computes quantity itself. What happens to a request already in flight when the dual-write deploy lands, and does it matter?
- Give two queries that prove the backfill is complete: one cheap enough to run every minute, one authoritative.
Explain how you would diagnose and resolve sudden server slowdowns cau…
Explain how you would diagnose and resolve sudden server slowdowns caused by a massive spike in incoming HTTP requests.
Approach
- Design the error taxonomy before the success shape; callers branch on it.
- Separate accepted, pending, failed and confirmed; they are different facts.
- Define the identity of a request so a retry cannot double-apply it.
Follow-up
- What does a partial failure look like to the caller?
- How does a client discover it is on an old version of this contract?
Design and implement a REST API for a multi-page web application, incl…
Design and implement a REST API for a multi-page web application, including data modeling, pagination, and JWT-based authentication/authorization.
Approach
- Say who the caller is and what they do when the call fails halfway.
- State how the contract changes without breaking existing clients.
- Separate accepted, pending, failed and confirmed; they are different facts.
Follow-up
- How does a client discover it is on an old version of this contract?
- What happens if the caller retries after a timeout?
Design a machine-readable error contract for the gateway
The edge gateway serves roughly 30k requests/second to SDKs and CI pipelines that retry automatically. Today every failure returns 500 with a prose message that clients string-match on. Design the error contract: the response body fields, and the status code for a malformed body, a revoked credential, a scope the credential lacks, a row belonging to another tenant, a reused idempotency key sent with a different body, an exceeded rate limit, and an unreachable dependency. For each, state whether the client may retry and on what schedule. Deliverable: the envelope schema plus the status-to-retry table.
Approach
- Split the envelope by audience: a stable
codestring for programs, amessagedocumented as human-only and free to change, arequest_idthat joins to gateway logs, and adetailsarray for per-field problems. The code list is an enum that only ever grows. - Assign status by who has to change something: 400/422 for the caller's bytes, 401 for a credential that no longer authenticates, 403 for a scope or entitlement, 404 rather than 403 for a row in another tenant because 403 confirms the identifier exists, 409 for an idempotency conflict, 429 for a limit, 503 for a dependency.
- Derive retryability from the method and the idempotency key rather than from the status: a 5xx or a timeout is an unknown outcome, not a failure, so GET/PUT/DELETE may be retried under HTTP semantics and POST only when it carries an idempotency key.
- Put the schedule in the response: Retry-After on 429 and 503 overrides the client's own backoff; otherwise capped exponential backoff with full jitter, sleeping uniformly in [0, min(cap, base * 2^attempt)], bounded by a total attempt budget so retries expire before the caller's deadline.
- Write the negative rules into the published contract: clients must never parse
message, must tolerate unknowncodevalues by falling back to the status class, and a code's meaning is never redefined once shipped.
Worked solution 20 min
- Write the envelope as a JSON schema with four top-level fields and say which are guaranteed present on every error.
- Fill a seven-row table: condition, status, code string, retryable yes/no, and the schedule or the reason retrying cannot help.
- For each non-retryable row, write the one thing the caller must change (bytes, credential, plan, key) so nothing is marked non-retryable without a remedy.
- Add the unknown-outcome row for timeouts and 5xx separately from the other rows, and give it an action other than 'treat as failed'.
- Write two sentences of client guidance: honour Retry-After when present, apply full jitter otherwise, and stop at the attempt budget.
Follow-up
- A customer reports they retried a 500 from POST /v1/runs and ended up with two sandboxes billed. Whose bug is it, and what in your contract permits their reading?
- You need to add a new error code next quarter without a version bump. What did the v1 contract have to say for that to be non-breaking?
Invoice detail latency triples after an ORM relationship refactor
An invoice detail endpoint returned in 40 ms at p99 last week. After a refactor replaced a hand-written join with ORM relationship access it returns in 1.4 s, and the regression grows with the number of invoice_line_item rows on the invoice. Database CPU rose, but no statement in the slow-query log exceeds 3 ms. You have request traces with per-span SQL, the ORM statement log, and a staging copy of the data. Produce an ordered diagnostic checklist, the measurement that confirms the cause before any code change, and the fix.
Approach
- Count statements per request before reading any statement duration. A slow-query log hides this class by construction, because every individual query is fast and only their number is wrong; take one trace and count SQL spans.
- Establish proportionality rather than asserting it: sample invoices with 5, 20, 60 and 200 line items and plot statements per request against line count. A straight line of slope 1 through an intercept of one or two identifies a lazy relationship load, and no index or cache would move that line.
- Locate the emitting attribute access in the refactored code and check whether the same shape repeats one level deeper, for instance a tax or adjustment collection hanging off each line, which turns the cost quadratic.
- Fix with a bounded statement count: either one join that fetches invoice and lines together, or two statements where the second is WHERE invoice_id = $1 AND tenant_id = $2. Keep tenant_id in the predicate so the read stays tenant-scoped even though invoice_id already implies it.
- Choose between the two deliberately: the join duplicates the wide parent row across N children on the wire, the two-statement form avoids that for one extra round trip. Prefer the join for narrow parents and the split for wide ones.
- Pin it with a per-request statement-count assertion in a test that varies line count, because a latency assertion passes on a small fixture and would not have caught this.
Follow-up
- The endpoint now also needs per-line tax rows. Show the shape that keeps statement count constant instead of reintroducing the same defect one level down.
- How does this change if a transaction-pooling proxy sits between the service and the database, so each statement may land on a different backend session?
- The same page paginates invoices with LIMIT and OFFSET. Why is that a second, independent defect, and what replaces it?
Day one measures instead of guessing, under a fixed rubric, and the remaining hours are allocated in proportion to the gaps before any studying begins. The allocation is deliberately not renegotiated midweek, because the area that feels worst on day three is usually the one that is moving.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Scheduling and interval logic
- Solve the reported appointment-slot question: given sorted appointments, find a slot of a given duration between 8 AM and 5 PM, using half-open intervals
- Extend it to find the earliest start after merging overlapping busy intervals, then to choose the narrowest gap that fits the duration
- Write tests for an empty calendar, a fully booked day, touching appointments and a gap exactly equal to the duration, and run them
Deliverable: A tested slot-finder with a written list of the boundary cases it handles.
Practice prompt ↗Practice prompt ↗Worked solution ↗02Parsing, simulation and fixing existing code
- Implement the reported markdown-to-HTML converter for nested lists, combined elements and line breaks, starting from a list of input cases
- Write the reported chance_of_personal_best simulation and state how the number of trials affects the estimate's stability
- Take code you did not write, add a deliberate bug to per-obstacle minimum logic like the reported best_of_bests question, then find and fix it by reproducing it first
- Work through the invoice-latency N+1 query diagnosis practice question, counting queries before looking at their durations
Deliverable: Three working functions with tests, plus a one-paragraph bug report: symptom, smallest repro, cause, fix.
Practice prompt ↗Practice prompt ↗03Rate limiting and traffic spikes
- Build middleware that enforces a global limit of 100 requests per minute and a per-endpoint limit of 3 requests per second, admitting a request only when both windows allow it
- Test it with an injectable clock across window boundaries, and confirm that rejected requests do not use up the global quota
- Answer the reported traffic-spike question out loud: how you would diagnose and fix sudden slowdowns under a large spike in HTTP requests, from the first metric you check to shedding load
- Complete the error-contract worked exercise, focusing on 429, Retry-After and which requests are safe to retry
Deliverable: A tested dual-window rate limiter and a written diagnosis runbook for a request spike.
Practice prompt ↗Practice prompt ↗04REST API design and data correctness
- Design the reported REST API question: an API for a multi-page web app with its data model, cursor or offset pagination and JWT-based authentication and authorization, then write out the endpoints and status codes
- Sketch a therapist-booking API and schema that rejects two concurrent bookings for the same slot, and name the constraint or statement that enforces it
- Complete the SQL concurrency worked exercise on concurrent quota enforcement, and connect its check-then-insert race to double-booking
Deliverable: An endpoint list with pagination and auth rules, plus a schema that prevents booking conflicts, with the enforcing SQL.
Practice prompt ↗Practice prompt ↗Worked solution ↗05System design: search, capacity and pipelines
- Design a rental or hotel search optimised for location and date range: write the queries first, then indexes, caching and how the design changes from startup scale to a larger one
- Estimate server and database capacity for fast user growth from stated access patterns, saying every assumption out loud
- Outline a pipeline that ingests and reconciles billing claims from several payers, covering idempotent ingestion and how mismatches get resolved
- Work through the dependency-graph worked exercise to practise explaining complexity and trade-offs
Deliverable: One search design diagram with its query list, and a one-page capacity estimate.
Practice prompt ↗Practice prompt ↗06Project walkthrough and career chapters
- Diagram a backend system you built in the last two years: data flows, storage choices, scaling bottlenecks, and what you would change, with no confidential details
- Write your career as chronological chapters, each with intent, what you owned, a result, a mistake and a lesson, for the reported career-chapters prompt
- Prepare the reported architectural trade-off or mistake story, plus one conflict, one explanation to a non-technical stakeholder, and one mentoring or tooling example
- Prepare the fail-open design disagreement practice prompt as practice for arguing against a design with numbers rather than opinion
Deliverable: A system diagram and a one-page career-chapters outline, with one story linked to each reported behavioral prompt.
Practice prompt ↗Practice prompt ↗07Mock loop and reference prep
- Run a mock that covers one coding question from days 1-3, one design question from days 4-5, and the career-chapters walkthrough, with a partner who interrupts and asks follow-ups
- After each part, write down the edge case, assumption or trade-off you did not say out loud, and redo that part
- Confirm your references, choosing people who worked with you on the projects you plan to discuss, and tell them about the role
Deliverable: A list of gaps from the mock with each one redone, and a confirmed list of references.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Candidates describe the behavioral interview as a structured 'career chapters' (WHO) round, where you present your work history in chronological order and explain the intent, results, mistakes and lessons at each stage. Prepare a few projects you know in depth, including the decisions you made and the trade-offs you weighed, and make your own role in each one clear and consistent every time you tell it.
Walk through your technical career history as a multi-chapter narrativ…
Walk through your technical career history as a multi-chapter narrative, highlighting your intent, key accomplishments, and lessons learned at each stage.
Approach
- Pick a story where you made the decision, not one where you watched it.
- State the situation in two sentences and spend the rest on the reasoning.
- Close with what you would do differently, concretely.
Follow-up
- What did you decide not to do, and why?
- How did you know your change caused the improvement?
Tell me about a time you made a major architectural trade-off or techn…
Tell me about a time you made a major architectural trade-off or technical mistake. What was the impact, and how did you rectify it?
Approach
- Give the blast radius: what could have broken, and what you measured.
- Pick a story where you made the decision, not one where you watched it.
- State the situation in two sentences and spend the rest on the reasoning.
Follow-up
- What would you do differently if you ran that again?
- How did you know your change caused the improvement?
Argue against failing open when the control plane is unreachable
The gateway caches credential-to-context decisions with a sixty-second TTL. A design proposal says that when control-plane reads fail, pods should keep serving from expired entries indefinitely so a control-plane outage never becomes a product outage. You believe that converts every revocation into an unbounded one. Describe a design you argued against while it was still a live proposal: what you measured or modelled to make the case, what you conceded, who decided, and what happened afterwards. Say what would have changed your mind before the decision, not after it.
Approach
- Reframe it from a values argument into a bounded-staleness argument. Both sides already accept the cache; the disagreement is only about the ceiling on how long a revoked credential keeps authorising. Put a number on the table — serve stale for up to fifteen minutes, then fail closed — and make the other side argue against a number rather than against a principle.
- Bring arithmetic rather than adjectives: the rate of revocations with revoked_reason in ('suspected_leak','auth_version_bump'), the observed distribution of control-plane unavailability, and the product of the two, which is expected requests served by revoked credentials per outage-hour. At 30k requests/second the unbounded version is not a subtle exposure and the number says so.
- Concede the strong half of the opposing case first, because that is what buys you the room: failing closed turns one service's outage into a total outage across three regions, and a control plane doing tens of writes per second is not engineered to the gateway's availability target. A proposal you have not steelmanned reads as reflex.
- Propose the asymmetry that usually resolves this: stale entitlements cost bounded money (a quota fifteen minutes out of date over-serves by a computable amount), while a stale revocation costs unbounded access. Split the cached decision by what it authorises, give the two halves different staleness ceilings, and let the entitlement half fail open while the revocation half fails closed.
- State the propagation dependency plainly, since it is the part that is missed: validity is also derived from the principal's auth_version, so password reset and sign-out-everywhere flow through this same cache. A design that bounds staleness for explicit revocation and not for auth_version bumps has only fixed half of it.
- Say who decided, and what you did afterwards in either outcome: write the decision down with its number and a review date, and instrument the exposure you were worried about so the next round of the argument is settled by data instead of by seniority.
Follow-up
- Publish-subscribe invalidation is lossy under a partition, and a TTL is the only hard bound. What TTL do you pick, and what does it cost you at 30k requests/second?
- The key was revoked because it was found in a public repository. Does your answer change, and where does that urgency live in the design?
- You lost the argument and six weeks later the failure you predicted happens. What do you say in the review, and what do you not say?
- 01
Walk through your technical career history as a multi-chapter narrative, highlighting your intent, key accomplishments, and lessons learned at each stage.
- 02
Describe a significant technical or interpersonal conflict you experienced during a project and detail how you resolved it.
- 03
Tell me about a time you made a major architectural trade-off or technical mistake. What was the impact, and how did you rectify it?
- 04
How do you approach explaining complex technical constraints or architectural trade-offs to non-technical stakeholders?
- 05
Give an example of how you have mentored junior team members or driven improvements to internal developer tooling and engineering processes.
Is this an official Headway interview guide?
No. It is PracHub's own research and practice material for the Software Engineer role at Headway. Rounds and questions reflect what candidates have reported, not a process Headway has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗What is the typical timeline for the Headway interview process?
Candidate reports vary, from about three to four weeks up to four to six weeks, across five stages: recruiter screen, technical screening, virtual onsite, behavioral interview and reference checks. Ask your recruiter for the expected schedule after the screen, and arrange your references early so the last step does not delay an offer.
PracHub interview research ↗Are the Headway coding interviews focused on obscure algorithm puzzles?
The reported coding questions are practical. They include finding an open appointment slot between 8 AM and 5 PM, converting markdown to HTML with nested lists, fixing and extending existing code that computes best times per obstacle, simulating the chance of a personal best, and rate-limiting middleware with global and per-endpoint limits. Prepare interval and time logic, parsing, simulation and debugging someone else's code, and check edge cases out loud.
PracHub interview research ↗What makes the "WHO" behavioral interview different from standard behavioral rounds?
Candidates describe it as a structured walkthrough of your career. You lead by presenting your work in chronological "chapters" and explain the intent, successes, technical failures and lessons at each stage. Prepare it as a narrative with a mistake and a lesson in each chapter, not as isolated STAR answers.
PracHub interview research ↗Are reference checks part of the Headway process?
Yes. Candidates report reference checks with past managers and colleagues as the final step before an offer. Choose people who can speak to your technical abilities and work ethic, ideally people who worked with you on the projects you plan to discuss.
PracHub interview research ↗Can I use my preferred programming language?
Candidates report that general coding screens accept any major language (Python, Java, C++), while some specialised or framework-specific rounds may require Python or JavaScript/TypeScript. Confirm with your recruiter before each technical round.
PracHub Software Engineer practice ↗What does the Headway virtual onsite cover?
It is reported to combine coding, system architecture and behavioral assessments. It runs either as one multi-hour block or split over two consecutive days. Candidates who pass the technical screen reportedly receive preparation materials for the final loop, so use them to adjust your plan.
PracHub Software Engineer practice ↗What system design topics should I prepare for Headway?
Reported design questions include a rental or hotel search optimised by location and date range, a REST API with data modelling, pagination and JWT authentication, diagnosing slowdowns during a large request spike, estimating server and database capacity during fast growth, and walking through the architecture of a production system you built. Start each one from the data model and the queries it has to serve.
PracHub Software Engineer practice ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-24 - 02PracHub Software Engineer practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-24 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-24