As a Software Engineer at D-Wave Quantum, you are at the frontier of commercial quantum computing. This role is not merely about writing code; it is about building the infrastructure and platforms that enable quantum-classical hybrid applications. You will be contributing to the evolution of the Leap quantum cloud service, AI platform development, or DevOps pipelines that sustain a high-performance, mission-critical environment.
Your work directly impacts how researchers and commercial enterprises interact with quantum hardware. Whether you are optimizing AI-driven workflows or architecting robust deployment systems, your contributions directly influence the scalability and reliability of quantum computing as a service. This position requires a unique blend of high-level systems thinking and an ability to navigate the complexities of emerging technology, making it a challenging but highly rewarding career step.
The work at D-Wave Quantum often involves bridging the gap between classical high-performance computing and quantum processing, so emphasize your ability to handle complex system integrations in your interviews.
Initial Screening
reportedHalf of this call is the part candidates treat as small talk: start date, notice period, work authorisation and its timing, location and time zone, on-call, and the number. Those are what kill offers late, after several engineers have each spent a day. Surfacing a hard constraint now costs you nothing and occasionally buys you something, since a loop compressed to fit a competing deadline can usually only be arranged if it is asked for early. The common failure is deflecting the compensation question twice, then discovering at offer stage that the band never reached your number.
What to demonstrate
- Whether your hard constraints are compatible with the role before a loop gets booked: earliest start, notice period, what authorisation you hold and when it needs action, days on site, willingness to carry a pager
- Whether you give a compensation range with something behind it, such as current total compensation or a competing timeline, rather than leaving the band untested
- Whether your stated timeline is real, since a competing deadline raised now is something scheduling can sometimes work around and the same deadline raised at offer stage usually is not
How to prepare
- Write each constraint down in one line before the call and state them as facts rather than negotiating them live under a question you were not expecting
- Set your range from two or three current data points for that level and location, and name the structure you are quoting in, so the number is comparable to the one they are holding
- If another process is running, say where it stands and by when, and ask directly whether this loop can be scheduled inside that window
Technical Deep-Dive
reportedMost of the time lost in this format is not lost to thinking. It goes to a standard-library call you half-remember, an off-by-one in a loop bound, and a debugging loop that mutates code at random until something passes. When output is wrong, stop re-reading the whole function: take the smallest input that reproduces it and walk the state through by hand, printing intermediates if the environment allows. Guessing at a fix without a failing case you understand is how a five-minute bug becomes twenty, and the clock does not pause while you do it.
What to demonstrate
- Whether you reach the right structure without a detour, and can write it from memory rather than only recall that one exists
- Whether overflow is considered where the language has fixed-width integers, since a signed 32-bit value stops at 2,147,483,647 and then wraps in Java, is undefined behaviour in C++, and does not arise in Python, whose integers grow instead
- Whether recursion depth is treated as a constraint on large inputs, given that CPython's default limit is 1000 frames and a deep recursion can exhaust the stack in any language where an iterative version would not
- Whether a failing case is isolated and explained before any edit is made to the code
How to prepare
- From an empty file and with no references open, implement the pieces you lean on most: a heap push and pop, an iterative DFS with an explicit stack, and a binary search whose midpoint is written lo + (hi - lo) / 2, which avoids the overflow that (lo + hi) / 2 can hit in a fixed-width integer type
- Time yourself on the ten library calls you look up most, such as sorting with a custom comparator, splitting and joining strings, and finding the next key at or above a value in an ordered map, until the lookup is gone
- Take a solution you know is broken and, before touching it, write one sentence naming the input, the expected value and the actual value. Repeat until you do it without deciding to.
Behavioral Assessment
reportedMany of these questions are about something that went wrong, and the grading sits mostly in the hours after you knew. Who found out first, whether that was you or an alert or a user, how long it took you to say it out loud, and whether the people who needed the news got it while they could still act on it. Engineers under-tell this part because it feels like confessing. The pattern it is looking for is the opposite: the quiet fix, an incident absorbed without telling anyone, after which nothing changed and the same failure is still available.
What to demonstrate
- How the problem was found, and whether that route was one you had built or one that happened to you, since a user reporting it first means your instrumentation did not cover that failure
- Whether time-to-detect and time-to-tell are separate numbers in your account and whether you know both, because a fast fix that nobody heard about until the retro is a different answer from a slow one that was announced immediately
- Whether the resolution left something durable behind, a check that fires or a default that changed, rather than depending on people remembering to be careful
- Whether you can say what the failure cost without either inflating it or waving it away
How to prepare
- Reconstruct one incident you were part of as a timeline with clock times: first bad request, first signal, first person who knew, first message outside the team, mitigation, permanent fix. The gaps between those entries are what gets asked about
- Look up the configuration of the signal that caught it, including its evaluation window and threshold. An alert defined on a five-minute aggregate cannot fire until the condition holds across that window, which puts a floor under time-to-detect that has nothing to do with how severe the failure was. Be able to say what that floor was and whether anyone had chosen it deliberately
- Prepare one story where you escalated early and the severity turned out to be smaller than you thought, including what it cost the people you pulled in. Without it, every answer you give about raising alarms is unfalsifiable
PracHub editorial advice for the preparation topics above.
Representing prices and quantities as binary floating point.
IEEE-754 binary64 cannot represent 0.01 exactly, so a cost basis accumulated in doubles drifts, and an equality test against a tick boundary fails unpredictably. The fix is an integer at a fixed scale - a count of ticks, or a scaled integer at the instrument's price_scale - which also turns the tick-size check into a modulus rather than a tolerance comparison. The nuance that lets this trap survive code review: binary64 represents every integer exactly up to 2^53, about 9.0e15, so a scaled integer carried in a double is exact right up until it is not, and the first failure tends to be a large notional in a low-priced instrument, in production.
Retrying an order send after a timeout, on the assumption that the send failed.
A timeout says nothing about whether the venue received, matched and acknowledged the order - only that a reply did not arrive in time. Retrying turns an unknown into a duplicate: two live orders, twice the intended position, and a hedge computed against a position record that is now wrong. The correct move is to treat the client order id as an idempotency key, persist it before the send, and on recovery query state (order status request, or the drop-copy stream) rather than resend. It is the same shape as a double-captured payment, except the second order can move the price against you while you work out what happened.
Sharing mutable state with no stated owner
Say which thread, request or task owns each mutable structure, and what protects it when the answer is more than one: a lock, a queue that hands ownership across, or an immutable copy per reader. A structure documented as safe for concurrent reads is usually not safe for a concurrent write alongside those reads.
Not asking what the system looks like if it dies halfway through
For any multi-step write, say what state remains if the process stops between step two and step three, and what brings it back: a single transaction, a saga with compensating actions, an outbox, or a reconciliation job. Partial failure is routine at any real call volume, so 'that shouldn't happen' is an answer with nothing behind it.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Choose a book structure for constant-time top-of-book reads
You maintain a price-level book for one instrument on one venue. Deltas arrive as (side, price, new_aggregate_qty), where price is an integer at the instrument's price_scale and tick_size is constant within the session; new_aggregate_qty of zero deletes the level. Sustained 10^5 to 10^6 deltas per second, over 90 percent landing within ten ticks of the touch, and top of book is read after every delta. Choose the representation, then give apply(delta) and best_bid()/best_ask() with their expected cost, their worst case, and the memory per instrument per side. Say what happens when the price leaves your band.
Approach
- Derive the structure from the access pattern rather than from the abstract operation set: reads are overwhelmingly of one element (the touch), writes cluster within ten ticks of it, and the price domain is a dense integer lattice because every valid price is a multiple of tick_size. That argues for a flat array of aggregate quantities indexed by (price - base_px) / tick_size, with cached best_bid_idx and best_ask_idx, over any comparison-based container.
- apply(delta) writes one slot in O(1). An insert at a price better than the current best updates the cached index in O(1) with no scan. A delete of a non-best level is O(1). Only a delete of the best level requires finding the next occupied level inward, which is where the worst case lives.
- Make that scan cheap by shadowing the array with a bitset, one bit per level, 64 levels per 64-bit word: the next occupied level is a find-first-set from the current index, so a 4096-tick band costs at most 64 word reads and typically one. Memory is 4096 x 8 bytes plus 512 bytes of bitset per side, about 32.5 KiB per side, 65 KiB per instrument.
- Price the alternative honestly. A sorted map is O(log L) per update with correct semantics and no band, but it allocates a node per new level on the decode path and chases pointers on every read; the disqualifier is the allocation and the cache misses, not the logarithm. Keep the flat array for the actively quoted set (1,000 instruments is about 65 MB) and a map for the long tail, because 10^5 instruments at 65 KiB each is 6.5 GB.
- Handle the band exit explicitly: when a price falls outside [base_px, base_px + width x tick], rebase by memmoving the live window and clearing the rest, an O(width) operation that must be counted and alarmed rather than hidden, since it is a latency spike correlated with fast markets.
- Tie it to the gap invariant: the book carries a staleness flag, and an unrecovered sequence gap sets it so best_bid()/best_ask() report unusable rather than returning the last well-formed state.
Follow-up
- Move from price-level to order-level (each resting order tracked individually). What does that cost you, and what does queue-position modelling need from it?
- The venue sends a snapshot plus increments after a gap. Where do you buffer the increments, and at which sequence number do you start applying them?
- How do you verify this book against an independent full-depth snapshot mid-session without stalling the decode thread?
Find every overlapping effective window before adding the constraint
instrument_version holds about 10^7 rows of (instrument_id, effective_from, effective_to nullable). Versions for one instrument_id must never overlap, and at most one may have effective_to NULL. Adding the EXCLUDE constraint aborts on the first conflicting pair it meets, so you need the whole picture first. Produce every row that overlaps another row for its instrument, per-instrument overlap counts, and every instrument holding more than one open-ended row. The median instrument has about ten versions; the worst has 10^5 after repeated vendor reloads. Explain why the pairwise comparison is correct but cannot ship.
Approach
- Cost the naive version properly rather than dismissing it. Grouped by instrument_id it is the sum of g^2 over groups, and at a median g of ten that is trivial — the blow-up lives entirely in the tail, where one 10^5-version group contributes 5 x 10^9 comparisons. Worse, its output is also quadratic: if that group is a repeated reload, nearly every pair overlaps, so the self-join emits billions of rows describing a few thousand bad versions. The output size, not the runtime, is the reason it cannot ship.
- Restate the deliverable as rows rather than pairs, which is what makes a linear output possible, then sort by (instrument_id, effective_from, effective_to nulls last) for O(n log n) and sweep each group once.
- In the sweep, carry running_max_end = MAX(COALESCE(effective_to, infinity)) over the rows already seen in this group, plus the row that set it. A row overlaps iff its effective_from is strictly less than running_max_end; flag both that row and the max-setter, so both sides of every overlap appear in the output. O(1) extra state per group.
- Do not shortcut to comparing adjacent rows only. It is sufficient to answer the boolean does any overlap exist — if every adjacent pair is disjoint then end_i <= start_(i+1) <= start_j for all j > i, so nothing overlaps — but it misses nested intervals when you must list every offending row, and a long open-ended version nesting later ones is the single most common real shape.
- If you write it in SQL, the window-function form is MAX(COALESCE(effective_to, 'infinity'::timestamptz)) OVER (PARTITION BY instrument_id ORDER BY effective_from ROWS BETWEEN UNBOUNDED PRECEDING AND 1 PRECEDING). The COALESCE is load-bearing, and in two distinct ways: MAX ignores NULLs, so an earlier still-open version contributes nothing to the running maximum, and where it is the only preceding row the maximum is NULL outright, which makes effective_from < running_max evaluate to NULL rather than true. Either way the comparison is not true and the overlap disappears from the report.
- Handle the open-row rule as its own aggregate in the same pass — COUNT(*) FILTER (WHERE effective_to IS NULL) > 1 per instrument — since a second open row is a violation whether or not your range model makes it overlap, and since it is the only part of the report that survives the COALESCE being dropped. Only then add EXCLUDE USING gist (instrument_id WITH =, tstzrange(effective_from, effective_to) WITH &&), which needs btree_gist for equality on a bigint, and expect to iterate because each failed build names one pair.
Follow-up
- tstzrange(a, a) is empty and empty ranges never overlap, so a zero-width version passes the constraint. Is that acceptable, and what would catch it?
- The fix requires closing thousands of rows. How do you do that without editing history that a replay depends on?
- How would you run this check continuously, so the next bad vendor load is caught at write time instead of at constraint-build time?
Fold a duplicated execution stream into per-order state
A day's execution reports arrive as one unordered array of tuples (venue_mic, venue_exec_id, order_id, exec_type, last_qty, cum_qty, leaves_qty). The same execution appears up to three times, once from the trading session, once from the drop copy, once from the clearing file. Assume 10^7 rows over 10^6 orders, and ignore trade_correct and trade_cancel for now. Given a map order_id to order_qty, return each order's final cum_qty and leaves_qty, plus the list of orders whose final cum_qty exceeds order_qty. One pass, O(n) expected time; justify your space and say what you would give up to shrink it.
Approach
- Notice what the deliverable needs before reaching for a dedupe set: final cum_qty and leaves_qty are a max-fold over the venue's own cumulative field, and max is idempotent, so a resent report changes nothing. That reduces state to one record per order_id in an open-addressed hash map keyed by a 64-bit id, roughly 24 bytes each, about 24 MB at 10^6 orders.
- Fold under an explicit total order rather than arrival order: keep the report with the greatest cum_qty, breaking ties toward the smaller leaves_qty. The tie-break is what lets a canceled report (same cum_qty, leaves_qty 0) displace the partial fill that preceded it without consulting any clock, which matters because the three sources arrive interleaved.
- Reintroduce deduplication only for the fields that are sums rather than maxima: commission_amt, fee_amt, rebate_amt, and any fill count. Those need a set on (venue_mic, venue_exec_id), and at 10^7 executions with roughly 50-byte keys that is about a gigabyte of state. Hashing the key to 64 bits shrinks it to around 256 MB but accepts a collision probability near n^2/2^65, about 3e-6 per run, where a collision silently drops a real fill.
- Emit the violation list in the same pass: any order whose winning cum_qty exceeds its order_qty. The fold exists to protect that invariant, so the check belongs inside the job rather than in a dashboard that reads its output later.
- State the bound: O(n) expected time, O(orders) space, and the job is I/O bound on reading 10^7 rows rather than bound by the fold itself.
Worked solution 20 min
- Build the fixture: order 1001 with order_qty 1000 has E1 (partial_fill, last 300, cum 300, leaves 700) and E2 (fill, last 700, cum 1000, leaves 0); order 1002 with order_qty 500 has E3 (partial_fill, last 200, cum 200, leaves 300) and E4 (canceled, last 0, cum 200, leaves 0).
- Duplicate E2 twice and E3 once, so the array holds 7 rows for 4 distinct executions, then shuffle it with a fixed seed.
- Run the max-fold with the (cum_qty desc, leaves_qty asc) tie-break and record both per-order outputs and the violation list.
- Add a corrupt row E5 for order 1001 with cum 1300, leaves 0, rerun, and confirm 1001 now appears in the violation list rather than being silently accepted.
- Swap the max-fold for a sum of last_qty on the same shuffled array and record what order 1001 reports.
Follow-up
- Now admit trade_correct and trade_cancel, which arrive later and reference an earlier venue_exec_id via corrects_exec_id. What breaks in the max-fold, and what replaces it?
- The clearing file disagrees with the session on cum_qty for 40 orders. Which side do you write into position_snapshot, and what do you do with the difference?
- How would you make this fold restartable partway through a 10^7-row file without recomputing from the beginning?
A desk report that silently multiplies filled quantity
A report joins order_event (many rows per order_id: submit, ack, each partial_fill, fill), execution_report (one row per venue execution, keyed (venue_mic, venue_exec_id)) and risk_limit (PRIMARY KEY (limit_id, limit_version), one row per version per effective window) to show, per strategy_run_id, filled quantity, filled notional, and the max_order_notional in force. Filled quantity comes back roughly five times too large and notional shifts run to run. Identify both fan-outs, give the diagnostic that proves each one, and rewrite the query so the totals are correct.
Approach
- Name the mechanism rather than the symptom: an inner join multiplies rows, and SUM over a multiplied row set multiplies the measure. Each execution matches every lifecycle row of its order, and each limit scope matches every stored version, so the two factors compound.
- Prove each fan-out with a cardinality probe on the key alone, before touching the aggregate: GROUP BY order_id HAVING count() > 1 on order_event, and GROUP BY limit_id HAVING count() > 1 on risk_limit. Then compare COUNT() of the joined set against COUNT() of execution_report restricted the same way; the ratio is the multiplier.
- Fix the execution side by aggregating before joining: a subquery summing last_qty and the notional per order_id, joined one-to-one against a genuine one-row-per-order source such as the submit event, instead of joining the event log directly.
- Fix the limit side with a LATERAL that picks the one version in force: effective_from <= the order's sent_ts AND (effective_to IS NULL OR effective_to > sent_ts) ORDER BY effective_from DESC LIMIT 1. Picking by MAX(limit_version) is a different answer and usually the wrong one, because the newest version may not have been in force when the order was sent.
- Scale notional correctly while you are in there: last_px is an integer at the instrument's price_scale and derivatives carry a contract_multiplier, so notional is last_qty * last_px * contract_multiplier / 10^price_scale, with both attributes resolved from the instrument version as of the execution.
- Reject the reflex fixes explicitly: SELECT DISTINCT and SUM(DISTINCT last_qty) both change the answer by collapsing two legitimately equal fills into one.
Follow-up
- A LEFT JOIN to risk_limit would keep orders that matched no limit. What does that do to your totals against an inner join, and which do you actually want?
- How would you catch this class of bug automatically before the report ships to a desk?
- Which of these joins would you push into a materialized view, and what event invalidates it?
Current order state from an append-only event log
order_event is append-only: event_id BIGINT PRIMARY KEY, order_id BIGINT, strategy_run_id UUID, seq_no INT (1-based and contiguous per order, UNIQUE (order_id, seq_no)), status_after, cum_qty, leaves_qty, received_ts TIMESTAMPTZ(6). The table holds 10^8 rows across 10^7 orders. Write the query returning current status_after, cum_qty and leaves_qty for every order in a supplied list of order_ids, then the variant returning current state for all orders of one strategy_run_id. State the index each needs, and why ordering by received_ts instead of seq_no is wrong.
Approach
- For the id list use SELECT DISTINCT ON (order_id) ... ORDER BY order_id, seq_no DESC, which takes the first row of each group and stops. The portable equivalent, ROW_NUMBER() OVER (PARTITION BY order_id ORDER BY seq_no DESC) = 1, ranks every row of every group before discarding all but one.
- Index (order_id, seq_no DESC). A btree can be scanned forwards or backwards as a whole, not per key, so the mixed-direction ORDER BY (order_id ASC, seq_no DESC) is not satisfied by (order_id, seq_no) and the planner inserts a sort.
- For the strategy_run_id variant the leading column changes to (strategy_run_id, order_id, seq_no DESC); otherwise the run's orders are found by a scan and grouped afterwards, which is the difference between 10^4 index descents and 10^8 row reads.
- Justify seq_no over received_ts concretely: received_ts is neither unique nor monotonic across reports. Two reports can share a microsecond, and a resend after a reconnect can arrive after a report it precedes, so MAX(received_ts) picks a non-deterministic row and can return an older state than the one the system acted on.
- Name the cost that motivates a maintained order_current table: this query is one index descent per order in scope, fine for a list and unacceptable for a dashboard refreshing every live order in a session. The trade-off is a second write on the order path and a crash window between the append and the update.
Worked solution 20 min
- Seed one order with six events, giving two of them the same received_ts and inserting the final event with a lower event_id than its predecessor.
- Run the DISTINCT ON query and a MAX(received_ts) variant against that order and diff the two rows.
- EXPLAIN the query under both index shapes, (order_id, seq_no) and (order_id, seq_no DESC), and look for the sort node.
- Repeat for the strategy_run_id variant over a run holding 10^4 orders.
Follow-up
- How do you assert seq_no is contiguous per order, and what should the system do when it is not?
- Write the query listing orders whose latest row violates cum_qty + leaves_qty = order_qty while the order is live.
- How would you keep order_current correct if the writer can crash between appending the event and updating the summary row?
Can you explain your process for troubleshooting a complex bottleneck …
Can you explain your process for troubleshooting a complex bottleneck in a distributed system?
Approach
- Say what you would check first and why it is the highest-information step.
- State your assumptions explicitly before working the problem.
- Clarify what is being asked and what a complete answer contains.
Follow-up
- What assumption would you test first?
- How would you know your answer was wrong?
How do you design for scalability in a cloud-native environment?
How do you design for scalability in a cloud-native environment?
Approach
- Work from the requirement backwards to the design.
- Clarify what is being asked and what a complete answer contains.
- State your assumptions explicitly before working the problem.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
Resumable position stream with a replayable watermark
Risk dashboards and an external allocator subscribe to intraday position updates over long-lived connections. position_snapshot is keyed by (account_id, instrument_id, business_date, snapshot_kind) and carries net_qty, avg_cost_px, mark_px, mark_source ENUM('last_trade','mid','settlement','vendor','stale') and fills_applied_through_exec_id as the watermark. Design the subscribe and resume contract: how a client disconnected for 90 seconds catches up without replaying the day, what happens when its resume point has aged out of retention, and how a consumer that must not double-count handles at-least-once delivery.
Approach
- Assign a server-side monotonic sequence per partition (account_id is the natural partition, since positions are folded per account) that is independent of any wall clock, and make the resume token (partition, last_seq, epoch). The epoch changes whenever the sequence is reset or the partition is rebuilt, so a stale token from a previous session cannot alias a new sequence and quietly skip a day.
- Define resume as 'send from last_seq+1, at-least-once, possibly redelivered', and publish the retention floor as part of the contract so a client knows how long it may be away.
- Make an expired resume point a typed error, not a silent restart from the head: the client must then do snapshot-then-increments, exactly as a market data book resynchronises. It fetches the snapshot with its sequence, buffers live increments arriving meanwhile, discards those at or below the snapshot's sequence, and applies the rest. Skipping the buffering step is what produces a book, or a position, that is wrong and well-formed.
- Prefer state-carrying messages over deltas: a message carrying net_qty, avg_cost_px and fills_applied_through_exec_id is idempotent under redelivery, whereas one carrying '+100' is not. Where a delta is unavoidable, carry the resulting sequence so the consumer can order and deduplicate.
- Put exactly-once at the sink, not on the wire: the consumer persists last_applied_seq per partition in the same transaction as the effect it applies, so a redelivered message is a no-op on restart. No transport gives exactly-once, and a design that assumes one has simply moved the bug.
- Carry staleness to the wire: mark_source='stale' must reach the client rather than being smoothed into a last known mark, so a dashboard can show a stale mark instead of an old number that looks current.
Worked solution 40 min
- Write the subscribe request, the resume request and the message envelope, with the sequence and epoch fields explicit.
- Trace a 90-second disconnect inside retention: list the messages sent on resume and the consumer's state before and after.
- Trace the same disconnect past the retention floor: write the error response and the full snapshot-then-increments sequence, including the buffer and the discard rule.
- Write the consumer's apply transaction showing last_applied_seq and the effect committed together.
- Write what the dashboard renders when mark_source is 'stale' and confirm nothing upstream can overwrite it with a fresher-looking value.
Follow-up
- A slow consumer falls behind the retention floor while still connected. Do you disconnect it, drop messages, or buffer, and what does each choice cost the risk dashboard?
- The allocator needs a consistent cross-instrument view for one account at a point in time. Does per-instrument sequencing give it that, and what would?
- How do you prove after the fact that a client saw a given position value, given that the stream is ephemeral?
Instrument cache flush degrades every book at once
At 13:05 a vendor correction reaches the reference service. Within two seconds its request rate rises forty-fold, p99 goes from 3 ms to 2 s, callers time out and retry, and the market data gateway marks books degraded because it cannot resolve symbology. Every service caches instrument_version with a five-minute TTL and flushes the entire cache on a correction notice. Give an ordered checklist that distinguishes a capacity problem from a stampede, and a fix that preserves as-of resolution for replay.
Approach
- Decide whether demand rose or supply fell. Compare distinct keys requested against total requests through the spike window. Forty times the requests over roughly the same key set is a stampede compounded by retries; forty times the distinct keys is genuine new demand and wants a different fix entirely.
- Check whether misses are correlated across processes. A global flush makes every instance miss the same key in the same instant, so the service receives one request per instance per key rather than one request per key. A uniform TTL does the same thing more slowly by aligning expiry across instances that started together.
- Quantify what the callers add before touching the server. A client that retries a timeout without backoff and without jitter multiplies an overload it is already inside; count retries separately from first attempts and state the amplification factor.
- Narrow the invalidation. A correction affects specific instrument_ids, and flushing discards 10^4 to 10^6 unrelated entries to fix a handful. Invalidate by key, which requires the notice to carry the affected instrument_id list rather than being a bare signal.
- Coalesce and stagger. One in-flight fetch per key per process with concurrent callers awaiting that single fetch bounds fan-out to the process count; TTL jitter stops survivors re-aligning; a bounded retry budget with backoff stops the client side recreating the burst.
- Handle the degraded case honestly. Serving a stale definition is only acceptable where the caller records which version it served, because a replay must resolve symbology exactly as the live session did; where it cannot, the correct behaviour is the one the gateway already took - mark the book degraded rather than price off a guess.
Follow-up
- Stale-while-revalidate would have kept the gateway quoting through this. When is that the right call, and when does it become a correctness bug?
- The correction notice is a topic every service subscribes to. What would you change about the notice itself?
- What signal would have made this visible as a design flaw before it was an incident?
For someone fluent in a dynamic language who has shipped real work but has never had to say what the runtime is doing underneath. The week is built on measuring and deliberately breaking things, because the questions that expose this background are the ones where the interviewer asks why a second time.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Measure before reasoning
- Take a slow piece of your own code, write down in advance where you believe the time goes, then profile it and record how wrong the guess was. The cost is usually an allocation you did not notice or an accidental quadratic membership test.
- Replace one list membership test inside a loop with a set and measure at a thousand, ten thousand and a hundred thousand elements, confirming the shape of the curve rather than only that it got faster.
- Write down the three quantities you can now measure instead of assert: wall time, peak memory, and call count for the function you suspected.
Deliverable: A before-and-after profile of real code plus a written note on the size of the gap between the guess and the measurement.
Practice prompt ↗Practice prompt ↗Worked solution ↗02References, copies, and the bugs they produce
- Write the function with a mutable default argument, call it three times, and explain the accumulating result: the default is evaluated once when the function is defined, so every call shares one object.
- Build a nested structure, take a shallow copy, mutate an inner element, and show that both views changed, because a shallow copy duplicates the container and not the elements. Then fix it with a deep copy and state the cost you just accepted.
- Write two functions, one mutating its argument in place and one rebinding the local name, and predict the caller's view of each before running it. That single distinction produces most of the bugs that pass their tests.
Deliverable: Three small programs whose output you predicted correctly before running, each with a one-line statement of the rule underneath.
Practice prompt ↗Practice prompt ↗03Types, once, in a language that checks them
- Port one module you have already written, roughly a hundred lines, into a statically typed language, and record every place the compiler demanded an answer your original had left implicit: a value that can be absent, a numeric width, a case never handled.
- Write the same signature in both languages and state what the static one guarantees before the program runs and what it does not, since it will not save you from a wrong algorithm or an index out of range.
- Write the difference between an interface satisfied by declaration and one satisfied structurally, with one case each where the other approach would miss the mistake.
Deliverable: One module in two languages plus a list of the questions the type checker forced you to answer.
Practice prompt ↗Practice prompt ↗04Concurrency, starting with what actually runs at the same time
- Run the same CPU-bound function across four threads and four processes and measure both. Under the default CPython build the threaded version will not speed up, because only one thread executes bytecode at a time; the process version will. Check which build you are on first, since free-threaded builds remove that lock and change the result.
- Then run a blocking I/O workload across four threads and measure it speeding up, because the interpreter releases that lock around blocking calls, which is why treating threads as useless is wrong as a general claim.
- Build the lost update: two threads each incrementing a shared counter a hundred thousand times, and show a final value below the expected sum, because an increment is a load, an add and a store and the thread can be suspended between them. Fix it with a lock and then measure what the lock costs.
Deliverable: Three measurements, threads against processes on CPU work, threads on I/O work, and a demonstrated lost update, each with the mechanism written underneath.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Debugging as a procedure rather than an instinct
- Work one real failure as a bisection: find a revision or an input size where it is good and one where it is bad, halve repeatedly, and state the two assumptions bisection needs, that the property changes exactly once across the range and that the test is reliable.
- Minimise one failing input to the smallest version that still fails, and record how many rounds it took.
- Keep a hypothesis log for one bug in three columns, what I believe, what would disprove it, what I observed, and stop yourself the first time you are about to change two things at once.
Deliverable: One bug worked to root cause with a written hypothesis log and a minimised reproducing input.
Practice prompt ↗Practice prompt ↗06Tests that catch the bug you are about to write
- Implement an LRU cache with a capacity bound, then write the three test cases that would catch an off-by-one in eviction: insert exactly capacity items and assert nothing was evicted, insert one more and assert the least recently used key is the one gone, and read an old key just before that insert so the eviction victim changes.
- Add a property test comparing your implementation against a deliberately slow reference, an ordered list scanned linearly, over a few thousand random operation sequences, because a slow reference finds the cases you would not have thought to write.
- Write one numeric test that fails under exact equality and passes with a tolerance, and state why the tolerance has to be relative rather than absolute once the magnitudes grow.
Deliverable: An LRU implementation with three boundary tests, one property test against a slow reference, and one tolerance-based numeric test.
Practice prompt ↗Practice prompt ↗07Debug something broken, out loud
- Have someone plant three defects in a two-hundred-line program, an off-by-one, a shared mutable state bug, and a wrong error-handling path, then find them while narrating, under a fixed rule: state the hypothesis before touching anything.
- Time each one and record which tool found it, reading, a printed value, a debugger, or a test, because the question asked in interviews is how you would find it rather than what it was.
- Write the sentence you will use when you do not yet know the cause, one that names the next measurement instead of offering a guess.
Deliverable: A recorded debugging session with time-to-find per defect and the method that found each.
Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
For anything that touched live traffic, be ready to say how you would have undone it: a flag, a staged rollout, dual writes with the old path still authoritative. Once the old column is dropped or the source rows are overwritten there is no reverse, so name what you kept a copy of and for how long.
How do you mentor junior developers or contribute to team culture?
How do you mentor junior developers or contribute to team culture?
Approach
- Pick a story where you made the decision, not one where you watched it.
- State the situation in two sentences and spend the rest on the reasoning.
- Close with what you would do differently, concretely.
Follow-up
- What did you decide not to do, and why?
- How did you know your change caused the improvement?
Describe a time you had to explain a highly technical concept to a non…
Describe a time you had to explain a highly technical concept to a non-technical stakeholder.
Approach
- Name the disagreement and how you resolved it with evidence.
- State the situation in two sentences and spend the rest on the reasoning.
- Pick a story where you made the decision, not one where you watched it.
Follow-up
- How did you know your change caused the improvement?
- What did you decide not to do, and why?
How do you handle disagreements with team members regarding technical …
How do you handle disagreements with team members regarding technical architecture?
Approach
- Close with what you would do differently, concretely.
- Name the disagreement and how you resolved it with evidence.
- Give the blast radius: what could have broken, and what you measured.
Follow-up
- What did you decide not to do, and why?
- What would you do differently if you ran that again?
Describe a time you had to optimize a piece of code for high-performan…
Describe a time you had to optimize a piece of code for high-performance computing needs.
Approach
- State the situation in two sentences and spend the rest on the reasoning.
- Name the disagreement and how you resolved it with evidence.
- Give the blast radius: what could have broken, and what you measured.
Follow-up
- What would you do differently if you ran that again?
- How did you know your change caused the improvement?
- 01
How do you mentor junior developers or contribute to team culture?
- 02
Describe a time you had to explain a highly technical concept to a non-technical stakeholder.
- 03
How do you handle disagreements with team members regarding technical architecture?
- 04
Describe a time you had to optimize a piece of code for high-performance computing needs.
Is this an official D-Wave Quantum interview guide?
No. It is PracHub's own research and practice material for the Software Engineer role at D-Wave Quantum. Rounds and questions reflect what candidates have reported, not a process D-Wave Quantum has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How much preparation time is typical for these interviews?
Most successful candidates spend 2–4 weeks preparing, focusing on both their core coding skills and their ability to explain complex system designs.
PracHub interview research ↗Is the interview process strictly technical?
No, while technical rigor is high, D-Wave Quantum places significant weight on how you work within a team, how you solve problems, and your alignment with the company’s vision.
PracHub interview research ↗Are there remote work opportunities?
Yes, many roles at D-Wave Quantum offer remote or hybrid options, though you should verify the specific requirements for the team you are applying to.
PracHub interview research ↗What differentiates a successful candidate?
Successful candidates typically demonstrate a "builder" mindset—they don't just solve the problem at hand, but consider how their solution impacts the long-term health and scalability of the platform.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Software Engineer practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22