As a Software Engineer at Quantifind, you play a pivotal role in building and scaling advanced software systems that transform complex data into actionable financial crime and risk intelligence. This position sits at the intersection of heavy data processing, modern web architecture, and robust infrastructure, directly empowering analysts and enterprise users to uncover critical insights. You will contribute to products that handle massive scale, demanding high performance, extreme reliability, and clean engineering principles.
The impact of this role extends across multiple domains, from core backend services built with Scala to critical infrastructure managed via OpenTofu and modern web experiences. Whether you are optimizing data pipelines, designing fault-tolerant distributed systems, or streamlining infrastructure, your code directly influences how financial institutions and government agencies mitigate risk. The work is technically challenging and intellectually stimulating, requiring a balance of rigorous algorithmic thinking and practical systems design.
You will find yourself collaborating with passionate, highly knowledgeable engineering teams in an environment that values autonomy, collaboration, and continuous learning. While expectations are high regarding code quality and architectural foresight, the culture emphasizes personable interactions and mentorship. Expect to engage deeply with the data industry ecosystem, pushing the boundaries of what automated risk discovery tools can achieve.
Initial Screening
reportedThe person on this call usually cannot evaluate your code and does not need to. They write a short paragraph, and that paragraph is what a hiring manager skims when deciding who to put on your loop. So the test is not whether your work was hard, it is whether a non-engineer can repeat it correctly. Name systems by what they did rather than by their internal codename, give each project a shape (what was breaking, what you changed, what happened after), and keep the whole walkthrough near ninety seconds. Depth that cannot survive a paraphrase reads as vagueness.
What to demonstrate
- Whether a non-engineer can restate your projects without distorting them, since their paraphrase is what travels to the hiring manager, not your sentences
- Whether each project has a shape rather than a stack list: the failure or constraint, the change you made, the result and how it was measured
- Whether you can say what was yours inside a team project without either inflating it or disappearing into the plural
How to prepare
- Rewrite each headline project as two sentences with no internal system names and no acronyms outside your company, then say them to someone outside engineering and have them repeat them back. Fix whatever came back wrong
- Attach one measured number to each project: the baseline, the change, and the window it was measured over. Where nothing was ever measured, say that plainly rather than reaching for a plausible percentage
- Time the background walkthrough against a clock. If it runs past two minutes, compress the earliest role to a single clause and spend the recovered time on the most recent one
Technical Interviews
reportedInput bounds are the part of the prompt most often skimmed, and they usually contain the answer. They tell you which complexity class is admissible, which narrows the search before you have thought about the problem itself. As a rough planning figure, a compiled language does on the order of 10^8 simple operations per second and an interpreted one roughly an order of magnitude less. So n up to about twenty admits enumerating subsets, a few thousand admits a quadratic pass, and a million admits neither: you need near-linear, or linear with a log factor. If the bounds are missing, ask for them.
What to demonstrate
- Whether the approach is justified by the stated input size rather than by whichever pattern you recognised first
- Whether you ask about the properties that change the algorithm: whether the input arrives sorted, whether duplicates occur, whether values are bounded integers, whether it all fits in memory
- Whether you can name the bottleneck in your own solution and what would remove it, even when you deliberately leave it in place
- Whether a claimed speedup is real, since memoising a recursion only helps when subproblems genuinely overlap and the state can be keyed cheaply
How to prepare
- For each algorithm you rely on, write down the largest n it handles in roughly a second, then check two of those figures by timing them in the language you will actually type in
- For two weeks, write one line naming your target complexity and the bound that justifies it before you write any code, then compare that line with what you ended up submitting
- Practise the conversion backwards: given a required O(n log n), list the mechanisms that get you there (sorting, a heap, an ordered map, divide and conquer) and choose by what the problem needs to query, not by what you used last
Behavioral Interviews
reportedWhat you say here is written down by each interviewer and compared afterwards, so the unit of evaluation is a claim someone else could check, not a well-told narrative. Two things make a story checkable: detail only a participant would hold, and a clean line around which part was yours. Vague ownership is the usual failure and it is usually accidental, because engineers say we about the team's work and we about their own, so the thing they personally built disappears into the plural. Name the part you wrote, and name who did the rest.
What to demonstrate
- Whether your details are ones a participant would hold and an observer would not: the constraint that ruled out the obvious approach, the first attempt that failed, the person who objected and on what grounds
- Whether ownership survives a direct question, since a follow-up to we decided is routinely who decided, and an answer that stays plural at that point is read as the work belonging to someone else
- Whether the numbers you quote are ones you would say identically to a former colleague with the dashboard open
How to prepare
- Go through each story replacing every we with either I or a named role (the on-call engineer, the reviewer, the other team) and check the story still holds together. Wherever it stops making sense you have found a part you cannot actually speak to
- Open the artefacts for two of your stories, the pull request, the design doc, the incident notes, and read them for dates and figures you have been rounding in the retelling. Correct your version to match
- For each story write the single sentence you would least want repeated to a former teammate, then either make it accurate or take it out
Final Interviews
reportedNobody in the room with you decides this. Interviewers typically write their rounds up separately, often before seeing anyone else's, and the outcome is settled later from those write-ups. A split panel gets resolved by whichever note carries specific evidence, so what you want out of each room is one concrete thing that person could write down: a bug you caught yourself, a trade-off you named, a decision you owned. The rest is arithmetic. The project you describe in a behavioural conversation is often the same system you sketched an hour earlier, and the two accounts have to agree.
What to demonstrate
- Whether the scale, team size and timeline you attach to a project hold steady when that project resurfaces in a different round
- Whether each interviewer leaves with a specific thing to cite rather than a general impression of competence
- Whether a trade-off you defended in one round survives a challenge in another, instead of being quietly swapped for the answer the new interviewer seemed to want
- Whether a question you have already answered earlier in the day gets the same answer at the same depth, without visible impatience
How to prepare
- Write a one-page sheet per project fixing the figures you will quote — request volume, data size, team size, elapsed time, what broke — and say them aloud from the sheet until they come out identical every time
- For each round on the schedule, decide in advance the one sentence you want in that person's notes, then check in a mock that you said it outright instead of leaving it to be inferred
- Have someone ask you the same project question twice, an hour apart, and diff the two answers for numbers that moved or a trade-off that reversed
PracHub editorial advice for the preparation topics above.
Comparing timestamps from two different venues, or subtracting a venue timestamp from a local one and calling the result latency.
Each venue stamps on its own clock, at its own point in its own pipeline, at its own resolution, so a difference between two of them mixes elapsed time with clock offset and with unstated stamping conventions. Subtracting a venue timestamp from a local receipt timestamp yields network delay plus the offset between two clocks, and unless both are disciplined to a common source the offset can exceed the quantity being measured. Use one local clock for any latency you intend to act on, discipline it properly - PTP with hardware timestamping reaches sub-microsecond, while plain NTP over a LAN is typically sub-millisecond, three orders of magnitude too coarse here - and keep venue and receipt timestamps in separate columns so nobody is tempted to conflate them.
Representing prices and quantities as binary floating point.
IEEE-754 binary64 cannot represent 0.01 exactly, so a cost basis accumulated in doubles drifts, and an equality test against a tick boundary fails unpredictably. The fix is an integer at a fixed scale - a count of ticks, or a scaled integer at the instrument's price_scale - which also turns the tick-size check into a modulus rather than a tolerance comparison. The nuance that lets this trap survive code review: binary64 represents every integer exactly up to 2^53, about 9.0e15, so a scaled integer carried in a double is exact right up until it is not, and the first failure tends to be a large notional in a low-priced instrument, in production.
Sorting when the problem never required a total order
Match the algorithm to the guarantee actually needed: the top k comes from a size-k heap in O(n log k) time and O(k) space, distinctness needs a set rather than an ordering, and a small bounded integer key range admits a linear counting pass. A full O(n log n) sort is the right default only when you genuinely need everything in order.
Abandoning working code to chase the optimal solution
Get the straightforward version correct, state its complexity, and only then optimise, keeping the working version until the faster one passes the same cases. A correct quadratic solution with a stated path to linear beats a half-written optimal one that never ran.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Explain the time and space complexity of your implemented solution and…
Explain the time and space complexity of your implemented solution and discuss potential optimizations.
Approach
- State the target complexity and say which constraint rules the naive version out.
- Walk one small example through your approach before writing the whole thing.
- Name the brute-force solution and its complexity before improving on it.
Follow-up
- Which test case would catch an off-by-one here?
- How does this change if the input no longer fits in memory?
Walk through an implementation of a caching layer with expiration poli…
Walk through an implementation of a caching layer with expiration policies.
Approach
- Choose the data structure from the access pattern, not from familiarity.
- State the target complexity and say which constraint rules the naive version out.
- Restate the input: its shape, its size, and what is guaranteed about it.
Follow-up
- Which test case would catch an off-by-one here?
- How does this change if the input no longer fits in memory?
Write a function to reverse a linked list or traverse a binary tree.
Write a function to reverse a linked list or traverse a binary tree.
Approach
- Walk one small example through your approach before writing the whole thing.
- Restate the input: its shape, its size, and what is guaranteed about it.
- Name the brute-force solution and its complexity before improving on it.
Follow-up
- Which test case would catch an off-by-one here?
- What is the worst case, and how likely is it on real data?
Publish top instruments by traded notional every second
Fills stream in as (instrument_id, last_qty, last_px, contract_multiplier), with last_px an integer at the instrument's price_scale. Roughly 10^7 fills a day across 10^5 instruments. Publish the fifty instruments with the highest cumulative traded notional for the session, refreshed once a second, using exact integer arithmetic. Per-fill work must be O(1), and the per-second emit must not rescan all 10^5 instruments in the common case. Trade cancels and corrections can reduce an instrument's total. Give the accumulator width you need, and the argument that your pruning cannot drop a qualifying instrument.
Approach
- Size the accumulator before choosing the algorithm. A single fill of 10^6 lots at a price of 100 carried at price_scale 9 is 10^6 x 10^11 = 10^17 in scaled units, so an int64 session total overflows after about 92 such fills (9.22e18 / 1e17). Accumulate in int128, or rescale to a coarser notional unit at ingest and document the rounding; a double is out on the usual grounds, since it is exact only to 2^53.
- Per fill: one open-addressed hash map lookup on instrument_id, one 128-bit add, and an insert into a dirty set — O(1) expected, no allocation, no ordering work at all between emits.
- Per emit: maintain tau, the current fiftieth-largest total. Candidates are the previous top fifty plus the dirty instruments whose total exceeds tau. Select with a bounded min-heap of size 50, which is O(c log 50) for c candidates, and c is a few thousand per second rather than 10^5.
- State the precondition that makes the pruning sound: while totals only increase, an untouched instrument below tau cannot have crossed it, because nothing changed its value. Cancels and corrections break exactly that precondition, so a corrected instrument must enter the dirty set and, if it falls out of the top fifty, the replacement may be an instrument the pruning skipped.
- Close that hole rather than ignoring it: keep a reserve of the top 200 and refill from it when a member drops, falling back to a full O(m) quickselect over all 10^5 totals when the reserve is exhausted. The fallback runs rarely, is bounded, and is the reason the output is correct rather than usually correct.
- Report both costs: amortised O(1) per fill, O(c log k) per emit with an O(m) worst case, and note that the emit runs on its own thread off a snapshot so it never sits in the fill path.
Worked solution 30 min
- Compute the overflow point by hand for the worst realistic fill (qty 10^6, px 100 at price_scale 9) and confirm the count at which an int64 total wraps.
- Build a fixture of 20 instruments with k = 3, feed fills so totals are 100, 90, 80, 70, ... and record tau after the first emit.
- Send a fill of size 5 to the instrument at 70, taking its running total to 75, and confirm it is pruned without entering the heap because 75 is still below tau; then send a further fill of size 20 to that same instrument, taking it to 95, and confirm it now enters the heap and displaces the instrument at 80.
- Send a trade cancel that removes 30 from the instrument at 100 and confirm the reserve supplies the replacement rather than the pruned set being consulted.
- Drain the reserve deliberately with a run of cancels and confirm the quickselect fallback fires and returns the same answer as a brute-force sort.
Follow-up
- Make it a rolling five-minute window rather than a session total. What state does each instrument now need, and what does that do to memory at 10^5 instruments?
- Two instruments tie exactly at the fiftieth place. What is your tie-break, and why does an unstable one make the output non-reproducible in replay?
- The totals are in mixed currencies. Where does the FX conversion go, and what happens to the ranking when a rate updates mid-second?
As-of symbology resolution over an effective-dated instrument table
instrument_version holds one row per instrument_id per effective-dated version: instrument_id BIGINT, venue_mic CHAR(4), venue_symbol VARCHAR(32), price_scale SMALLINT, effective_from TIMESTAMPTZ NOT NULL, effective_to TIMESTAMPTZ NULL. A vendor correction currently arrives as UPDATE ... SET venue_symbol = ... on the open row, and retired instruments carry an is_deleted BOOLEAN flag. Write the query that resolves (venue_mic, venue_symbol) to an instrument_id as of an arbitrary past instant, state the constraint that makes overlapping versions impossible, and explain why the in-place update and the is_deleted flag each break a replay.
Approach
- Fix the interval convention before writing anything: treat the window as half-open [effective_from, effective_to) with a NULL upper bound meaning open-ended, so every instant matches at most one version, an instant in a coverage gap matches none, and a boundary instant is not claimed by two.
- Write the predicate as effective_from <= $ts AND (effective_to IS NULL OR effective_to > $ts). Both halves are load-bearing, and dropping the upper bound in favour of ORDER BY effective_from DESC LIMIT 1 is not an equivalent query: at any instant inside a coverage gap, where a symbol has been released and not yet reassigned, the correct answer is zero rows, and the lower-bound-only form returns the last closed version instead, which is the issuer that used to hold the symbol. Keep both bounds. A btree on (venue_mic, venue_symbol, effective_from) still serves the full predicate with one descent and a backward walk, and INCLUDE (effective_to, instrument_id) lets the upper-bound test be evaluated from the index tuple, holding Heap Fetches at zero for as long as the visibility map is current for those pages. LIMIT 1 then buys a stopping condition rather than the semantics: under the non-overlap constraint at most one row can qualify anyway, and in a gap the walk reads every earlier version of that symbol before correctly returning nothing.
- Enforce non-overlap in the database rather than in the loader: EXCLUDE USING gist (instrument_id WITH =, tstzrange(effective_from, effective_to) WITH &&), which needs the btree_gist extension to get an equality operator class for a bigint. UNIQUE (instrument_id, effective_from) alone happily admits two windows that overlap.
- Replace the in-place update with close-the-current-row-then-insert-the-new-one inside one transaction, and replace is_deleted with a new version carrying trading_status = 'delisted'. A row that was true yesterday has to stay readable exactly as it was, because a replay reads it again.
- State the consequence that motivates all of it: venue symbols are reassigned to new issuers, so resolving today's symbol against today's open row can return a different instrument_id than the session being replayed actually traded.
Follow-up
- A version was loaded three weeks ago with the wrong effective_from. How do you correct it without editing history, and what second time axis does that force into the table?
- What does the lookup return during the instant a corporate action closes one version and opens the next, and what guarantees those two writes are one transaction?
- A replay resolves a million symbol lookups per run. What changes about the query shape at that volume?
Changing a price column's type on a live four-billion-row table
execution_report.last_px is NUMERIC(18,6) and must become a BIGINT integer at the instrument's price_scale, which lives on the effective-dated instrument_version row. The table holds 4 x 10^9 rows, takes writes continuously through the session, and is replicated. Produce the migration plan with no write downtime: the exact DDL steps and the lock each takes, how the backfill is batched and throttled, how you verify before cutting reads over, and the rollback point at every stage. Name the correctness bug specific to resolving price_scale during the backfill.
Approach
- Add the column cheaply and defensively. ALTER TABLE ... ADD COLUMN last_px_scaled BIGINT is a catalog-only change in PostgreSQL 11 and later for a nullable column or a non-volatile default, so there is no table rewrite. It still takes ACCESS EXCLUSIVE briefly, and that lock queues ahead of every new query, so set lock_timeout to a couple of seconds and retry rather than letting one long-running transaction stall the entire table behind the ALTER.
- Start dual-writing before backfilling, from the application or a BEFORE INSERT trigger, so the backfill only ever has to catch up on a closed set of old rows instead of chasing a moving tail.
- Backfill in bounded batches over a key range, one commit per batch, throttled on replication lag and autovacuum progress. Every UPDATE writes a new row version, so a full-table backfill roughly doubles the physical size until vacuum reclaims it; unthrottled, it bloats the table and lags the replicas at the same time, and the second symptom hides the first.
- Resolve price_scale as of the execution's venue_ts, not as of today. instrument_version is effective-dated, so a scale change since the trade makes a today-resolved backfill mis-scale every historical row of that instrument by a power of ten, silently, because the values stay plausible. Assert per row that the stored NUMERIC carries no more decimals than the target scale, and fail the batch rather than round away real precision.
- Verify before cutting reads over: a full pass returning zero rows for last_px_scaled IS DISTINCT FROM round(last_px * power(10::numeric, price_scale)) (numeric power, not the float8 ^ operator), plus a zero count of NULLs, run twice with the second run after further writes have landed. Then move reads behind a flag, soak, and only then DROP COLUMN last_px, which is catalog-only with space reclaimed on the next rewrite.
- State the rollback at each stage: before dual-write, drop the column; during backfill, stop and the old column is still authoritative; after read cutover, flip the flag back because both columns remain populated; after the drop there is no rollback, which is why the drop is a separate change weeks later.
Worked solution 45 min
- On a copy of the table, time ADD COLUMN with and without lock_timeout while a long-running SELECT holds a conflicting lock, and watch the queue form behind the ALTER.
- Write the batch backfill keyed on a range, joining instrument_version as of venue_ts, and measure rows per second and table size growth over ten batches.
- Run the verification query over the backfilled range, then corrupt one row deliberately and confirm it is caught.
- Add the NOT VALID check constraint and VALIDATE it, timing both and confirming writes continue throughout.
- Write the rollback sentence for each stage and identify the one stage that has none.
Follow-up
- Where do you add NOT NULL, and what does SET NOT NULL cost on this table compared with a NOT VALID check constraint you validate afterwards?
- The backfill needs an index that does not exist yet. How do you add it without blocking writes, and what do you do when it fails halfway through?
- A replica serves reporting and cannot tolerate an hour of lag. What changes about the batch size and the schedule?
Explain how you would secure sensitive data at rest and in transit acr…
Explain how you would secure sensitive data at rest and in transit across cloud infrastructure.
Approach
- Name the read and write paths separately; they rarely have the same bottleneck.
- Name the failure you are designing for, then the recovery path.
- State the consistency you need, and where you are willing to be stale.
Follow-up
- What would you drop to keep the system up under load?
- How does this behave when that dependency is down for an hour?
Design a distributed logging and monitoring system for microservices.
Design a distributed logging and monitoring system for microservices.
Approach
- Fix the scope first: who calls this, how often, and what they do when it fails.
- Name the failure you are designing for, then the recovery path.
- State the consistency you need, and where you are willing to be stale.
Follow-up
- What would you drop to keep the system up under load?
- What breaks first when traffic grows ten times?
How would you architect a data ingestion pipeline that handles sudden …
How would you architect a data ingestion pipeline that handles sudden traffic spikes?
Approach
- Fix the scope first: who calls this, how often, and what they do when it fails.
- State the consistency you need, and where you are willing to be stale.
- Choose a partition key and say what query it makes expensive.
Follow-up
- What breaks first when traffic grows ten times?
- How does this behave when that dependency is down for an hour?
How do you manage infrastructure as code and maintain declarative conf…
How do you manage infrastructure as code and maintain declarative configurations?
Approach
- Clarify what is being asked and what a complete answer contains.
- Say what you would check first and why it is the highest-information step.
- Work from the requirement backwards to the design.
Follow-up
- What assumption would you test first?
- How would you know your answer was wrong?
Fold execution reports into positions exactly once
execution_report has PRIMARY KEY (venue_mic, venue_exec_id) and carries exec_type, last_qty, last_px, cum_qty, leaves_qty, corrects_exec_id, source ('trading_session', 'drop_copy' or 'clearing_file'), is_possible_duplicate, venue_ts, received_ts and applied_ts. position_snapshot holds one row per (account_id, instrument_id, business_date, snapshot_kind) with net_qty, avg_cost_px and fills_applied_through_exec_id. Cost basis is weighted average: an increasing fill re-weights avg_cost_px, a reducing fill leaves avg_cost_px unchanged and realizes the difference between last_px and avg_cost_px on the quantity closed (signed so that closing a long above its average and closing a short below its average are both gains), and a fill crossing through zero realizes against the whole old position and opens a new basis at that fill's price. Volume is 100,000 to 10,000,000 executions a day, the same execution legitimately arrives from all three sources, and drop copy can deliver one fill before the session delivers an earlier one. Design the fold: what makes reapplication a no-op, what order the cost fold replays in and what a row arriving out of that order costs, how trade corrections are represented, and how an intraday reader gets a position no more than a few seconds stale.
Approach
- Separate identity from arithmetic. Deduplicate at insert with INSERT ... ON CONFLICT (venue_mic, venue_exec_id) DO NOTHING so three sources carrying one execution produce one row, then fold rows rather than messages. That places exactly-once at the sink, where it can actually be enforced, instead of assuming it of the transport.
- Split the derived state by its algebra, because only part of it is order-free. net_qty is a sum of signed last_qty over a set of exec ids, so arrival order cannot change it. cum_qty and leaves_qty are the venue's statement about an order, taken from the report carrying that order's latest venue state and never accumulated locally; conflating the two is the classic overfill bug. avg_cost_px and realized_pnl are neither, they are a sequential fold with carry: a reducing fill realizes against the running average, an increasing fill re-weights it, and a fill through zero discards the basis entirely. The one order-free case is pure same-side accumulation, where the average collapses to sum(last_qty * last_px) / sum(last_qty).
- Give the cost fold a total order derived from the rows themselves: the venue's own monotonic execution sequence where it publishes one, otherwise (venue_ts, venue_mic, venue_exec_id), whose tail is the primary key so the order is total. The tiebreak inside one venue_ts is arbitrary but fixed, which is all reproducibility requires; it is not a claim about the venue's true matching sequence. The internal ingest sequence is an arrival artifact, the right resume watermark and the wrong fold order, because a session resend, or a drop-copy row that beats the session copy of an earlier fill, moves a reducing fill relative to the average it realizes against while leaving net_qty, the row count and every duplicate count intact.
- Treat a row that sorts below rows already folded as a rewind rather than an append: reopen that (account_id, instrument_id, business_date) from its opening snapshot and replay the day's rows for that key. Scoping the replay to the key bounds it at thousands of rows rather than the day's millions, and the rewind rate is worth counting, since a rising one is a feed or drop-copy problem rather than a fold problem. Trade corrections are the benign case, since trade_correct and trade_cancel arrive as new rows referencing corrects_exec_id with their own later venue_ts and therefore append in order. Fold them as a negating entry for the original plus, for a correction, the replacement, and never UPDATE the original row; it is the record of what the desk acted on at the time, and realized P&L attribution needs both sides.
- Fix the watermark and the read path separately. Resume from the internal monotonic ingest sequence assigned at insert, since venue_exec_id is an opaque vendor string with no ordering guarantee within a venue let alone across venues, and keep fills_applied_through_exec_id for audit. Serve intraday reads from the snapshot row rather than from execution history, so a read is a primary-key lookup instead of an O(n) fold. One writer per account partition keeps the fold single-threaded per row; if several writers can touch one position row, a check-then-act under READ COMMITTED reads a value that is stale before it acts, so use SELECT ... FOR UPDATE, a version-guarded conditional UPDATE, or SERIALIZABLE with a retry on serialization failure.
Worked solution 35 min
- Create execution_report with its primary key and load a recorded stream of 50,000 reports covering 500 orders across 40 (account_id, instrument_id) pairs, including 2,000 deliberate duplicates, 50 trade_correct rows, and the two shapes the cost fold is sensitive to: at least 200 reducing fills and 20 fills that cross through zero.
- Implement apply(report) as dedup on (venue_mic, venue_exec_id) followed by a fold that replays in (venue_ts, venue_mic, venue_exec_id) order with avg_cost_px and realized_pnl in exact decimal, then run the stream three ways: in order, shuffled, and with every row duplicated.
- Compare net_qty, avg_cost_px, realized_pnl and per-order cum_qty across the three runs.
- Run the shuffled stream once more through a control fold that replays in ingest order instead, and diff its per-pair avg_cost_px and realized_pnl against the canonical result.
- Assert cum_qty + leaves_qty = order_qty on every live order_event row the fold produced.
Follow-up
- The clearing file at T+1 says net_qty is 300 and the derived value is 295. Which row do you write, and which do you leave alone?
- Primary and replica agree on net_qty for every account and disagree on avg_cost_px for three of them. What is the smallest difference between their inputs that explains that, and which side is wrong?
- Two writers own the same account partition for 30 seconds during a failover. What breaks, and what does not?
Instrument cache flush degrades every book at once
At 13:05 a vendor correction reaches the reference service. Within two seconds its request rate rises forty-fold, p99 goes from 3 ms to 2 s, callers time out and retry, and the market data gateway marks books degraded because it cannot resolve symbology. Every service caches instrument_version with a five-minute TTL and flushes the entire cache on a correction notice. Give an ordered checklist that distinguishes a capacity problem from a stampede, and a fix that preserves as-of resolution for replay.
Approach
- Decide whether demand rose or supply fell. Compare distinct keys requested against total requests through the spike window. Forty times the requests over roughly the same key set is a stampede compounded by retries; forty times the distinct keys is genuine new demand and wants a different fix entirely.
- Check whether misses are correlated across processes. A global flush makes every instance miss the same key in the same instant, so the service receives one request per instance per key rather than one request per key. A uniform TTL does the same thing more slowly by aligning expiry across instances that started together.
- Quantify what the callers add before touching the server. A client that retries a timeout without backoff and without jitter multiplies an overload it is already inside; count retries separately from first attempts and state the amplification factor.
- Narrow the invalidation. A correction affects specific instrument_ids, and flushing discards 10^4 to 10^6 unrelated entries to fix a handful. Invalidate by key, which requires the notice to carry the affected instrument_id list rather than being a bare signal.
- Coalesce and stagger. One in-flight fetch per key per process with concurrent callers awaiting that single fetch bounds fan-out to the process count; TTL jitter stops survivors re-aligning; a bounded retry budget with backoff stops the client side recreating the burst.
- Handle the degraded case honestly. Serving a stale definition is only acceptable where the caller records which version it served, because a replay must resolve symbology exactly as the live session did; where it cannot, the correct behaviour is the one the gateway already took - mark the book degraded rather than price off a guess.
Follow-up
- Stale-while-revalidate would have kept the gateway quoting through this. When is that the right call, and when does it become a correctness bug?
- The correction notice is a topic every service subscribes to. What would you change about the notice itself?
- What signal would have made this visible as a design flaw before it was an incident?
Four days sample coding, design, fundamentals and the practical rounds at deliberately shallow depth, which is enough to surface the topics you did not know were in scope. That map, rather than a guess made on day one, decides where the last three days go.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Coding, one pass at shallow depth
- Solve one problem from each of six families, an array with two pointers, hash counting, binary search, a tree traversal, a graph traversal and one dynamic program, under a hard twenty-minute cap with no extensions, marking each finished, late, or stalled.
- For every stall, write the exact move you could not make rather than the subject, so the note reads could not turn the recurrence into a loop rather than bad at dynamic programming.
- Fix nothing today. The value of the pass is the unfixed record.
Deliverable: Six timed attempts marked finished, late or stalled, each stall carrying a named blocking move.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Design, one pass at shallow depth
- Spend twenty minutes each on three different shapes, a read-heavy feed, a write-heavy ingest path, and something needing a transaction across two entities, stopping each at requirements, interface and data model.
- After each, write the first question you could not answer, which is usually a number you could not estimate or a failure mode you had no vocabulary for.
- Mark which of the three you would be most relieved not to be asked, and treat that as data rather than as a preference.
Deliverable: Three shallow designs, each with the first unanswerable question written at the bottom.
Practice prompt ↗Practice prompt ↗03Fundamentals and the practical rounds
- Answer eight short questions in writing at four minutes each, covering the material that fills the gaps between the big rounds: what happens between a URL and a rendered page, what an index costs on write, when a process is preferable to a thread, and what conditions a deadlock requires.
- Do one thirty-minute practical task of the kind a take-home compresses: read an unfamiliar two-hundred-line file and write what it does, what you would change, and the one thing you remain unsure of.
- Score every answer fluent, correct but slow, or absent, and keep the absent ones visible.
Deliverable: Eight scored short answers and one written reading of unfamiliar code.
Practice prompt ↗Practice prompt ↗04The rounds that are about you, and the map
- Deliver three behavioural answers aloud against a timer, a conflict, a failure you owned, and a decision made without enough information, marking any that ran past three minutes or contained no number.
- Assemble the map: every marked item from days one to three on a single page, sorted by how likely it is to appear in your loop rather than by how uncomfortable it felt.
- Choose exactly two areas for the remaining three days and write down what you are deliberately abandoning.
Deliverable: A one-page scored map of the whole surface area with two areas chosen and the rest explicitly abandoned.
Practice prompt ↗Practice prompt ↗Worked solution ↗05First chosen area, to the depth you skipped
- Work the higher-ranked area in four focused blocks, choosing items one level above where you stalled rather than repeating what already works.
- After each block write the rule you extracted in one sentence with its precondition attached, since a rule carrying no precondition is exactly what fails under a variation.
- Re-attempt the day-one or day-two item that exposed this area and compare against the original timing.
Deliverable: Four worked blocks, a timed re-attempt against the original, and three one-sentence rules with preconditions.
Practice prompt ↗Practice prompt ↗06Second chosen area, where the gap is coverage rather than speed
- Treat the second area differently from the first. Day five drilled something you could already half-do; this one is usually a topic you had simply never met, so build one worked reference example end to end and keep it, rather than attempting six problems badly.
- Write down the vocabulary you were missing on day two or three, five terms at most, each with the one sentence that makes it usable in an answer rather than the textbook definition.
- Redo the shallow attempt that exposed this area and note whether you now fail later in the problem, because moving the failure point is the realistic gain from a single day and is worth more than a score that did not change.
Deliverable: One worked reference example for the newly covered area, a five-term vocabulary list, and a note on where the failure point moved.
Practice prompt ↗Practice prompt ↗07Reassemble the loop
- Sit two rounds back to back with no gap, ordering them so the area you chose second comes last, because the map was built from rested, isolated attempts and the loop will reach your weaker area when you are already spent.
- Write where the second round suffered from the first, which is normally the point at which structure collapses into narration.
- Reduce the week to one page holding only the rules you can state without reading them.
Deliverable: Mock notes on cross-round carryover plus a one-page card of rules you can recite from memory.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Team size, service count and tickets closed say very little. Seniority shows in the decision you owned: what you chose not to build, which constraint you traded away, whose objection you had to resolve before anything could move. A large project where you executed someone else's plan is a small story.
Describe a time when you had to debug a critical production issue unde…
Describe a time when you had to debug a critical production issue under pressure. How did you resolve it?
Approach
- State the situation in two sentences and spend the rest on the reasoning.
- Pick a story where you made the decision, not one where you watched it.
- Close with what you would do differently, concretely.
Follow-up
- What did you decide not to do, and why?
- How did you know your change caused the improvement?
How do you handle concurrency and race conditions in multi-threaded ap…
How do you handle concurrency and race conditions in multi-threaded applications?
Approach
- Close with what you would do differently, concretely.
- Name the disagreement and how you resolved it with evidence.
- State the situation in two sentences and spend the rest on the reasoning.
Follow-up
- What would you do differently if you ran that again?
- How did you know your change caused the improvement?
Resolve a review disagreement over floating-point prices
A pull request stores avg_cost_px as a double and checks it against a tick boundary with a tolerance of 1e-9. The author notes that every test passes and the values are small. You think it should be an integer at the instrument's price_scale. Describe a code review disagreement you resolved. State the argument you made, the evidence you produced inside the review rather than after it, how long it took, and how it ended — including the case where you were the one who changed position.
Approach
- Lead with a reproduction, not a principle. Four lines that accumulate 0.01 a hundred times and print a value that is not 1.0 settle more than a paragraph, because IEEE-754 binary64 has no exact representation of 0.01.
- Explain the author's own evidence rather than dismissing it. Binary64 represents every integer exactly up to 2^53, about 9.0e15, so a scaled value carried in a double is exact until magnitudes grow — which is why the tests pass and why the first failure is a large notional in a low-priced instrument, in production, months later.
- Name the alternative and what it buys beyond exactness: an integer at the instrument's price_scale, with tick_size expressed at that same scale, turns the boundary check into px % tick_size == 0 and deletes the 1e-9 constant that nobody in the room can derive.
- Scope the change honestly instead of expanding the review. It is not one column — it is every boundary that reads it and a migration for rows already written. Say whether you asked for it in this pull request or filed it with an owner.
- State your stopping rule. A review is not a place to win: an unresolved correctness point goes to a third person or a written decision, not to another round of comments, and a strong answer names the round at which it escalates.
Follow-up
- The author counters that NUMERIC in PostgreSQL is exact, so why not use that end to end?
- Where in this system is a double still the right choice?
- How would you detect the drift in rows already written before you migrate them?
- 01
Describe a time when you had to debug a critical production issue under pressure. How did you resolve it?
- 02
How do you handle concurrency and race conditions in multi-threaded applications?
- 03
A pull request stores avg_cost_px as a double and checks it against a tick boundary with a tolerance of 1e-9. The author notes that every test passes and the values are small. You think it should be an integer at the instrument's price_scale. Describe a code review disagreement you resolved. State the argument you made, the evidence you produced inside the review rather than after it, how long it took, and how it ended — including the case where you were the one who changed position.
Is this an official Quantifind interview guide?
No. It is PracHub's own research and practice material for the Software Engineer role at Quantifind. Rounds and questions reflect what candidates have reported, not a process Quantifind has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult are the technical interviews at Quantifind?
The technical interviews strike a balance between rigorous problem-solving and practical engineering discussions. While coding questions test your core algorithmic competence, the overall process prioritizes real-world system design and collaboration over trick questions or hostile grilling.
PracHub interview research ↗How much preparation time should I plan for?
Most candidates benefit from dedicating two to four weeks of focused preparation. This allows sufficient time to review core data structures, practice live coding problems, brush up on system design principles, and review your past project history.
PracHub interview research ↗What differentiates successful candidates during the loop?
Successful candidates stand out by communicating their thought process clearly, asking clarifying questions when faced with ambiguity, and demonstrating genuine curiosity about the data ecosystem. Interviewers value engineers who collaborate well and treat the interview as a two-way technical discussion.
PracHub interview research ↗What is the typical timeline from initial screen to offer?
The entire interview process is generally organized and fast-moving, often spanning two to three weeks from your initial recruiter conversation to the final decision. Correspondence is typically quick, reflecting the team's organized and respectful approach to candidate experience.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Software Engineer practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22