This guide covers what a Software Engineer at Nimble is expected to do and how to prepare for the interview.
Take-Home Assignment
reportedThe README is read before the code, and a follow-up conversation is usually built from it, so treat every sentence you put there as a question you have agreed to answer. It needs the command that runs the thing, the assumptions you made where the prompt was ambiguous, and the limits of what you built stated with the preconditions that make them true. Overclaiming is the expensive mistake here. Writing that something is thread-safe, or constant-time, or handles files larger than memory invites a reader to check that exact line, and a claim the code cannot support costs more than silence would have.
What to demonstrate
- Whether the run instructions work from a clean clone, naming the exact commands, the language version you tested on, and any environment variable the program expects
- Whether ambiguities in the prompt are resolved in writing, with the interpretation you picked and the reason, rather than settled silently in the code
- Whether documented limits match the implementation, so a stated input bound is one the code enforces or at least does not contradict
- Whether the trade-offs you list come with the condition that would make you choose the other way, instead of reading as a list of alternatives you happened to consider
How to prepare
- Write the README before the final hour, then read the code against it claim by claim and correct or delete every statement the implementation does not back
- For each ambiguity in the prompt, write one sentence fixing your interpretation and keep it; those sentences become the assumptions section and your answer when someone asks why you did it that way
- Give the repository to someone who has not seen the prompt and ask them to run it using only what is written down, treating every question they have to ask you as a gap in the document
Technical Discussions
reportedThe same problem is scored by two different mechanisms depending on the format, and preparing for one does not cover the other. With a person watching, partial progress is visible and a hint is a correction you can absorb; silence is the expensive failure, because nobody can read a half-written function. With an automated grader there is no partial credit for what you were about to do, nobody to ask, and the worked examples in the prompt are the entire specification. Read them as a contract, down to whether an empty result should be an empty list or no output at all.
What to demonstrate
- In a live session, whether your commentary tracks what your hands are doing, and whether a hint redirects you or gets defended against
- In an automated one, whether you cover the cases the examples do not show, since the hidden cases are where the score moves
- Whether you manage the clock on purpose: abandoning an approach that is not converging while there is still time to write something simpler that finishes
How to prepare
- Have someone hand you a problem and feed you one deliberately wrong hint. Practise testing it against a concrete case instead of accepting or rejecting it on authority.
- Do one timed run a week in a plain browser editor with autocomplete, linting and your own snippets switched off, which is closer to what these environments give you
- For the automated format, write the harness before the solution: a main that feeds the worked examples plus an empty and a single-element case and prints expected against actual, so a wrong submission is caught by you first
Leadership Interview
reportedAn unlabelled round is first an information problem, and the cheapest information is free. Whoever schedules it can usually tell you how long it runs, who will be in the room and what they work on, whether you will be writing code and in what environment, and whether anything is being sent beforehand. Ask in writing so the answer is on record, then prepare for the two or three formats those answers still leave open instead of betting on one. What separates a strong candidate is not guessing right; it is having an opening that works whichever one it turns out to be.
What to demonstrate
- Whether you can start work from an ambiguous brief, since tolerating a vague scope without stalling is the same thing the job asks for
- Whether the questions you asked beforehand were ones that change your preparation, such as duration, medium and who is joining, rather than ones whose answers you could not have acted on
- Whether you adapt when the round turns out to be something other than what you were told, instead of spending the first ten minutes visibly recalibrating
How to prepare
- Send one short scheduling message asking four things: how long, who is joining and what they work on, whether you will be writing code and where, and whether to prepare anything in advance. Treat a vague reply as real information, since it means the round is loosely structured and you will be shaping it yourself.
- Write one opening that works in any of the formats still open: restate in your own words what you have been asked to do, then ask which of two directions is more useful to them. Say it aloud until it stops sounding recited.
- Set up for the two most likely formats before the call starts, with a blank editor in the language you would choose and a shared document you can type into, so a format surprise costs you nothing in the first minutes
1 candidate reports. Individual accounts describe a particular role and hiring cycle.
Nimble Technical Screen Interview Experience — A TTL-LRU Cache Nobody Could Agree On, Same-Day Rejection
View report detailsPracHub editorial advice for the preparation topics above.
Holding a database transaction open across a physical or third-party operation
It is natural to open a transaction, lock the position, call the rating or tendering API, and commit on the response, and it works perfectly until the partner's p99 goes from 200 milliseconds to thirty seconds. At that point every request holding a lock on a hot item-node pair queues behind it, the connection pool fills with transactions that are waiting on the network rather than on the database, and an unrelated service sharing the pool fails at the same moment. In this domain the effect is amplified because demand concentrates on a few hot rows during a promotion or a seasonal peak, exactly when partner latency is also degraded. The structural fix is to keep transactions short and local - commit the state change together with an outbox row, let a relay perform the external call, and reconcile asynchronously - accepting at-least-once delivery and making the effect idempotent rather than trying to stretch a transaction over something the database cannot roll back.
Floating-point quantities and implicit unit-of-measure conversions
Binary floating point cannot represent most decimal fractions exactly, so a chain of pallet-to-case-to-each conversions accumulates a residue that a later rounding turns into a unit that was created or destroyed, breaking conservation with no concurrency involved at all. The bug is slow and non-local: it surfaces as a cycle-count variance weeks later, at a node nobody changed, and it is untraceable because the arithmetic looked correct at every individual step. The fix is to hold quantities as integers in the smallest transacting unit, carry the UoM explicitly on every row that carries a number, store conversion factors as integers that are versioned in time, and reject a conversion that does not divide exactly instead of rounding it. The same discipline applies to money on the freight side, where an accessorial in floating-point dollars produces invoice reconciliation differences of a cent that cost more to investigate than the freight.
Assuming the input fits in memory
Ask how large the input is in bytes before committing to an in-memory algorithm; beyond that point the options are a single streaming pass, an external sort with bounded buffers, or a sketch that trades exactness for constant memory. An algorithm that assumes random access to the whole input is a different algorithm from one that sees each element once.
Choosing a schema before the access patterns are known
Write the queries first, with their filters, sort orders, cardinalities and which ones sit on the latency-critical path, then design tables and indexes to serve them. An index nothing queries still costs write throughput and storage, and a hot query with no supporting index becomes a full scan that only hurts once the table is big.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Trace a suspect lot forward to every shipped order line
lot_edge(parent_lot_id, child_lot_id) has 400 million rows with btree indexes on both (parent_lot_id, child_lot_id) and (child_lot_id, parent_lot_id). lot_shipment(lot_id, order_line_id) maps lots to shipped lines. Given one suspect received lot, return every downstream order_line_id. The reachable subgraph can reach 2 million lots and 12 levels; the full edge table does not fit in memory. Give the traversal, its complexity in terms of the reachable subgraph, and what you do if the data contains a cycle.
Approach
- Genealogy is a directed acyclic graph, not a tree: a production run consumes several lots and one lot splits across many shipments, so a node is reached by several paths and a tree traversal revisits it combinatorially. Breadth-first from the suspect lot over the
(parent_lot_id, ...)index, with a hash set of visited lot ids, is O(V' + E') where V' and E' count only the reachable subgraph — 2 million 8-byte ids is 16 MB of payload and a few tens of MB with open-addressing overhead, which is why the visited set is affordable and the edge table's 400 million rows never enter memory. - Batch each level into one indexed read. Issue
WHERE parent_lot_id = ANY($1)with a few thousand parents per call instead of a query per node: at 12 levels and 2 million nodes, per-node round trips dominate the wall clock entirely, and the recall's deadline is a real-world response time, not a query-plan aesthetic. The batch size is the tuning knob between round-trip count and per-statement planning cost. - Keep the visited set even though the schema says DAG. Rework and repack loops — material returned into an earlier lot — do occur, and an unguarded traversal on a cyclic edge set does not terminate; a set membership test is O(1) and turns a hang into a correct answer plus a data-quality alert. Also cap the traversal by level and by node count, and fail loudly on the cap rather than returning a partial set silently.
- Join to
lot_shipmentonce at the end over every visited lot id, not over the leaves and not per node. A lot is routinely both shipped and consumed: half a lot goes out on an order line while the other half is repacked into a child lot, so it has alot_shipmentrow and outgoing edges at the same time and is not a leaf. Restricting the join to nodes with no outgoing edges drops exactly those order lines, which is under-recall — the one failure mode this trace is not allowed to have. One bulk join against an index onlot_idover the whole visited set is a single pass; per-node lookups multiply the round-trip problem by the number of visited lots. - State the semantics honestly: reachability answers which shipped lines could contain material from the suspect lot, and without propagating edge quantities it is a superset of physical containment. For a recall that is the correct direction of error — over-recall, never under-recall — and saying so is the point, because a trace that returns quickly and incompletely will be believed. The upstream question, which received lots are in a suspect shipped unit, is the same traversal over the reverse index, which is why both indexes exist.
Worked solution 30 min
- Draw a six-lot fixture containing a merge (two parents into one child), a split (one parent into three children), and one interior lot that is partly shipped on an order line and partly consumed into a child lot, then write down the expected reachable set and the expected order lines by hand.
- Implement level-batched BFS with an explicit frontier list, a visited hash set, and a per-level batch read.
- Add a cycle to the fixture and confirm the traversal terminates and flags it rather than hanging.
- Bulk-join
lot_shipmentonce over every visited lot id — the whole visited set, not the leaves — and count the queries issued, confirming it is one per level plus one. - Re-run with the frontier order reversed to confirm the output set is unchanged.
Follow-up
- A recursive CTE would express this in ten lines. What is its cost here if the index on the traversal column is missing, and how would you notice?
- You need the quantity of suspect material in each downstream line, not just reachability. What has to be on the edge, and what does that do to the traversal?
- The traversal must run while a production run is writing new edges. What does a consistent answer mean here, and what do you cut it at?
Split an order across the fewest nodes that can fill it
A fulfilment order has L lines, L up to 25. After filtering by capability, service time and stock, N candidate nodes remain; N is usually under 14 but can reach 200. Each node covers a known subset of the lines in full. Return a minimum-size set of nodes covering every line, tie-broken by total distance. Give an exact algorithm for the small case with its exact cost in operations and bytes, say precisely why enumerating node subsets stops working, and give what you ship when N is 200.
Approach
- Name it: this is minimum-cardinality set cover, which is NP-hard, so the useful move is choosing which parameter you are exponential in rather than hunting for a polynomial exact algorithm. Enumerating subsets of nodes is 2^N and correct — at N = 14 that is 16,384 subsets and finishes instantly, which is exactly why it survives review and then fails in production at N = 200, where 2^200 is about 1.6e60.
- Be exponential in L instead, because L is the parameter you bound. Index the DP by the set of lines still uncovered:
dp[rem]is the cheapest way to cover the lines inrem, withdp[0] = (0, 0)and the value a(node_count, total_distance)pair compared lexicographically. Letbbe the lowest set bit ofrem; any cover must contain some node that stocks lineb, so the transition ranges only over those nodes:dp[rem] = min over v with b in cover(v) of dp[rem & ~cover(v)] + (1, dist(v)). Worst case O(2^L * N) — at L = 25 and N = 14 that is 33,554,432 states times 14, about 4.7e8 transitions — and the lowest-bit restriction cuts the real branching factor to the nodes stocking one specific line. - Quote the memory for what you actually store, not for a count array alone. Over 2^25 masks you need three dense arrays:
uint8node count (33.5 MB),uint32total distance (134.2 MB, integer metres so the tie-break compares exactly and the array stays fixed-width), anduint8chosen node (33.5 MB, since N <= 200 fits a byte) — about 201 MB, roughly six times the 33.5 MB that a bare count array would suggest. The same layout at L = 20 is 6.3 MB, which is why this is comfortable up to about 20 lines and needs a deliberate decision at 25. - Reconstruction is free in this formulation and needs no parent-mask array: because every transition eliminates the lowest uncovered line, the predecessor of
remunder the stored choicevis exactlyrem & ~cover(v), so walkremfrom the full mask down to 0 emittingchoice[rem]. If the reachable state count is far below 2^L — coverage sets are usually highly correlated — a top-down memoized recursion over a hash map trades the 201 MB for roughly 40-60 bytes per visited state; measure the visited count first, because a hash map that ends up touching most masks is strictly worse than the dense arrays. Add a branch-and-bound cut ofceil(popcount(rem) / max_v popcount(cover(v) & rem)), which prunes long before the theoretical bound bites. - When exact is out of reach, ship greedy with a stated guarantee: repeatedly take the node covering the most currently-uncovered lines, using bitset popcounts, O(N * ceil(L/64)) word operations per pick and at most min(N, L) picks. Greedy is within H(L) <= ln L + 1 of optimal (Chvatal), and no polynomial algorithm achieves (1 - o(1)) ln L unless P = NP (Dinur and Steurer, 2014) — so effort belongs in the pre-filter and the tie-break, not in a cleverer heuristic. At L = 25 the bound is H(25), about 3.82x worst case, while the realistic gap on correlated coverage sets is zero or one node.
- Close on the objective, because fewest nodes is not cheapest. Two shipments from distant nodes routinely cost more than three from near ones, and a split also costs a second box, a second carrier pickup and a worse customer experience. If cost is the real objective it is weighted set cover — greedy by cost divided by newly-covered lines, same logarithmic bound — and if a node can cover a line only partially, the line mask is no longer sufficient state, since you must track remaining quantity per line; that is the point at which you stop hand-rolling and either restrict splits or hand an MILP to a solver with a time budget.
Follow-up
- The pre-filter is what keeps N small. What is in it, and what happens to your exact path the day someone loosens it?
- One node can cover a line only partially. Show where the bitmask formulation breaks and what state replaces it.
- You have 40 milliseconds inside a checkout call. Which of these runs there, and what runs asynchronously afterwards?
Size dock doors from overlapping appointment windows
For one node and one local calendar day you have up to 200,000 planned dock appointments from shipment_leg: leg_id, planned_arrive_at, planned_depart_at (both TIMESTAMPTZ), plus the node's IANA zone. A door is occupied over [arrive, depart). Return the minimum number of doors that lets every appointment start on time, and the maximal interval over which that peak is sustained. Target O(n log n). Say how you derive the day's boundaries and how you break ties between an arrival and a departure at the same instant.
Approach
- Name the result before computing it: with one interchangeable door type, the minimum door count equals the maximum number of simultaneously occupied intervals. That equality is not a heuristic — interval graphs are perfect, so their chromatic number equals their clique number, and greedy left-to-right assignment achieves it. Say this, because it is what licenses solving a scheduling question with a counter.
- Emit 2n endpoints,
(t, +1)at each arrival and(t, -1)at each departure, sort byt, and at equaltorder the-1before the+1. Half-open occupancy means a trailer leaving at 10:00 frees the door for one arriving at 10:00; ordering arrivals first inflates the answer by exactly the number of back-to-back handoffs, which on a well-packed schedule is most of them. - Scan once with a running counter, tracking the maximum and the endpoint index where it was first reached. The peak is attained on a half-open interval between two consecutive endpoints,
[t_i, t_{i+1}), not at an instant — report it that way or the operations team cannot act on it. Extend the interval while the counter stays at the maximum. - Complexity: O(n log n) dominated by the sort, O(n) space. If rows already arrive ordered by
planned_arrive_atfrom an index, use the min-heap variant instead — push each departure, pop all departures at or before the current arrival, and the heap size is the current occupancy. Same time bound, O(peak) space rather than O(n), which matters when peak occupancy is 40 and n is 200,000. - Derive the day's boundaries from the zone, not from arithmetic on instants. Convert local midnight and the next local midnight to instants in the node's IANA zone; on a transition day that span is 23 or 25 hours, so adding 86,400 seconds silently drops or duplicates an hour of appointments. Decide explicitly what a zero-length dwell means — under half-open semantics it occupies nothing — and state it.
Follow-up
- Doors are typed: some break pallets, some are parcel-only. Does max-overlap still equal the door count?
- Appointments have a tolerance — a trailer may start up to 20 minutes late without penalty. How does that change the objective, and is it still solvable by a sweep?
- Half the appointments are actuals and half are plans. Which do you sweep for tomorrow's staffing, and which for last week's utilisation report?
Stop two allocators promising the same unit of stock
inventory_position is keyed (item_id, node_id, lot_id, state) and holds qty BIGINT, which is on-hand stock and is moved only by the movement ledger; available_qty BIGINT, maintained on the same row; and version. The declared invariant is available_qty = qty - SUM(allocation.qty) over that key in status held or committed, enforced by CHECK (available_qty >= 0) and CHECK (available_qty <= qty). allocation holds allocation_id, order_line_id, item_id, node_id, lot_id, qty, tier, status (held, committed, consumed, released, expired) and expires_at. Twelve allocator threads try to claim the last eight units of one item at one node inside the same 50 ms. Give the exact statements that let at most eight units be claimed, then say what changes if identical code runs under PostgreSQL REPEATABLE READ versus InnoDB REPEATABLE READ, and what a retried allocate for the same order line must do.
Approach
- Collapse the check and the write into one statement whose WHERE clause is the business rule: UPDATE inventory_position SET available_qty = available_qty - :q, version = version + 1 WHERE item_id = :i AND node_id = :n AND lot_id = :l AND state = 'on_hand' AND available_qty >= :q. qty is deliberately untouched: a hold is a promise, not a physical move, so only a movement may change on-hand and the ledger-equals-position invariant survives. Rowcount zero is a refusal, not an error, and the allocation row is inserted in the same transaction only when the rowcount is one.
- If the deployment carries no available_qty column and availability must be derived, the alternative is SELECT ... FOR UPDATE on the position row followed by the SUM over held and committed allocations inside the same transaction. That is correct but holds the row lock for two statements instead of one, which matters on a key taking hundreds of writes a second.
- Isolation is not a substitute for the atomic step, and it changes the error contract. Under PostgreSQL READ COMMITTED the losing UPDATE blocks, then re-evaluates its predicate against the newly committed row version and returns rowcount 0, which the handler already handles. Under PostgreSQL REPEATABLE READ the same statement instead raises SQLSTATE 40001 and the whole transaction must be retried, so a handler that only tests rowcount fails closed on an exception it never expected. InnoDB REPEATABLE READ does neither: a locking read or UPDATE performs a current read of the latest committed row, so the loser blocks and then sees rowcount 0, with a deadlock rather than a serialization failure as its exceptional case.
- Make the retry a no-op with a partial unique index: UNIQUE (order_line_id, node_id, lot_id) WHERE status IN ('held','committed'). Then a retried allocate conflicts instead of double-claiming, and the handler reads the existing allocation and returns it.
- Never source the availability number from a replica or a cache for this path. Replication lag is largest during the write burst that made the key hot, so the guarantee fails precisely when it is load-bearing.
- For the hottest item-node pairs, shed load without overselling: route allocations for that key through a single writer with a bounded queue and refuse with a conservative not-available when the queue is full. Splitting the position into N sub-rows cuts contention N-fold but strands the remainder across sub-rows and produces false out-of-stock when the last units are spread thin.
Follow-up
- One order line needs two units from two different nodes, atomically. What do you do when those positions live in different shards and no distributed transaction is available?
- Soft-tier allocations expire. Show the sweeper statement, and say what stops it from releasing an allocation that is being committed at that instant.
- Under sustained contention your CAS retries. What is the retry budget, and what does the caller see when it is exhausted?
Attribute multi-stop freight cost without fanning it out
shipment_leg(leg_id, shipment_id, sequence_no, linehaul_cost_cents, accessorial_cost_cents, actual_arrive_at) joins to fulfilment_order_line through leg_order_line(leg_id, order_line_id, allocated_weight_g). One multi-stop truckload leg carries 40 order lines. The current report SUMs linehaul_cost_cents after joining all three tables and returns roughly forty times the real freight spend. Write the query that returns cost per order for one month, explain the fan-out in terms of the join's row count, and make the per-order amounts sum exactly to the leg totals in integer cents.
Approach
- Name the cardinality: joining shipment_leg to leg_order_line produces one row per (leg, line) pair, so the result has SUM over legs of lines-per-leg rows rather than one row per leg. SUM(linehaul_cost_cents) over that result counts each leg's cost once per line it carries, which is why 40 lines gives 40 times the spend. The bug is not the SUM, it is aggregating a parent measure across a child-expanded row set.
- Compute a share before aggregating anything. In one pass: leg.linehaul_cost_cents * b.allocated_weight_g / SUM(b.allocated_weight_g) OVER (PARTITION BY b.leg_id), which keeps the denominator local to the leg and needs no self-join.
- Do the division in integer cents and settle the remainder deliberately. Floor each share, then distribute the leftover cents one each by descending remainder, ordered by a stable tiebreak such as order_line_id. Floating-point shares leave a residue that makes the freight accrual miss by a cent per leg, which costs more to investigate than the freight.
- Aggregate to the order only after the share exists: GROUP BY order_id over the per-line allocated cents. Restrict the month on shipment_leg.actual_arrive_at, not on the order date, so a leg is attributed in the period it actually ran.
- Index for the shape of the query: leg_order_line needs (leg_id) for the window partition and (order_line_id) for the reverse lookup, and shipment_leg needs an index on actual_arrive_at to bound the month without a full scan.
- Apply the same treatment to accessorial_cost_cents separately. Accessorials such as a detention charge often belong to one stop rather than to the whole leg, so weight-sharing them across every line on the load misattributes cost even though the total ties out.
Worked solution 30 min
- Build a fixture with one leg at 100,003 cents carrying three lines weighted 1000 g, 1000 g and 1 g.
- Write the naive join and observe the total reported as three times the leg cost.
- Rewrite using the SUM(...) OVER (PARTITION BY leg_id) share, floor to integer cents, and add a largest-remainder pass ordered by remainder then order_line_id.
- Aggregate to order level and compare the grand total against SELECT SUM(linehaul_cost_cents + accessorial_cost_cents) FROM shipment_leg for the same window.
Follow-up
- A line is cancelled after the leg ran. Does its share of the freight disappear, move to the remaining lines, or stay as a cost with no revenue?
- Weight is the wrong basis for a light bulky pallet. What would you use instead, and does the reconciliation check still hold?
- The carrier invoice arrives three weeks later and differs from the rated cost. Where does the correction land, and does it re-run this allocation?
Describe your process for ensuring safety and reliability in a distrib…
Describe your process for ensuring safety and reliability in a distributed software system.
Approach
- Work from the requirement backwards to the design.
- Clarify what is being asked and what a complete answer contains.
- State your assumptions explicitly before working the problem.
Follow-up
- What assumption would you test first?
- How would you know your answer was wrong?
How do you approach debugging a race condition in a high-performance s…
How do you approach debugging a race condition in a high-performance system?
Approach
- State your assumptions explicitly before working the problem.
- Say what you would check first and why it is the highest-information step.
- Work from the requirement backwards to the design.
Follow-up
- What assumption would you test first?
- How would you know your answer was wrong?
Walk me through your design process when building a system from scratc…
Walk me through your design process when building a system from scratch.
Approach
- Say what you would check first and why it is the highest-information step.
- Work from the requirement backwards to the design.
- State your assumptions explicitly before working the problem.
Follow-up
- How would you know your answer was wrong?
- What assumption would you test first?
Keep a distribution centre picking through a six-hour WAN outage
A site's link to the rest of the system fails for six hours mid-shift. The Warehouse Execution Service runs on site and holds the work already released to it; the Availability Service, the Fulfilment Orchestrator and every other node are on the far side of the partition. Decide what the floor may keep doing, what central sourcing may promise from that node meanwhile, and how buffered movements and a cycle count taken during the outage are reconciled on reconnect. Name the invariant you relax and for how long.
Approach
- Name both sides and what each still knows. The site holds released work with hard allocations and can still generate movements; it cannot see new demand or any other node. Central holds demand and the ledger but is blind to picks. The relaxed invariant is the freshness of that node's derived position, bounded by the outage. Conservation itself is not relaxed: movements are durable locally with their idempotency keys and will be applied exactly once on replay.
- Let the floor continue, but only on work already released. Buffer movements to local durable storage in occurred_at order, each carrying its idempotency_key, the device clock offset and the reference that caused it. The floor stopping costs labour immediately and visibly, while a six-hour-stale record costs promise accuracy that can be bounded on the other side - so the trade is worth making, and stating that reasoning is the answer, not the choice itself.
- Bound central's exposure instead of pretending it has data. After a staleness threshold, freeze that node's available quantity at last-known on_hand minus everything already released to it; after a longer threshold, drop the node out of sourcing entirely. The error direction is under-promising: a missed sale rather than a short ship, a re-source and a second shipment's freight.
- Defend against the real hazard, which is the double ship. Central will want to re-source a line the site is actively picking. Make a released line's hard allocation unexpiring and require an acknowledged cancel from the site before re-sourcing; with the site unreachable no ack is possible, so the line stays with the site and is re-sourced only after reconnect confirms it unpicked. An allocation expiry sweeper that fires during a partition is precisely how two trucks leave with the same unit.
- Write the reconnect protocol as an order, not a hope: replay buffered movements by occurred_at with their idempotency keys, so a replay attempted twice is a no-op; then run the reconciliation re-sum for that node bounded by each row's last_movement_id, so writes still in flight do not read as variances; then release allocations the site reports it never consumed; then re-enable the node for sourcing.
- Apply the cycle count as evidence, not as a command. Record counted_qty with its count time and derive the adjustment delta at apply time against the position as of that moment. Replaying a precomputed set-to-40 overwrites every pick that happened after the count was taken, which is the version of this mistake that survives review because the number looks authoritative.
Worked solution 40 min
- Draw the partition, list what each side can still read and write, and write the relaxed invariant as one sentence with an explicit time bound.
- Write the site buffer's record format and the exact fields a buffered movement must carry to be replayable months later.
- Write central's degradation rule: the frozen availability formula, both thresholds as numbers, and the direction of error with its cost on each side.
- Write the reconnect sequence in order, then work the count-versus-picks case concretely - a count of 40 at 14:10 followed by twelve picks before the link returns - and show the resulting position.
- State the residual damage the design accepts and the signal that detects it.
Follow-up
- The site's local buffer is lost during the outage. What is still recoverable, and from which records?
- Two sites partition at once and both were candidate sources for the same lines. What does sourcing do differently from the single-site case?
- How do you choose the two staleness thresholds, and what measurement tells you afterwards that they were wrong?
Availability p99 spikes in a sawtooth locked to the cache TTL
Available-to-promise reads sit inside checkout with a 12 ms p99 behind a 60-second per-key cache. During a promotion, p99 spikes to 900 ms in a sawtooth exactly on 60-second boundaries and is normal between spikes. The inventory database shows bursts of thousands of identical single-row position reads inside 50 ms across about 40 item-node keys, and the connection pool saturates during each burst. Request throughput is unchanged. Diagnose, fix, and state which callers may read a cached availability and which may not.
Approach
- Separate a stampede from saturation by the shape. Latency that rises smoothly with request rate is capacity; a sawtooth whose period equals the TTL, with normal service between peaks, is expiry-driven. Confirm by overlaying spike timestamps on key expiry times and watching the hit ratio fall to near zero for a few milliseconds each cycle.
- Size the herd with Little's law rather than adjectives. Concurrency equals arrival rate times service time, so a key taking 1,200 requests per second with a 30 ms origin read admits about 36 simultaneous origin reads at each expiry, and 40 keys that expire together produce roughly 1,440 against a 100-connection pool. That fourteen-fold oversubscription is the 900 ms, and the arithmetic matching is what confirms the diagnosis.
- Explain why the 40 keys expire in lockstep: they were all populated in the same second when promotion traffic arrived, so their TTLs are synchronised and stay synchronised. Jitter of plus or minus twenty percent on the TTL desynchronises them; it does not reduce the per-key herd.
- Bound the per-key herd with single-flight. One in-flight origin read per cache key, with concurrent callers awaiting the same result, takes origin load to one read per key per TTL regardless of request rate. Add serve-stale-while-revalidate so the waiters get the previous value immediately and staleness is bounded and measured rather than incidental.
- Draw the correctness boundary explicitly, because it is not a performance question. Browse and product pages may read a cached or stale availability. The allocation path may not: the check and the write must be one atomic step against the primary, so a cached number can inform a display but must never authorise a reservation. Checkout sits in between and should either read through with a short TTL or accept that its number is advisory and can fail at allocation.
- Decide the degradation direction in advance. If the origin is saturated, the read path should shed to a conservative answer, showing limited availability rather than a precise count and never rounding upward, because an optimistic stale read converts a latency incident into oversells and short-ships.
Follow-up
- What does your single-flight do when the origin read times out: fail the waiters, or serve them the stale value, and what bounds the staleness?
- One key gets hot enough that even one origin read per TTL per replica is too much across 200 application instances. What changes?
- How would you invalidate a key on a movement write without opening a second stampede on the invalidation itself?
Roughly ninety minutes on weeknights with one longer weekend block. The plan cuts scope rather than compressing everything, on the assumption that one thing finished per night beats four half-started.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Fix the scope and take a cold baseline
- Read the role description and write the three things the loop will almost certainly test, then write an explicit not-doing list and keep it visible all week.
- Take one twenty-five-minute coding problem and one fifteen-minute design prompt cold, and write the single sentence naming what blocked each, because those two sentences decide where the remaining evenings go.
- Set the week's rule: one thing finished every night, including the night you only have forty minutes.
Deliverable: A one-page scope with a not-doing list and two cold attempts, each carrying one sentence on what blocked it.
Practice prompt ↗Practice prompt ↗Worked solution ↗02One pattern, written three times from blank
- Choose the single pattern most likely to appear in your loop and write it three times from an empty file rather than editing the previous attempt.
- On the third pass, write the invariant as a comment before the loop body and the complexity before the first line of code.
- Stop at ninety minutes even if the third version is imperfect, and write the one thing you would fix given another hour.
Deliverable: Three independent implementations of the same pattern plus a note on what changed between them.
Practice prompt ↗Practice prompt ↗03One design, only to the depth you can defend
- Take one system shape and go only as far as requirements, interface and data model, refusing to draw a box you could not survive a follow-up about.
- Attach one number to each non-functional requirement, deriving it rather than asserting it, and write the assumption the number rests on.
- Write the one tradeoff you are choosing against and the observation that would make you reverse it.
Deliverable: One design at interface-and-schema depth with derived numbers and one written reversible tradeoff.
Practice prompt ↗Practice prompt ↗04Only the fundamentals you will have to defend
- Write, in under two hundred words each, the answers to the two questions that follow almost any implementation: why this structure and not the obvious alternative, and what happens to this code at a hundred times the input.
- Write what an index actually costs: faster lookups on the indexed columns against a write that now maintains a second structure, plus the cases where the planner declines to use it anyway, low selectivity, or a predicate wrapping the column in a function.
- Delete any answer you cannot deliver aloud in under a minute, since an answer that needs reading is not an answer you have.
Deliverable: Three written answers, each under two hundred words and each timed aloud.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Your own work, timed
- Write a ninety-second and a four-minute version of your main project and time both aloud rather than reading them.
- Prepare the two follow-ups that always come: what you would do differently, and how you knew it worked.
- Put one number in the first sentence and be ready to say exactly where it came from and what it excludes.
Deliverable: Two timed narratives with one defensible number in the opening line.
Practice prompt ↗Practice prompt ↗06The one full rehearsal, in the weekend block
- Run a sixty-minute mock covering a coding round and a design round in one sitting with no break, because sustained attention is the thing evenings have not trained.
- Immediately afterwards, and before hearing any feedback, write the three moments you lost the thread.
- Spend the rest of the block only on those three moments, and on nothing you merely feel shaky about.
Deliverable: Mock notes naming three failure moments with a specific fix written under each.
Practice prompt ↗Practice prompt ↗07Taper
- Write the twenty-minute warm-up you will actually do on the morning: one problem you can already solve from a blank file, one design you can narrate, and nothing you have never seen.
- Re-read only your own notes from this week and open no new material.
- Write the logistics down: the editor or shared document you will be working in, whether execution and lookups are permitted, and the sentence you will use when you do not know something.
Deliverable: A one-page card holding the design structure, the project numbers, and the logistics.
Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Every story you tell gets read for blast radius and judgement: what could have broken, who else it touched, what you knew at the moment you decided. Nobody can audit your code in an hour, so they audit your reasoning instead. Pick work where the call was genuinely yours and the consequences were real enough to remember.
Tell me about a "what-if" scenario where you had to pivot your technic…
Tell me about a "what-if" scenario where you had to pivot your technical strategy under pressure.
Approach
- Pick a story where you made the decision, not one where you watched it.
- State the situation in two sentences and spend the rest on the reasoning.
- Close with what you would do differently, concretely.
Follow-up
- What did you decide not to do, and why?
- How did you know your change caused the improvement?
What is your personal mission statement, and how does it align with wh…
What is your personal mission statement, and how does it align with what we are doing at Nimble?
Approach
- Name the disagreement and how you resolved it with evidence.
- Close with what you would do differently, concretely.
- State the situation in two sentences and spend the rest on the reasoning.
Follow-up
- How did you know your change caused the improvement?
- What would you do differently if you ran that again?
Ship freight cost attribution with named, dated debt
You have eight working days to get freight cost onto order margin. The correct model allocates each shipment_leg cost across the orders on that leg through a bridge, with a stated basis such as billable weight, and attaches accessorials to the stop that incurred them. The version that fits the window copies linehaul_cost_cents onto each line, split evenly. Describe a time you shipped known debt: exactly what you cut, the guard that made the shortcut visible rather than silent, who agreed to carry it, the date it was paid down, and what it cost while it stood.
Approach
- State the accepted failure concretely rather than as a category: on a multi-stop load, an even split charges a two-pallet order the same freight as a twenty-pallet order, and a detention accessorial incurred waiting at one stop is smeared across orders that were not there. Margin on small orders is understated and on large ones overstated, in a way that is systematic rather than random, so it will not average out.
- Make the shortcut loud. Emit a counter on legs with more than one order, tag every attributed row with the basis used so downstream consumers can filter, and report the share of cost that went through the even split. A shortcut that reports its own frequency converts an unknown into a number you can take to the paydown conversation.
- Bound the exposure before shipping: count historical legs with more than one order and the weight dispersion within them. If ninety percent of legs are single-order parcel, the debt is small and quantified; if a third are multi-stop truckload, the plan does not survive contact and you learned that in a query rather than in a finance review.
- Handle the arithmetic correctly even in the shortcut, because integer minor units do not divide evenly. Splitting 100,003 cents across three orders must assign the residue deterministically, to the largest order or to the first by id, and the sum of the parts must be asserted equal to the leg cost; rounding each share independently is how a reconciliation against the carrier invoice ends up off by cents nobody can explain.
- Name the owner, the date and where both are written, then report what actually happened: how much cost flowed through the shortcut, whether the paydown date held, and if it slipped, by how long and what finally forced it.
- Distinguish debt you can pay from debt that compounds. An attribution you can recompute from shipment_leg and the bridge later is recoverable; one that overwrites the leg cost or discards the stop detail destroys the inputs and cannot be unwound, and a strong answer says it would have refused that version at any deadline.
Follow-up
- The paydown date arrives and the team has a new deadline. What is different about this conversation from the first one?
- Finance has been reporting margin off this for six weeks. What do you tell them, and what do you restate?
- Which version of this shortcut would you have refused outright, and why that one?
- 01
Tell me about a "what-if" scenario where you had to pivot your technical strategy under pressure.
- 02
What is your personal mission statement, and how does it align with what we are doing at Nimble?
- 03
You have eight working days to get freight cost onto order margin. The correct model allocates each shipment_leg cost across the orders on that leg through a bridge, with a stated basis such as billable weight, and attaches accessorials to the stop that incurred them. The version that fits the window copies linehaul_cost_cents onto each line, split evenly. Describe a time you shipped known debt: exactly what you cut, the guard that made the shortcut visible rather than silent, who agreed to carry it, the date it was paid down, and what it cost while it stood.
Is this an official Nimble interview guide?
No. It is PracHub's own research and practice material for the Software Engineer role at Nimble. Rounds and questions reflect what candidates have reported, not a process Nimble has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult is the interview process?
The process is widely considered difficult and rigorous. It is designed to challenge your technical depth and your problem-solving speed, so ensure you are well-practiced in both coding and architectural design.
PracHub interview research ↗How much time should I set aside for the take-home challenges?
Treat these challenges as a significant commitment. They are a core part of the evaluation process, and you should dedicate several days of focused effort to ensure your solutions are robust, clean, and well-documented.
PracHub interview research ↗What is the culture like at Nimble?
The culture is mission-driven and focused on solving hard, real-world problems. You will be surrounded by engineers who value technical excellence and a proactive, collaborative approach to complex system challenges.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-24 - 02PracHub Software Engineer practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-24 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-24