As a Software Engineer at Delivery Hero, you are at the heart of a global ecosystem that powers food delivery, logistics, and quick-commerce services for millions of users across the world. You will contribute to high-scale platforms that handle complex, real-time challenges, such as demand forecasting, logistics optimization, and massive transaction volumes. Whether you are working on the Consumer app, Fintech infrastructure, or AdTech platforms, your code directly impacts the speed and reliability of services that people depend on daily.
This role requires more than just technical proficiency; it demands an ability to thrive in a fast-paced, distributed environment where system performance and reliability are paramount. You will collaborate with cross-functional teams to solve sophisticated engineering problems—ranging from microservices architecture and database optimization to machine learning integration. Successful engineers here are those who view technical challenges as opportunities to improve the user experience and who communicate effectively to navigate the complexities of a large, global organization.
Recruiter Screening
reportedBefore anything technical happens, someone has to decide which rung of the ladder your loop is calibrated to, and that decision sets the bar for every round after it. It comes from how you describe scope, not from your title, because titles do not convert cleanly between companies. The weak version of the answer is team size and years. The strong version names the largest change you shipped where nobody reviewed the design, what would have broken if you had been wrong, and what you were paged for. Get the level said out loud on this call, because the range and the loop both follow from it.
What to demonstrate
- Whether the scope in your own account maps onto a level the team actually has an opening at, so a mismatch ends the process cheaply rather than after four interviewers have spent a day
- Whether your title needs re-mapping: the same word describes very different amounts of independent decision-making at a twenty-person company and a ten-thousand-person one
- Whether your compensation expectation can be filled at that level in the structure the role pays in, which is why the number gets asked for before any engineer is scheduled
How to prepare
- Write down two changes from the last two years: the largest one you designed with nobody reviewing the design, and the largest one where someone more senior did. Lead with the first when scope comes up, and be ready to say which parts of the second were yours
- Ask which level the loop is calibrated to and what changes at the level above it, then plan your weeks from that answer rather than from the posting
- Settle a total-compensation range beforehand with the split named, base against bonus against equity and its vesting period, so a question about numbers gets a number instead of the word market
Technical Rounds
reportedMost of the time lost in this format is not lost to thinking. It goes to a standard-library call you half-remember, an off-by-one in a loop bound, and a debugging loop that mutates code at random until something passes. When output is wrong, stop re-reading the whole function: take the smallest input that reproduces it and walk the state through by hand, printing intermediates if the environment allows. Guessing at a fix without a failing case you understand is how a five-minute bug becomes twenty, and the clock does not pause while you do it.
What to demonstrate
- Whether you reach the right structure without a detour, and can write it from memory rather than only recall that one exists
- Whether overflow is considered where the language has fixed-width integers, since a signed 32-bit value stops at 2,147,483,647 and then wraps in Java, is undefined behaviour in C++, and does not arise in Python, whose integers grow instead
- Whether recursion depth is treated as a constraint on large inputs, given that CPython's default limit is 1000 frames and a deep recursion can exhaust the stack in any language where an iterative version would not
- Whether a failing case is isolated and explained before any edit is made to the code
How to prepare
- From an empty file and with no references open, implement the pieces you lean on most: a heap push and pop, an iterative DFS with an explicit stack, and a binary search whose midpoint is written lo + (hi - lo) / 2, which avoids the overflow that (lo + hi) / 2 can hit in a fixed-width integer type
- Time yourself on the ten library calls you look up most, such as sorting with a custom comparator, splitting and joining strings, and finding the next key at or above a value in an ordered map, until the lookup is gone
- Take a solution you know is broken and, before touching it, write one sentence naming the input, the expected value and the actual value. Repeat until you do it without deciding to.
Behavioral Rounds
reportedYour first answer is not really what is scored. It buys the follow-up questions, and those decide the round. An interviewer with fifteen minutes takes one thread and pushes on it four or five times, so a story you can only tell at a single level of detail collapses under the third why. That is an argument for fewer stories known deeply rather than one prepared per prompt. Four or five pieces of work you can still explain down to the code you changed and the argument you had about it will cover nearly anything asked in this round.
What to demonstrate
- Whether a story holds as the questioning moves from what you did to why that instead of the alternative, and then to what you would change knowing what you know now
- Whether you can re-cut a project to answer the question actually asked rather than delivering a rehearsed block that answers an adjacent one
- Whether your level of detail is chosen rather than habitual: going down to the schema when the question is about the data model, staying out of it when the question is about the person who disagreed with you
How to prepare
- Pick four projects and write the chain out four levels deep for each: what you did, why that, why not the alternative, and what would have to be true for the alternative to have won. Where you cannot reach the fourth level, you have a placeholder rather than a story
- Have someone ask why three times in a row on a single thread with nothing else added, and mark the point where you start repeating a sentence you already said. That point is where the interviewer stops learning anything
- Build a one-page index instead of an answer bank: the common prompts in this round (disagreement, a failure that was yours, thin requirements, a deadline you missed, work you inherited) mapped to which of your four projects you would use for each, so the choosing is done now rather than while an interviewer waits
Bar Raiser Interview
reportedWhen a round has no standard shape, it is often there because something is still open: an area no earlier conversation reached, a round where the signal came out mixed, or a decision someone is not ready to make alone. Work out which by going back over what each earlier round actually covered rather than how it felt, and arrive able to give evidence on that point without being asked twice. Weak answers replay the loop's earlier material at the same depth. Strong ones go a level deeper and stay consistent with what you already said.
What to demonstrate
- Whether your account of a project matches the one you gave earlier in the loop, since what you said before may be available to whoever runs this round
- Whether you can go a level deeper on something already covered, reaching the decision and its alternatives rather than repeating the summary
- Whether you state your own uncertainty accurately, including parts of a system you did not build and decisions you inherited, instead of claiming even ownership across all of it
- Whether you can answer a question you handled poorly earlier by naming what you missed, rather than delivering a polished second version as if the first had not happened
How to prepare
- Reconstruct the loop on one page: for each round, the questions you were asked and the answer you actually gave, not the better one you thought of afterwards. The gaps on that page are your best available guess at why this round exists.
- Take the two claims you made earlier that carry the most weight and assemble the backing for each: the measurement, the date, what broke, the decision you would make differently now.
- Write down the three facts about your work that must not drift between tellings, such as team size, timeline and your own role, and check your stories against that list rather than trusting recall under pressure
2 candidate reports. Individual accounts describe a particular role and hiring cycle.
Delivery Hero Software Engineer interview: backend-heavy technical rounds
The process began with an HR screen, moved through technical interviews, and ended with a hiring-manager discussion. I was not offered the role. The technical rounds were the hardest part. They covered Python fundamentals, backend development, system design, APIs, databases, concurrency, and general problem solving. The questions were practical and tied to engineering scenarios rather than stayin…
Read full experienceDelivery Hero Software Engineer Interview Experience: Three quick and kind rounds
After the recruiter call, I went through three rounds: a technical interview, a hiring-manager round, and a bar raiser. The process moved quickly and matched what I expected going in. Everyone I spoke with was kind and understanding, which made the pace easier to handle. I did not get an offer, but the process did not feel hostile or chaotic. It was a straightforward, quick evaluation path. Locat…
Read full experiencePracHub editorial advice for the preparation topics above.
Filtering providers with a bounding box or a distance function over raw lat/lon columns
A WHERE clause computing great-circle distance per row cannot use a B-tree index and degenerates to a scan of every provider in the table, which is fine at 5k providers in staging and falls over at 100k in a dense market at peak. A lat/lon bounding box is index-assisted but returns a square that over-selects badly near the poles and still needs a second-pass distance filter. The working answers are a spatial index — PostGIS GiST with ST_DWithin (which is index-assisted, unlike ST_Distance in a predicate) or the KNN <-> operator for ordered nearest-neighbour — or cell-based bucketing in a key-value store. Cell bucketing has its own edge: two points metres apart can sit in different cells, so a prefix-only lookup silently misses the closest provider unless the query also covers the eight neighbouring cells.
Publishing an event and committing the state change as two separate operations
Whichever order they are attempted in, a crash between them breaks the system in a specific way: publish-then-commit emits an event for a transition that never happened, so downstream consumers charge or notify for a phantom booking; commit-then-publish silently drops the notification, so the provider never receives an offer that the database believes was sent. Neither is fixed by retrying harder, because the failure is the gap, not the operation. The transactional outbox closes it — the event row is written in the same local transaction as the state change and a relay publishes it afterwards — at the cost of at-least-once delivery, which pushes the idempotency requirement onto every consumer.
Never running a concrete value through the code
Trace one small input and one edge input by hand, index by index, out loud. Re-reading your own code catches design mistakes; walking a real value through it catches the off-by-one, the uninitialised accumulator and the loop that never advances.
Treating a network call as though it were a local function call
A remote call can be slow, fail, or return after you stopped waiting, so name the timeout, the retry policy, and what the caller sees while the dependency is down. A call with no timeout turns one slow dependency into an exhausted thread or connection pool in every service upstream of it.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Maximum offers to one provider in any rolling window
Given up to 20 million dispatch_offer rows as (provider_id, offered_at), sorted ascending by offered_at and stamped by the server rather than the device, compute for each provider the maximum number of offers falling inside any rolling 15-minute window, and return the providers whose maximum exceeds 20. Treat the window as half-open: (t - 900s, t]. Memory must be proportional to the offers active in a 15-minute span, not to the length of the input. State your time and space complexity.
Approach
- Keep one FIFO deque of timestamps per provider. On each arrival at time t, pop from the front while front <= t - 900; the deque's size is then the window count at t, and the per-provider maximum is the running max of that size. Each timestamp is pushed once and popped once, so the work is amortised O(1) per event and O(n) overall.
- Exploit the given ordering: because the global stream is sorted by offered_at, each provider's subsequence inherits that order and no sort is needed. If the input were unsorted, sorting would dominate at O(n log n) and would also force the whole stream into memory, defeating the space bound.
- Bound memory by construction: a deque can only hold events from a 15-minute span, so total memory is the peak offer count in any 15 minutes rather than n. Evict a provider's entry entirely once its deque empties and no offer has arrived for a full window.
- Be explicit about the boundary. With (t - 900, t], two offers exactly 900 seconds apart are never in the window together. Use the same convention in this alert and in the enforcement path, or the alert fires on a count the limiter never observed.
- Say why rolling rather than tumbling. Fixed 15-minute calendar buckets split any burst straddling a boundary, so a 30-offer burst centred on the boundary reports as roughly 15 and 15 and never trips a threshold of 20 — which is precisely the pattern the check exists to find.
- If only the boolean 'did it ever exceed 20' is needed, a ring buffer of the last 20 timestamps per provider suffices: the threshold is breached exactly when t minus the oldest of the last 20 is less than 900. That is O(1) space per provider and loses the actual peak value.
Follow-up
- Extend this to a 24-hour window over 200 million events — what breaks and what replaces the deque?
- Events now arrive within seconds of real time and you must alert live; what does the restart story look like for this job?
- A provider's device clock is three minutes fast and the client stamps the event — exactly where does that corrupt this computation?
Find overlapping reservations and the largest free gap
You have up to 2,000,000 booking rows with booking_id, listing_id, status, and reserved_during stored as a half-open UTC range [start, end). Ignoring rows whose status is 'cancelled', report every pair of overlapping reservations on the same listing, and for each listing return the largest free gap inside a supplied horizon. The ranges were materialised from local wall-clock check-in and check-out times in the listing's own timezone. Target O(n log n), and state the exact overlap predicate you use.
Approach
- Bucket rows by listing_id in a hash map first. Overlap is only ever possible within a listing, so you sort many small groups instead of one large one; the bound stays O(n log n) but the constant falls and the work parallelises per listing.
- Within a listing, sort by start ascending with end descending as the tiebreak, then sweep carrying running_max_end. Report a conflict when current.start < running_max_end. Comparing only against the immediately preceding end is the common wrong version: [1,10), [2,3), [4,5) passes that check even though [4,5) is nested inside [1,10).
- State the predicate for half-open ranges explicitly: a and b overlap iff a.start < b.end AND b.start < a.end. Ranges that touch, where a.end == b.start, do not overlap — a same-day turnover is legal. PostgreSQL's && on a tstzrange with the default [) bounds is exactly this predicate, which is why the EXCLUDE constraint in the schema and this offline check agree.
- Compute free gaps from the same sweep: merge while start <= running_max_end, then take max(next.start - prev.end) over consecutive merged ranges, including the leading gap from the horizon start and the trailing gap to the horizon end, clipped to the horizon.
- Bound the output. Reporting all overlapping pairs is O(n log n + p) and p is quadratic on a pathological listing, so either cap p per listing or report only the first conflict per listing when the caller just needs a repair signal.
- Handle the timezone precondition: a local calendar day is 23 or 25 hours across a DST transition, so reserved_during must be materialised at write time by converting local wall-clock times in the listing's zone to UTC instants. Reconstructing it later from a stored local date plus a fixed 24-hour length silently shifts one night per year in each direction.
Follow-up
- This audit found conflicts that a live EXCLUDE constraint on (listing_id WITH =, reserved_during WITH &&) should have made impossible — what could have produced them?
- How do you run this incrementally over only the bookings written since the last pass without missing a conflict with an older row?
- One listing has 40,000 reservations and the pair count explodes — what do you return to the caller instead?
Why the naive proximity scan fails at market scale
provider_presence holds 120,000 rows for one dense market, with lat, lon, cell_id, status and expires_at. The dispatch loop runs every 500 ms and, for each of up to 400 open requests, needs the ten nearest idle providers within 3 km. The obvious implementation computes a great-circle distance for every row and sorts. It returns the correct answer. Quantify why it cannot be shipped, give a design that holds a p99 under 30 ms, and state precisely what that design gives up.
Approach
- Cost it in numbers, not adjectives: 400 requests times 120,000 rows is 48 million distance evaluations per cycle, and at a 500 ms cadence that is 96 million per second, plus 400 sorts of 120,000 elements. The defect is not that the haversine formula is slow; it is that the work per request is proportional to fleet size while the whole dispatch budget is a couple of seconds end to end.
- Explain why no B-tree rescues it. The predicate is a function of two columns, so an index on lat, or a composite on (lat, lon), can restrict only the leading column and the rest is a filter. A degree bounding box is index-assisted on that leading column but over-selects: the square circumscribing a circle of radius r has area 4r^2 against the circle's pi*r^2, so under locally uniform density about 21 percent of the rows that survive the box fall outside the radius, (4 - pi)/4, and still need an exact second pass. Keep the two ratios apart: 4/pi - 1, about 27 percent, is how much more area the box covers than the circle — the extra work done, not the false-positive share of what comes back. Converting metres to a longitude delta divides by cos(latitude), which inflates the box toward the poles.
- Give two designs that work. In PostgreSQL: a geography column with a GiST index and ST_DWithin(pos, point, 3000), which is index-assisted, plus ORDER BY pos <-> point LIMIT 10 for ordered nearest-neighbour. ST_Distance(...) < 3000 written as a predicate is not index-assisted and is the version written by accident. Outside PostgreSQL: bucket on cell_id in an in-memory store and read the block of cells covering the radius, which keeps the candidate set in the low thousands.
- Argue the storage split with the write rate rather than by preference: 100,000 providers heartbeating every 4 seconds is about 25,000 writes per second, which is presence traffic competing for the same WAL as bookings, for state that is worthless 30 seconds later. Redis matches the durability this data actually needs: GEOADD stores members in a sorted set scored by a 52-bit interleaved geohash, and GEOSEARCH ... BYRADIUS reads it. Expiry has to be built rather than assumed, because a sorted set has no per-member TTL — EXPIRE applies to the whole key, and per-field expiry exists only for hashes, via HEXPIRE from Redis 7.4. Have each heartbeat also ZADD presence:seen <epoch_seconds> <provider_id>, sweep with ZRANGEBYSCORE presence:seen -inf (now - 30) and ZREM those members from both keys in one pipeline, and re-check the stored heartbeat at offer time so a member that outlived a missed sweep is still discarded. Snapshot periodically for analytics.
- State what is given up. Great-circle distance is a lower bound on road distance, so a radius prefilter on it never drops a provider whose road distance is inside the radius — the shortlist is admissible. Ranking on it is not defensible: it ignores rivers, one-way systems and the direction of travel. Shortlist by distance, then rank the shortlist by a routing ETA, paying that call on 50 candidates rather than 120,000.
- Name the residual failure modes so the design is not oversold: cell bucketing misses a provider just across a boundary unless the neighbouring cells are queried, and a radius that returns nothing must widen rather than fail — with a bounded number of widenings, because a request that expires unmatched is a first-class outcome, not an error.
Worked solution 30 min
- Write the naive cost: 400 x 120,000 = 48M distance evaluations per cycle, 96M per second at a 500 ms cadence, before the sorts.
- Assume the market spans about 40 km by 40 km, giving a density of 120,000 / 1,600 = 75 providers per square kilometre.
- Size a geohash-6 cell: 360 / 2^15 degrees of longitude is about 1.22 km and 180 / 2^15 degrees of latitude is about 0.61 km, so a 3 km radius needs about 3 rings east-west and 5 rings north-south, a 7 x 11 block of 77 cells.
- Compute the candidate set: 77 cells x 0.744 square kilometres each is about 57 square kilometres, so about 4,300 candidates, against the exact circle's pi x 9 = 28.3 square kilometres and about 2,120 providers.
- Compare: 120,000 / 4,300 is roughly a 28x reduction in rows scanned, with about half the survivors inside the true radius and needing the exact distance filter.
- Redo step 1 with a fleet of 12,000 to see where the naive plan becomes acceptable.
Follow-up
- At what fleet size does the naive version start meeting the budget again — show the arithmetic rather than guessing.
- The 3 km radius returns zero idle providers at 03:00; what does the loop do next, and when does it stop trying?
- Two dispatch partitions read the same idle provider from the index within one cycle — which layer prevents the double assignment, and why not this one?
Explain why the nearest-idle-provider query never uses its index
provider_presence holds 180,000 rows for one market and carries a btree on (lat, lon). Dispatch filters status='idle' and earth_distance(ll_to_earth(lat, lon), ll_to_earth($1, $2)) < 3000, orders by that same distance and takes 20. EXPLAIN reports a sequential scan and p99 is 240 ms against a 30 ms dispatch budget. Explain precisely why the btree cannot serve that predicate, then give two index-assisted rewrites — one in PostgreSQL, one using cell buckets in a key-value store — and name where each degrades.
Approach
- Name the sargability failure exactly: a btree on
(lat, lon)orders rows by latitude then longitude, and the predicate is a function of both columns, so the planner cannot derive a range on the leading column and must evaluate the expression per row. An expression index on the distance does not rescue it either, because the anchor point is a query parameter, so the expression is not constant across calls. - Dispose of the bounding-box half-fix:
lat BETWEEN a AND b AND lon BETWEEN c AND duses the index only for thelatrange and filterslonafterwards, reading a latitude band across the entire index. It also over-selects a square against a circle, and worsens with latitude because a degree of longitude shrinks bycos(lat). - PostgreSQL rewrite: store a
geographycolumn and buildCREATE INDEX CONCURRENTLY ... USING gist (geog) WHERE status = 'idle', then queryWHERE ST_DWithin(geog, $point, 3000).ST_DWithinis index-assisted (bounding-box search followed by an exact recheck) whileST_Distance(...) < 3000in a predicate is not, and ordered nearest-k comes from the KNN operator,ORDER BY geog <-> $point LIMIT 20. It degrades on writes: GiST maintenance at tens of thousands of heartbeats per second is the wrong workload for an OLTP index, which is the argument for presence living outside the transactional store entirely. - Cell rewrite: store
cell_idand look up the cell plus its eight neighbours as exact keys, then apply an exact distance filter and ranking as a second pass. It degrades at the boundary, and the arithmetic is unforgiving: the 3x3 block only covers the search radius when the cell edge is at least that radius. A geohash-7 cell is about 153 m on a side, so nine of them span under 500 m and a 3 km search silently misses nearly every candidate; geohash-5 at roughly 4.9 km is the right resolution for this radius. - Finish on the ranking, because the fastest correct filter still returns the wrong order: straight-line distance is a candidate filter, not a dispatch ranking. A provider 400 m away across a river is further in travel time than one 1.2 km away on the same road, so rank a bounded candidate set by ETA and keep the radius purely as the bound that makes ETA computation affordable.
Worked solution 30 min
- Generate 180,000 synthetic presence rows and reproduce the sequential scan with
EXPLAIN (ANALYZE, BUFFERS). - Add the geography column and a partial GiST index, re-run with
ST_DWithin, and record the plan node plus the row counts before and after the exact recheck. - Compute the cell coverage for a 3 km radius on paper and choose the resolution before writing any cell lookup code.
- Diff the two result sets: the cell version must return the same provider set as the spatial version for the same radius.
Follow-up
- The
cubeandearthdistanceextensions can be made index-assisted with a GiST index onll_to_earth(lat, lon)and anearth_box(...) @>predicate. Why is that still only a bounding volume, and what second-pass predicate do you still need? - At 25,000 presence writes per second in this market, is a GiST index on the transactional table defensible at all? If presence moves out, what do you lose?
- The filter returns 300 providers inside 3 km. How do you rank them, and what does that ranking cost per dispatch cycle?
Model provider deactivation, audit history and erasure together
provider(provider_id, email CITEXT UNIQUE, compliance_state, payout_account_ref, legal_name, phone) is referenced by booking.provider_id and, through it, by ledger_entry. Three requirements land in the same sprint: support must deactivate and later reactivate a provider; a dispute six months later must show which compliance state was in force on a given date; and an erasure request must remove personal data while payout history stays reconcilable. A deleted_at column plus query filters is proposed. Give the schema, and say precisely what deleted_at alone breaks.
Approach
- Let the erasure requirement decide the layout, because it is the one that constrains the others. Split identity from personal data:
providerkeeps the surrogate key, compliance state,eligibility_versionand the payout reference, whileprovider_pii(provider_id, email, legal_name, phone)holds everything erasable. Erasure deletes or crypto-shreds that row;bookingandledger_entrycarry onlyprovider_id, so every historical sum stays reproducible and no ledger row is ever touched. - Enumerate what
deleted_atalone breaks, concretely rather than as a style objection.UNIQUE (email)now blocks the same person re-registering, so it must becomeCREATE UNIQUE INDEX ... ON provider_pii (email) WHERE deleted_at IS NULL. Every query silently includes deleted rows unless it remembers the filter, and the queries that forget are the aggregates nobody re-reads. Foreign keys still resolve, which is correct and is exactly why the row cannot be hard-deleted — and whyON DELETE CASCADEanywhere on this chain would destroy financial history. - Treat deactivation as a state, not a deletion: a value in the compliance state machine with legal transitions (
approved -> suspended -> deactivated -> approved), andeligibility_versionbumped on every change so dispatch re-reads eligibility at dispatch time instead of trusting a flag cached at login. A boolean cannot express reactivation at all, and a nullable timestamp can only express it by being nulled, which erases the history you were asked to keep. - Record the audit trail as an append-only table written in the same transaction as the state change:
provider_compliance_event(event_id, provider_id, from_state, to_state, actor_type, actor_id, reason, occurred_at).provider.compliance_statethen becomes a cache of the latest event rather than an independent fact. - Make 'what was in force on date D' a temporal query with a temporal structure:
provider_compliance_period(provider_id, state, valid_during TSTZRANGE)withCREATE EXTENSION btree_gistandEXCLUDE USING gist (provider_id WITH =, valid_during WITH &&). The open period has an unbounded upper bound, and the dispute query isWHERE provider_id = $1 AND valid_during @> $d, which the GiST index serves. Closing one period and opening the next must happen in a single transaction, or the constraint rejects the overlap — which is the constraint working, not an obstacle. - State the cost honestly: three artefacts now describe the same state, so nominate one writer. Either the application writes the event and a trigger maintains the period and the cached column, or the period table is derived by a job from the event log; say which, and say what a crash between two of them leaves behind.
Follow-up
- A partial unique index's predicate must be immutable. Why can it not be
WHERE deleted_at > now() - interval '30 days', and how would you express a 30-day grace period instead? - Erasure removed the email, and a chargeback then arrives on a two-year-old booking. What can you still do, and what can you no longer do?
- Who writes
provider_compliance_period— the application, a trigger on the event table, or a batch job? Defend the choice against a crash mid-transition.
Specify the idempotency contract for booking creation from a quote
POST /bookings binds a consumer to an existing quote: the body carries quote_id, request_id and a tokenised instrument. The server authorises against the payment gateway and inserts a booking row; booking has UNIQUE (request_id) and quote rows are immutable with an expiry. The caller is a mobile client with a 10 s timeout that retries twice. Specify the idempotency contract: where the key comes from, what is persisted and at what point, the response when the same key arrives with a different body, the response while the first call is still in flight, and the retention floor.
Approach
- Tie the key to the user's intent, not to the HTTP attempt: the client mints a UUID when the consumer taps confirm, stores it, and replays it on every retry of that tap. A key generated per request defeats the whole mechanism while leaving it visibly in place, which is the failure mode that survives code review.
- Persist the key before doing anything external. Insert an idempotency record with a UNIQUE constraint on the key and a status of in_progress, and commit that before calling the gateway - if it is written after the call, a crash in between leaves a charge with nothing pointing at it. Store a fingerprint alongside it: a hash over the canonicalised fields that must not change, here quote_id, request_id and the instrument token.
- Define the three arrival cases precisely. A new key runs the operation. A known key whose fingerprint matches and whose record is completed replays the stored status code and body byte-identically, including when that status was a 4xx. A known key whose fingerprint differs is refused with 422 and a distinct code, without executing - the client has reused a key for a different intent, and guessing which one it meant is how you double-book.
- Handle the in-flight case explicitly rather than by waiting: return 409 with Retry-After: 1 and a code the client can act on. A 202 invites the client to move on as though a booking exists, and blocking the second request until the first finishes just moves the timeout.
- Set retention from the client's behaviour, not from taste: the record must outlive the longest retry horizon the client can produce plus the gateway reconciliation lag for unknown outcomes, so 24 hours is a floor here, held in a TTL-partitioned table. Note the interaction with quote expiry: a replay must return the stored response and must not re-validate the quote, because the booking already exists and re-checking an expired quote would turn a successful retry into a spurious failure. UNIQUE (request_id) is the backstop for a client that loses its key entirely - map that violation to the existing booking's representation rather than a 500.
Worked solution 30 min
- Write the key provenance rule and the exact moment the client generates and discards a key.
- Implement the reservation insert with its UNIQUE constraint, committed before the gateway call, and the read-back path on conflict.
- Define the fingerprint over the canonicalised triple and the 422 refusal on mismatch.
- On completion, write the status code, response body and booking_id in the same transaction that flips the record to completed.
- Exercise three interleavings: a sequential retry after a timeout, two concurrent requests sharing a key, and a replay issued after the quote has expired.
Follow-up
- The gateway call times out and the record stays in_progress forever because the process died. What sweeps it, and what does the endpoint return in the meantime?
- The consumer taps confirm twice deliberately, wanting two bookings. How does the contract tell that apart from a retry?
- Where does the idempotency record live relative to the booking row, and what breaks if they are in different databases?
Calendar holds that cannot double-book a listing
The rental vertical reserves a listing for a range: booking(booking_id, listing_id, consumer_id, status in accepted, en_route, arrived, in_progress, completed, cancelled, disputed, reserved_during TSTZRANGE, version). Two consumers can check out the same weekend at the same moment, check-in is a local wall-clock time in the listing's timezone, and a hold lasts 10 minutes before expiring. Design the hold-then-confirm flow and the storage-level guarantee that two confirmed reservations on one listing can never overlap. State how you represent the range and what you do about daylight saving transitions.
Approach
- Rule out the application-level check first and say why precisely: SELECT conflicting rows followed by INSERT is check-then-act, and at READ COMMITTED or REPEATABLE READ neither transaction can see the other's uncommitted insert, so both commit a valid-looking overlap. Only SERIALIZABLE or a storage-level constraint closes it.
- Enforce it declaratively: CREATE EXTENSION btree_gist, then EXCLUDE USING gist (listing_id WITH =, reserved_during WITH &&) WHERE (status <> 'cancelled'). An overlap becomes SQLSTATE 23P01, which the handler turns into a user-visible refusal. Two preconditions worth stating: every pre-existing overlap must be resolved before the index will build, and there is no CONCURRENTLY path for an exclusion constraint (the UNIQUE ... USING INDEX trick does not apply), so adding it to a large live table holds ACCESS EXCLUSIVE for the build.
- Get the interval semantics right. Use half-open '[)' bounds so an 11:00 checkout and an 11:00 check-in do not overlap, and store absolute instants built by converting the listing's local wall time through the listing's IANA timezone rather than the server's or the consumer's. A local day is 23 or 25 hours across a transition, so computing a range as nights times 24 hours is wrong by an hour twice a year, in the direction that silently shortens or extends a stay.
- Model the hold as a real constrained row (a 'hold' status in the same table, or a separate table covered by an equivalent constraint) with its own expires_at, so a hold genuinely blocks a competing confirm. A hold kept in a cache or an in-process lock is invisible to the constraint and reintroduces exactly the race the constraint exists to close.
- Make confirm a guarded transition like every other one in the system: UPDATE ... WHERE status = 'hold' AND version = $v AND expires_at > now(), so the sweeper that cancels expired holds and the confirm that promotes one resolve by affected-row count rather than by a lock the application has to remember to take.
Follow-up
- The listing owner blocks dates for maintenance. Same table or a different one, and what does the constraint see in each case?
- A guest extends by one night while in_progress. Write that transition and say what the constraint rejects.
- How do you add this constraint to a 40-million-row table that already contains overlaps, without a long write outage?
Cancellations fail hourly while the outbox relay falls minutes behind
Every hour on the hour, 2-4% of cancellations return 500 and the database log records deadlock detected (SQLSTATE 40P01); outbox_event lag climbs to eight minutes over the same window and then recovers. The cancellation transaction updates booking, then payment_attempt, then inserts ledger_entry and outbox_event. An hourly reconciliation job updates payment_attempt rows whose status is 'unknown', then updates the matching booking rows with UPDATE ... WHERE booking_id IN (...). Give the ordered diagnostic checklist and the fix.
Approach
- Read the deadlock report in the server log before anything else. PostgreSQL logs both processes, the statement each was running, and what each held and waited for; the detector fires after deadlock_timeout, one second by default, so a report exists for every occurrence. That report usually names the two lock orders outright and saves the entire investigation.
- Confirm the ordering conflict it describes. The cancellation path locks a booking row and then a payment_attempt row; the reconciliation job locks payment_attempt rows and then booking rows. Two transactions taking the same pair in opposite orders deadlock by construction, and the hourly schedule is why the rate is periodic rather than correlated with traffic.
- Note the detail that makes the batch worse than it looks: UPDATE ... WHERE booking_id IN (...) acquires row locks in whatever order the plan produces rows, not in the order of the IN list. Two overlapping batches can therefore deadlock with each other while both developers believe they used the same order. Imposing an order requires an explicit SELECT booking_id FROM booking WHERE booking_id = ANY($1) ORDER BY booking_id FOR UPDATE before the update.
- Account for the lag separately instead of assuming one cause. A deadlock victim rolls its whole transaction back, so the outbox rows it wrote disappear and are rewritten on retry; the lag spike is that retry backlog, not a relay defect. Verify by checking whether relay sessions are actually blocked -- a lock wait event in the activity view -- or merely behind, and confirm its claim query is still using the partial index on created_at where published_at is null with FOR UPDATE SKIP LOCKED so workers do not serialize.
- Fix in two parts. Give every multi-row path a canonical lock order, ascending booking_id imposed by an explicit ordered SELECT ... FOR UPDATE, and chunk the job so lock hold time is bounded by a chunk rather than by the whole run. Then shorten the cancellation transaction so it does not hold a booking row lock across payment work at all: moving the money work to its own consumer off the outbox removes the pairing instead of ordering it, which is the durable version of the fix.
- Keep retrying 40P01, but correctly. The transaction rolled back completely, so a retry is safe and required; bound it, add jitter, and enable log_lock_waits so contention that has not yet escalated into a deadlock becomes visible before the next incident.
Follow-up
- The job now runs in 200-row chunks. What is the worst-case lock hold time, and what does chunking do to the deadlock rate compared with one large transaction?
- If you switched to SERIALIZABLE, which of these failures disappears and which new error would every caller have to handle?
- How would you distinguish an outbox gap caused by rolled-back transactions from a relay that crashed mid-batch?
For someone fluent in a dynamic language who has shipped real work but has never had to say what the runtime is doing underneath. The week is built on measuring and deliberately breaking things, because the questions that expose this background are the ones where the interviewer asks why a second time.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Measure before reasoning
- Take a slow piece of your own code, write down in advance where you believe the time goes, then profile it and record how wrong the guess was. The cost is usually an allocation you did not notice or an accidental quadratic membership test.
- Replace one list membership test inside a loop with a set and measure at a thousand, ten thousand and a hundred thousand elements, confirming the shape of the curve rather than only that it got faster.
- Write down the three quantities you can now measure instead of assert: wall time, peak memory, and call count for the function you suspected.
Deliverable: A before-and-after profile of real code plus a written note on the size of the gap between the guess and the measurement.
Practice prompt ↗Practice prompt ↗Worked solution ↗02References, copies, and the bugs they produce
- Write the function with a mutable default argument, call it three times, and explain the accumulating result: the default is evaluated once when the function is defined, so every call shares one object.
- Build a nested structure, take a shallow copy, mutate an inner element, and show that both views changed, because a shallow copy duplicates the container and not the elements. Then fix it with a deep copy and state the cost you just accepted.
- Write two functions, one mutating its argument in place and one rebinding the local name, and predict the caller's view of each before running it. That single distinction produces most of the bugs that pass their tests.
Deliverable: Three small programs whose output you predicted correctly before running, each with a one-line statement of the rule underneath.
Practice prompt ↗Practice prompt ↗03Types, once, in a language that checks them
- Port one module you have already written, roughly a hundred lines, into a statically typed language, and record every place the compiler demanded an answer your original had left implicit: a value that can be absent, a numeric width, a case never handled.
- Write the same signature in both languages and state what the static one guarantees before the program runs and what it does not, since it will not save you from a wrong algorithm or an index out of range.
- Write the difference between an interface satisfied by declaration and one satisfied structurally, with one case each where the other approach would miss the mistake.
Deliverable: One module in two languages plus a list of the questions the type checker forced you to answer.
Practice prompt ↗Practice prompt ↗04Concurrency, starting with what actually runs at the same time
- Run the same CPU-bound function across four threads and four processes and measure both. Under the default CPython build the threaded version will not speed up, because only one thread executes bytecode at a time; the process version will. Check which build you are on first, since free-threaded builds remove that lock and change the result.
- Then run a blocking I/O workload across four threads and measure it speeding up, because the interpreter releases that lock around blocking calls, which is why treating threads as useless is wrong as a general claim.
- Build the lost update: two threads each incrementing a shared counter a hundred thousand times, and show a final value below the expected sum, because an increment is a load, an add and a store and the thread can be suspended between them. Fix it with a lock and then measure what the lock costs.
Deliverable: Three measurements, threads against processes on CPU work, threads on I/O work, and a demonstrated lost update, each with the mechanism written underneath.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Debugging as a procedure rather than an instinct
- Work one real failure as a bisection: find a revision or an input size where it is good and one where it is bad, halve repeatedly, and state the two assumptions bisection needs, that the property changes exactly once across the range and that the test is reliable.
- Minimise one failing input to the smallest version that still fails, and record how many rounds it took.
- Keep a hypothesis log for one bug in three columns, what I believe, what would disprove it, what I observed, and stop yourself the first time you are about to change two things at once.
Deliverable: One bug worked to root cause with a written hypothesis log and a minimised reproducing input.
Practice prompt ↗06Tests that catch the bug you are about to write
- Implement an LRU cache with a capacity bound, then write the three test cases that would catch an off-by-one in eviction: insert exactly capacity items and assert nothing was evicted, insert one more and assert the least recently used key is the one gone, and read an old key just before that insert so the eviction victim changes.
- Add a property test comparing your implementation against a deliberately slow reference, an ordered list scanned linearly, over a few thousand random operation sequences, because a slow reference finds the cases you would not have thought to write.
- Write one numeric test that fails under exact equality and passes with a tolerance, and state why the tolerance has to be relative rather than absolute once the magnitudes grow.
Deliverable: An LRU implementation with three boundary tests, one property test against a slow reference, and one tolerance-based numeric test.
Practice prompt ↗07Debug something broken, out loud
- Have someone plant three defects in a two-hundred-line program, an off-by-one, a shared mutable state bug, and a wrong error-handling path, then find them while narrating, under a fixed rule: state the hypothesis before touching anything.
- Time each one and record which tool found it, reading, a printed value, a debugger, or a test, because the question asked in interviews is how you would find it rather than what it was.
- Write the sentence you will use when you do not yet know the cause, one that names the next measurement instead of offering a guess.
Deliverable: A recorded debugging session with time-to-find per defect and the method that found each.
Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Conflict answers where you were right and everyone came round are the weakest ones. Stronger: the evidence you went and collected, what would have changed your mind, and what you did in the weeks after the call went against you. Implementing a design you argued against, properly, is a specific and checkable behaviour.
These assess how you approach teamwork, conflict, and professional gro…
These assess how you approach teamwork, conflict, and professional growth. Describe an achievement you are particularly proud of and the impact it had. Tell me about a time you faced a major conflict; how did you handle it? How do you approach code reviews and provide feedback to peers? What motivates you to join Delivery Hero? How do you prioritize tasks when working on a high-pressure, diverse team?
Approach
- Pick a story where you made the decision, not one where you watched it.
- Name the disagreement and how you resolved it with evidence.
- Give the blast radius: what could have broken, and what you measured.
Follow-up
- How did you know your change caused the improvement?
- What did you decide not to do, and why?
Walk an outage you owned from first signal to permanent fix
Pick an incident where you were the primary owner, not a helper. Tell it as a timeline: the first signal and what threshold fired it, what you ruled out and how, the mitigation you shipped and how long it took to reach production, and the permanent fix that followed. Quantify the blast radius in a countable unit - bookings affected, duplicate captures, minutes of degraded dispatch - and say how you bounded that number instead of extrapolating it. Close with the one change that made the class of failure impossible rather than merely unlikely.
Approach
- Lead with impact in one sentence - who was affected, for how long, in what unit - before any chronology, because that sentence is what the listener calibrates seniority against.
- Give the detection path honestly. 'A consumer emailed support' and 'the duplicate-capture alert fired at 14 per minute against a baseline of zero' describe very different systems, and claiming the second when it was the first collapses on the first follow-up.
- Run two clocks: time to mitigate and time to fix. Mitigation is whatever stops the bleeding within minutes (a flag, a rate cap, draining one partition); the fix is the structural change that lands days later. Collapsing them into one story hides whether you can triage under pressure.
- Bound the affected set with a query you can state, not a rate times a duration - for example successful captures grouped by booking_id having count(*) > 1 across the incident window, cross-checked against the gateway's own record. An exact list survives scrutiny; an estimate invites it.
- End on the structural change and say what it does not catch. A partial unique index on the live-offer status, an EXCLUDE constraint on overlapping reservations, or a guarded UPDATE whose affected-row count decides the winner each convert a silent corruption into a loud error, and each has a boundary worth naming.
Follow-up
- What did you believe mid-incident that turned out to be wrong, and what made you drop it?
- How did you verify impact had stopped, without relying on the alert clearing?
- Who else on the team could have shipped the same bug that quarter, and what stops them now?
A code review disagreement you carried past one round
Recall a review where you and the author still disagreed after the first round of comments. Say what the change did, what you believed was wrong with it, and how you argued - a failing test, a two-transaction interleaving written out, a prior incident, or authority. Say how it ended: merged as written, revised, escalated, or abandoned, and who decided. Then classify your own objection: correctness bug, maintainability opinion, or style. A strong answer treats those three as deserving different amounts of pushback.
Approach
- Classify before you narrate. Correctness bugs justify blocking, maintainability arguments justify one round and a follow-up ticket, and style belongs in a formatter rather than in a human's comment.
- For a correctness objection, describe the cheapest convincing artifact you produced. Usually that is a failing test or four lines showing two workers both reading status = 'offered' under READ COMMITTED, the default isolation level in PostgreSQL, and both proceeding to write an accept.
- Say who converged and how. 'They agreed once the test went red' and 'we escalated to the tech lead and I lost' are both credible; 'I approved it to avoid friction and it broke in production' is also credible if you follow it with what you would do now.
- Account for the cost you imposed. A blocked review costs the author momentum and sometimes a release slot, and a senior reviewer prices that in rather than treating rigour as free.
- Close with the systemic change if there was one - a test in CI, a constraint moved into the schema, a lint rule. The same objection raised by hand twice is a process failure, not a reviewing triumph.
Follow-up
- When did you last approve something you disagreed with, and how did it turn out?
- Where does that class of bug get caught today, if not in human review?
- How do you run the same objection when the author is considerably more senior than you?
- 01
These assess how you approach teamwork, conflict, and professional growth. Describe an achievement you are particularly proud of and the impact it had. Tell me about a time you faced a major conflict; how did you handle it? How do you approach code reviews and provide feedback to peers? What motivates you to join Delivery Hero? How do you prioritize tasks when working on a high-pressure, diverse team?
- 02
Pick an incident where you were the primary owner, not a helper. Tell it as a timeline: the first signal and what threshold fired it, what you ruled out and how, the mitigation you shipped and how long it took to reach production, and the permanent fix that followed. Quantify the blast radius in a countable unit - bookings affected, duplicate captures, minutes of degraded dispatch - and say how you bounded that number instead of extrapolating it. Close with the one change that made the class of failure impossible rather than merely unlikely.
- 03
Recall a review where you and the author still disagreed after the first round of comments. Say what the change did, what you believed was wrong with it, and how you argued - a failing test, a two-transaction interleaving written out, a prior incident, or authority. Say how it ended: merged as written, revised, escalated, or abandoned, and who decided. Then classify your own objection: correctness bug, maintainability opinion, or style. A strong answer treats those three as deserving different amounts of pushback.
Is this an official Delivery Hero interview guide?
No. It is PracHub's own research and practice material for the Software Engineer role at Delivery Hero. Rounds and questions reflect what candidates have reported, not a process Delivery Hero has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How long does the interview process usually take?
The process typically moves quickly, often spanning 3 to 6 weeks, though this can vary based on team availability. You can expect regular updates from your recruiter throughout the stages.
PracHub interview research ↗What is the "Bar Raiser" round?
The Bar Raiser is an interview conducted by someone outside of the team you are interviewing for. Their goal is to ensure a high and consistent standard of hiring across the entire company, focusing on both technical competence and long-term cultural fit.
PracHub interview research ↗Is there feedback provided if I am not selected?
While the company strives to provide feedback, it is not always guaranteed for every candidate. If you are not selected, focus on the specific areas where you felt less confident during the technical rounds.
PracHub interview research ↗Is the work environment remote or hybrid?
Expectations vary by location and team, but Delivery Hero generally operates with a hybrid approach. Be sure to clarify the specific expectations for your role during the initial recruiter screen.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Software Engineer practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22