Delivery Hero · Software Engineer
Updated · 2026-09-24

Delivery Hero Software Engineer
Interview Questions & Guide 2026

THE 60-SECOND BRIEF

As a Software Engineer at Delivery Hero, you are at the heart of a global ecosystem that powers food delivery, logistics, and quick-commerce services for millions of users across the world. You will contribute to high-scale platforms that handle complex, real-time challenges, such as demand forecasting, logistics optimization, and massive transaction volumes. Whether you are working on the Consumer app, Fintech infrastructure, or AdTech platforms, your code directly impacts the speed and reliability of services that people depend on daily.

For a design discussion, a catalogue of architectures is worth less than the ability to turn a vague requirement into a data model and an API contract. A box diagram with no schema under it collapses at the first follow-up question.

Delivery Hero candidates report 4 rounds · ≈ 3-5 weeks. The stages below are what candidates describe, not a published process.

Index provider locations for bounded proximity queriesMake offer acceptance idempotent under expiry racesEnforce a single active assignment per provider

45 min read

Practice 11 Software Engineer prompts
2Candidate experiences ↗Read their reports
11Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

As a Software Engineer at Delivery Hero, you are at the heart of a global ecosystem that powers food delivery, logistics, and quick-commerce services for millions of users across the world. You will contribute to high-scale platforms that handle complex, real-time challenges, such as demand forecasting, logistics optimization, and massive transaction volumes. Whether you are working on the Consumer app, Fintech infrastructure, or AdTech platforms, your code directly impacts the speed and reliability of services that people depend on daily.

This role requires more than just technical proficiency; it demands an ability to thrive in a fast-paced, distributed environment where system performance and reliability are paramount. You will collaborate with cross-functional teams to solve sophisticated engineering problems—ranging from microservices architecture and database optimization to machine learning integration. Successful engineers here are those who view technical challenges as opportunities to improve the user experience and who communicate effectively to navigate the complexities of a large, global organization.

01

Recruiter Screening

reported

Before anything technical happens, someone has to decide which rung of the ladder your loop is calibrated to, and that decision sets the bar for every round after it. It comes from how you describe scope, not from your title, because titles do not convert cleanly between companies. The weak version of the answer is team size and years. The strong version names the largest change you shipped where nobody reviewed the design, what would have broken if you had been wrong, and what you were paged for. Get the level said out loud on this call, because the range and the loop both follow from it.

What to demonstrate

  • Whether the scope in your own account maps onto a level the team actually has an opening at, so a mismatch ends the process cheaply rather than after four interviewers have spent a day
  • Whether your title needs re-mapping: the same word describes very different amounts of independent decision-making at a twenty-person company and a ten-thousand-person one
  • Whether your compensation expectation can be filled at that level in the structure the role pays in, which is why the number gets asked for before any engineer is scheduled

How to prepare

  • Write down two changes from the last two years: the largest one you designed with nobody reviewing the design, and the largest one where someone more senior did. Lead with the first when scope comes up, and be ready to say which parts of the second were yours
  • Ask which level the loop is calibrated to and what changes at the level above it, then plan your weeks from that answer rather than from the posting
  • Settle a total-compensation range beforehand with the split named, base against bonus against equity and its vesting period, so a question about numbers gets a number instead of the word market
PracHub interview research ↗
02

Technical Rounds

reported

Most of the time lost in this format is not lost to thinking. It goes to a standard-library call you half-remember, an off-by-one in a loop bound, and a debugging loop that mutates code at random until something passes. When output is wrong, stop re-reading the whole function: take the smallest input that reproduces it and walk the state through by hand, printing intermediates if the environment allows. Guessing at a fix without a failing case you understand is how a five-minute bug becomes twenty, and the clock does not pause while you do it.

What to demonstrate

  • Whether you reach the right structure without a detour, and can write it from memory rather than only recall that one exists
  • Whether overflow is considered where the language has fixed-width integers, since a signed 32-bit value stops at 2,147,483,647 and then wraps in Java, is undefined behaviour in C++, and does not arise in Python, whose integers grow instead
  • Whether recursion depth is treated as a constraint on large inputs, given that CPython's default limit is 1000 frames and a deep recursion can exhaust the stack in any language where an iterative version would not
  • Whether a failing case is isolated and explained before any edit is made to the code

How to prepare

  • From an empty file and with no references open, implement the pieces you lean on most: a heap push and pop, an iterative DFS with an explicit stack, and a binary search whose midpoint is written lo + (hi - lo) / 2, which avoids the overflow that (lo + hi) / 2 can hit in a fixed-width integer type
  • Time yourself on the ten library calls you look up most, such as sorting with a custom comparator, splitting and joining strings, and finding the next key at or above a value in an ordered map, until the lookup is gone
  • Take a solution you know is broken and, before touching it, write one sentence naming the input, the expected value and the actual value. Repeat until you do it without deciding to.
PracHub interview research ↗
03

Behavioral Rounds

reported

Your first answer is not really what is scored. It buys the follow-up questions, and those decide the round. An interviewer with fifteen minutes takes one thread and pushes on it four or five times, so a story you can only tell at a single level of detail collapses under the third why. That is an argument for fewer stories known deeply rather than one prepared per prompt. Four or five pieces of work you can still explain down to the code you changed and the argument you had about it will cover nearly anything asked in this round.

What to demonstrate

  • Whether a story holds as the questioning moves from what you did to why that instead of the alternative, and then to what you would change knowing what you know now
  • Whether you can re-cut a project to answer the question actually asked rather than delivering a rehearsed block that answers an adjacent one
  • Whether your level of detail is chosen rather than habitual: going down to the schema when the question is about the data model, staying out of it when the question is about the person who disagreed with you

How to prepare

  • Pick four projects and write the chain out four levels deep for each: what you did, why that, why not the alternative, and what would have to be true for the alternative to have won. Where you cannot reach the fourth level, you have a placeholder rather than a story
  • Have someone ask why three times in a row on a single thread with nothing else added, and mark the point where you start repeating a sentence you already said. That point is where the interviewer stops learning anything
  • Build a one-page index instead of an answer bank: the common prompts in this round (disagreement, a failure that was yours, thin requirements, a deadline you missed, work you inherited) mapped to which of your four projects you would use for each, so the choosing is done now rather than while an interviewer waits
PracHub interview research ↗
04

Bar Raiser Interview

reported

When a round has no standard shape, it is often there because something is still open: an area no earlier conversation reached, a round where the signal came out mixed, or a decision someone is not ready to make alone. Work out which by going back over what each earlier round actually covered rather than how it felt, and arrive able to give evidence on that point without being asked twice. Weak answers replay the loop's earlier material at the same depth. Strong ones go a level deeper and stay consistent with what you already said.

What to demonstrate

  • Whether your account of a project matches the one you gave earlier in the loop, since what you said before may be available to whoever runs this round
  • Whether you can go a level deeper on something already covered, reaching the decision and its alternatives rather than repeating the summary
  • Whether you state your own uncertainty accurately, including parts of a system you did not build and decisions you inherited, instead of claiming even ownership across all of it
  • Whether you can answer a question you handled poorly earlier by naming what you missed, rather than delivering a polished second version as if the first had not happened

How to prepare

  • Reconstruct the loop on one page: for each round, the questions you were asked and the answer you actually gave, not the better one you thought of afterwards. The gaps on that page are your best available guess at why this round exists.
  • Take the two claims you made earlier that carry the most weight and assemble the backing for each: the measurement, the date, what broke, the decision you would make differently now.
  • Write down the three facts about your work that must not drift between tellings, such as team size, timeline and your own role, and check your stories against that list rather than trusting recall under pressure
PracHub interview research ↗

2 candidate reports. Individual accounts describe a particular role and hiring cycle.

Software Engineer

Delivery Hero Software Engineer interview: backend-heavy technical rounds

HR Screen → Technical Screen

The process began with an HR screen, moved through technical interviews, and ended with a hiring-manager discussion. I was not offered the role. The technical rounds were the hardest part. They covered Python fundamentals, backend development, system design, APIs, databases, concurrency, and general problem solving. The questions were practical and tied to engineering scenarios rather than stayin…

Read full experience
Software Engineer

Delivery Hero Software Engineer Interview Experience: Three quick and kind rounds

Technical Screen → HR Screen → Other

After the recruiter call, I went through three rounds: a technical interview, a hiring-manager round, and a bar raiser. The process moved quickly and matched what I expected going in. Everyone I spoke with was kind and understanding, which made the pace easier to handle. I did not get an offer, but the process did not feel hostile or chaotic. It was a straightforward, quick evaluation path. Locat…

Read full experience

PracHub editorial advice for the preparation topics above.

01

Filtering providers with a bounding box or a distance function over raw lat/lon columns

A WHERE clause computing great-circle distance per row cannot use a B-tree index and degenerates to a scan of every provider in the table, which is fine at 5k providers in staging and falls over at 100k in a dense market at peak. A lat/lon bounding box is index-assisted but returns a square that over-selects badly near the poles and still needs a second-pass distance filter. The working answers are a spatial index — PostGIS GiST with ST_DWithin (which is index-assisted, unlike ST_Distance in a predicate) or the KNN <-> operator for ordered nearest-neighbour — or cell-based bucketing in a key-value store. Cell bucketing has its own edge: two points metres apart can sit in different cells, so a prefix-only lookup silently misses the closest provider unless the query also covers the eight neighbouring cells.

02

Publishing an event and committing the state change as two separate operations

Whichever order they are attempted in, a crash between them breaks the system in a specific way: publish-then-commit emits an event for a transition that never happened, so downstream consumers charge or notify for a phantom booking; commit-then-publish silently drops the notification, so the provider never receives an offer that the database believes was sent. Neither is fixed by retrying harder, because the failure is the gap, not the operation. The transactional outbox closes it — the event row is written in the same local transaction as the state change and a relay publishes it afterwards — at the cost of at-least-once delivery, which pushes the idempotency requirement onto every consumer.

03

Never running a concrete value through the code

Trace one small input and one edge input by hand, index by index, out loud. Re-reading your own code catches design mistakes; walking a real value through it catches the off-by-one, the uninitialised accumulator and the loop that never advances.

04

Treating a network call as though it were a local function call

A remote call can be slow, fail, or return after you stopped waiting, so name the timeout, the retry policy, and what the caller sees while the dependency is down. A call with no timeout turns one slow dependency into an exhausted thread or connection pool in every service upstream of it.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

8 technical prompts3 include a worked solution

Maximum offers to one provider in any rolling window

medium
sliding windowtwo pointersevent streamrate limiting

Given up to 20 million dispatch_offer rows as (provider_id, offered_at), sorted ascending by offered_at and stamped by the server rather than the device, compute for each provider the maximum number of offers falling inside any rolling 15-minute window, and return the providers whose maximum exceeds 20. Treat the window as half-open: (t - 900s, t]. Memory must be proportional to the offers active in a 15-minute span, not to the length of the input. State your time and space complexity.

Approach
  1. Keep one FIFO deque of timestamps per provider. On each arrival at time t, pop from the front while front <= t - 900; the deque's size is then the window count at t, and the per-provider maximum is the running max of that size. Each timestamp is pushed once and popped once, so the work is amortised O(1) per event and O(n) overall.
  2. Exploit the given ordering: because the global stream is sorted by offered_at, each provider's subsequence inherits that order and no sort is needed. If the input were unsorted, sorting would dominate at O(n log n) and would also force the whole stream into memory, defeating the space bound.
  3. Bound memory by construction: a deque can only hold events from a 15-minute span, so total memory is the peak offer count in any 15 minutes rather than n. Evict a provider's entry entirely once its deque empties and no offer has arrived for a full window.
  4. Be explicit about the boundary. With (t - 900, t], two offers exactly 900 seconds apart are never in the window together. Use the same convention in this alert and in the enforcement path, or the alert fires on a count the limiter never observed.
  5. Say why rolling rather than tumbling. Fixed 15-minute calendar buckets split any burst straddling a boundary, so a 30-offer burst centred on the boundary reports as roughly 15 and 15 and never trips a threshold of 20 — which is precisely the pattern the check exists to find.
  6. If only the boolean 'did it ever exceed 20' is needed, a ring buffer of the last 20 timestamps per provider suffices: the threshold is breached exactly when t minus the oldest of the last 20 is less than 900. That is O(1) space per provider and loses the actual peak value.
Follow-up
  • Extend this to a 24-hour window over 200 million events — what breaks and what replaces the deque?
  • Events now arrive within seconds of real time and you must alert live; what does the restart story look like for this job?
  • A provider's device clock is three minutes fast and the client stamps the event — exactly where does that corrupt this computation?

Find overlapping reservations and the largest free gap

medium
intervalssweep linesortingtimezones

You have up to 2,000,000 booking rows with booking_id, listing_id, status, and reserved_during stored as a half-open UTC range [start, end). Ignoring rows whose status is 'cancelled', report every pair of overlapping reservations on the same listing, and for each listing return the largest free gap inside a supplied horizon. The ranges were materialised from local wall-clock check-in and check-out times in the listing's own timezone. Target O(n log n), and state the exact overlap predicate you use.

Approach
  1. Bucket rows by listing_id in a hash map first. Overlap is only ever possible within a listing, so you sort many small groups instead of one large one; the bound stays O(n log n) but the constant falls and the work parallelises per listing.
  2. Within a listing, sort by start ascending with end descending as the tiebreak, then sweep carrying running_max_end. Report a conflict when current.start < running_max_end. Comparing only against the immediately preceding end is the common wrong version: [1,10), [2,3), [4,5) passes that check even though [4,5) is nested inside [1,10).
  3. State the predicate for half-open ranges explicitly: a and b overlap iff a.start < b.end AND b.start < a.end. Ranges that touch, where a.end == b.start, do not overlap — a same-day turnover is legal. PostgreSQL's && on a tstzrange with the default [) bounds is exactly this predicate, which is why the EXCLUDE constraint in the schema and this offline check agree.
  4. Compute free gaps from the same sweep: merge while start <= running_max_end, then take max(next.start - prev.end) over consecutive merged ranges, including the leading gap from the horizon start and the trailing gap to the horizon end, clipped to the horizon.
  5. Bound the output. Reporting all overlapping pairs is O(n log n + p) and p is quadratic on a pathological listing, so either cap p per listing or report only the first conflict per listing when the caller just needs a repair signal.
  6. Handle the timezone precondition: a local calendar day is 23 or 25 hours across a DST transition, so reserved_during must be materialised at write time by converting local wall-clock times in the listing's zone to UTC instants. Reconstructing it later from a stored local date plus a fixed 24-hour length silently shifts one night per year in each direction.
Follow-up
  • This audit found conflicts that a live EXCLUDE constraint on (listing_id WITH =, reserved_during WITH &&) should have made impossible — what could have produced them?
  • How do you run this incrementally over only the bookings written since the last pass without missing a conflict with an older row?
  • One listing has 40,000 reservations and the pair count explodes — what do you return to the caller instead?

Why the naive proximity scan fails at market scale

hardWorked solution
spatial indexquery planningcomplexityscale

provider_presence holds 120,000 rows for one dense market, with lat, lon, cell_id, status and expires_at. The dispatch loop runs every 500 ms and, for each of up to 400 open requests, needs the ten nearest idle providers within 3 km. The obvious implementation computes a great-circle distance for every row and sorts. It returns the correct answer. Quantify why it cannot be shipped, give a design that holds a p99 under 30 ms, and state precisely what that design gives up.

Approach
  1. Cost it in numbers, not adjectives: 400 requests times 120,000 rows is 48 million distance evaluations per cycle, and at a 500 ms cadence that is 96 million per second, plus 400 sorts of 120,000 elements. The defect is not that the haversine formula is slow; it is that the work per request is proportional to fleet size while the whole dispatch budget is a couple of seconds end to end.
  2. Explain why no B-tree rescues it. The predicate is a function of two columns, so an index on lat, or a composite on (lat, lon), can restrict only the leading column and the rest is a filter. A degree bounding box is index-assisted on that leading column but over-selects: the square circumscribing a circle of radius r has area 4r^2 against the circle's pi*r^2, so under locally uniform density about 21 percent of the rows that survive the box fall outside the radius, (4 - pi)/4, and still need an exact second pass. Keep the two ratios apart: 4/pi - 1, about 27 percent, is how much more area the box covers than the circle — the extra work done, not the false-positive share of what comes back. Converting metres to a longitude delta divides by cos(latitude), which inflates the box toward the poles.
  3. Give two designs that work. In PostgreSQL: a geography column with a GiST index and ST_DWithin(pos, point, 3000), which is index-assisted, plus ORDER BY pos <-> point LIMIT 10 for ordered nearest-neighbour. ST_Distance(...) < 3000 written as a predicate is not index-assisted and is the version written by accident. Outside PostgreSQL: bucket on cell_id in an in-memory store and read the block of cells covering the radius, which keeps the candidate set in the low thousands.
  4. Argue the storage split with the write rate rather than by preference: 100,000 providers heartbeating every 4 seconds is about 25,000 writes per second, which is presence traffic competing for the same WAL as bookings, for state that is worthless 30 seconds later. Redis matches the durability this data actually needs: GEOADD stores members in a sorted set scored by a 52-bit interleaved geohash, and GEOSEARCH ... BYRADIUS reads it. Expiry has to be built rather than assumed, because a sorted set has no per-member TTL — EXPIRE applies to the whole key, and per-field expiry exists only for hashes, via HEXPIRE from Redis 7.4. Have each heartbeat also ZADD presence:seen <epoch_seconds> <provider_id>, sweep with ZRANGEBYSCORE presence:seen -inf (now - 30) and ZREM those members from both keys in one pipeline, and re-check the stored heartbeat at offer time so a member that outlived a missed sweep is still discarded. Snapshot periodically for analytics.
  5. State what is given up. Great-circle distance is a lower bound on road distance, so a radius prefilter on it never drops a provider whose road distance is inside the radius — the shortlist is admissible. Ranking on it is not defensible: it ignores rivers, one-way systems and the direction of travel. Shortlist by distance, then rank the shortlist by a routing ETA, paying that call on 50 candidates rather than 120,000.
  6. Name the residual failure modes so the design is not oversold: cell bucketing misses a provider just across a boundary unless the neighbouring cells are queried, and a radius that returns nothing must widen rather than fail — with a bounded number of widenings, because a request that expires unmatched is a first-class outcome, not an error.
Worked solution 30 min
  1. Write the naive cost: 400 x 120,000 = 48M distance evaluations per cycle, 96M per second at a 500 ms cadence, before the sorts.
  2. Assume the market spans about 40 km by 40 km, giving a density of 120,000 / 1,600 = 75 providers per square kilometre.
  3. Size a geohash-6 cell: 360 / 2^15 degrees of longitude is about 1.22 km and 180 / 2^15 degrees of latitude is about 0.61 km, so a 3 km radius needs about 3 rings east-west and 5 rings north-south, a 7 x 11 block of 77 cells.
  4. Compute the candidate set: 77 cells x 0.744 square kilometres each is about 57 square kilometres, so about 4,300 candidates, against the exact circle's pi x 9 = 28.3 square kilometres and about 2,120 providers.
  5. Compare: 120,000 / 4,300 is roughly a 28x reduction in rows scanned, with about half the survivors inside the true radius and needing the exact distance filter.
  6. Redo step 1 with a fleet of 12,000 to see where the naive plan becomes acceptable.
EXPECTED RESULTThe naive plan does about 96 million distance evaluations per second. A geohash-6 ring query at the stated density returns about 4,300 candidates per request, roughly 28 times fewer rows, of which about half fall inside the 3 km circle. The cost becomes proportional to local density rather than to fleet size, which is the property that matters as the fleet grows.
Follow-up
  • At what fleet size does the naive version start meeting the budget again — show the arithmetic rather than guessing.
  • The 3 km radius returns zero idle providers at 03:00; what does the loop do next, and when does it stop trying?
  • Two dispatch partitions read the same idle provider from the index within one cycle — which layer prevents the double assignment, and why not this one?

For someone fluent in a dynamic language who has shipped real work but has never had to say what the runtime is doing underneath. The week is built on measuring and deliberately breaking things, because the questions that expose this background are the ones where the interviewer asks why a second time.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Measure before reasoning
  • Take a slow piece of your own code, write down in advance where you believe the time goes, then profile it and record how wrong the guess was. The cost is usually an allocation you did not notice or an accidental quadratic membership test.
  • Replace one list membership test inside a loop with a set and measure at a thousand, ten thousand and a hundred thousand elements, confirming the shape of the curve rather than only that it got faster.
  • Write down the three quantities you can now measure instead of assert: wall time, peak memory, and call count for the function you suspected.

Deliverable: A before-and-after profile of real code plus a written note on the size of the gap between the guess and the measurement.

Practice prompt ↗Practice prompt ↗Worked solution ↗
02References, copies, and the bugs they produce
  • Write the function with a mutable default argument, call it three times, and explain the accumulating result: the default is evaluated once when the function is defined, so every call shares one object.
  • Build a nested structure, take a shallow copy, mutate an inner element, and show that both views changed, because a shallow copy duplicates the container and not the elements. Then fix it with a deep copy and state the cost you just accepted.
  • Write two functions, one mutating its argument in place and one rebinding the local name, and predict the caller's view of each before running it. That single distinction produces most of the bugs that pass their tests.

Deliverable: Three small programs whose output you predicted correctly before running, each with a one-line statement of the rule underneath.

Practice prompt ↗Practice prompt ↗
03Types, once, in a language that checks them
  • Port one module you have already written, roughly a hundred lines, into a statically typed language, and record every place the compiler demanded an answer your original had left implicit: a value that can be absent, a numeric width, a case never handled.
  • Write the same signature in both languages and state what the static one guarantees before the program runs and what it does not, since it will not save you from a wrong algorithm or an index out of range.
  • Write the difference between an interface satisfied by declaration and one satisfied structurally, with one case each where the other approach would miss the mistake.

Deliverable: One module in two languages plus a list of the questions the type checker forced you to answer.

Practice prompt ↗Practice prompt ↗
04Concurrency, starting with what actually runs at the same time
  • Run the same CPU-bound function across four threads and four processes and measure both. Under the default CPython build the threaded version will not speed up, because only one thread executes bytecode at a time; the process version will. Check which build you are on first, since free-threaded builds remove that lock and change the result.
  • Then run a blocking I/O workload across four threads and measure it speeding up, because the interpreter releases that lock around blocking calls, which is why treating threads as useless is wrong as a general claim.
  • Build the lost update: two threads each incrementing a shared counter a hundred thousand times, and show a final value below the expected sum, because an increment is a load, an add and a store and the thread can be suspended between them. Fix it with a lock and then measure what the lock costs.

Deliverable: Three measurements, threads against processes on CPU work, threads on I/O work, and a demonstrated lost update, each with the mechanism written underneath.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05Debugging as a procedure rather than an instinct
  • Work one real failure as a bisection: find a revision or an input size where it is good and one where it is bad, halve repeatedly, and state the two assumptions bisection needs, that the property changes exactly once across the range and that the test is reliable.
  • Minimise one failing input to the smallest version that still fails, and record how many rounds it took.
  • Keep a hypothesis log for one bug in three columns, what I believe, what would disprove it, what I observed, and stop yourself the first time you are about to change two things at once.

Deliverable: One bug worked to root cause with a written hypothesis log and a minimised reproducing input.

Practice prompt ↗
06Tests that catch the bug you are about to write
  • Implement an LRU cache with a capacity bound, then write the three test cases that would catch an off-by-one in eviction: insert exactly capacity items and assert nothing was evicted, insert one more and assert the least recently used key is the one gone, and read an old key just before that insert so the eviction victim changes.
  • Add a property test comparing your implementation against a deliberately slow reference, an ordered list scanned linearly, over a few thousand random operation sequences, because a slow reference finds the cases you would not have thought to write.
  • Write one numeric test that fails under exact equality and passes with a tolerance, and state why the tolerance has to be relative rather than absolute once the magnitudes grow.

Deliverable: An LRU implementation with three boundary tests, one property test against a slow reference, and one tolerance-based numeric test.

Practice prompt ↗
07Debug something broken, out loud
  • Have someone plant three defects in a two-hundred-line program, an off-by-one, a shared mutable state bug, and a wrong error-handling path, then find them while narrating, under a fixed rule: state the hypothesis before touching anything.
  • Time each one and record which tool found it, reading, a printed value, a debugger, or a test, because the question asked in interviews is how you would find it rather than what it was.
  • Write the sentence you will use when you do not yet know the cause, one that names the next measurement instead of offering a guess.

Deliverable: A recorded debugging session with time-to-find per defect and the method that found each.

Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

Conflict answers where you were right and everyone came round are the weakest ones. Stronger: the evidence you went and collected, what would have changed your mind, and what you did in the weeks after the call went against you. Implementing a design you argued against, properly, is a specific and checkable behaviour.

These assess how you approach teamwork, conflict, and professional gro…

medium
behavioural and engineering judgement

These assess how you approach teamwork, conflict, and professional growth. Describe an achievement you are particularly proud of and the impact it had. Tell me about a time you faced a major conflict; how did you handle it? How do you approach code reviews and provide feedback to peers? What motivates you to join Delivery Hero? How do you prioritize tasks when working on a high-pressure, diverse team?

Approach
  1. Pick a story where you made the decision, not one where you watched it.
  2. Name the disagreement and how you resolved it with evidence.
  3. Give the blast radius: what could have broken, and what you measured.
Follow-up
  • How did you know your change caused the improvement?
  • What did you decide not to do, and why?

Walk an outage you owned from first signal to permanent fix

medium
incident responseblast radiuspostmortem

Pick an incident where you were the primary owner, not a helper. Tell it as a timeline: the first signal and what threshold fired it, what you ruled out and how, the mitigation you shipped and how long it took to reach production, and the permanent fix that followed. Quantify the blast radius in a countable unit - bookings affected, duplicate captures, minutes of degraded dispatch - and say how you bounded that number instead of extrapolating it. Close with the one change that made the class of failure impossible rather than merely unlikely.

Approach
  1. Lead with impact in one sentence - who was affected, for how long, in what unit - before any chronology, because that sentence is what the listener calibrates seniority against.
  2. Give the detection path honestly. 'A consumer emailed support' and 'the duplicate-capture alert fired at 14 per minute against a baseline of zero' describe very different systems, and claiming the second when it was the first collapses on the first follow-up.
  3. Run two clocks: time to mitigate and time to fix. Mitigation is whatever stops the bleeding within minutes (a flag, a rate cap, draining one partition); the fix is the structural change that lands days later. Collapsing them into one story hides whether you can triage under pressure.
  4. Bound the affected set with a query you can state, not a rate times a duration - for example successful captures grouped by booking_id having count(*) > 1 across the incident window, cross-checked against the gateway's own record. An exact list survives scrutiny; an estimate invites it.
  5. End on the structural change and say what it does not catch. A partial unique index on the live-offer status, an EXCLUDE constraint on overlapping reservations, or a guarded UPDATE whose affected-row count decides the winner each convert a silent corruption into a loud error, and each has a boundary worth naming.
Follow-up
  • What did you believe mid-incident that turned out to be wrong, and what made you drop it?
  • How did you verify impact had stopped, without relying on the alert clearing?
  • Who else on the team could have shipped the same bug that quarter, and what stops them now?

A code review disagreement you carried past one round

easy
code reviewconflictcorrectness

Recall a review where you and the author still disagreed after the first round of comments. Say what the change did, what you believed was wrong with it, and how you argued - a failing test, a two-transaction interleaving written out, a prior incident, or authority. Say how it ended: merged as written, revised, escalated, or abandoned, and who decided. Then classify your own objection: correctness bug, maintainability opinion, or style. A strong answer treats those three as deserving different amounts of pushback.

Approach
  1. Classify before you narrate. Correctness bugs justify blocking, maintainability arguments justify one round and a follow-up ticket, and style belongs in a formatter rather than in a human's comment.
  2. For a correctness objection, describe the cheapest convincing artifact you produced. Usually that is a failing test or four lines showing two workers both reading status = 'offered' under READ COMMITTED, the default isolation level in PostgreSQL, and both proceeding to write an accept.
  3. Say who converged and how. 'They agreed once the test went red' and 'we escalated to the tech lead and I lost' are both credible; 'I approved it to avoid friction and it broke in production' is also credible if you follow it with what you would do now.
  4. Account for the cost you imposed. A blocked review costs the author momentum and sometimes a release slot, and a senior reviewer prices that in rather than treating rigour as free.
  5. Close with the systemic change if there was one - a test in CI, a constraint moved into the schema, a lint rule. The same objection raised by hand twice is a process failure, not a reviewing triumph.
Follow-up
  • When did you last approve something you disagreed with, and how did it turn out?
  • Where does that class of bug get caught today, if not in human review?
  • How do you run the same objection when the author is considerably more senior than you?
  • 01

    These assess how you approach teamwork, conflict, and professional growth. Describe an achievement you are particularly proud of and the impact it had. Tell me about a time you faced a major conflict; how did you handle it? How do you approach code reviews and provide feedback to peers? What motivates you to join Delivery Hero? How do you prioritize tasks when working on a high-pressure, diverse team?

  • 02

    Pick an incident where you were the primary owner, not a helper. Tell it as a timeline: the first signal and what threshold fired it, what you ruled out and how, the mitigation you shipped and how long it took to reach production, and the permanent fix that followed. Quantify the blast radius in a countable unit - bookings affected, duplicate captures, minutes of degraded dispatch - and say how you bounded that number instead of extrapolating it. Close with the one change that made the class of failure impossible rather than merely unlikely.

  • 03

    Recall a review where you and the author still disagreed after the first round of comments. Say what the change did, what you believed was wrong with it, and how you argued - a failing test, a two-transaction interleaving written out, a prior incident, or authority. Say how it ended: merged as written, revised, escalated, or abandoned, and who decided. Then classify your own objection: correctness bug, maintainability opinion, or style. A strong answer treats those three as deserving different amounts of pushback.

PracHub interview preparation framework ↗
Is this an official Delivery Hero interview guide?

No. It is PracHub's own research and practice material for the Software Engineer role at Delivery Hero. Rounds and questions reflect what candidates have reported, not a process Delivery Hero has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
How long does the interview process usually take?

The process typically moves quickly, often spanning 3 to 6 weeks, though this can vary based on team availability. You can expect regular updates from your recruiter throughout the stages.

PracHub interview research ↗
What is the "Bar Raiser" round?

The Bar Raiser is an interview conducted by someone outside of the team you are interviewing for. Their goal is to ensure a high and consistent standard of hiring across the entire company, focusing on both technical competence and long-term cultural fit.

PracHub interview research ↗
Is there feedback provided if I am not selected?

While the company strives to provide feedback, it is not always guaranteed for every candidate. If you are not selected, focus on the specific areas where you felt less confident during the technical rounds.

PracHub interview research ↗
Is the work environment remote or hybrid?

Expectations vary by location and team, but Delivery Hero generally operates with a hybrid approach. Be sure to clarify the specific expectations for your role during the initial recruiter screen.

PracHub interview research ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.