As a Machine Learning Engineer at United Airlines, you will be at the intersection of massive-scale logistics and cutting-edge data science. You are not just building models; you are solving the complex optimization challenges inherent in global aviation, from dynamic pricing and flight scheduling to predictive maintenance and passenger experience personalization.
The impact of your work is tangible and immediate. By refining algorithms that process vast amounts of operational data, you directly contribute to the efficiency of the world’s most interconnected airline. You will collaborate with cross-functional teams to turn raw data into strategic insights, ensuring that United Airlines maintains its competitive edge through technical innovation and operational excellence.
Preparation focus
editorialNo round sequence has been reported for this company, so work the categories below and confirm the format with your recruiter.
What to demonstrate
- Breadth across SQL, experimentation and product reasoning
- Ability to state assumptions before choosing a method
How to prepare
- Drill the practice exercises below and time yourself
- Prepare three quantified stories about decisions you drove
PracHub editorial advice for the preparation topics above.
Running a schema change as though the lock lasts as long as the statement
In PostgreSQL an ALTER TABLE that needs an ACCESS EXCLUSIVE lock must first wait for every open transaction touching that table, and while it waits, later queries needing a conflicting lock queue behind it rather than overtaking it. A DDL statement that would execute in milliseconds, issued while a thirty-second analytics query is open, therefore stalls all traffic on that table for thirty seconds: the outage length is set by the longest open transaction, not by the change. The defences are specific and worth knowing by name - set lock_timeout low and retry rather than queue, add columns without a volatile default so no table rewrite occurs (from version 11 a non-volatile default is a metadata-only change), build indexes with CREATE INDEX CONCURRENTLY while accepting that it cannot run inside a transaction block and leaves an invalid index behind if it fails, and add constraints as NOT VALID followed by a separate VALIDATE CONSTRAINT, which takes a weaker lock.
Assuming an isolation level prevents the anomaly you actually have
Isolation levels are named by the SQL standard but implemented differently, so any claim about one is only true of a named engine. PostgreSQL defaults to READ COMMITTED, where every statement takes a fresh snapshot, so two statements inside one transaction can legitimately disagree about the same row. Its REPEATABLE READ is snapshot isolation: it removes non-repeatable and phantom reads but permits write skew, where two transactions each read a set, each conclude their own write is safe, both commit, and the combined result violates a constraint that no single row expresses. Only SERIALIZABLE closes that, and it closes it by aborting a transaction with a serialization failure (SQLSTATE 40001), which means the guarantee is theoretical unless the application has a retry loop. InnoDB's REPEATABLE READ is a different mechanism again - plain SELECTs read a consistent snapshot while locking reads and writes see the latest committed row - so a read-modify-write inside one transaction can act on a value that the transaction's own earlier SELECT never returned.
Optimising an axis nobody named
Ask which resource is actually scarce here: wall-clock latency, throughput, memory footprint, cost per request, or engineering time. Shaving a constant factor off an in-memory step is wasted effort when the same function makes a blocking remote call inside the loop.
Saying 'eventually consistent' without naming the anomaly a user would see
Describe the concrete symptom you are choosing to accept: the author reloads and their own comment is missing for two seconds, or two devices show different balances for a minute. The class of consistency model is a technical label; the tolerable anomaly is the actual product decision.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Explain the process of feature engineering for a time-series forecasti…
Explain the process of feature engineering for a time-series forecasting model.
Approach
- Pick the metric from the cost of each error type, not from habit.
- Say how you would validate it, and where leakage could enter the split.
- Name the simplest model that could work and what would make you move past it.
Follow-up
- What changes if the classes are heavily imbalanced?
- How would you know the model is overfitting?
How do you validate the performance of a model before deploying it int…
How do you validate the performance of a model before deploying it into a live environment?
Approach
- State the learning problem: the label, the unit of prediction and how the model is used.
- Name the simplest model that could work and what would make you move past it.
- Say how you would validate it, and where leakage could enter the split.
Follow-up
- Where could label leakage enter this setup?
- What changes if the classes are heavily imbalanced?
Merge partitioned event streams into one ordered feed with bounded lateness
The read-model service consumes 64 log partitions carrying about 4,000 events per second in total. Each partition is ordered within itself, but partitions drift by up to 30 seconds, and the activity feed must present a tenant's events in occurred_at order. Produce the merge. State its complexity, the buffer it requires in events and in bytes, what happens when one partition is idle, and what you do with an event that arrives after you have already emitted its position. Payloads average 1 KB.
Approach
- Merge with a min-heap over the 64 partition heads keyed on (occurred_at, event_id): O(log P) per event and O(n log P) overall. The tie-break on event_id is what makes the output deterministic when two partitions carry the same millisecond, which matters because the feed is paginated and a non-deterministic order reorders pages under the reader.
- Emitting the heap head is only correct once every partition has produced everything up to that timestamp, so the emit condition is a watermark: the minimum across partitions of the highest occurred_at seen, less the allowed lateness. Events are held until the watermark passes them, which is what turns individually ordered streams into a jointly ordered one.
- Size the buffer from the lateness rather than guessing: 4,000 events per second times 30 seconds is 120,000 buffered events, and at 1 KB each about 120 MB of heap. That number is the real price of the ordering guarantee and belongs in front of whoever asked for it.
- Handle the idle partition explicitly, because it fails the feed rather than corrupting it: a partition with no traffic never advances its own maximum, so the watermark freezes and output stops entirely. Either every partition emits a periodic idle marker carrying the broker's current time, or the watermark falls back to wall clock for a partition silent beyond a threshold.
- Choose the late-event policy from what the projection is keyed on. The projection upserts on (aggregate_id, aggregate_version) and discards a version it has already applied, so a late event is safe to apply out of order and correctness never depended on the merge at all. Apply it, recompute the affected feed page, and count lateness so the 30-second budget can be re-derived from data rather than folklore.
- Say what the merge does not buy: ordering is guaranteed within one aggregate by the log's partitioning, and no watermark makes the cross-aggregate order authoritative. Two events from different aggregates in the same millisecond have no true order, so the feed's order is a presentation choice that must be stable rather than correct.
Worked solution 35 min
- Write the heap comparator on (occurred_at, event_id) and the per-partition head refill.
- Write the watermark computation and the emit-loop condition, then list which buffered events are held at a chosen instant.
- Compute the buffer at 4,000 events per second, 30 seconds and 1 KB per event, and state what fraction of a worker's heap that represents.
- Add the idle-partition marker and trace the watermark with one silent partition, both with and without the marker.
- Write the late-event path and name the key that makes applying it safe.
Follow-up
- The lateness budget is raised to five minutes. What is the new buffer, and what besides memory changes?
- The consumer restarts. Where does it resume from, and what does the feed look like for the first 30 seconds?
- One partition is ten minutes behind because its producer is slow. Do you stall the feed or emit without it?
What are the trade-offs between different join types in SQL when deali…
What are the trade-offs between different join types in SQL when dealing with large-scale datasets?
Approach
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Say which index the query would use, and what makes it unusable.
- Name the grain you start from and join outward from it.
Follow-up
- What happens to this when the table is ten times larger?
- How does the query change if that join becomes one-to-many?
Can you describe a time you had to optimize a slow-running SQL query?
Can you describe a time you had to optimize a slow-running SQL query?
Approach
- Check whether any join is one-to-many before aggregating, or the sums inflate.
- Say which index the query would use, and what makes it unusable.
- State the isolation you are assuming and the anomaly it still allows.
Follow-up
- How does the query change if that join becomes one-to-many?
- How would you run this migration without downtime?
Write the update path that detects a concurrent edit
resource carries version INT NOT NULL DEFAULT 1. resource_revision holds revision_id, resource_id, version, actor_user_id, change_kind, patch JSONB, request_id, created_at with UNIQUE (resource_id, version). outbox_event holds aggregate_type, aggregate_id, aggregate_version, event_type, payload, status. A PUT carries the version the client read. Write the exact statements for the single transaction that applies the edit, records the revision and enqueues 'resource.updated', and give the handler's branch on zero affected rows. Then say what PostgreSQL 16 does under READ COMMITTED when two of these updates hit one row at once.
Approach
- One transaction, three writes, no network call inside it: UPDATE resource SET title = $3, version = version + 1, updated_at = now() WHERE resource_id = $1 AND tenant_id = $4 AND version = $2; then INSERT the resource_revision row at version $2 + 1; then INSERT the outbox_event row at the same aggregate_version. The event goes to a table rather than a broker because no transaction spans both.
- Branch on the affected-row count before doing anything else. Zero has three causes — stale version, wrong tenant, row gone — so re-read once and map to 409 carrying the current version, or 404 for an id outside the caller's tenant, which also stops the endpoint confirming that another tenant's id exists.
- State the engine behaviour instead of assuming it. Under READ COMMITTED the second UPDATE blocks on the row lock, and when the first commits PostgreSQL re-evaluates the WHERE clause against the newly committed row, so the version predicate now fails and the statement reports zero rows. Under REPEATABLE READ the identical collision raises SQLSTATE 40001 instead, so the handler must fold both shapes into one conflict response.
- Keep UNIQUE (resource_id, version) even though the predicate already serialises writers. It is what makes a lost update unwritable if any other path ever reaches the revision table, and it converts a logic bug into 23505 rather than into a silently missing history row.
- Refuse to auto-retry the whole PUT. A retry re-reads the winner's state and reapplies an intent formed against data that no longer exists — the silent overwrite the version token was added to detect. Return the conflict; merge field-wise only if the patches are provably disjoint.
- Note that now() is the transaction timestamp in PostgreSQL, so resource.updated_at, the revision's created_at and the outbox row share one instant, which is what later makes reconciliation between the three tables unambiguous.
Worked solution 25 min
- Write the three statements plus the rowcount branch and confirm they sit in one BEGIN/COMMIT with no outbound call between them.
- Run two clients that both read version 7 and apply their updates 5 ms apart; assert one 200 and one 409.
- Assert resource.version = 8, exactly one resource_revision row at version 8, and one outbox_event row at aggregate_version 8.
- Repeat at REPEATABLE READ and record the different failure shape (SQLSTATE 40001) the handler must also map to 409.
- Delete the version predicate and re-run: both writes commit and the first edit disappears with no error raised anywhere.
Follow-up
- A client sends the version it read ten minutes ago and the resource has moved three versions. What is in your 409 so it can resolve the conflict without a full re-fetch?
- Two editors, two disjoint fields, no overlap. Does your answer still refuse the second write, and should it?
- Every write now touches a second hot table. How do you keep the outbox insert and its partial index from becoming the write bottleneck at 1.2k writes/second?
Design the async export contract a client can resume safely
A tenant asks for a CSV of every resource. The work runs for minutes on the worker fleet through a job_run row carrying a lease, an attempt count and a dedupe_key, far past the edge's 400 ms budget. Callers are a browser that polls and a script that walks away and checks later. Specify what the submit call returns, the operation resource and its states, how a duplicate submit is handled, how a client learns about completion, what cancellation means given that a lease can expire mid-run, and how the result is fetched and when it expires.
Approach
- Split the API in two. Submit returns 202 with an operation id and a location to poll, and never blocks on the work. The operation is a real resource with its own lifecycle - queued, running, succeeded, failed, cancelled - plus attempt, a monotonic progress figure, and a terminal error drawn from the same code taxonomy the synchronous endpoints use, so a client needs one error vocabulary rather than two.
- Deduplicate at submit using job_run.dedupe_key, unique over (job_type, dedupe_key) while status is 'queued' or 'running': a repeat submit of the same logical export returns 200 with the existing operation instead of 202 with a new one, and the partial index deliberately permits a legitimate re-run once the first has finished. Pair it with the request's idempotency key so an HTTP-level retry of the submit is exact rather than merely similar.
- Tell the poller how to poll: Retry-After on the polling response, a minimum interval enforced at the edge, and a documented maximum lifetime after which an operation is reaped. Polling is the contract of record; the webhook is the fast path, and both must lead to the same terminal state, so a client that receives the completion event and then polls anyway sees no contradiction.
- Be exact about cancellation. A cancel request records intent; it cannot stop work already executing. The handler reads the flag at checkpoints, and because a lease expires on a clock that cannot distinguish a dead worker from a slow one, a second copy may start after the cancel was recorded - so the handler re-reads the flag immediately after claiming the lease. 'cancelled' becomes terminal only when no lease is outstanding; reporting it earlier shows a client a stopped job while a worker is still writing output.
- Make the handler safe to run twice, because the lease guarantees that it will be. Write output to a deterministic object key derived from the operation id so a second copy overwrites its own work instead of appending a second file, and record completion with a conditional update that only the copy holding the current lease can win.
- Treat result fetch as a separate authorised read: a short-lived signed URL, the tenant checked when it is issued rather than only when the file was produced, and a documented retention after which the operation remains terminal but the bytes are gone - a state the client must be able to tell apart from a failure.
Follow-up
- An operation has said 'running' for 40 minutes and the worker is gone. What does the client see, and which columns in job_run decide that?
- Two tenants each submit 50 exports at once. What in this contract stops one of them delaying the other?
- The customer wants the export emailed instead. What changes, and what becomes harder to make exactly-once?
Design the bulk write endpoint a migration script retries blindly
A customer's migration script pushes 2 million resources through POST /v1/resources:batch, up to 500 items per call, and retries any call that errors or times out. Within a call some items fail validation, some collide with rows that already exist, and some succeed. Specify the request and response shape, whether a batch is atomic or per-item, how idempotency works for the call and for each item, the status code for a mixed outcome, the size and item-count limits with their error codes, and exactly what the script does after a timeout mid-batch.
Approach
- Choose atomicity deliberately and price it. All-or-nothing means one transaction holding locks for the whole batch on a primary already absorbing about 1.2k writes/second, which bounds batch size by lock duration, and it turns one bad row into 499 rejections the script must resend. Per-item partial success is the right default for a migration, and the contract's job is then to make a partial outcome impossible to miss.
- Require a client-supplied id on every item and echo it in every result. Deriving a per-item key from the array index breaks the first time the script resends a batch with the failures removed: the indices shift, previously-succeeded items acquire new keys, and they are created a second time.
- Key the effects at two levels. The call's Idempotency-Key covers an exact resend of the same bytes; per-item keys of (tenant_id, client_item_id) make a partially-applied batch safe to resend whole. Resending an identical batch must reproduce the same per-item results, not 500 conflicts the script has to interpret.
- Answer a mixed outcome with one status plus per-item detail: 200, or 207 borrowed from WebDAV if you prefer it - document whichever you pick - carrying an array of client_item_id, per-item status, and either resource_id or an error code from the same taxonomy the single-item endpoint uses. Reserve 4xx for the request as a whole: unparseable body, too many items, payload over the limit (413). Put a failed count at the top level so that even a script checking only the cheapest thing cannot conclude success while rows were dropped.
- Cap the request before doing any of it: item count, total bytes, and concurrent batches per tenant, since one tenant's migration otherwise consumes write capacity everyone shares. Anything that cannot finish inside the request deadline belongs on job_run behind a 202, not in a synchronous call that will time out halfway.
- Script behaviour after a timeout: the outcome is unknown and no results were received, so resend the identical batch with the same keys and read the results. Never resend 'only the items I have no result for' - a timeout yields no results at all, and that rule silently means resend everything anyway.
Worked solution 30 min
- Write the request schema with the per-item client id, and the response schema with per-item status and a top-level failed count.
- Write the atomicity decision and the sentence of justification that names the lock cost or the resend cost.
- Define both key levels and trace a resend of a half-applied batch through them, item by item.
- List the whole-request rejections with their status codes and limits.
- Write the script's timeout rule and the pacing it should apply between calls.
Follow-up
- At 500 items per call, what is the wall-clock time for 2 million rows, and what pacing do you publish so the migration does not become an incident?
- One item in every batch fails with the same code. How does the script discover that without a human reading logs?
Edge instances grow 400 MB per hour until the nightly restart
Edge API instances start at 700 MB resident and grow about 400 MB/hour; a nightly rolling restart has hidden it for weeks. Growth continues unchanged when request rate halves overnight, p99 degrades in the last hours before an instance is recycled, and heap used immediately after a forced full GC rises monotonically. The service holds no product state. Name the discriminating measurement that separates the plausible causes, give the most likely cause, and give the fix and how you would verify it.
Approach
- Separate resident memory from live heap first, because they fail differently. Resident size can grow from fragmentation, native buffers or thread stacks while the heap is flat; heap used after a full GC rising monotonically is the measurement that says objects are reachable and not being released. You already have it, so this is retention, not fragmentation, and that closes off half the candidate list.
- Use the rate's independence from traffic as the discriminator. Growth that continues at half the request rate rules out per-request objects that are merely slow to collect and points at a structure that grows with distinct values observed rather than with call volume. Write the candidates that have that property: a metrics registry keyed on a high-cardinality label, an unevicted cache, an interner, a per-key lock map.
- Take two heap snapshots an hour apart and diff by retained size, reading the dominator tree, not by allocation count or instance count. Expect one root holding a map with millions of entries, then follow the reference chain to the code that inserts and never removes. Allocation profilers point at churn, which is the wrong signal here.
- The candidate that fits this service is an observability label carrying an identifier, such as a request path recorded before templating so that /v1/resources/48213 becomes its own metric series. That grows with distinct ids seen, is independent of rate, and explains the late p99 degradation, since GC cost rises with the size of the live set.
- Fix by bounding cardinality at the source: template the path to /v1/resources/{id} before it becomes a label, move tenant id from a label to a log field or an exemplar, and cap the registry with a bounded map that evicts. Add a cardinality ceiling that fails loudly in a lower environment rather than growing quietly in production.
- Verify with a soak rather than a restart. Hold one instance out of the nightly recycle for 48 hours with the fix and compare post-GC heap and series count against an unfixed control taking the same traffic.
Follow-up
- Post-GC heap is now flat but resident size still creeps. What are you looking at, and does it matter?
- How would you have detected this before an OOM, given the nightly restart masked the trend?
- That label is what makes one dashboard useful. How do you keep the dashboard and lose the leak?
For someone who has spent the last few years shipping features and reading other people's code, and who has not solved a timed problem from a blank file in a long time. Five days rebuild the primitives and the patterns that sit on them, working from invariants rather than remembered solutions, and the last two attach that back to the rest of the loop.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Rebuild the primitives by implementing them
- Implement a dynamic array with doubling growth and an operation counter, then change the growth rule to add a fixed sixteen slots instead, and time both for n of ten thousand, a hundred thousand and a million. The fixed-increment version resizes n/16 times at O(n) each, so its total work is quadratic; doubling is what makes append amortised constant.
- Implement a hash map with separate chaining and a load-factor resize, then insert ten thousand keys engineered to land in one bucket and record what happens to lookup time, so that average-case O(1) becomes a claim with a stated precondition rather than a reflex.
- For dynamic-array append and hash-map insert, write down which cost is amortised rather than worst-case, which single operation pays the whole bill, and what a system with a hard per-operation deadline would have to do instead.
Deliverable: Two working implementations plus a timing table showing the input at which each structure's advertised complexity stops holding.
Practice prompt ↗Practice prompt ↗Worked solution ↗02Arrays under an invariant: two pointers, sliding window, binary search
- Solve longest-subarray-with-sum-at-most-K using a sliding window, then run it on an input containing negative numbers and watch it return the wrong answer: extending the window only moves the sum monotonically when every element is non-negative, and that precondition is the whole reason the technique works.
- Write the binary search that finds the first index satisfying a predicate rather than an exact value, put the loop invariant above the loop in a comment, and verify termination on the two inputs that break careless versions: the empty range, and a range where every element satisfies the predicate.
- Compute the midpoint as lo + (hi - lo) / 2 and write one line on why the obvious (lo + hi) / 2 is a genuine defect in a fixed-width integer type and a non-issue in a language with arbitrary-precision integers.
Deliverable: Three solved problems, each with its invariant written above the loop, plus one recorded input on which the sliding window is provably wrong.
Practice prompt ↗Practice prompt ↗03Sorting, heaps, and the greedy argument that has to be proved
- Solve one top-k problem three ways, by full sort, by a size-k heap, and by quickselect, then write the values of n and k at which each becomes the right choice, along with quickselect's quadratic worst case and why a randomised pivot makes that unlikely rather than impossible.
- Implement bottom-up heapify and count sift-down steps to confirm it does linear work rather than n log n, because most nodes sit near the bottom of the tree and therefore move only a short distance.
- Take interval scheduling by earliest finishing time and write the exchange argument out in full: given any optimal schedule, swapping in the earliest-finishing interval keeps it feasible and no smaller. Then construct the weighted variant where that same greedy fails and name what has to replace it.
Deliverable: A three-way top-k comparison with measured crossover points, one written exchange argument, and one counterexample to a greedy rule that looks almost identical.
Practice prompt ↗Practice prompt ↗04Recursion, memoisation, and the step to a table
- Take one problem with overlapping subproblems, such as edit distance or coin change, instrument the plain recursion with a call counter to show the blow-up, then add memoisation and re-count.
- Convert the memoised version to a bottom-up table and state the two properties you relied on: each subproblem's result depends only on its arguments, and the dependencies form a DAG you can enumerate in order.
- Rewrite one deep recursion with an explicit stack, then find the input length at which the original hits the interpreter's frame limit, which defaults to about a thousand frames in CPython, so you know when the rewrite is required rather than decorative.
Deliverable: One problem in three forms, naive, memoised and tabulated, with call counts for each and the input length at which recursion depth becomes the binding constraint.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Graphs, where most of the work is choosing the traversal
- Implement BFS and DFS over one adjacency list, then answer for each which finds a shortest path in an unweighted graph and which you would use to detect a cycle in a directed graph, including why the in-progress versus finished distinction matters for the second.
- Implement topological sort by in-degree, feed it a graph containing a cycle, and confirm the failure signature is that fewer than V nodes come out rather than an exception, then note that the order it produces is one of several valid ones.
- Run a shortest-path search on a graph with a single negative edge weight and show the wrong answer, then write the precondition Dijkstra actually needs, non-negative weights, because it finalises a node's distance the first time that node is popped, and name the algorithm you would switch to and its own limit.
Deliverable: A small graph library with BFS, DFS and topological sort, plus two inputs that produce documented wrong answers under the wrong algorithm choice.
Practice prompt ↗Practice prompt ↗06One day for everything that is not an algorithm
- Sketch one system only to the depth a coding-heavy loop tends to reach: the endpoints, what the service stores, and the single query pattern that decides the schema. Stop at twenty-five minutes.
- Prepare the project answer for an interviewer who codes, which means rehearsing the two levels they push to: the specific thing you built, and why you chose that approach over the alternative they will name. Open with a number and be ready to say what it excludes.
- Prepare the answer to what you would do differently, choosing a real technical mistake with a specific fix rather than a complaint about process or staffing.
Deliverable: One design sketch at endpoint-and-schema depth, plus a project answer rehearsed to two levels of follow-up.
Practice prompt ↗07Solve out loud, under time
- Do three timed problems at twenty-five minutes each in a plain editor with no autocomplete and no execution until the end, then tally separately the failures that were syntax and the ones that were approach, because those two numbers call for different fixes.
- Narrate one solution from the first sentence, stating the approach and its complexity before writing any code, and rehearse the sentence you will use when you realise mid-solution that the approach is wrong.
- Re-solve from blank the two problems you were slowest on this week and compare the times against the day they first appeared.
Deliverable: A recording of one fully narrated solution and a tally that separates syntax failures from approach failures.
Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
For anything that touched live traffic, be ready to say how you would have undone it: a flag, a staged rollout, dual writes with the old path still authoritative. Once the old column is dropped or the source rows are overwritten there is no reverse, so name what you kept a copy of and for how long.
How do you handle missing or noisy data in a production-level machine …
How do you handle missing or noisy data in a production-level machine learning pipeline?
Approach
- State the situation in two sentences and spend the rest on the reasoning.
- Pick a story where you made the decision, not one where you watched it.
- Close with what you would do differently, concretely.
Follow-up
- How did you know your change caused the improvement?
- What would you do differently if you ran that again?
Narrate an outage you owned from page to postmortem
Pick an incident you personally drove, ideally one where writes were affected rather than reads. In six to eight minutes: state the symptom as it first appeared on a dashboard, the blast radius you established before you knew the cause, the mitigation you applied and when, the mechanism you eventually proved, and the follow-up that would prevent a repeat. Bring numbers: error rate, tenants affected, minutes to mitigate, minutes to resolve. If you cannot name what you measured, choose a different incident.
Approach
- Open on the signal rather than the cause: which metric at which percentile moved, on which service, at what time, so the listener follows the same evidence you had rather than a conclusion you already reached.
- Separate mitigation from diagnosis out loud. State what you did to stop the bleeding (flag off, shed traffic, drain a lease, roll back a deploy) and say plainly that you did it before the mechanism was known, because those are two jobs with different deadlines.
- Establish blast radius in countable terms: how many tenants, how many writes, and crucially whether the effect was loss or only delay. An append-only revision table or a pending outbox row means the change survived and the projection was merely behind, which is a repair rather than a data-loss incident.
- Prove the mechanism instead of asserting it. Name the trace span that grew, the plan that flipped to a sequential scan, the lease that expired, plus one alternative you ruled out and the signal that stayed flat while you ruled it out.
- Close on the durable fix and its cost, distinguishing what landed that week from what needed an expand-and-contract migration across several deploys, and say which of the two you actually finished.
Follow-up
- What would you do differently in the first five minutes, given the same dashboard and no more information?
- Which follow-up action did you deliberately not take, and why was dropping it the right call?
- How did you convince yourself the mitigation was safe to apply while the cause was still unknown?
Turn a code review disagreement into a decision
A colleague's change updates a row with UPDATE resource SET version = version + 1 WHERE resource_id = $1 AND version = $2 and treats an affected-row count of zero as a successful no-op. You read that as a silently lost update; they think returning 200 is friendlier to clients than returning a conflict. Describe how you have handled a review disagreement of this shape: what goes in the comment, when you leave the thread, and who decides. Then write the comment you would leave here, in under 80 words.
Approach
- Sort the disagreement before writing anything. A silently discarded write is a correctness claim about data; the choice between 409 and 412 is taste. Only the first justifies blocking a merge, and saying which one you are doing is most of the value of the comment.
- Make the claim reproducible in the comment itself with an interleaving rather than a principle: A reads version 7, B reads version 7, B commits version 8, A's predicate matches zero rows, A is told it succeeded and A's edit is gone.
- Offer the alternative with its cost attached: return 409 carrying the current version and the revision that won, so the client can re-read and re-apply. Note that automatic retry is not the fix, because a retry re-reads the winner's state and reapplies an intent formed against data that no longer exists.
- Apply an escalation rule you can state: two round trips on the thread, then a call, and the service's owner decides rather than the reviewer. A reviewer who cannot be overruled is a bottleneck with extra steps.
- Close in writing wherever the decision lands, so the next reader finds the reasoning in the code or the ticket instead of in a collapsed review thread.
Follow-up
- Where would you put the test that fails if someone reintroduces the swallowed zero rowcount?
- The author says clients cannot handle a 409. How do you check whether that is true?
- How do you handle the same review comment when the author is more senior than you and in a hurry?
- 01
How do you handle missing or noisy data in a production-level machine learning pipeline?
- 02
Pick an incident you personally drove, ideally one where writes were affected rather than reads. In six to eight minutes: state the symptom as it first appeared on a dashboard, the blast radius you established before you knew the cause, the mitigation you applied and when, the mechanism you eventually proved, and the follow-up that would prevent a repeat. Bring numbers: error rate, tenants affected, minutes to mitigate, minutes to resolve. If you cannot name what you measured, choose a different incident.
- 03
A colleague's change updates a row with UPDATE resource SET version = version + 1 WHERE resource_id = $1 AND version = $2 and treats an affected-row count of zero as a successful no-op. You read that as a silently lost update; they think returning 200 is friendlier to clients than returning a conflict. Describe how you have handled a review disagreement of this shape: what goes in the comment, when you leave the thread, and who decides. Then write the comment you would leave here, in under 80 words.
Is this an official United Airlines interview guide?
No. It is PracHub's own research and practice material for the Machine Learning Engineer role at United Airlines. Rounds and questions reflect what candidates have reported, not a process United Airlines has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult is the technical assessment?
The assessment is of average difficulty for a mid-level engineer. It focuses on practical application rather than "trick" questions; if you are comfortable with intermediate to advanced SQL, you will be well-prepared.
PracHub interview research ↗What is the company culture like?
The culture is described as professional, collaborative, and highly focused on operational success. You will find that interviewers are personable and interested in your potential to contribute to the team.
PracHub interview research ↗How long does the hiring process take?
While timelines vary, you can expect a cadence that includes a two-week window between the initial interviews and the follow-up technical assessment.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-30 - 02PracHub Machine Learning Engineer practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-30 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-30