A Software Engineer at Optum Health Services plays a critical role in designing, developing, and maintaining high-scale technology solutions that power the modern healthcare system. As part of UnitedHealth Group (UHG), Optum is dedicated to improving the health system for everyone. In this position, you will work on platforms that process massive volumes of clinical data, manage complex pharmacy benefits, and streamline provider-patient interactions. Your work directly impacts millions of lives by making healthcare more accessible, efficient, and secure.
The engineering culture at Optum Health Services balances the agility of modern software development with the rigor and security required for healthcare infrastructure. You will be tasked with modernizing legacy systems, building robust microservices, and ensuring seamless data integration across various health networks. Whether you are developing secure APIs, optimizing database queries, or deploying cloud-native applications, your technical contributions will solve real-world problems that have a profound human impact.
To succeed in this role, you must possess strong technical foundations, a passion for clean code, and a collaborative mindset. You will work closely with cross-functional teams, including product managers, data scientists, and clinical experts, to deliver high-quality software. Optum values engineers who are not only technically proficient but also deeply aligned with the company's mission to help people live healthier lives and make the health system work better for everyone.
Recruiter Screening Call
reportedThe person on this call usually cannot evaluate your code and does not need to. They write a short paragraph, and that paragraph is what a hiring manager skims when deciding who to put on your loop. So the test is not whether your work was hard, it is whether a non-engineer can repeat it correctly. Name systems by what they did rather than by their internal codename, give each project a shape (what was breaking, what you changed, what happened after), and keep the whole walkthrough near ninety seconds. Depth that cannot survive a paraphrase reads as vagueness.
What to demonstrate
- Whether a non-engineer can restate your projects without distorting them, since their paraphrase is what travels to the hiring manager, not your sentences
- Whether each project has a shape rather than a stack list: the failure or constraint, the change you made, the result and how it was measured
- Whether you can say what was yours inside a team project without either inflating it or disappearing into the plural
How to prepare
- Rewrite each headline project as two sentences with no internal system names and no acronyms outside your company, then say them to someone outside engineering and have them repeat them back. Fix whatever came back wrong
- Attach one measured number to each project: the baseline, the change, and the window it was measured over. Where nothing was ever measured, say that plainly rather than reaching for a plausible percentage
- Time the background walkthrough against a clock. If it runs past two minutes, compress the earliest role to a single clause and spend the recovered time on the most recent one
Asynchronous Assessment
reportedThe same problem is scored by two different mechanisms depending on the format, and preparing for one does not cover the other. With a person watching, partial progress is visible and a hint is a correction you can absorb; silence is the expensive failure, because nobody can read a half-written function. With an automated grader there is no partial credit for what you were about to do, nobody to ask, and the worked examples in the prompt are the entire specification. Read them as a contract, down to whether an empty result should be an empty list or no output at all.
What to demonstrate
- In a live session, whether your commentary tracks what your hands are doing, and whether a hint redirects you or gets defended against
- In an automated one, whether you cover the cases the examples do not show, since the hidden cases are where the score moves
- Whether you manage the clock on purpose: abandoning an approach that is not converging while there is still time to write something simpler that finishes
How to prepare
- Have someone hand you a problem and feed you one deliberately wrong hint. Practise testing it against a concrete case instead of accepting or rejecting it on authority.
- Do one timed run a week in a plain browser editor with autocomplete, linting and your own snippets switched off, which is closer to what these environments give you
- For the automated format, write the harness before the solution: a main that feeds the worked examples plus an empty and a single-element case and prints expected against actual, so a wrong submission is caught by you first
Technical Panel Interview
reportedCoding rounds mostly set a floor. They decide whether you clear the bar, not where you land on the ladder. Level tends to come out of the design discussion and the ownership stories, so the question worth auditing beforehand is whether the scope you describe matches the scope of the job. Work that stops at your own service, or a story whose hard part was writing the code rather than getting several people to agree on an interface, reads a level below where you think you are interviewing, and that gap is usually resolved downwards.
What to demonstrate
- Whether the largest thing you describe owning ran end to end — the decision, the migration path, the rollout, and what you did when it went wrong — or stopped at the change you merged
- Whether design answers include what you would not build, what you would defer, and what you would measure before committing, rather than only what the boxes are
- Whether a disagreement in a story was settled with something checkable — a benchmark, a prototype, a written proposal — instead of by seniority or by waiting it out
- Whether you can say which calls you made alone and which you escalated, and why the line sat where it did
How to prepare
- Write your largest piece of owned work as a timeline of decisions — who decided what, when, and what you did when the plan broke — then delete every sentence whose subject is "we" and see how much survives
- Take one system you know well and drill the migration answer: how old and new paths run side by side under live traffic, how you compare their outputs, what the rollback is once writes are going to both, and which step you would not automate
- Map each line of the ladder in the job posting to a specific thing you have done, find the line you cannot support, and prepare the closest evidence you have plus an honest account of the gap
Live Coding Exercise
reportedMost of the time lost in this format is not lost to thinking. It goes to a standard-library call you half-remember, an off-by-one in a loop bound, and a debugging loop that mutates code at random until something passes. When output is wrong, stop re-reading the whole function: take the smallest input that reproduces it and walk the state through by hand, printing intermediates if the environment allows. Guessing at a fix without a failing case you understand is how a five-minute bug becomes twenty, and the clock does not pause while you do it.
What to demonstrate
- Whether you reach the right structure without a detour, and can write it from memory rather than only recall that one exists
- Whether overflow is considered where the language has fixed-width integers, since a signed 32-bit value stops at 2,147,483,647 and then wraps in Java, is undefined behaviour in C++, and does not arise in Python, whose integers grow instead
- Whether recursion depth is treated as a constraint on large inputs, given that CPython's default limit is 1000 frames and a deep recursion can exhaust the stack in any language where an iterative version would not
- Whether a failing case is isolated and explained before any edit is made to the code
How to prepare
- From an empty file and with no references open, implement the pieces you lean on most: a heap push and pop, an iterative DFS with an explicit stack, and a binary search whose midpoint is written lo + (hi - lo) / 2, which avoids the overflow that (lo + hi) / 2 can hit in a fixed-width integer type
- Time yourself on the ten library calls you look up most, such as sorting with a custom comparator, splitting and joining strings, and finding the next key at or above a value in an ordered map, until the lookup is gone
- Take a solution you know is broken and, before touching it, write one sentence naming the input, the expected value and the actual value. Repeat until you do it without deciding to.
Behavioral Discussions
reportedThis round is deciding whether a change you make without supervision can be allowed to reach production. It is scored on what you knew at the moment you decided, not on how it turned out, so a story that opens with the result and works backwards reads as luck retold as judgement. Say what the options were, what you did not know, what you did to shrink the unknown before committing, and what you accepted as the worst plausible case. The detail that separates answers is a bound: how many users, how much data, and for how long, if you had been wrong.
What to demonstrate
- Whether the reasoning you give was available at the time you decided rather than after the result came in, since a story whose deciding evidence arrived later describes an outcome and not a judgement
- Whether you can put units on the exposure (users, rows, minutes of degraded service) and whether the containment you chose actually bounded it: a canary bounds the request path it fronts, while a background job writing to a shared table reaches every user regardless of which version served their requests
- Whether the reversal path existed before you shipped or was improvised during the incident, and whether it restores state or only stops further damage
How to prepare
- For your three largest changes, write down the one thing you would have had to be wrong about for it to fail, and what your best estimate of it was on the day you shipped. If you never held an estimate, that is the gap the follow-up questions will find
- Write the undo procedure for one of those changes as it existed at the time, then mark which steps restore data and which only stop new damage. Turning a flag off or reverting a deploy ends the new writes; rows already written come back only from a copy you kept, and a dropped column comes back empty unless something outside the schema holds the values
- Rehearse one story from the decision point forward and stop before the outcome, then have someone ask what you would do next. If the story only works with the ending attached, it is an anecdote rather than a decision you can defend
PracHub editorial advice for the preparation topics above.
Upserting on the order or result identifier, so a correction overwrites the original row.
It makes the question 'what did the clinician see at 14:02' unanswerable, which is exactly what an incident review or a legal hold asks. It also leaves downstream consumers that already acted on the preliminary value with no correction event to react to, because the state transition was collapsed into a single mutated row and never emitted.
Treating a medical record number or a member ID as a globally unique key and joining on it directly.
These identifiers are unique only within the authority that issued them. Two facilities in one network routinely have the same medical record number for different people, and member IDs get reissued when someone changes plans. Joining on the bare value merges two patients' records, which is the most damaging failure available in this domain, and it passes every test written against a single-facility fixture because the collision only appears once a second source is connected.
Saying 'eventually consistent' without naming the anomaly a user would see
Describe the concrete symptom you are choosing to accept: the author reloads and their own comment is missing for two seconds, or two devices show different balances for a minute. The class of consistency model is a technical label; the tolerable anomaly is the actual product decision.
Abandoning working code to chase the optimal solution
Get the straightforward version correct, state its complexity, and only then optimise, keeping the working version until the faster one passes the same cases. A correct quadratic solution with a stated path to linear beats a half-written optimal one that never ran.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Walk me through how you would debug a performance bottleneck in a prod…
Walk me through how you would debug a performance bottleneck in a production application.
Approach
- Restate the input: its shape, its size, and what is guaranteed about it.
- Walk one small example through your approach before writing the whole thing.
- Name the brute-force solution and its complexity before improving on it.
Follow-up
- Which test case would catch an off-by-one here?
- What is the worst case, and how likely is it on real data?
How do you ensure data security and compliance (such as HIPAA standard…
How do you ensure data security and compliance (such as HIPAA standards) when designing APIs that handle sensitive patient information?
Approach
- Walk one small example through your approach before writing the whole thing.
- Name the brute-force solution and its complexity before improving on it.
- State the target complexity and say which constraint rules the naive version out.
Follow-up
- Which test case would catch an off-by-one here?
- What is the worst case, and how likely is it on real data?
Group transfer-chained encounters into episodes with out-of-order arrival
You have up to 5 million encounter rows for a reporting month, in arbitrary order: encounter_id, enterprise_person_id, facility_id, encounter_class, admit_ts, discharge_ts (nullable), status, prior_encounter_id (nullable, set on transfer), source_event_ts. prior_encounter_id may reference a row that arrives later or never arrives. Group rows into episodes, one per maximal transfer chain, returning episode root, person, earliest admit_ts, latest discharge_ts (null if any member is still open) and member count. Report cycles instead of following them. Target O(N) expected time.
Approach
- Recognise the shape: every row has at most one prior_encounter_id, so this is a forest of parent pointers, not a general graph. A hash map from encounter_id to row is the whole index you need, and the work is O(N) expected rather than anything sort-shaped.
- Resolve roots by iterative pointer-following with write-back memoisation: from each node walk prior_encounter_id until you reach a node with no parent, a parent absent from the input, or a node whose root is already known, then stamp the discovered root onto every node on that path. Each node is finalised once, so total work stays O(N) even when one chain is very long.
- Do not recurse. A malformed feed can produce a chain tens of thousands deep, and a recursive resolver dies on stack depth taking the whole batch with it. An explicit loop with write-back is both faster and bounded.
- Detect cycles with three-state marking: unvisited, on the current path, finalised. On reaching a node that is on the current path, emit a data-quality record naming every encounter_id in the cycle and exclude that component from the episode output. Choosing an arbitrary root instead would produce an episode that looks plausible and will be believed.
- Keep a dangling prior_encounter_id distinct from a genuine root. It means the chain starts outside the window or has not arrived, so flag the episode left-truncated and carry that flag into the output: an episode with an unknown start cannot be used for length of stay or readmission counting without biasing both downward.
- Compute aggregates during the same pass that assigns roots: minimum admit_ts, maximum discharge_ts with null propagation so any open member leaves the episode open, and exclude cancelled and entered_in_error members from the aggregates while keeping them retrievable.
Worked solution 25 min
- Load all rows into a hash map keyed by encounter_id, with a colour array alongside.
- For each unfinalised node, walk parents onto an explicit path stack until you hit a root, a dangling reference, a finalised node or an on-path node.
- Stamp the resolved root, or the cycle marker, back onto every node on the path.
- Fold each node into its root's aggregate: min admit_ts, max discharge_ts with null propagation, member count, left-truncated flag.
- Emit episodes and the cycle records as two separate outputs.
Follow-up
- Two encounters name each other as prior_encounter_id. What does your detector emit, and who needs to see it?
- One episode spans two facilities whose clocks differ by four seconds. Do you order on source_event_ts or admit_ts, and what does the choice cost?
- A chain is cut in half by the month boundary. What does this month's report say, and how do you make next month's agree with it?
Add a backfilled NOT NULL column to a live encounter table
encounter holds 900 million rows and is written continuously by the admit/discharge/transfer consumer. You must add admission_source_code text, backfill it per row from a lookup on the source message archive, constrain it to a fixed code set, make it NOT NULL, and add an index on (facility_id, admit_ts) — with no write outage and no statement blocked longer than one second. Give the ordered steps, the lock level each takes, why the obvious single ALTER is unavailable here, and the session setting that stops a DDL statement from stalling everything behind it.
Approach
- Say why the shortcut does not apply. Since PostgreSQL 11, ADD COLUMN with a constant default is metadata-only and cheap, but this value is derived per row from an archive lookup, so there is no constant to store and the table must be written row by row regardless.
- Add the column nullable first: catalog-only, ACCESS EXCLUSIVE for microseconds. The danger is not the statement's duration but the lock queue — a pending ACCESS EXCLUSIVE blocks every reader that arrives behind it, so a long-running report turns a microsecond DDL into a multi-minute outage. SET lock_timeout = '1s' before the ALTER and retry on failure; that converts a potential outage into a retried statement.
- Deploy dual-write before backfilling. The consumer must populate the column on every insert and transfer from that moment, or the backfill chases a moving tail forever.
- Backfill in bounded batches over primary-key ranges, a few thousand rows per transaction with a pause between, restricted to WHERE admission_source_code IS NULL so it is resumable and idempotent. One giant UPDATE holds a transaction open for hours, pins the xmin horizon so autovacuum cannot clean anything, and bloats the table by a full row version per updated row.
- Add the value constraint as CHECK (...) NOT VALID first — ACCESS EXCLUSIVE, no scan — then ALTER TABLE ... VALIDATE CONSTRAINT, which takes only SHARE UPDATE EXCLUSIVE and scans without blocking reads or writes.
- Reach NOT NULL without a blocking scan: add CHECK (admission_source_code IS NOT NULL) NOT VALID, VALIDATE it, then SET NOT NULL, which on PostgreSQL 12 and later uses the validated constraint as proof and skips the full scan; drop the now-redundant CHECK afterwards. Build the index with CREATE INDEX CONCURRENTLY, which takes SHARE UPDATE EXCLUSIVE, makes two passes plus a wait for concurrent transactions, cannot run inside a transaction block, and leaves an INVALID index behind on failure that must be dropped and rebuilt.
Follow-up
- CREATE INDEX CONCURRENTLY has been running for four hours and you need to cancel it. What state is the index left in, how do you detect it, and what do you run next?
- Your backfill is at 40% and the replica is 90 seconds behind. What do you change, and what do you measure to decide the new batch size?
- The code set gains a value two weeks later. Does your CHECK constraint or a lookup table with a foreign key make that change cheaper, and what does each cost on the write path?
Return the displayable version of each analyte with a window function
observation_result is append-only: observation_id, enterprise_person_id, encounter_id, accession_id, placer_order_id, filler_order_id, loinc_code, value_numeric, unit_ucum, result_status in ('registered','preliminary','final','amended','corrected','entered_in_error'), collected_ts, issued_ts, version, supersedes_observation_id. A chart view needs the currently displayable value for each (accession_id, loinc_code) for one person over the last 90 days, excluding analytes whose newest version is entered_in_error but keeping analytes whose older versions were. Write the query, give the index that serves it, and say honestly which part of the plan the index cannot remove.
Approach
- Recognise the shape: this is a top-1-per-group over a version chain, not a filter. Pick the newest version first, then apply the status predicate to the winner — the order is the whole exercise.
- Use DISTINCT ON (accession_id, loinc_code) with ORDER BY accession_id, loinc_code, version DESC, observation_id DESC on PostgreSQL; the ORDER BY must start with the DISTINCT ON expressions or it is a syntax error. Use ROW_NUMBER() OVER (PARTITION BY ... ORDER BY version DESC, observation_id DESC) with an outer WHERE rn = 1 if you need portability, since a window function cannot be referenced in the WHERE of its own query level.
- Push only person and time into the inner WHERE. Pushing result_status <> 'entered_in_error' inside promotes a retracted result's predecessor back onto the chart, which is the bug the prompt is testing for.
- Index (enterprise_person_id, collected_ts DESC) so the 90-day slice is an index range scan rather than a scan of the person's whole history.
- Be straight about the limit: because collected_ts carries a range predicate, no b-tree index can also deliver rows pre-ordered by (accession_id, loinc_code, version DESC), so the sort stays in the plan. That is acceptable because one person's 90-day slice is hundreds to low thousands of rows, and it is a quicksort in work_mem rather than an external merge — verify that rather than assume it.
Worked solution 30 min
- Write the DISTINCT ON form with the full sort key and only person plus time in the inner WHERE.
- Wrap it in an outer query and apply the entered_in_error filter there, so it can only eliminate winners.
- Write the ROW_NUMBER equivalent and confirm both return identical rows against a fixture containing one preliminary, one final, one corrected and one retracted chain.
- Add the index, then read EXPLAIN (ANALYZE, BUFFERS) and confirm an index scan feeding a Sort node, with Sort Method reported as quicksort and no Disk usage.
- Construct the adversarial fixture: an analyte whose version 1 is entered_in_error and whose version 2 is final. It must appear in the output; an analyte whose version 2 is entered_in_error must not.
Follow-up
- A correction arrives carrying the same filler_order_id and the same version number as the row it corrects. What breaks, and what do you add to the sort key?
- The same query for a cohort of 50,000 people instead of one person: does DISTINCT ON still hold up, and at what point do you materialise a latest-version table with a partial index?
- How would you show the clinician that the displayed value was amended, given the chart query returns only the winner?
Describe a scenario where you had to implement a DevOps pipeline. How …
Describe a scenario where you had to implement a DevOps pipeline. How did you ensure continuous integration and deployment security?
Approach
- Fix the scope first: who calls this, how often, and what they do when it fails.
- Name the failure you are designing for, then the recovery path.
- Choose a partition key and say what query it makes expensive.
Follow-up
- What would you drop to keep the system up under load?
- What breaks first when traffic grows ten times?
What are ICD-10 CM and PCS codes, and how do they function within a he…
What are ICD-10 CM and PCS codes, and how do they function within a healthcare billing or clinical software system?
Approach
- Choose a partition key and say what query it makes expensive.
- Name the read and write paths separately; they rarely have the same bottleneck.
- Fix the scope first: who calls this, how often, and what they do when it fails.
Follow-up
- How does this behave when that dependency is down for an hour?
- What would you drop to keep the system up under load?
How would you design a high-throughput system to ingest and process re…
How would you design a high-throughput system to ingest and process real-time health data from wearable devices?
Approach
- Name the failure you are designing for, then the recovery path.
- Name the read and write paths separately; they rarely have the same bottleneck.
- Choose a partition key and say what query it makes expensive.
Follow-up
- What breaks first when traffic grows ten times?
- What would you drop to keep the system up under load?
Set rate limits across registration, batch ingest and bulk export
Three callers share the identity resolution and longitudinal record APIs: a registration desk making 1-3x10^3 interactive match calls a second against a 200ms budget, a nightly ingest replaying millions of messages in a burst, and partner bulk exports walking whole panels. One global limit either starves registration during the batch window or stretches the batch past morning. Define the limiting scheme - the key, the algorithm, the reserved capacity, the response when a caller is limited - and state precisely what each of the three callers does when it is limited.
Approach
- Reject a single global limit, then reject per-IP: the batch runs from a handful of hosts and the desk sits behind a shared egress, so IP is simultaneously too coarse and too fine. Key on the authenticated principal plus a traffic class bound to the credential, never a class the caller declares in a header it controls.
- Pick the algorithm from the traffic shape. A fixed-window counter admits up to twice the limit across a window boundary, which is exactly the top-of-hour moment the desk bursts; a token bucket with a burst allowance, or a sliding-window counter, does not. The desk needs burst tolerance; the batch needs a steady ceiling.
- Give the classes different treatment rather than different numbers of the same thing: a reserved floor of capacity the interactive class cannot be pushed below, the remainder shared by batch and export, and the batch class shedding first under pressure. This is a scheduling decision - the class that can wait is the one that waits.
- Answer with 429 plus Retry-After in delta-seconds and the limit, remaining and reset headers, so the caller need not invent a backoff. Keep 429 distinct from 503: 429 means try later and the request did no work, 503 means the service is unhealthy. Conflating them makes correct client behaviour impossible to write.
- Write each caller's behaviour, because a limit without a documented client response only relocates the failure. The desk degrades to deterministic-only matching and flags the registration for review rather than blocking a patient; the ingest applies backpressure to its consumer, since the messages are durable and delay is free while loss is not; the export sleeps Retry-After and resumes from its cursor. Then state where the counter lives: a shared store with one atomic increment-and-expire per decision, sized so the limiter is not the bottleneck at tens of thousands of decisions a second, and with a stated fail-open or fail-closed behaviour when it is unreachable.
Worked solution 25 min
- Write the three traffic profiles as numbers: request rate, burst shape, latency budget, and what each caller can tolerate on refusal.
- Choose the key, show where the class comes from on the credential, and write the check that stops a caller self-declaring.
- Work the fixed-window boundary arithmetic that admits twice the limit, then the token-bucket or sliding-window version that does not.
- Write the full 429 response with every header, and the exact backoff each of the three callers implements.
- Write the reserved-capacity rule and trace all three classes at 150 percent of total demand.
Follow-up
- The batch runs under the same credential as an interactive tool. How do you separate them?
- The limiter's shared store goes down. Does traffic fail open or closed, and what is the argument for your choice here specifically?
- One partner's export is now 40 percent of read load while staying inside its limit. What do you change?
Morning eligibility spike saturates the outbound payer pool
The eligibility service caches answers keyed by (person, plan, service date) with a 24-hour TTL. Every weekday at 07:00 local, for roughly 90 seconds, p99 goes from 45ms to the 6s outbound timeout, the connection pool to external payers saturates, and payers begin returning 429. Cache hit rate dips to 55%. Payer-reported latency for the requests they do serve is unchanged. Name the mechanism, the evidence that separates it from a slow dependency, and the fix — including what you would not do to the TTL.
Approach
- Separate many distinct keys missing from one key missed many times concurrently. Log (key, outcome) and, over a ten-second window inside the spike, compute misses divided by distinct keys missed. A ratio near one is a cold cache and a capacity problem; a ratio of twenty or fifty is a stampede, where concurrent requests for the same key each perform their own origin fetch.
- Confirm the expiry is synchronised. Histogram the remaining TTL across live keys: a single absolute 24-hour TTL set during yesterday's 07:00 peak produces a spike in that histogram exactly 24 hours later, which is why the incident is punctual rather than proportional to load.
- Locate the saturation. Payer-side latency flat while client-observed latency climbs to the timeout means the queueing is on your side. Measure in-flight outbound requests against the pool maximum and the time spent waiting to acquire a connection. That is the evidence that rules out a slow dependency, and the 429s are a consequence of your fan-out rather than its cause.
- Remove the duplication first with single-flight coalescing keyed on (enterprise_person_id, plan_id, service_date): concurrent misses for one key make one origin call and the rest wait on its result. This alone collapses the misses-to-keys ratio and is independent of every other change.
- Then fix the synchronisation and bound the exposure. Use a TTL of minutes with plus or minus 20% jitter so expiries spread, plus a bounded stale-while-revalidate. State the worst case explicitly — a 300s TTL with a 10s stale window is at most 310s of staleness — and return the computed-as-of timestamp to the caller, because coverage terminates retroactively and a long TTL is a correctness decision, not a performance knob.
- Bulkhead the outbound path per payer so one slow or rate-limiting payer cannot consume the concurrency budget of requests destined for the others, and define what the service returns when it cannot reach the origin at all.
Follow-up
- What does the service return when the payer is unreachable and the cache is empty? Does registration proceed with an unverified flag or block, and who owns that decision?
- How would you warm the cache before 07:00 without the warm-up itself becoming the stampede — and which keys are even warmable given the service date is part of the key?
Roughly ninety minutes on weeknights with one longer weekend block. The plan cuts scope rather than compressing everything, on the assumption that one thing finished per night beats four half-started.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Fix the scope and take a cold baseline
- Read the role description and write the three things the loop will almost certainly test, then write an explicit not-doing list and keep it visible all week.
- Take one twenty-five-minute coding problem and one fifteen-minute design prompt cold, and write the single sentence naming what blocked each, because those two sentences decide where the remaining evenings go.
- Set the week's rule: one thing finished every night, including the night you only have forty minutes.
Deliverable: A one-page scope with a not-doing list and two cold attempts, each carrying one sentence on what blocked it.
Practice prompt ↗Practice prompt ↗Worked solution ↗02One pattern, written three times from blank
- Choose the single pattern most likely to appear in your loop and write it three times from an empty file rather than editing the previous attempt.
- On the third pass, write the invariant as a comment before the loop body and the complexity before the first line of code.
- Stop at ninety minutes even if the third version is imperfect, and write the one thing you would fix given another hour.
Deliverable: Three independent implementations of the same pattern plus a note on what changed between them.
Practice prompt ↗Practice prompt ↗03One design, only to the depth you can defend
- Take one system shape and go only as far as requirements, interface and data model, refusing to draw a box you could not survive a follow-up about.
- Attach one number to each non-functional requirement, deriving it rather than asserting it, and write the assumption the number rests on.
- Write the one tradeoff you are choosing against and the observation that would make you reverse it.
Deliverable: One design at interface-and-schema depth with derived numbers and one written reversible tradeoff.
Practice prompt ↗Practice prompt ↗04Only the fundamentals you will have to defend
- Write, in under two hundred words each, the answers to the two questions that follow almost any implementation: why this structure and not the obvious alternative, and what happens to this code at a hundred times the input.
- Write what an index actually costs: faster lookups on the indexed columns against a write that now maintains a second structure, plus the cases where the planner declines to use it anyway, low selectivity, or a predicate wrapping the column in a function.
- Delete any answer you cannot deliver aloud in under a minute, since an answer that needs reading is not an answer you have.
Deliverable: Three written answers, each under two hundred words and each timed aloud.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Your own work, timed
- Write a ninety-second and a four-minute version of your main project and time both aloud rather than reading them.
- Prepare the two follow-ups that always come: what you would do differently, and how you knew it worked.
- Put one number in the first sentence and be ready to say exactly where it came from and what it excludes.
Deliverable: Two timed narratives with one defensible number in the opening line.
Practice prompt ↗Practice prompt ↗06The one full rehearsal, in the weekend block
- Run a sixty-minute mock covering a coding round and a design round in one sitting with no break, because sustained attention is the thing evenings have not trained.
- Immediately afterwards, and before hearing any feedback, write the three moments you lost the thread.
- Spend the rest of the block only on those three moments, and on nothing you merely feel shaky about.
Deliverable: Mock notes naming three failure moments with a specific fix written under each.
Practice prompt ↗Practice prompt ↗07Taper
- Write the twenty-minute warm-up you will actually do on the morning: one problem you can already solve from a blank file, one design you can narrate, and nothing you have never seen.
- Re-read only your own notes from this week and open no new material.
- Write the logistics down: the editor or shared document you will be working in, whether execution and lookups are permitted, and the sentence you will use when you do not know something.
Deliverable: A one-page card holding the design structure, the project numbers, and the logistics.
Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
A slipped date is only a bad story if you sat on it. What matters is what you believed when you gave the number, the signal that told you it was wrong, how many days passed before you said so, and what you cut rather than asking for more time. Scope you defended counts as much as scope you dropped.
Tell me about a time you made a mistake on a project. What did you lea…
Tell me about a time you made a mistake on a project. What did you learn, and how did you rectify it?
Approach
- Give the blast radius: what could have broken, and what you measured.
- Name the disagreement and how you resolved it with evidence.
- State the situation in two sentences and spend the rest on the reasoning.
Follow-up
- What would you do differently if you ran that again?
- What did you decide not to do, and why?
What are the core principles of RESTful API design, and how do you han…
What are the core principles of RESTful API design, and how do you handle error states in a microservices architecture?
Approach
- Pick a story where you made the decision, not one where you watched it.
- Give the blast radius: what could have broken, and what you measured.
- Name the disagreement and how you resolved it with evidence.
Follow-up
- What would you do differently if you ran that again?
- How did you know your change caused the improvement?
Ship an eligibility cache you already know is too stale
You shipped the eligibility service against a fixed go-live with a 24-hour cache keyed on person and plan, no service date in the key and no computed-as-of field in the response, because the real-time payer inquiry path was two sprints behind. Enrolment files routinely terminate coverage retroactively — a file received on the fifth can terminate coverage effective the first. Describe shipping a compromise of this shape: what you wrote down, who accepted the risk, the bound you could honestly state, the measurement that sized the debt in money, and whether it was ever paid.
Approach
- The probe is whether you can state a compromise as a bounded risk rather than as a known limitation. Start by separating two staleness sources, because they have different fixes: cache staleness is bounded by the TTL, while business-time staleness is not bounded by anything in your system, since the source itself asserts changes retroactively.
- Follow that distinction to its uncomfortable conclusion — shortening the TTL to five minutes leaves the second unchanged. A zero-TTL cache still serves an answer that was false when computed, so anyone who proposes TTL tuning as the fix has misread which of the two is biting.
- Write the debt as a sentence with numbers: the service can answer 'covered' for a person whose coverage terminated up to the enrolment cadence plus its retroactive window earlier, with up to 24 hours of cache on top. That sentence is what a reviewer can accept or refuse; a ticket title is not.
- Name the mitigation that costs days rather than sprints: put service_date in the cache key and computed_as_of in the response body so consumers can see an answer's age, and re-check eligibility at claim submission rather than trusting the encounter-time answer.
- Give the measurement that sizes it in money — the weekly count and dollar value of claim_line rows denied for terminated coverage where the answer served at encounter time said active, filtered to the surviving claim version. That number is what buys the sprint later; without it the fix competes against features on opinion.
- Be specific about the acceptance. Name the artefact the decision lives in and the person with budget authority who signed it, because a compromise nobody accountable acknowledged is a decision you made alone and later described as a trade-off.
Follow-up
- The payer inquiry path ships and the cache stays. Which of the two staleness bounds has actually moved?
- A service was delivered against a stale 'covered' answer. Who absorbs the cost, and does your design record enough to answer that?
- Registration wants a 50ms p99 on this call. What do you give up, and how do you say that to them in one sentence?
- 01
Tell me about a time you made a mistake on a project. What did you learn, and how did you rectify it?
- 02
What are the core principles of RESTful API design, and how do you handle error states in a microservices architecture?
- 03
You shipped the eligibility service against a fixed go-live with a 24-hour cache keyed on person and plan, no service date in the key and no computed-as-of field in the response, because the real-time payer inquiry path was two sprints behind. Enrolment files routinely terminate coverage retroactively — a file received on the fifth can terminate coverage effective the first. Describe shipping a compromise of this shape: what you wrote down, who accepted the risk, the bound you could honestly state, the measurement that sized the debt in money, and whether it was ever paid.
Is this an official Optum Health Services interview guide?
No. It is PracHub's own research and practice material for the Software Engineer role at Optum Health Services. Rounds and questions reflect what candidates have reported, not a process Optum Health Services has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How technical is the HireVue stage?
The HireVue assessment typically consists of a mix of behavioral questions and a basic coding or technical challenge. It is designed to filter for foundational coding ability and communication skills before you speak with a live engineer.
PracHub interview research ↗What programming languages are preferred at Optum?
Optum uses a variety of technologies, but Java is highly prevalent across many of their backend systems. Demonstrating strong object-oriented design principles in Java or C# is highly beneficial.
PracHub interview research ↗How long does the entire hiring process take?
The timeline can vary. While some candidates receive offers within a few weeks, others report a more drawn-out process spanning several months. It is important to stay patient and follow up periodically.
PracHub interview research ↗Is healthcare domain knowledge required to get hired?
No, general software engineering excellence is the primary focus. However, showing an interest in the healthcare industry and having a basic understanding of compliance and data security will give you a distinct advantage.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-22 - 02PracHub Software Engineer practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-22 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-22