Software Engineers at Visa, as the role is described, build and run payment infrastructure: transaction processing, real-time payment routing, fraud prevention and checkout. The teams named for the role include Processing & Products Technology (PPT), Developer Platform and Core Infrastructure. The work covers distributed backend microservices, event-driven architectures, APIs and developer-facing tooling, and Java, Kafka, Docker and Kubernetes are the listed technologies.
The reported design prompts center on retries, idempotency and zero data loss. They include a global payment system with a state machine, retries and zero data loss, a payment gateway API that handles concurrency, rate limiting, caching and idempotency, a distributed scheduler, and an asynchronous pipeline that ingests transaction events into a persistent store, with Kafka named for a question about moving monolithic payment routing to an event-driven architecture.
Reported questions fall into four groups. Data structures and algorithms: longest increasing subsequence, an LRU cache, linked-list cycle detection, longest substring without repeating characters, and a structure that tracks string counts with O(1) max and min. System design: the payment system, payment gateway API, distributed scheduler and transaction-event ingest prompts from the previous paragraph. CS fundamentals: JVM internals and garbage collection, a top-three-salaries-per-department window-function query, concurrency and deadlocks, SQL versus NoSQL, and TCP/IP, REST, Docker and Kubernetes. Behavioral: project walkthroughs, deadlines, disagreements and code quality under pressure.
The practical takeaway is to prepare fundamentals as a category of their own. Candidates report being asked short technical questions alongside coding, so a clear two-sentence answer on garbage collection or on how ROW_NUMBER and DENSE_RANK treat ties is worth rehearsing just as much as an algorithm.
Technical Screen
reportedCandidates report that this first stage is an online coding assessment on a platform such as HackerRank or CodeSignal. It has three to four problems, from easy to medium-hard, covering arrays, hash maps, sliding windows, two pointers, dynamic programming and trees. Nobody is there to answer questions, so the worked examples in each prompt are the whole specification. Read them for the exact output format and for what an empty result should look like. Solve the easier problems first to secure them, then put the remaining time into the hardest one.
What to demonstrate
- Whether each submission handles the cases the examples leave out: empty input, a single element, duplicates, and values near the integer limits
- Whether you split the time across problems deliberately instead of sinking it into the first hard one
- Whether you spot the standard pattern (sliding window, two pointers, prefix sums, dynamic programming) fast enough to leave time for testing
How to prepare
- Do two timed sets of three to four problems in a plain browser editor with autocomplete off, using the easy-to-medium-hard mix candidates describe
- Before each submission, run a small harness that feeds in the prompt's examples plus an empty case and a single-element case, and prints expected output next to actual output
- Drill the reported patterns: longest substring without repeating characters with a sliding window, longest increasing subsequence in O(n log n) with a sorted tails array, and Floyd's cycle detection that returns the node where the cycle starts
Technical Phone Screen
reportedCandidates describe a phone or video interview with an engineering manager or a senior engineer that includes coding exercises and technical questions. Prepare to code while talking, and be ready for a technical question drawn from the reported fundamentals categories, such as Java internals, SQL, concurrency or networking, since any of them may come up. Because a person is watching, partial progress is visible and counts. Explain your approach, state the complexity before you write, and walk a small example through the code out loud before you say you are done.
What to demonstrate
- Whether what you say matches the code you are writing, and whether a hint changes your direction
- Whether you can answer a technical question in two or three precise sentences with an example, then go deeper when asked
- Whether you clarify inputs and edge cases such as null, empty collections and overflow before coding
How to prepare
- Run a mock where the interviewer switches from a coding problem to a fundamentals question partway through, for example concurrency versus asynchrony, or what happens when you enter a URL
- Prepare a short walkthrough of your most relevant project in case the call opens with your background, naming what you built rather than internal codenames
- Implement the reported LRU cache out loud and explain why a hash map plus a doubly linked list gives O(1) get and put
Virtual Onsite Interview
reportedNobody in the room with you decides this. Interviewers typically write their rounds up separately, often before seeing anyone else's, and the outcome is settled later from those write-ups. A split panel gets resolved by whichever note carries specific evidence, so what you want out of each room is one concrete thing that person could write down: a bug you caught yourself, a trade-off you named, a decision you owned. The rest is arithmetic. The project you describe in a behavioural conversation is often the same system you sketched an hour earlier, and the two accounts have to agree.
What to demonstrate
- Whether the scale, team size and timeline you attach to a project hold steady when that project resurfaces in a different round
- Whether each interviewer leaves with a specific thing to cite rather than a general impression of competence
- Whether a trade-off you defended in one round survives a challenge in another, instead of being quietly swapped for the answer the new interviewer seemed to want
- Whether a question you have already answered earlier in the day gets the same answer at the same depth, without visible impatience
How to prepare
- Write a one-page sheet per project fixing the figures you will quote — request volume, data size, team size, elapsed time, what broke — and say them aloud from the sheet until they come out identical every time
- For each round on the schedule, decide in advance the one sentence you want in that person's notes, then check in a mock that you said it outright instead of leaving it to be inferred
- Have someone ask you the same project question twice, an hour apart, and diff the two answers for numbers that moved or a trade-off that reversed
Coding and Algorithms Round
reportedMost of the time lost in this format is not lost to thinking. It goes to a standard-library call you half-remember, an off-by-one in a loop bound, and a debugging loop that mutates code at random until something passes. When output is wrong, stop re-reading the whole function: take the smallest input that reproduces it and walk the state through by hand, printing intermediates if the environment allows. Guessing at a fix without a failing case you understand is how a five-minute bug becomes twenty, and the clock does not pause while you do it.
What to demonstrate
- Whether you reach the right structure without a detour, and can write it from memory rather than only recall that one exists
- Whether overflow is considered where the language has fixed-width integers, since a signed 32-bit value stops at 2,147,483,647 and then wraps in Java, is undefined behaviour in C++, and does not arise in Python, whose integers grow instead
- Whether recursion depth is treated as a constraint on large inputs, given that CPython's default limit is 1000 frames and a deep recursion can exhaust the stack in any language where an iterative version would not
- Whether a failing case is isolated and explained before any edit is made to the code
How to prepare
- From an empty file and with no references open, implement the pieces you lean on most: a heap push and pop, an iterative DFS with an explicit stack, and a binary search whose midpoint is written lo + (hi - lo) / 2, which avoids the overflow that (lo + hi) / 2 can hit in a fixed-width integer type
- Time yourself on the ten library calls you look up most, such as sorting with a custom comparator, splitting and joining strings, and finding the next key at or above a value in an ordered map, until the lookup is gone
- Take a solution you know is broken and, before touching it, write one sentence naming the input, the expected value and the actual value. Repeat until you do it without deciding to.
System Design Round
reportedCandidates report that this round depends on seniority. Mid-level and senior candidates get full system design. Entry-level candidates are more often asked low-level design, database schema questions or the architecture of a past project. The reported prompts center on payments: a global payment system with a state machine, retries and zero data loss; a payment gateway API with concurrency, rate limiting, caching and idempotency; a distributed scheduler; and a pipeline that ingests transaction events. Whatever the prompt, explain early how a retried or duplicated request is made safe, because each of these problems depends on it.
What to demonstrate
- Whether the payment lifecycle is modelled as explicit states with defined transitions, and what a retry does in each state
- Whether idempotency is enforced by a key and a unique constraint rather than by checking first and then inserting
- Whether you handle partial failure, such as a process dying between calling a processor and recording the result, or a consumer receiving events twice or out of order
- Whether you argue trade-offs such as SQL versus NoSQL for transactional data, or REST versus gRPC between services, from the requirements
How to prepare
- Work through this guide's webhook-intake design exercise from start to finish, then redo it with Kafka as the transport and say where the deduplication key lives
- Sketch the payment gateway API: endpoints, the idempotency header, the stored response, where the rate limiter sits, and what a duplicate request gets back
- For entry-level loops, prepare a schema walkthrough of a past project and one low-level design, such as an LRU cache or a scheduler interface
Behavioral and Cultural Fit Round
reportedMany of these questions are about something that went wrong, and the grading sits mostly in the hours after you knew. Who found out first, whether that was you or an alert or a user, how long it took you to say it out loud, and whether the people who needed the news got it while they could still act on it. Engineers under-tell this part because it feels like confessing. The pattern it is looking for is the opposite: the quiet fix, an incident absorbed without telling anyone, after which nothing changed and the same failure is still available.
What to demonstrate
- How the problem was found, and whether that route was one you had built or one that happened to you, since a user reporting it first means your instrumentation did not cover that failure
- Whether time-to-detect and time-to-tell are separate numbers in your account and whether you know both, because a fast fix that nobody heard about until the retro is a different answer from a slow one that was announced immediately
- Whether the resolution left something durable behind, a check that fires or a default that changed, rather than depending on people remembering to be careful
- Whether you can say what the failure cost without either inflating it or waving it away
How to prepare
- Reconstruct one incident you were part of as a timeline with clock times: first bad request, first signal, first person who knew, first message outside the team, mitigation, permanent fix. The gaps between those entries are what gets asked about
- Look up the configuration of the signal that caught it, including its evaluation window and threshold. An alert defined on a five-minute aggregate cannot fire until the condition holds across that window, which puts a floor under time-to-detect that has nothing to do with how severe the failure was. Be able to say what that floor was and whether anyone had chosen it deliberately
- Prepare one story where you escalated early and the severity turned out to be smaller than you thought, including what it cost the people you pulled in. Without it, every answer you give about raising alarms is unfalsifiable
12 candidate reports. Individual accounts describe a particular role and hiring cycle.
Visa Software Engineer Interview Experience — Three-Round Onsite With Longest Increasing Path, a Greedy Problem, and an Inventory Service Design
Round 1: the original LeetCode problem, longest-increasing-path-in-a-matrix. Debugging wasn't very convenient. I also had to write my own test cases and my own main, so that held me up for a long time and I didn't have enough time. At the end I brought up the optimal solution, using memoization to record results, but I was out of time. > Given an m x n integers matrix, return the length of the lo…
Read full experienceVisa New Grad Software Engineer Interview Experience — Three-Question OA Ending in a Shortest-Prefix Permutation Problem
Visa's latest OA questions. Three questions in total: the first two were fundamental questions, and the third was an intermediate question. Question 1 I don't remember the problem statement clearly. It was fairly basic; I remember you needed a hashmap optimization to avoid TLE. Question 2 Given an integer array (a vector). For each integer in the array, find the minimum k such that: It is a posit…
Read full experienceVisa Software Engineer interview: short Java and algorithms round
My Visa interview lasted about 30 minutes and moved quickly. I began with a coding question and had around five minutes to produce a solution. That was followed by a logical puzzle. The interviewer then tested object-oriented programming and Java fundamentals, including implementing inheritance. The last part was a binary-search problem: I started with a brute-force approach, explained how I woul…
Read full experienceSoftware Engineer interview at Visa
Visa began in a standard way with an OA and then technical rounds. In CodeSignal Round 1, I had a medium array problem, a deep resume discussion, Docker and Kubernetes questions, and basic database-query material. Round 2 was more coding-heavy and asked me to implement an LRU cache. I got working LRU-cache code, but the interview ended around the 50-minute mark, leaving expected material uncovere…
Read full experienceVisa Software Engineer interview experience
After finishing the OA, I was selected for the next step. The interview combined a resume walkthrough with coding and fundamentals. They focused on my projects and hackathons, then asked at least one DSA question. The coding question was cycle detection in a linked-list-like structure, requiring me to reason through how to identify a loop. We also discussed CS fundamentals, situational questions,…
Read full experiencePracHub editorial advice for the preparation topics above.
Spending the online assessment on the hardest problem first
Candidates report three to four problems, from easy to medium-hard, on HackerRank or CodeSignal. Read all of them first, solve the easier ones, then spend what is left on the hardest. Before each submission, run the code on an empty input, a single element and the largest value the constraints allow. An automated grader scores the hidden cases and gives no credit for what you meant to write.
Writing the LRU cache or the O(1) count structure from a half-remembered pattern
For the reported data-structure questions, naming the pattern is not the hard part. The pointer bookkeeping is: an LRU cache with O(1) get and put, and a structure that increments and decrements string counts and returns the max and min keys in O(1). Use sentinel head and tail nodes so that insert and remove never need a special case for an empty list. Before you type, say which map points to which node, and agree on the approach with the interviewer.
Designing a payment flow with retries but no idempotency key
The reported design prompts (a global payment system, a payment gateway API, an event-ingest pipeline) all depend on what happens when a request is retried. Name the payment states first. Then name the idempotency key that makes a retry safe and the unique constraint that enforces it. Finally, say what the caller sees if the process dies after the processor call but before the local commit. A retry policy without that key creates duplicate charges.
Treating Java and SQL fundamentals as warm-up trivia
Reported questions include JVM garbage collection, interface versus abstract class, deadlock prevention and a top-three-salaries-per-department query. For the query, name your ranking function and explain why. ROW_NUMBER with rn <= 3 returns at most three rows per department (fewer when a department has fewer than three employees) and breaks ties arbitrarily, while DENSE_RANK returns every employee whose salary is among the three highest distinct values. If you skip the tie behaviour, the interviewer has to ask. Prepare a two-sentence answer with one example for each fundamentals topic.
Describing the same project differently in the phone screen, design and behavioral rounds
The resume walkthrough, the tight-deadline prompt and the technical-disagreement prompt often draw on the same project. Before the loop, put its figures on one page: traffic, data size, team size, timeline and your own part. Quote them the same way every time, and say which decisions were yours and which were the team's.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Find the longest strictly increasing subsequence in an array using opt…
Find the longest strictly increasing subsequence in an array using optimal dynamic programming or binary search approaches.
Approach
- Restate the input: its shape, its size, and what is guaranteed about it.
- Choose the data structure from the access pattern, not from familiarity.
- Walk one small example through your approach before writing the whole thing.
Follow-up
- Which test case would catch an off-by-one here?
- What is the worst case, and how likely is it on real data?
Detect a cycle in a Linked List and return the node where the cycle be…
Detect a cycle in a Linked List and return the node where the cycle begins.
Approach
- Choose the data structure from the access pattern, not from familiarity.
- Walk one small example through your approach before writing the whole thing.
- State the target complexity and say which constraint rules the naive version out.
Follow-up
- How does this change if the input no longer fits in memory?
- What is the worst case, and how likely is it on real data?
Design a data structure that supports incrementing/decrementing string…
Design a data structure that supports incrementing/decrementing string counts and retrieving maximum/minimum count keys in (O(1)) time.
Approach
- Choose the data structure from the access pattern, not from familiarity.
- Restate the input: its shape, its size, and what is guaranteed about it.
- State the target complexity and say which constraint rules the naive version out.
Follow-up
- How does this change if the input no longer fits in memory?
- What is the worst case, and how likely is it on real data?
Implement an LRU Cache in (O(1)) time complexity for get and put opera…
Implement an LRU Cache in (O(1)) time complexity for get and put operations using a doubly linked list and hash map.
Approach
- State the target complexity and say which constraint rules the naive version out.
- Restate the input: its shape, its size, and what is guaranteed about it.
- Walk one small example through your approach before writing the whole thing.
Follow-up
- How does this change if the input no longer fits in memory?
- Which test case would catch an off-by-one here?
Answer as-of balance queries over an append-only entry log
Given 400 million ledger_entry rows (entry_id, account_id, direction, amount_minor, currency, business_date) and 2 million queries of (account_id, currency, as_of_date) asking for the balance at the end of that business date, produce every answer. The obvious solution — per query, sum that account's entries with business_date <= as_of_date — is correct. Say precisely why it will not finish, then give one that will, with time and space complexity. Corrections are posted as new entries carrying their own business_date.
Approach
- Cost the naive version in numbers before rejecting it. Spread uniformly over 20 million accounts, each query touches about 20 rows behind a per-account index and 2 million queries is 4e7 row touches — perfectly fine. The problem is skew: one pooled clearing or merchant settlement account holding 3e7 entries, taking 10% of the queries, is 6e12 row touches. Name the skew; 'n is large' is not the reason.
- The structural fact that buys a cheap answer: entries are append-only and never updated, so a prefix sum over an account's entries ordered by
(business_date, entry_id)is stable — nothing behind position i can change. No mutable-balance design offers that, and it is why the storage is worth paying for. - Offline sweep, when all queries are known up front: externally sort entries by
(account_id, currency, business_date, entry_id)and queries by(account_id, currency, as_of_date), then merge-walk both with a running sum, emitting each query's answer as the sweep passes its date. O((n + q) log(n + q)) dominated by the sort, O(1) beyond sort buffers, one sequential pass over each input instead of 2 million random seeks. - Online alternative: materialise end-of-day snapshots — one row per
(account_id, currency, business_date)that had activity, holding the cumulative total. A query becomes one index seek for the latest snapshot at or beforeas_of_date, O(log n) per query, over far fewer rows than n. Use snapshots when queries arrive singly and the sweep when they arrive as a batch. - Corrections are the subtlety: an entry posted today but dated back changes historical answers, so every snapshot for that account from that date forward is stale. Either keep a Fenwick tree over dates per account (O(log D) update and prefix query) or recompute that account's snapshots from the corrected date onward. Then be precise about what reproducibility means — yesterday's statement is reproducible as of a stated snapshot time, not identical forever.
- Bound the resources: int64 sums throughout, no float; 400 million rows at roughly 48 bytes of the columns you actually need is about 19 GB, so the sort is external and its fan-out is chosen from the sort buffer, not from the row count.
Worked solution 40 min
- Compute both costs explicitly: the uniform case at about 4e7 row touches, and the skewed case at about 6e12. Showing that arithmetic is the answer to 'why'.
- Implement the offline sweep on a 10-million-row, 50,000-query fixture, merging on
(account_id, currency, business_date, entry_id). - Implement the naive version as the reference answer and assert both agree on every fixture query.
- Add a correction entry dated 30 days back, re-run, and assert that exactly the queries with
as_of_dateon or after that date move, all by the same signed amount. - Measure rows touched and wall time for each at 10 million rows, then extrapolate to 400 million and state the assumption that makes the extrapolation valid — sequential I/O, no random seeks.
Follow-up
- One account holds 30% of all entries. What does the external sort do with it, and what would you do for that one key instead?
- Queries now arrive online at 500 per second. Which design survives, and what does keeping the other one warm cost?
- A correction lands with a
business_date90 days back. Which snapshots are now wrong, and how does a reader find out?
Make a charge endpoint safe under concurrent duplicate retries
idempotency_key holds id, scope, key, request_fingerprint (SHA-256 over the canonicalised body), status (in_progress, completed, failed), response_status, response_body, locked_at, completed_at, expires_at, created_at. Fifty identical create-payment requests carrying the same scope and key reach four application instances inside the same 20 ms. Give the DDL constraint and the exact statements the handler runs so that exactly one payment_intent is created and all fifty callers receive the same response body. State what you return when that key arrives with a different fingerprint, and what an arrival after expires_at means.
Approach
- Put the concurrency control in the schema: UNIQUE (scope, key). A SELECT-then-INSERT cannot work because both transactions can read nothing before either commits, so the check passes twice and the constraint then surfaces as an error on a payment that succeeded.
- Claim the key with INSERT ... ON CONFLICT (scope, key) DO NOTHING RETURNING id. A conflict returns zero rows rather than the existing row, so branch on rowcount: the winner proceeds, the loser reads the stored row.
- Keep that path on READ COMMITTED deliberately. The loser's follow-up SELECT takes a fresh statement snapshot and therefore sees the winner's committed row; under REPEATABLE READ the transaction snapshot predates that commit, the row stays invisible and the loser concludes the key does not exist.
- Split the work across two transactions because the processor call cannot sit inside one: commit the in_progress row with locked_at first so losers can see a claim, perform the effect, then write payment_intent plus status=completed with response_status and response_body in a single second transaction.
- Handle the crash window explicitly: a row stuck in_progress past its lease is an unknown outcome, not a failure, so the reaper queries the processor for that key before deciding. A loser that sees in_progress returns 409 and retries rather than repeating the effect.
- Compare request_fingerprint before replaying anything. Same key with a different body is 409, never the cached response, because replaying confirms a payment the caller did not request; and set expires_at beyond the client's and the processor's maximum retry horizon, since a replay after it is a genuinely new request.
Follow-up
- The handler dies after the processor call and before the local commit. What does the next retry with that key observe, and how does the system converge on exactly one charge?
- Does the downstream processor honour an idempotency key of its own? Who mints it, and what breaks if a fresh one is generated per attempt?
- How do you purge rows past expires_at without the delete contending with the insert path?
Version a loan schedule instead of soft-deleting posted instalments
loan_instalment is keyed by (loan_id, schedule_version, instalment_no) and carries due_date, principal_minor, interest_minor, fee_minor, paid_principal_minor, paid_interest_minor, status, days_past_due, effective_from (date) and superseded_at (timestamptz). A borrower defers two payments on 2026-03-14, and instalments 1 to 6 already have allocations posted against them. Model exactly what the deferral writes, and write the query that returns the schedule as the borrower saw it on an arbitrary date. Say why stamping the old rows with deleted_at, or updating them in place, fails an audit.
Approach
- Treat the deferral as an insert, not an edit: write a complete new schedule_version with effective_from = the deferral instant on 2026-03-14, and in the same transaction stamp superseded_at on every row of the outgoing version with that identical instant, so the two version windows are half-open and adjacent rather than overlapping. Instalments 1 to 6 are reproduced unchanged, because they are what the borrower was told and what was posted to the ledger.
- Write the as-of read as a version selection, not a row filter: WHERE loan_id = $1 AND effective_from <= $2 AND (superseded_at IS NULL OR superseded_at > $2), then assert that exactly one schedule_version comes back - COUNT(DISTINCT schedule_version) = 1, not one row - so an overlapping window raises rather than silently returning two interleaved schedules under one instalment_no.
- Fix the type mismatch before writing that predicate: effective_from is declared date and superseded_at timestamptz, so comparing them casts the date at the session TimeZone and two readers in different zones select different versions near midnight. Put both columns in one domain, timestamptz, keep the window half-open with effective_from inclusive and superseded_at exclusive, and resolve a bare as-of date to an instant once, in the loan's booking timezone, at the edge of the system.
- Say what deleted_at loses. It records that a row stopped being current but not what replaced it or from when, it leaves every downstream query obliged to remember deleted_at IS NULL, and one query that forgets double-counts the schedule. Versioning puts the same information in the primary key where it cannot be forgotten.
- Close the arithmetic: under the loan's stated day-count convention the new version must still sum to outstanding principal plus scheduled interest to the minor unit, with the per-period rounding residual placed in one named instalment, conventionally the last, rather than smeared across the tail.
- Index (loan_id, effective_from DESC) for the as-of lookup, keep superseded versions online rather than archiving them, and enforce with a trigger that no UPDATE touches a row whose paid_principal_minor or paid_interest_minor is non-zero.
Worked solution 25 min
- Insert a 12-instalment version 1 with effective_from at origination, then allocate payments against instalments 1 to 6.
- Apply the deferral: insert version 2 effective at the 2026-03-14 deferral instant, reproducing instalments 1 to 6 byte for byte and re-amortising 7 to 12, and set superseded_at on all twelve rows of version 1 to that same instant in the same transaction.
- Write the as-of query with the version-selection predicate and run it for instants resolved from 2026-03-01 and 2026-03-20 in the loan's booking zone, and for the deferral instant itself.
- Compare SUM(principal_minor) per version against the original principal and locate the rounding residual.
Follow-up
- What does days_past_due mean for an instalment that exists in two versions with different due_dates?
- A payment arrives allocated to an instalment_no that exists only in the superseded version. What do you do with it?
- How do you prove the ledger postings made under version 1 still reconcile once version 2 exists?
Design a payment gateway API handling high concurrency, rate limiting,…
Design a payment gateway API handling high concurrency, rate limiting, distributed caching, and strict idempotency.
Approach
- Choose a partition key and say what query it makes expensive.
- Name the read and write paths separately; they rarely have the same bottleneck.
- State the consistency you need, and where you are willing to be stale.
Follow-up
- How does this behave when that dependency is down for an hour?
- What would you drop to keep the system up under load?
Architect an asynchronous event-driven system that ingests high-volume…
Architect an asynchronous event-driven system that ingests high-volume transaction events and writes them safely to a persistent data store.
Approach
- Name the failure you are designing for, then the recovery path.
- Choose a partition key and say what query it makes expensive.
- Fix the scope first: who calls this, how often, and what they do when it fails.
Follow-up
- What would you drop to keep the system up under load?
- What breaks first when traffic grows ten times?
Explain concurrency vs. asynchrony, race conditions, thread synchroniz…
Explain concurrency vs. asynchrony, race conditions, thread synchronization, and how to prevent deadlocks in high-throughput applications.
Approach
- State your assumptions explicitly before working the problem.
- Say what you would check first and why it is the highest-information step.
- Clarify what is being asked and what a complete answer contains.
Follow-up
- What assumption would you test first?
- How would you know your answer was wrong?
Webhook intake under duplicate, unordered delivery and ledger outage
A processor POSTs 5,000 signed events/s to one endpoint, at least once, unordered, retrying for 24 hours until it gets a 2xx, and treating a response slower than 10 seconds as a failure. Each event carries a provider event id, an object id and an object version. Your ledger is occasionally unavailable for minutes at a time. Design intake: signature verification, deduplication, how a succeeded event arriving before the processing event it follows is applied, and what status you return while the ledger is down. Justify that status and bound the backlog.
Approach
- Verify before parsing. Compute HMAC-SHA256 over the timestamp concatenated with the raw request bytes, compare in constant time, and reject a timestamp outside a tolerance of a few minutes so a captured request cannot be replayed indefinitely. Parsing to an object and re-serialising before verification is how a signature check silently stops working the day a JSON library reorders keys or renormalises a number.
- Make 'accept' mean 'durably recorded', not 'applied'. Insert (provider_id, provider_event_id, raw body, received_at) under a UNIQUE constraint on the provider's event id and return 2xx in single-digit milliseconds; apply asynchronously. The endpoint doing no business logic is what keeps it inside the provider's timeout at 5,000/s, and the constraint handles the at-least-once repeats without application logic.
- Order by version, never by arrival: UPDATE payment_intent SET status = $s, version = $v WHERE intent_id = $1 AND version < $v. A stale succeeded-then-processing pair updates zero rows on the second event and is recorded as discarded. Arrival time and the provider's own created timestamp are both unusable for ordering, because retries reorder both.
- Split the failure case by which store is down, because the answer differs. If acceptance storage is healthy and only the ledger is down, keep returning 2xx: you hold the durable record and version-ordered application makes the replay safe, so availability costs you nothing but lag. If acceptance storage itself is down, return 5xx and let the provider's 24-hour retry schedule be your queue; a 2xx you cannot honour is an event the provider will never send again.
- Bound the backlog explicitly in rows or in lag-minutes, alarm on it, and shed to 503 past the bound rather than accepting work you cannot drain. The health metric that matters is per-object version gaps, not queue depth: a queue can be empty while an object is stuck three versions behind.
Worked solution 25 min
- Build the receiver: constant-time HMAC over timestamp plus raw body, a tolerance window, then an INSERT under UNIQUE (provider_id, provider_event_id) and an immediate 200 with no application.
- Replay a captured request outside the tolerance window and assert rejection; replay it inside the window and assert one stored row rather than two.
- Deliver succeeded at version 7 before processing at version 5 and apply both through the version-conditional UPDATE.
- Stop the ledger, drive 5,000 events/s for five minutes, and watch endpoint p99 and backlog growth.
Follow-up
- After a four-hour outage the provider replays everything. What does your consumer do with 70 million duplicates?
- One hot merchant object receives 50 events/s. Does the version-conditional update starve or livelock, and what do you change?
- You detect a permanent gap at version 12 for one object. How do you close it?
Duplicate captures appear only in production, roughly weekly
About once a week one payment is captured twice. The idempotency path is: SELECT id, response_body FROM idempotency_key WHERE scope = $1 AND key = $2; if no row, call the processor; then INSERT. The table has UNIQUE (scope, key). Logs for each duplicate show one successful capture pair and one HTTP 500 carrying SQLSTATE 23505. A 200-iteration sequential test passes, and a 50-thread version passes on a laptop but fails on the production-sized cluster. Explain why, and give the fix.
Approach
- Read the 23505 as evidence, not as noise. A unique violation on the INSERT proves two requests both passed the SELECT and both reached the INSERT, which means both had already called the processor. The duplicate charge happened before the constraint fired. The constraint is reporting the race; it is not causing it, and anyone who treats the 500 as the bug fixes the wrong thing.
- Name the interleaving precisely. Under READ COMMITTED each statement takes a fresh snapshot, so two concurrent requests with the same key can both run the SELECT before either INSERTs and both see zero rows. Raising the isolation level does not fix check-then-act by itself, because at SELECT time the first transaction has written nothing to conflict with; SERIALIZABLE only converts the race into a 40001 abort that the code must then retry.
- Explain the reproduction gap rather than calling the bug rare. The window is the duration of the processor call: milliseconds against a stub on a laptop, hundreds of milliseconds against a real processor. Production retries are also correlated, since a client timeout produces a second request at a predictable delay, while a thread-pool test fires all 50 within microseconds and lands them on the same side of the window. The local test is not exercising the window at all.
- Restructure so the database picks the winner before any side effect. INSERT the key first with ON CONFLICT (scope, key) DO NOTHING RETURNING id. A returned id means this request owns the effect and may call the processor. No returned row means another request owns it, and note that RETURNING yields nothing on conflict, so the loser must then SELECT the existing row explicitly. Mutual exclusion now lives in one atomic statement and the window is gone.
- Give the loser something to read. If the winner is still in_progress, the loser must neither error nor perform the effect: it polls the row inside the caller's timeout budget and replays response_status and response_body once status is completed, or returns 409 when request_fingerprint differs. Without this, deduplication turns a successful payment into a visible failure.
- Close the crash window separately, because the atomic insert does not cover it. If the winner dies after calling the processor and before writing completed, the row stays in_progress with a stale locked_at. Recovery must query the processor for that key or client reference rather than assume either outcome, which is why the processor's own idempotency key has to be the same value, generated once by the caller and reused on every attempt.
Follow-up
- The same key arrives with a different request_fingerprint. What do you return, and why is returning the cached response wrong?
- What is your locked_at staleness threshold, and what does the sweeper do when it finds an expired one?
- Write the test that fails on the laptop. What do you have to inject to make the window observable there?
Day one measures instead of guessing, under a fixed rubric, and the remaining hours are allocated in proportion to the gaps before any studying begins. The allocation is deliberately not renegotiated midweek, because the area that feels worst on day three is usually the one that is moving.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Technical Screen: online assessment simulation
- Sit a timed set of three to four problems in a plain browser editor, matching the easy-to-medium-hard mix candidates report. Include longest substring without repeating characters, the start node of a linked-list cycle, and one dynamic programming problem.
- Solve the problems from easiest to hardest and log the minutes spent on each.
- Before each submission, run the prompt's examples plus an empty input, a single element and a maximum-size input through a harness.
Deliverable: A timing log for each problem and a list of every edge case your first submissions missed.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Coding round: data-structure design from a blank file
- Implement the reported LRU cache with a hash map and a doubly linked list with sentinel nodes, and state why get and put are O(1).
- Implement the reported structure that increments and decrements string counts and returns the max and min keys in O(1), using a linked list of count buckets. Write down the invariant that keeps both ends correct.
- Implement longest increasing subsequence in O(n log n) and explain why the tails array stays sorted.
- Work through the coding exercise on as-of balance queries over an append-only log, and compare the cost of the naive approach with the offline sweep.
Deliverable: Three implementations written without references, each with its invariant in one sentence, plus the written cost comparison for the balance exercise.
Practice prompt ↗Practice prompt ↗03Fundamentals: SQL and data modelling
- Write the reported top-three-salaries-per-department query twice, once with ROW_NUMBER and once with DENSE_RANK, and show on sample data how each handles ties.
- Explain inner versus outer joins on a two-table example that includes unmatched rows.
- Work through the SQL exercise on versioning a loan schedule and run its checks.
- Answer the idempotent charge endpoint drill: the unique constraint, the ON CONFLICT claim, and what a concurrent loser returns.
Deliverable: Two ranking queries with a tie case that separates them, and a passing run of the loan-schedule checks.
Practice prompt ↗Practice prompt ↗04Fundamentals: Java, concurrency and networking
- Write two-sentence answers with one example each on JVM memory areas, garbage collection, interface versus abstract class, and one SOLID principle applied to real code.
- Explain concurrency versus asynchrony, a race condition you could reproduce, and two ways to prevent deadlock, such as a global lock order and lock timeouts.
- Walk through what happens when you enter a URL, from DNS through the TCP handshake to the HTTP response, then cover OAuth and JWT basics and what Docker and Kubernetes each manage.
- Work through the duplicate-captures debugging drill and explain why checking first and then inserting lets two requests reach the processor.
Deliverable: A one-page fundamentals card with a two-sentence answer for each topic, readable aloud without notes.
Practice prompt ↗Practice prompt ↗Worked solution ↗05System Design round: payments and idempotency
- Design the reported payment gateway API: endpoints, the idempotency header, the stored response, where the rate limiter sits, and where caching is safe.
- Model a payment as explicit states, and say what a retry does in each state and what happens if the process dies after calling the processor.
- Work through the webhook-intake design exercise from start to finish, then redo it with Kafka as the transport, stating the partition key and the deduplication key.
- If your loop is entry-level, also prepare a schema walkthrough of one past project.
Deliverable: One annotated design covering the API contract, states, idempotency, and failure handling for the ledger-down case.
Practice prompt ↗Practice prompt ↗06Phone screen and Behavioral round
- Write a one-page fact sheet for your main project: traffic, data size, team size, timeline, what broke and which decisions were yours.
- Prepare STAR answers for the reported prompts: a complex project walkthrough, a tight deadline or changing requirements, a technical disagreement, code quality under timelines, and why Visa.
- Rehearse the incident-postmortem drill with a timeline, how you sized the damage and the prevention change, including one thing you got wrong.
- Run a mock phone screen where a coding problem switches to a fundamentals question partway through.
Deliverable: A project fact sheet and five STAR answers, each ending in a measured result or an honest statement that nothing was measured.
Practice prompt ↗Practice prompt ↗07Virtual Onsite rehearsal
- Run mock rounds back to back: one coding problem from the reported DSA list, one design prompt (a distributed scheduler, or an event-driven pipeline that ingests transaction events), and one behavioral round.
- Afterwards, compare the project figures you quoted in each round against your fact sheet and correct any that changed.
- Redo the single problem from days 1 to 6 that you missed most badly, from a blank file.
Deliverable: A list of the remaining weak spots from the mock, each with the first thing you will say if it comes up.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Reported behavioral prompts for this role cover a complex project walkthrough, tight deadlines and changing requirements, a technical disagreement with a senior engineer or stakeholder, code quality under business timelines, and why Visa. Answer in STAR form. Spend most of the answer on your own actions and the trade-off you made, and end with a measured result or an honest statement that nothing was measured. Keep the project figures identical to the ones you quote in technical rounds.
Why do you want to work at Visa, and how does this role align with you…
Why do you want to work at Visa, and how does this role align with your long-term software engineering career goals?
Approach
- Name the disagreement and how you resolved it with evidence.
- Give the blast radius: what could have broken, and what you measured.
- Close with what you would do differently, concretely.
Follow-up
- How did you know your change caused the improvement?
- What did you decide not to do, and why?
How do you ensure code quality and system safety when pushing updates …
How do you ensure code quality and system safety when pushing updates to production under strict business timelines?
Approach
- Give the blast radius: what could have broken, and what you measured.
- State the situation in two sentences and spend the rest on the reasoning.
- Pick a story where you made the decision, not one where you watched it.
Follow-up
- How did you know your change caused the improvement?
- What did you decide not to do, and why?
Own the postmortem for a duplicate-capture incident
A processor slowed down, callers timed out and retried without reusing their idempotency key, and 412 captures were duplicated over 90 minutes before a reconciliation break report surfaced it. Take the on-call role. Describe an incident you owned of comparable blast radius: how it was detected, how you bounded the affected population, what you stopped first, and how customers were made whole. Give a wall-clock timeline, the query or metric that sized the damage, and the change that would have prevented it. Include what you got wrong during the response, not only after it.
Approach
- The probe is whether you can bound an unknown blast radius under time pressure. Open with the invariant that broke (at most one capture per authorisation attempt) rather than the symptom, because the invariant tells the listener what to count.
- Size the population with a stated query, not an adjective: duplicate captures are ledger_entry rows with source_type='capture' grouped by source_id having count(*) > 1, joined back to payment_intent for the affected merchants and amounts. Say how long that query took and whether you could run it against a replica while the incident was live.
- Separate mitigation from fix and say which you did first. Mitigation is usually cheap and blunt (disable the retry path, drop the caller's concurrency, hold captures behind a flag); the fix is a UNIQUE constraint plus a stored response, and it is not an incident-window change.
- State the remediation arithmetic explicitly: refunds are new customer-visible movements with their own fees and their own settlement lag, so the count of duplicates, the total minor units, the refund posting date and the customer notification are four separate numbers a strong answer has ready.
- Close on the prevention change and its cost. Naming one guard that would have caught it earlier (a break-age alert, a duplicate-capture counter on the ledger write path) beats listing five that nobody staffed.
- Name your own error inside the response window: a mitigation you tried that made it worse, or the 20 minutes you spent on the wrong hypothesis. Interviewers weight that heavily because it is the part candidates rehearse away.
Follow-up
- The retry came from a client you do not control. What do you change so a client that regenerates its key per attempt cannot cause this again?
- How would you have detected it in 5 minutes instead of 90, and what would that detector cost in false pages per week?
- A merchant disputes your count of affected transactions. What do you show them?
- 01
Walk through a complex project on your resume: your specific contribution, the technical trade-offs, and the architectural choices.
- 02
Describe a time you faced a tight deadline or changing requirements. How did you weigh technical debt against delivery speed?
- 03
Tell me about a technical disagreement with a senior engineer or stakeholder and how it was resolved.
- 04
How do you ensure code quality and system safety when pushing updates to production under strict business timelines?
- 05
Describe how you assessed, contained, communicated and resolved a critical production vulnerability under pressure.
- 06
Why do you want to work at Visa, and how does this role fit your long-term engineering goals?
Is this an official Visa interview guide?
No. It is PracHub's own research and practice material for the Software Engineer role at Visa. The rounds and questions reflect what candidates have reported, not a process Visa has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How difficult is the Online Assessment (OA) at Visa, and how should I prepare?
Candidates report three to four coding problems on HackerRank or CodeSignal, with a time limit of 70 to 90 minutes. Difficulty ranges from easy to medium-hard, and the problems cover array manipulation, hash maps, sliding windows, two pointers, dynamic programming and trees. Practise timed sets in a plain editor, solve the easier problems first, and test every submission on empty, single-element and maximum-size inputs before you submit.
PracHub interview research ↗What is the typical timeline from initial application to offer?
Candidate reports put the process at about three to five weeks, and some round summaries stretch it to around six. After you apply or clear the online assessment, a recruiter usually follows up within one to two weeks to schedule the next screens or the final rounds. Ask your recruiter for the schedule of your specific loop.
PracHub interview research ↗Are system design interviews required for junior or entry-level candidates?
Candidates report that full system design rounds are mostly for mid-level, senior and staff roles. Entry-level candidates may still get low-level design questions, database schema questions, or questions about the architecture of a personal or academic project. If you are early-career, prepare a schema walkthrough of one project and one low-level design, such as an LRU cache.
PracHub interview research ↗Does Visa allow you to choose your preferred programming language during technical interviews?
Candidates report being allowed to code in their language of choice, such as Java, Python, C++ or Go, in live algorithm rounds. For backend Java roles, team-specific questions can move into frameworks such as Spring Boot. If your target team is Java-heavy, review Spring Boot dependency injection and how you would size a database connection pool in a high-concurrency service.
PracHub interview research ↗What should a payment system design answer cover?
The reported prompts include a global payment system, a payment gateway API, a distributed scheduler and an event-driven ingest pipeline. A complete answer names the payment states and transitions, the idempotency key and the unique constraint that make retries safe, how duplicate or out-of-order events are handled, and what happens when a process dies halfway through a multi-step write. Also cover rate limiting and caching on the API path, and argue the SQL versus NoSQL choice from the consistency the data needs. The webhook-intake exercise in this guide rehearses most of these points.
PracHub Software Engineer practice ↗How much Java and CS fundamentals should I review?
Plan real time for it. Reported topics include JVM internals and garbage collection, interface versus abstract class, SOLID and design patterns, concurrency versus asynchrony, race conditions and deadlock prevention, SQL window functions and joins, SQL versus NoSQL key design, TCP/IP, REST, OAuth and JWT, and Docker and Kubernetes basics. For each topic, prepare a two-sentence answer with one concrete example, and be ready to go a level deeper when asked.
PracHub Software Engineer practice ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-24 - 02PracHub Software Engineer practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-24 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-24