Candidate reports describe Affirm Software Engineers working on transaction engines, credit decisioning, identity verification pipelines, streaming data infrastructure, consumer and merchant platforms, checkout integrations and direct-to-consumer card experiences. In your preparation, expect problems where the data has to stay consistent, every change has to be auditable, and money must never be double-applied or lost.
The reported questions reflect that domain. The coding prompts include a multi-player card game engine with extensions, an insert/delete/getRandom structure, and a memory-efficient aggregation of a large transaction CSV. The bank titles cover matching loans to transactions, duplicate transactions in a time window, and resolving a loan request's top-level parent. The design prompts include a real-time ledger, a payment-due notification service across SMS, email and push, a rate limiter for merchant APIs, and an identity decisioning system that combines credit bureau data with internal fraud models.
Candidates report that coding happens in a live, browser-based IDE where the code has to compile and run against test cases, and that problems start with a base requirement and then add one or two extensions. Prepare to write working, modular code in the language you know best, and to finish the first part quickly enough to reach the extension.
Recruiter Pre-Screen
reportedCandidates describe this as a first conversation with a recruiter to check fit and talk through the role. Use it to raise your hard constraints: start date, notice period, work authorisation, location and compensation range. These are the issues that sink an offer late. Also ask which format your initial technical screen will take. Candidates report it can be an online assessment or a live coding interview, and that some teams offer a take-home instead. Ask which languages are accepted too. Confirming the format now lets you practise in the right setting.
What to demonstrate
- Whether your constraints on start date, authorisation, location and compensation fit the role before a loop gets scheduled
- Whether you can explain in a sentence or two why this role and domain interest you, backed by something specific from your own work
- Whether your stated timeline is real, including any competing process you raise
How to prepare
- Write each constraint as a one-line fact before the call so you state it rather than negotiate it on the spot
- Ask whether the technical screen is an online assessment, live coding or a take-home, and which languages are accepted
- Prepare a short answer to why Affirm and why this role, tied to a system you have built that handled money, transactions or data consistency
Initial Technical Screen
reportedCandidates say this stage is either an online assessment or a live coding interview. Across the coding rounds they report a browser-based IDE where the code must compile and run against test cases, so pseudo-code will not carry you. They also report that coding problems generally start with a base requirement and then add one or two extensions. Most lost time goes to over-building part one, half-remembered library calls, and random edits made without a failing case you understand. Get a correct base version running, then refactor only as far as the next extension needs.
What to demonstrate
- Whether a clean, correct base solution compiles and passes its tests before you add abstractions
- Whether your classes and functions are split so an extension slots in without a rewrite, with data, entities and rules kept apart
- Whether you pick standard structures (hash maps, sets, queues, heaps) that fit the access pattern and can explain the complexity
- Whether you find a failing case and explain it before changing the code
How to prepare
- Take a multi-part coding problem, solve the base version, then add each extension in turn, timing how long each change takes and how much base code it touches
- Implement common O(1) building blocks from a blank file, such as a hash map of indices paired with an array for constant-time removal, until they come without thinking
- Drill your language's standard library (Python's collections or Java's java.util) in a browser editor with no autocomplete until you stop looking up basic calls
- Practise single-pass streaming solutions over large inputs that keep only running aggregates in memory, and run them early on small inputs
Hiring Manager Screen
reportedCandidates report that this conversation covers your background, technical interests and experience with high-level architecture. Expect to walk through a system you worked on at the architecture level. Explain the components and why each is there, where the data lives, what must not be lost, and what you would change now. The common stall is not having the numbers that justify the design. Pick one or two systems before the call and rehearse them until you can answer follow-ups without drifting into generalities.
What to demonstrate
- Whether you can describe a system you worked on end to end, with the reason each component exists
- Whether you know the load and data figures that justify your design choices, and can say what would break first as they grow
- Whether you separate what you personally decided and operated from what the team or another group owned
- Whether your stated technical interests connect to the kind of work described for the role, such as transaction processing, data pipelines or decisioning systems
How to prepare
- For one system you built, write down peak request rate, largest table size, the latency you were held to and what must survive a crash
- Practise explaining what fails first at ten times the load, and the smallest change that would buy headroom
- Prepare one decision where the simpler option was right, along with the figure that justified it
- List two or three technical interests and connect each to a project you can discuss in depth
Virtual Onsite
reportedCandidates describe the virtual onsite as several technical and behavioral rounds, and report that it always includes live, interactive coding. Which technical formats your onsite contains is not specified in candidate reports, so ask your recruiter and prepare for runnable coding with extensions plus any design discussion they confirm. For the behavioral rounds, prepare stories in STAR form on ownership, collaboration and feedback; the behavioral section of this guide lists the reported prompts.
What to demonstrate
- Whether coding answers stay runnable and modular across extensions, not just correct for the base case
- Whether any design answer states the API contract, the data model and how failures are handled, including idempotency and audit trails where money moves
- Whether you argue trade-offs, such as choosing a readable, modular design over a slightly more optimal but harder-to-maintain one
- Whether behavioral stories show your own actions and outcomes on ownership, disagreement and feedback
How to prepare
- Practise a payments design question such as a ledger: double-entry postings, an idempotency key on writes and an append-only audit history, then explain how a retried request is prevented from posting twice
- Practise writing a design out as text in a shared document, with endpoints, schema, per-channel delivery status and retry behaviour, as well as drawing it
- Run a mock of one multi-part coding problem in a browser editor with no autocomplete and check that you reach the extension with runnable code
- Rehearse a trade-off story, a feedback story and a cross-team production incident story, each with a measurable result
7 candidate reports. Individual accounts describe a particular role and hiring cycle.
Affirm Senior Data Analyst Interview Experience — A Downlevel After Passing the Take-home
In July, a kind younger alum referred me for Full Stack Analyst II. In early August, HR contacted me to say that role was gone and asked whether I'd consider another Senior Analyst role. After that, it was one step each week: HR call -> hiring manager interview -> take-home. The HM case: Analyze what could have caused the delinquency rate to jump in the past quarter. The take-home: What product a…
Read full experienceAffirm Software Engineer interview: finance-style coding and a timed technical screen
My online assessment felt like a finance-specific LeetCode problem. Afterward, I had a 30-minute recruiter screen about the role and logistics, including compensation, the actual work and cultural fit. It wasn't a particularly deep conversation, but it gave me an idea of what they wanted. Next came a technical screen where I shared my screen for a timed coding problem. I only saw the prompt when…
Read full experienceAffirm Software Engineer Interview Experience: payments design and poor recruiter communication
After an initial recruiter call, I entered a fintech-focused process. The roughly 30-minute screen covered my background, interest in fintech and BNPL, compensation, and location. A deeper hiring-manager conversation followed, covering past projects, technical decisions, and fit for the specific team. The coding portion was a 60-minute HackerRank session. Next came about an hour of system design…
Read full experienceAffirm Software Engineer Interview Experience — Fraud Detection Coding and a Payment System Design Round
Four rounds total: one hour of coding + half an hour of behavioral questions + one hour of system design + half an hour of behavioral questions. The two behavioral rounds were just the standard questions. Coding round (the prompt was long, but not hard): Part 1: Live Fraud Detector (Debug Existing Function) Description One of Affirm's competitive edges is our ability to do credit underwriting and…
Read full experienceAffirm Software Engineer Interview Experience — Five Onsite Rounds, One Shaky HM Round
View report detailsPracHub editorial advice for the preparation topics above.
Over-building part one of a multi-part coding problem and never reaching the extension
Candidates report that coding problems start with a base requirement and add one or two extensions, and that failing to reach the second part is a frequent cause of rejection. Get a correct, runnable base version first with only the separation the next step will obviously need, such as deck, player and scoring rules in the reported card game question. Then refactor when the extension arrives. Say out loud what you are deferring so the interviewer knows it was a choice.
Writing pseudo-code or code that never compiles in the live IDE
Candidates say the code must compile and run against test cases. Practise in a browser editor with no autocomplete, in the language you will use, until standard-library calls for sorting, grouping, string splitting and counters come from memory. Run the code early on a small input instead of writing the whole solution before the first execution.
Designing a ledger or payment flow without saying what stops a double post
For the reported ledger, notification and decisioning design questions, state the write path's identity before you draw components. Use an idempotency key generated once by the caller, a unique constraint that enforces it, and balanced double-entry postings stored as integer minor units, never floats. Say what a caller sees after a timeout: an unknown outcome to look up, not a failure to retry blindly. The capture-endpoint and money-parsing worked exercises drill exactly this.
Answering the equality-versus-equity question with a definition and nothing you have done
Candidates report this question comes up in the behavioral rounds, alongside questions on building an inclusive team. Give a short definition: equality gives everyone the same thing, equity adjusts support so people can reach the same outcome. Then follow at once with a specific workplace example: a process you noticed was uneven, such as how on-call load, code review attention or meeting airtime was distributed, what you changed and what happened next.
Describing past architecture in the hiring manager screen without numbers or personal ownership
Candidates say this screen covers your background and your experience with high-level architecture. Before the call, write down the request rate, data size, latency target and failure requirements for one system you worked on, and which decisions were yours. A walkthrough in which 'we' did everything and no figure justifies any component is hard for an interviewer to credit to you.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Implement a multi-player card game engine where a deck of cards is dis…
Implement a multi-player card game engine where a deck of cards is distributed among players, and the player with the highest score wins. Extend your solution to handle a variable number of players, custom deck sizes, and dynamic winning rules.
Approach
- Name the brute-force solution and its complexity before improving on it.
- Walk one small example through your approach before writing the whole thing.
- Restate the input: its shape, its size, and what is guaranteed about it.
Follow-up
- What is the worst case, and how likely is it on real data?
- How does this change if the input no longer fits in memory?
Design a custom data structure that supports insert, delete, and getRa…
Design a custom data structure that supports insert, delete, and getRandom operations in O(1) time complexity, incorporating specific rules for data persistence.
Approach
- Restate the input: its shape, its size, and what is guaranteed about it.
- Walk one small example through your approach before writing the whole thing.
- Choose the data structure from the access pattern, not from familiarity.
Follow-up
- Which test case would catch an off-by-one here?
- How does this change if the input no longer fits in memory?
Write a clean utility to parse and filter a large CSV file containing …
Write a clean utility to parse and filter a large CSV file containing transaction charges, aggregating the data by merchant category while optimizing for minimal memory usage.
Approach
- Name the brute-force solution and its complexity before improving on it.
- Choose the data structure from the access pattern, not from familiarity.
- State the target complexity and say which constraint rules the naive version out.
Follow-up
- What is the worst case, and how likely is it on real data?
- How does this change if the input no longer fits in memory?
Parse settlement amounts into minor units without floating point
A settlement file carries amounts as text: 1234.56, -0.07, 1,234.5, 1234, 12 345,67, and for three-exponent currencies 1.234. Each line also carries an ISO 4217 alphabetic code and the file's sign convention. Write the function converting one amount string plus its currency into a signed amount_minor int64, scaling by the currency's ISO 4217 exponent — 0 for JPY and KRW, 2 for USD and EUR, 3 for KWD and BHD. Reject anything not convertible exactly, with a typed reason. Never raise, never use a float.
Approach
- Fix the contract first: return a result union of
Ok(int64)orErr(reason)with reason drawn fromunknown_currency,malformed,too_many_fraction_digits,overflow. A parser on this path that throws turns one bad line into a failed fifty-million-line run. - Decide which character is the decimal mark from the file's format spec, not from a heuristic on the last separator, then strip grouping separators.
1,234is genuinely ambiguous between 1234 and 1.234 exactly when the exponent is 3, so an unstated spec is a reject, not a guess. - Split on the decimal mark. Right-pad the fraction with zeros to the currency's exponent e; if it is longer than e and any dropped digit is non-zero, return
too_many_fraction_digits. Silently truncating is how half a minor unit becomes a reconciliation break nobody can source. - Build the integer by digit accumulation on int64 with checked multiply and add, or equivalently concatenate the integer and padded-fraction substrings and parse once. Either way the value never passes through a float.
- Apply the sign last, from the convention the file declares — trailing
CR/DR, a trailing minus as in1234.56-, or a leading minus. Assuming a leading minus silently flips every credit in a file that uses the trailing form. - Complexity is O(len) per amount with no allocation beyond a digit buffer, so the whole file is one O(total bytes) pass.
Worked solution 20 min
- Write the exponent table
{JPY:0, KRW:0, USD:2, EUR:2, KWD:3, BHD:3}and the result type before any parsing code. - Implement integer-only scaling: split on the decimal mark, pad or reject the fraction, concatenate, parse once with checked arithmetic.
- Table-test: (
1234.56, USD) to 123456; (1234, JPY) to 1234; (1.234, KWD) to 1234; (1234.56, JPY) toErr(too_many_fraction_digits); (0.005, USD) toErr; (-0.07, USD) to -7. - Write the float implementation beside it and run both over ten thousand random amounts across all six currencies, printing every disagreement.
- Round-trip: format
amount_minorback to text at the currency's exponent and assert equality with the normalised input.
Follow-up
- Exponent 4 exists (CLF). Does the exponent come from a compiled-in table or from the file header, and what happens when the two disagree?
- A line has more fraction digits than the exponent allows. Is that a reject, a round, or a break — and who makes that call?
- How do you fuzz this so that every failure is a typed rejection rather than a wrong number?
Explain why the outbox relay stopped using its partial index
outbox_event holds event_id, aggregate_type, aggregate_id, aggregate_version, event_type, payload jsonb, published_at, attempts, last_error, created_at, with index ix_unpub ON outbox_event (created_at) WHERE published_at IS NULL. The relay runs SELECT ... WHERE published_at IS NULL ORDER BY created_at LIMIT 500 FOR UPDATE SKIP LOCKED, then marks each row by setting published_at. Unpublished rows hold steady near 400, but the query has gone from 3 ms to 900 ms. Explain what EXPLAIN (ANALYZE, BUFFERS) will show, why it happens, and the fix.
Approach
- Read the plan for the gap between rows returned and work done: an index scan on ix_unpub returning 500 rows while touching tens of thousands of buffers is the signature. Rows Removed by Filter and the buffer counts name it; wall-clock alone does not, because a warm cache hides it.
- Explain the mechanism: marking a row published is an UPDATE, which writes a new tuple version. The new version fails the index predicate and leaves ix_unpub, but the dead old version's index entry stays until vacuum removes it, so the scan walks dead entries and discards them. PostgreSQL can hint an entry LP_DEAD once a scan has proved it dead, which cheapens repeat visits, but the index pages themselves still have to be read and are not reclaimed.
- Ask why vacuum is not reclaiming. Anything holding the xmin horizon back prevents removal: a long-running query, an idle-in-transaction session, an abandoned prepared transaction, or an inactive replication slot. Check the oldest xact_start in pg_stat_activity, pg_replication_slots, pg_prepared_xacts, and n_dead_tup with last_autovacuum in pg_stat_all_tables.
- Fix in order of leverage: delete or archive published rows instead of leaving them in place, so a queue table stays a queue; keep the transaction horizon short and alert on it; then tune autovacuum on this one table with an aggressive scale factor rather than changing the global setting.
- Rule out the other failure with the same symptom: a partial index is usable only when the planner can prove the query predicate implies the index predicate, so rewriting the filter as coalesce(published_at, 'epoch') = 'epoch' or wrapping the column in a function disqualifies the index entirely and produces a sequential scan instead of a bloated index scan.
- Verify by re-running EXPLAIN (ANALYZE, BUFFERS) after the horizon is released and a VACUUM completes, comparing shared buffer reads rather than elapsed time, and confirm the relay keeps per-destination ordering after the change.
Worked solution 25 min
- Open a second session with BEGIN; SELECT 1; and leave it idle to hold the xmin horizon.
- Churn 500,000 events through insert and publish, then run EXPLAIN (ANALYZE, BUFFERS) on the relay query inside a transaction you roll back, so FOR UPDATE does not hold locks.
- Record the plan node, rows, Rows Removed by Filter and shared buffer counts.
- Close the idle transaction, VACUUM outbox_event, and re-run the identical EXPLAIN, then add the archive step and re-measure.
Follow-up
- SKIP LOCKED means two relay workers never block each other. What else does it change about ordering guarantees for a single destination?
- You archive published rows to a second table. What does that do to the relay's crash recovery and to duplicate delivery?
- The relay batches 500 rows and publishes them, then marks them. Where exactly can it crash, and what does the consumer see?
Add and backfill business_date on a live ledger table
ledger_entry holds 4 billion rows, is append-only, takes 10,000 inserts per second, and every reporting query currently derives the business date as posted_at::date. You must add business_date date NOT NULL, populated from the cutoff rule (17:00 in the account's own timezone), backfilled across all history, indexed, and cut over, with no write downtime and no long-held lock. Give the ordered migration steps with the lock each one takes, how you make the backfill restartable and throttled, and how you retire the old expression safely.
Approach
- Add the column nullable and with no default. ALTER TABLE ... ADD COLUMN takes ACCESS EXCLUSIVE but is a catalogue-only change held for microseconds. The hazard is the lock queue, not the statement: a blocked ALTER waits behind one long reader holding ACCESS SHARE, and every query arriving afterwards queues behind the ALTER's pending ACCESS EXCLUSIVE, so set lock_timeout to a couple of seconds and retry rather than letting a metadata change take the table down.
- Deploy the write path before the backfill, so new inserts populate business_date from the cutoff rule while reads stay on the old expression. The backfill then chases a closed set with a fixed upper bound instead of a moving target.
- Backfill in bounded batches keyed by entry_id range, on the order of 50,000 rows per statement, committing between batches and recording the high-water mark in its own table so a killed run resumes instead of restarting. WHERE business_date IS NULL makes each batch idempotent, and the pacing is set by replica lag and dead-tuple growth rather than CPU, since each UPDATE writes a new row version and the WAL volume is proportional to the rows touched.
- Install the constraint without a blocking scan: ALTER TABLE ... ADD CONSTRAINT ck_business_date CHECK (business_date IS NOT NULL) NOT VALID takes a brief ACCESS EXCLUSIVE and scans nothing, then VALIDATE CONSTRAINT takes SHARE UPDATE EXCLUSIVE and runs alongside reads and writes. On PostgreSQL 12 and later, SET NOT NULL can then use the validated CHECK and skip its own full scan; on 11 and earlier it always scans, so the CHECK is the migration on those versions.
- Build the index with CREATE INDEX CONCURRENTLY, which avoids ACCESS EXCLUSIVE at the cost of two table passes, cannot run inside a transaction block, and on failure leaves an INVALID index that must be dropped and rebuilt rather than reused.
- Cut over behind a flag: run the new and old expressions side by side for one reporting cycle and compare totals per day, since the cutoff rule will legitimately move entries near 17:00 across the boundary. Only once they reconcile do you retire the posted_at::date expression index, and you keep posted_at as the ordering key rather than repurposing it.
Follow-up
- Reporting now wants ledger_entry partitioned by business_date. Why can this not be another ALTER, and what is the migration instead?
- Nightly totals move for the days around the cutoff change. How do you tell a correct restatement from a backfill bug?
- The backfill is halfway done when a replica falls 20 minutes behind. What do you throttle, and what do you refuse to throttle?
Architect a real-time ledger system that records user transactions, en…
Architect a real-time ledger system that records user transactions, ensuring strict transaction consistency, auditability, and zero data loss.
Approach
- Design the error taxonomy before the success shape; callers branch on it.
- Say who the caller is and what they do when the call fails halfway.
- State how the contract changes without breaking existing clients.
Follow-up
- What does a partial failure look like to the caller?
- How does a client discover it is on an old version of this contract?
Design a notification service that alerts users of payment due dates a…
Design a notification service that alerts users of payment due dates across multiple channels (SMS, email, push notifications) while managing user preferences and high-throughput delivery queues.
Approach
- Separate accepted, pending, failed and confirmed; they are different facts.
- Say who the caller is and what they do when the call fails halfway.
- Define the identity of a request so a retry cannot double-apply it.
Follow-up
- What happens if the caller retries after a timeout?
- What does a partial failure look like to the caller?
Architect an identity decisioning system that aggregates data from ext…
Architect an identity decisioning system that aggregates data from external credit bureaus and internal fraud models to approve or deny loan applications in real-time.
Approach
- Define the identity of a request so a retry cannot double-apply it.
- Separate accepted, pending, failed and confirmed; they are different facts.
- Design the error taxonomy before the success shape; callers branch on it.
Follow-up
- What happens if the caller retries after a timeout?
- How does a client discover it is on an old version of this contract?
Capture endpoint that survives concurrent duplicate retries
POST /payments/{intent_id}/capture carries an Idempotency-Key header and an amount. Callers retry on timeout, and two retries can land concurrently on different instances. You have idempotency_key(id, scope, key, UNIQUE(scope,key), request_fingerprint bytea, status in_progress|completed|failed, response_status, response_body jsonb, locked_at, completed_at, expires_at). The processor capture takes 200 to 2,000 ms. Design the path so exactly one capture reaches the processor, every duplicate receives the identical response, and a crash between the processor call and your commit converges. Name the single statement that is the concurrency control.
Approach
- The concurrency control is INSERT INTO idempotency_key (...) VALUES (...) ON CONFLICT (scope, key) DO NOTHING RETURNING id, and it must commit before the processor call. A returned row means you own the effect; zero rows means you lost and must read the winner's result. A SELECT-then-INSERT check cannot substitute: two requests both read 'absent' and both proceed, and the window is exactly the concurrency you are defending against.
- Commit the in_progress row in its own short transaction. A concurrent duplicate insert blocks on an in-flight conflicting insert until that transaction ends, so wrapping the 2 s processor call in the same transaction turns every duplicate into a 2 s lock wait and a retry storm into pool exhaustion.
- Define loser behaviour per status rather than uniformly: completed replays response_status and response_body unchanged; in_progress returns 409 with Retry-After and performs nothing; failed splits by cause, since a terminal processor decline should replay but a transport failure should let the key be retried. Getting this wrong in the safe direction (replay a decline) is better than returning an error for a capture that succeeded.
- Fingerprint the canonicalised body with SHA-256 and compare on every hit. Same key with a different amount is a client bug and must return 409 or 422, never the cached response, because returning the cached body silently captures the old amount and looks successful.
- Pass the same key to the processor so deduplication holds end to end, and generate it once at the originating caller. A key regenerated per attempt leaves every line of idempotency code in place while disabling the mechanism entirely.
- Converge after a crash by querying the authoritative side rather than guessing: a reaper picks up rows in_progress past locked_at plus a bound, asks the processor for that key or the intent's processor_reference, and completes the row from the answer. Set expires_at longer than the caller's full retry schedule and document that a replay after expiry is a new request.
Worked solution 30 min
- Implement the endpoint with the ON CONFLICT DO NOTHING insert committed before the processor call, and a processor stub that counts calls and sleeps 1,500 ms.
- Drive 50 concurrent identical requests through two application instances.
- Repeat with the same key and the amount changed by one minor unit.
- Kill the instance between the stub's response and the local commit, restart, and run the reaper.
Follow-up
- The processor does not honour idempotency keys. What is the end-to-end design now, and what can you no longer promise?
- Two merchants send the same key value. What makes that safe?
- You keep keys for 24 hours at 3,000 requests/s. Size the table and the index, and say what expires them.
Authorisation p99 tripled overnight with no deploy
Payment orchestration serves about 3,000 authorisations per second against a 150 ms p99 budget. Since 02:00, p99 is 460 ms and rising about 8 ms per hour, while p50 is unchanged at 11 ms. There was no deploy, no traffic change and no processor degradation. The hot path does one INSERT into idempotency_key, which has UNIQUE (scope, key), then two UPDATEs on that row: locked_at, then status and response_body. A nightly reconciliation job started at 01:50 and is still running. Give the ordered diagnostic checklist and the cause.
Approach
- Read the shape first. p50 flat with p99 rising and no deploy is a resource or data-volume effect, not a code path, because a code change moves the median too. A tail that climbs monotonically at fixed workload means something monotonically grows.
- Ask what started at 01:50. In PostgreSQL an open transaction holds back the xmin horizon cluster-wide, so autovacuum can reclaim no dead tuple newer than that snapshot. Confirm with pg_stat_activity (state, now() - xact_start, backend_xmin) and with pg_stat_all_tables (n_dead_tup, last_autovacuum) for idempotency_key.
- Connect it to the write pattern. Three writes per key produce up to two dead tuples each, so at 3,000 rps the table sheds roughly 6,000 dead tuples per second. Heap-only tuple updates would keep those out of the index, but only when no indexed column changes and the page has room, and appending response_body grows the tuple enough to force a new page. So the unique index on (scope, key) grows too.
- Explain why only the tail suffers. A larger index means more pages per lookup and a rising fraction of them missing shared_buffers; the median request still hits cache while the tail pays physical I/O. This is exactly the p50-flat, p99-rising signature, and it is worth stating before acting.
- Verify before fixing rather than after. n_dead_tup in the millions and rising, last_autovacuum stale since about 01:50, and pg_relation_size on the unique index measured twice fifteen minutes apart showing growth at constant workload. Index bloat cannot be inferred from row count alone; use the size series or pgstattuple.
- Fix in two moves and prevent separately. End or chunk the long transaction so the reconciliation job commits per batch instead of holding one snapshot over 50M lines, then let autovacuum catch up or run REINDEX CONCURRENTLY. Add a transaction-age alert and a statement timeout on the reporting role. Collapsing the two UPDATEs into one and lowering fillfactor halves dead-tuple production, but that is an optimisation, not the cause.
Follow-up
- The reconciliation job legitimately needs a consistent view of 50M lines. How do you give it one without pinning the xmin horizon?
- Why did p50 not move at all?
- You also run an expires_at cleanup job that DELETEs old idempotency keys. During this incident, does running it help or hurt, and what design avoids the question entirely?
For someone who has spent the last few years shipping features and reading other people's code, and who has not solved a timed problem from a blank file in a long time. Five days rebuild the primitives and the patterns that sit on them, working from invariants rather than remembered solutions, and the last two attach that back to the rest of the loop.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Coding: multi-part problems in a runnable editor
- Solve the reported card game engine coding question in a browser editor: deal a deck to N players, score hands and return the winner, with tests that run.
- Add the reported extensions one at a time (variable player count, custom deck size, pluggable winning rules) and note how much base code each change touched.
- Solve the bank question 'Design a High-Card Game with a Persistent Tie Pot' as a second pass on the same shape.
Deliverable: A runnable card game engine with tests for the base case and each extension, plus a note on which design choice made the extensions cheap.
Practice prompt ↗Practice prompt ↗Worked solution ↗02Coding: data structures the reported questions lean on
- Implement the reported insert, delete and getRandom structure with an array plus a hash map of indices, and test delete on the last element and on a duplicate insert.
- Implement an LRU cache (a bank question) with a hash map and a doubly linked list, testing eviction order after a get.
- Write down the time and space complexity of each operation and the input that would break an incorrect delete.
Deliverable: Two tested data structures with every operation's complexity written above its method.
Practice prompt ↗Practice prompt ↗03Coding: transaction data processing
- Write the reported memory-efficient CSV aggregation as one streaming pass that keeps only a running total per merchant category, then test it on a file with a malformed row.
- Work through the 'Parse settlement amounts into minor units without floating point' worked exercise and run its table tests.
- Solve 'Duplicate Transactions in Time Window' and 'Match Transactions to Loans' from the bank, stating how you handle ties and unmatched records.
Deliverable: Three data-processing solutions that never use floats for money, each with an edge-case test.
Practice prompt ↗Practice prompt ↗04Coding and debugging: hierarchies, settlements and existing code
- Solve 'Top-Level Parent for Loan Request' or 'Aggregate Loans by Ultimate Parent Company', including a cycle or missing-parent case.
- Solve 'Compute Balances and Minimize Settlements' and explain why your settlement count is correct.
- Practise 'Solve the Existing Code Problem' and 'Implement and debug event filtering in Python': for each bug, write the input, expected value and actual value before editing any code.
- Read the 'Authorisation p99 tripled overnight' debugging drill and write the diagnostic order in your own words.
Deliverable: Two solved hierarchy and settlement problems, plus a one-line failing-case statement for every bug you fixed.
Practice prompt ↗Practice prompt ↗Worked solution ↗05System design: ledger and payment flows
- Design the reported real-time ledger: API contract, double-entry schema, idempotent writes, audit history, and how a retried request is prevented from posting twice.
- Work through the 'Capture endpoint that survives concurrent duplicate retries' worked exercise and name the one statement that provides the concurrency control.
- Sketch 'Design Installment-Loan Payment Processing' from the bank, stating what happens when a payment attempt times out.
Deliverable: A written ledger design with API, schema and failure handling, in the shared-document format candidates describe for the system design interview.
Practice prompt ↗Practice prompt ↗06System design and SQL: notifications, decisioning and data stores
- Design the reported payment-due notification service: user preferences, per-channel delivery status for SMS, email and push, queues, and retries that do not send twice.
- Design the reported identity decisioning system: how bureau and fraud-model results are combined, what happens when an external source is slow, and how each decision is recorded for audit.
- Work through the 'outbox relay stopped using its partial index' SQL worked exercise and read the ledger backfill drill, noting the lock each step takes.
Deliverable: Two design write-ups with explicit failure handling, plus notes on the SQL exercise's diagnosis and fix.
Practice prompt ↗Practice prompt ↗07Hiring manager screen and behavioral rounds, then a mock loop
- Prepare the hiring manager walkthrough of one system: components, load figures, what must not be lost, and which decisions were yours.
- Write STAR answers to the equality-versus-equity question, the reported trade-off and feedback prompts, and a cross-team production incident.
- Run a mock loop: one coding problem with an extension, one design prompt written in a shared document, and two behavioral questions, each answered aloud.
Deliverable: A one-page architecture walkthrough, four behavioral stories with measurable results, and notes from the mock loop on where you ran slow.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Candidates say Affirm's behavioral rounds put strong weight on diversity, equity and inclusion as well as ownership and collaboration, including a question on the difference between equality and equity. Prepare stories in STAR form (Situation, Task, Action, Result) that show what you personally did and what changed afterwards. For the inclusion questions, move quickly from a definition to a concrete action you took on a real team.
Describe a time when you had to make a difficult technical trade-off b…
Describe a time when you had to make a difficult technical trade-off between delivering a feature quickly and maintaining long-term code quality.
Approach
- Close with what you would do differently, concretely.
- Name the disagreement and how you resolved it with evidence.
- State the situation in two sentences and spend the rest on the reasoning.
Follow-up
- What did you decide not to do, and why?
- What would you do differently if you ran that again?
Explain your understanding of the difference between equality and equi…
Explain your understanding of the difference between equality and equity in a non-financial, collaborative workplace context.
Approach
- Give the blast radius: what could have broken, and what you measured.
- State the situation in two sentences and spend the rest on the reasoning.
- Close with what you would do differently, concretely.
Follow-up
- How did you know your change caused the improvement?
- What would you do differently if you ran that again?
Estimate a reconciliation rebuild you have never attempted
You are asked how long it takes to replace a reconciliation service matching 30 million settlement lines a day against the ledger, including a bounded fuzzy fallback for netted fees and an ageing model for breaks. You have never built one. Produce an estimate, the range around it, and the two or three unknowns that dominate that range. Then describe a time you estimated unfamiliar work: what you did in the first day to shrink the range, what you committed to publicly, how far off you were, and what you would tell the requester differently now.
Approach
- The probe is whether you can be useful under uncertainty without either refusing to estimate or inventing false precision. Give a number with an explicit range and the basis for both, then immediately name what would move it, rather than asking for two weeks of discovery first.
- Decompose into parts with different uncertainty profiles. The hash join on (external_reference, amount_minor, currency, business_date) over 30 million lines is well understood engineering and estimates tightly; the fuzzy fallback for netted and fee-adjusted lines does not, because its scope is defined by whatever the files actually contain; the ageing and break workflow is mostly operations-facing surface area, which estimates by counting screens and states.
- Name the dominating unknowns concretely: how many distinct file formats and cutoff conventions the sources use, what fraction of lines are netted rather than itemised, and whether business_date is derivable from any field in the file or must be reconstructed from the cutoff rule. Each is a factor on the fuzzy path, not a percentage on the whole.
- Describe the first-day range-shrinking work, which is the part that separates strong from generic: take one real file, count distinct formats, measure the netted fraction, and attempt the exact join on a single day of postings to see what the residual actually is. One day of that typically converts a 3x range into something near 1.5x.
- Commit in a form that survives being wrong: a range plus a checkpoint date at which you will replace it with a narrower one, and an explicit statement of what you will cut first if the range turns out to be optimistic.
- In the retrospective half, give the real numbers: the estimate, the actual, and the specific thing that consumed the difference. Answers that were within 10 percent are less informative than answers that were 2x off for a nameable reason.
Follow-up
- The requester wants one number, not a range, for a board deadline. What do you give them?
- Your one-day probe finds 40 percent netted lines instead of the 5 percent you assumed. What changes in the plan, not just the estimate?
- What do you cut first if you are at the deadline and the fuzzy fallback is not done?
- 01
Explain your understanding of the difference between equality and equity in a non-financial, collaborative workplace context.
- 02
Tell me about a time when you noticed an inequitable process or dynamic on your team, and what actions you took to address it.
- 03
Describe a time when you had to make a difficult technical trade-off between delivering a feature quickly and maintaining long-term code quality.
- 04
Talk about a project where you had to collaborate with stakeholders across multiple teams (such as Product, Risk, or Operations) to resolve a critical production issue.
- 05
Describe a situation where you received constructive feedback that significantly changed your approach to a technical problem or team dynamic.
- 06
Describe a past project that failed to meet its deadline or technical goals. What went wrong, and how did you manage the aftermath?
Is this an official Affirm interview guide?
No. It is PracHub's own research and practice material for the Software Engineer role at Affirm. The rounds and questions reflect what candidates have reported, not a process Affirm has published, and they change over time. Confirm the current format with your recruiter.
PracHub interview research ↗How long does the interview process typically take?
Candidates report about three to five weeks from the recruiter screen to a decision, across four stages: recruiter pre-screen, initial technical screen, hiring manager screen and virtual onsite. Your timeline can differ depending on scheduling and the team.
PracHub interview research ↗What programming languages can I use in the technical interviews?
Candidates describe the technical interviews as generally language-agnostic, using languages such as Python, Java, Kotlin or C++. Pick the one whose standard library you know best, because you will write runnable code, and confirm the accepted languages with your recruiter.
PracHub interview research ↗Does my code have to run, or is pseudo-code acceptable?
Candidates report that coding happens in a live, browser-based IDE where the code has to compile and run against test cases. Practise in a plain browser editor, run the code early on small inputs, and avoid leaving sections as pseudo-code.
PracHub Software Engineer practice ↗What should I expect from the coding problems' structure?
Candidates report that problems start with a base requirement and add one or two extensions. The reported card game engine question, for example, grows to handle variable player counts, custom deck sizes and changing win rules. Finish a correct base version quickly so you have time to reach the extension.
PracHub Software Engineer practice ↗Is the initial technical screen an online assessment or live coding?
Candidates report it can be either, and that some teams offer a take-home assessment instead. Ask your recruiter which format you will get and prepare in a matching environment. Candidates also report that the virtual onsite always includes live, interactive coding.
PracHub Software Engineer practice ↗How much weight do the behavioral rounds put on DEI?
Candidate reports describe diversity, equity and inclusion as a strong focus of the behavioral rounds. The reported questions include explaining the difference between equality and equity and describing how you have addressed an inequitable process on a team. Prepare a clear definition and at least one concrete example of something you changed.
PracHub interview research ↗What is the format of the system design interview?
Candidates describe a collaborative discussion over video, on a virtual whiteboard or in a shared document where you may write API contracts and schemas. Practise writing a design out as text (endpoints, schema, failure handling) as well as drawing it, and focus on data modelling and trade-offs.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-24 - 02PracHub Software Engineer practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-24 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-24