This guide covers what a Software Engineer at Chime is expected to do and how to prepare for the interview.
Recruiter Screening
reportedHalf of this call is the part candidates treat as small talk: start date, notice period, work authorisation and its timing, location and time zone, on-call, and the number. Those are what kill offers late, after several engineers have each spent a day. Surfacing a hard constraint now costs you nothing and occasionally buys you something, since a loop compressed to fit a competing deadline can usually only be arranged if it is asked for early. The common failure is deflecting the compensation question twice, then discovering at offer stage that the band never reached your number.
What to demonstrate
- Whether your hard constraints are compatible with the role before a loop gets booked: earliest start, notice period, what authorisation you hold and when it needs action, days on site, willingness to carry a pager
- Whether you give a compensation range with something behind it, such as current total compensation or a competing timeline, rather than leaving the band untested
- Whether your stated timeline is real, since a competing deadline raised now is something scheduling can sometimes work around and the same deadline raised at offer stage usually is not
How to prepare
- Write each constraint down in one line before the call and state them as facts rather than negotiating them live under a question you were not expecting
- Set your range from two or three current data points for that level and location, and name the structure you are quoting in, so the number is comparable to the one they are holding
- If another process is running, say where it stands and by when, and ask directly whether this loop can be scheduled inside that window
Technical Screen
reportedThe same problem is scored by two different mechanisms depending on the format, and preparing for one does not cover the other. With a person watching, partial progress is visible and a hint is a correction you can absorb; silence is the expensive failure, because nobody can read a half-written function. With an automated grader there is no partial credit for what you were about to do, nobody to ask, and the worked examples in the prompt are the entire specification. Read them as a contract, down to whether an empty result should be an empty list or no output at all.
What to demonstrate
- In a live session, whether your commentary tracks what your hands are doing, and whether a hint redirects you or gets defended against
- In an automated one, whether you cover the cases the examples do not show, since the hidden cases are where the score moves
- Whether you manage the clock on purpose: abandoning an approach that is not converging while there is still time to write something simpler that finishes
How to prepare
- Have someone hand you a problem and feed you one deliberately wrong hint. Practise testing it against a concrete case instead of accepting or rejecting it on authority.
- Do one timed run a week in a plain browser editor with autocomplete, linting and your own snippets switched off, which is closer to what these environments give you
- For the automated format, write the harness before the solution: a main that feeds the worked examples plus an empty and a single-element case and prints expected against actual, so a wrong submission is caught by you first
Virtual Onsite
reportedNobody in the room with you decides this. Interviewers typically write their rounds up separately, often before seeing anyone else's, and the outcome is settled later from those write-ups. A split panel gets resolved by whichever note carries specific evidence, so what you want out of each room is one concrete thing that person could write down: a bug you caught yourself, a trade-off you named, a decision you owned. The rest is arithmetic. The project you describe in a behavioural conversation is often the same system you sketched an hour earlier, and the two accounts have to agree.
What to demonstrate
- Whether the scale, team size and timeline you attach to a project hold steady when that project resurfaces in a different round
- Whether each interviewer leaves with a specific thing to cite rather than a general impression of competence
- Whether a trade-off you defended in one round survives a challenge in another, instead of being quietly swapped for the answer the new interviewer seemed to want
- Whether a question you have already answered earlier in the day gets the same answer at the same depth, without visible impatience
How to prepare
- Write a one-page sheet per project fixing the figures you will quote — request volume, data size, team size, elapsed time, what broke — and say them aloud from the sheet until they come out identical every time
- For each round on the schedule, decide in advance the one sentence you want in that person's notes, then check in a mock that you said it outright instead of leaving it to be inferred
- Have someone ask you the same project question twice, an hour apart, and diff the two answers for numbers that moved or a trade-off that reversed
Cross-Functional Interviews
reportedAn unlabelled round is first an information problem, and the cheapest information is free. Whoever schedules it can usually tell you how long it runs, who will be in the room and what they work on, whether you will be writing code and in what environment, and whether anything is being sent beforehand. Ask in writing so the answer is on record, then prepare for the two or three formats those answers still leave open instead of betting on one. What separates a strong candidate is not guessing right; it is having an opening that works whichever one it turns out to be.
What to demonstrate
- Whether you can start work from an ambiguous brief, since tolerating a vague scope without stalling is the same thing the job asks for
- Whether the questions you asked beforehand were ones that change your preparation, such as duration, medium and who is joining, rather than ones whose answers you could not have acted on
- Whether you adapt when the round turns out to be something other than what you were told, instead of spending the first ten minutes visibly recalibrating
How to prepare
- Send one short scheduling message asking four things: how long, who is joining and what they work on, whether you will be writing code and where, and whether to prepare anything in advance. Treat a vague reply as real information, since it means the round is loosely structured and you will be shaping it yourself.
- Write one opening that works in any of the formats still open: restate in your own words what you have been asked to do, then ask which of two directions is more useful to them. Say it aloud until it stops sounding recited.
- Set up for the two most likely formats before the call starts, with a blank editor in the language you would choose and a shared document you can type into, so a format surprise costs you nothing in the first minutes
5 candidate reports. Individual accounts describe a particular role and hiring cycle.
Chime Software Engineer interview: React and JavaScript component task
I faced a front-end-focused technical challenge centered on building React/JavaScript components and implementing the data logic behind them. In the first round, I worked in a shared virtual environment to create components in a React/JS app. The components had to render data, and I had to implement the logic that processed the data before rendering. The UI scaffold was already provided. The diff…
Read full experienceChime Software Engineer interview: basic SQL without window functions
I interviewed for a data engineering-oriented SDE role where the questions were unusually basic. They focused on my past work and simpler algorithm and SQL knowledge. The recruiter screen covered my background and how it matched the role. The technical portion covered SQL and algorithms. The SQL questions were very basic and explicitly did not include window functions. The algorithm problems lean…
Read full experienceChime Software Engineer interview with coding, system design, and manager rounds
I went through a fairly standard Software Engineer process at Chime, with a recruiter conversation followed by coding, system design, manager, and product-style rounds. The difficulty was mostly average, although I had mixed feelings about how clearly some interviewers ran their sessions. The recruiter screen was a phone call about my background, the role, and expectations. The technical intervie…
Read full experienceChime Software Engineer interview with a complex game-design question
I had a difficult start when a phone screen turned into an extremely complex game-design question that I couldn't get clear enough to understand. The recruiter call came first. During the technical phone screen, I was asked to design a game with complicated rules. The interviewer struggled to explain how the game worked, and the prompt felt unrealistic without prior knowledge. I didn't make it to…
Read full experienceChime Software Engineer Interview Experience — Two Back-to-Back Phone Screens, Rejected for Pipeline Fit
A recruiter reached out. I looked up the company online and felt the culture seemed okay, so I went ahead and interviewed. The recruiter call was just casual chatting, and we quickly landed on an interview time. I hadn't grinded LeetCode in a long time so I was a little worried, so I scheduled the phone screen slot for two weeks out and started grinding through interview experience posts. In the…
Read full experiencePracHub editorial advice for the preparation topics above.
Treating money as a decimal with two places
ISO 4217 exponents are 0 for currencies such as JPY and KRW, 2 for most, and 3 for BHD, KWD, JOD, OMR and TND, so a hard-coded multiply-by-100 is off by a factor of 100 or 10 depending on the currency, in opposite directions. Floating point is worse: IEEE 754 binary64 cannot represent 0.1 exactly, so repeated accrual accumulates drift that appears as a handful of minor units in the daily reconciliation and then gets 'fixed' by widening the match tolerance, which is how a genuine break becomes invisible. The only forms that survive a reconciliation are integer minor units with the exponent carried alongside the currency code, or a fixed-scale decimal type with exactly one documented rounding point.
Retrying a charge after a timeout
A timeout is not a failure; it is an unknown outcome, and the request may have been processed in full with only the response lost. Re-sending it without an idempotency key that the processor itself honours produces a duplicate charge, which is a customer-visible incident and usually a dispute. The correct handling is to treat the state as unknown, query the processor for that key or client reference, and only then decide. The mechanism also depends on the key being generated once by the caller and reused across every attempt — generating a fresh key per retry turns the whole scheme into a no-op while leaving all the code that appears to implement it in place.
Comparing floating-point values for equality, or holding money in them
Binary floating point cannot represent 0.1 exactly, so repeated addition drifts and an equality check fails on values that are mathematically equal. Store currency as integer minor units or a decimal type, and compare floats against a tolerance you chose for a stated reason.
Assuming the bug is in the framework
Suspect your own code first: read the stack trace top to bottom, check which versions are actually installed rather than which ones you believe are, and reproduce in isolation before blaming a library that thousands of people run daily. When the fault really is upstream, you need that minimal reproduction to say so credibly anyway.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Build a custom component for a frontend application using React to ren…
Build a custom component for a frontend application using React to render dynamically formatted ledger items while managing local state and asynchronous API calls.
Approach
- Restate the input: its shape, its size, and what is guaranteed about it.
- State the target complexity and say which constraint rules the naive version out.
- Name the brute-force solution and its complexity before improving on it.
Follow-up
- What is the worst case, and how likely is it on real data?
- Which test case would catch an off-by-one here?
Write a program to parse, aggregate, and analyze transaction log strin…
Write a program to parse, aggregate, and analyze transaction log strings to detect potential financial anomalies across user accounts.
Approach
- Name the brute-force solution and its complexity before improving on it.
- Restate the input: its shape, its size, and what is guaranteed about it.
- State the target complexity and say which constraint rules the naive version out.
Follow-up
- What is the worst case, and how likely is it on real data?
- Which test case would catch an off-by-one here?
Implement an in-memory transactional key-value data store that support…
Implement an in-memory transactional key-value data store that supports nested operations, rollback, and commit capabilities.
Approach
- Name the brute-force solution and its complexity before improving on it.
- Restate the input: its shape, its size, and what is guaranteed about it.
- Walk one small example through your approach before writing the whole thing.
Follow-up
- How does this change if the input no longer fits in memory?
- Which test case would catch an off-by-one here?
Design and code a game logic module with complex rules, handling state…
Design and code a game logic module with complex rules, handling state updates, boundary checks, and win conditions cleanly.
Approach
- Choose the data structure from the access pattern, not from familiarity.
- Restate the input: its shape, its size, and what is guaranteed about it.
- State the target complexity and say which constraint rules the naive version out.
Follow-up
- How does this change if the input no longer fits in memory?
- Which test case would catch an off-by-one here?
Match a settlement file to ledger postings under duplicate keys
You have one business date of ledger postings (about 12 million rows: transaction_id, source_id, amount_minor, currency, business_date) and the processor's settlement file (about 12 million lines: settlement_line_id, external_reference, amount_minor, currency, business_date), where source_id carries the external reference. Match one-to-one on (external_reference, amount_minor, currency, business_date). Duplicate keys occur legitimately — the same amount can appear twice. Emit every unmatched item classified ledger_only, file_only or duplicate_match, in linear expected time. Then say what you do when neither side fits in memory.
Approach
- Build the smaller side into
key -> deque of row ids, neverkey -> row id. A duplicate key is data, not corruption; a single-row map drops one of a legitimate pair and the break report then shows afile_onlythat does not exist. - Probe the larger side once, carrying one extra bit per bucket: whether that bucket was ever hit. Pop from the bucket on a match and set the bit. An absent key is a probe-side-only row. A present-but-empty bucket means the probe side holds more copies than the build side — surplus, so
duplicate_match. After the pass, a leftover non-empty bucket that was never hit is build-side-only; one that was hit is build-side surplus, so alsoduplicate_match. - That bit is what makes the classification a function of the per-key counts rather than of which side you happened to build. For a key with L ledger and F file copies: min(L, F) match, and the |L - F| surplus rows are
duplicate_matchtagged with the side that is over, degenerating toledger_onlyorfile_onlyexactly when min(L, F) is 0. Without the bit, surplus is only observable as a present-but-empty bucket, which can only ever happen on the probe side — so the same input reports different break classes depending on build order, and the smaller-side heuristic in bullet one silently decides which. - A genuine amount difference does not surface as
amount_mismatchhere, because the amount is inside the key — it surfaces as aledger_onlyand afile_onlysharing a reference. Promote those in a second, separate pass keyed on reference alone, recording signeddelta_minoras ledger minus file. Keep that promotion out of the exact pass. - Cost: O(N+M) expected time and O(min(N,M)) memory; the hit bit packs into the bucket header and changes neither bound. The constant is the hash map, roughly 60 to 100 bytes per entry in most runtimes, so 12 million rows is order 1 GB — measure it rather than assert it.
- When neither side fits, use a grace hash join: partition both sides with the same hash function into P spill files so a key lands in the same partition on both sides, then join partition by partition in memory. Cost is two extra sequential passes; skew inside one partition is the failure mode, handled by re-partitioning that partition under a second hash.
- Sort-merge is the alternative at O(N log N + M log M) with external sort, and it wins when the file already arrives sorted by reference or the output must be ordered. It also gets the surplus classification for free, since a merge sees L and F side by side. Whichever you pick, do not widen the amount comparison to make breaks disappear: a tolerance wide enough to absorb rounding is wide enough to absorb a real loss.
Worked solution 30 min
- Build the bucket map over the ledger side and assert that total bucket length equals row count — that single assertion catches the single-row-map bug immediately.
- Probe with the file side, popping on match and setting each touched bucket's hit bit, then drain the leftovers: hit means surplus, never hit means absent on the probe side.
- Fixture of 11 ledger rows and 11 file lines: 7 keys matching one-to-one, one key where the ledger has 2 copies and the file has 3, 2 ledger-only rows and 1 file-only line.
- Assert the counting identity
2*matched + ledger_only + file_only + duplicate_match == N + M; it holds under either build order, so it is necessary but not sufficient — it does not catch a misclassification that moves a row between the three break classes. - Swap build and probe sides and re-run, asserting the four counts and the surplus side are identical, not merely mirrored. Then delete the hit bit and re-run the swap to watch the surplus file copy get reclassified as an absence.
Follow-up
- The file nets three fee lines into one batch total. Which pass catches that, and what is its stopping rule?
- The processor's business date sits one cutoff behind yours for forty minutes of traffic. What does that do to the exact join, and what does it do to break ages?
- The same break recurs on the next run. Why must it link to the existing
reconciliation_breakrow rather than open a second one?
Add and backfill business_date on a live ledger table
ledger_entry holds 4 billion rows, is append-only, takes 10,000 inserts per second, and every reporting query currently derives the business date as posted_at::date. You must add business_date date NOT NULL, populated from the cutoff rule (17:00 in the account's own timezone), backfilled across all history, indexed, and cut over, with no write downtime and no long-held lock. Give the ordered migration steps with the lock each one takes, how you make the backfill restartable and throttled, and how you retire the old expression safely.
Approach
- Add the column nullable and with no default. ALTER TABLE ... ADD COLUMN takes ACCESS EXCLUSIVE but is a catalogue-only change held for microseconds. The hazard is the lock queue, not the statement: a blocked ALTER waits behind one long reader holding ACCESS SHARE, and every query arriving afterwards queues behind the ALTER's pending ACCESS EXCLUSIVE, so set lock_timeout to a couple of seconds and retry rather than letting a metadata change take the table down.
- Deploy the write path before the backfill, so new inserts populate business_date from the cutoff rule while reads stay on the old expression. The backfill then chases a closed set with a fixed upper bound instead of a moving target.
- Backfill in bounded batches keyed by entry_id range, on the order of 50,000 rows per statement, committing between batches and recording the high-water mark in its own table so a killed run resumes instead of restarting. WHERE business_date IS NULL makes each batch idempotent, and the pacing is set by replica lag and dead-tuple growth rather than CPU, since each UPDATE writes a new row version and the WAL volume is proportional to the rows touched.
- Install the constraint without a blocking scan: ALTER TABLE ... ADD CONSTRAINT ck_business_date CHECK (business_date IS NOT NULL) NOT VALID takes a brief ACCESS EXCLUSIVE and scans nothing, then VALIDATE CONSTRAINT takes SHARE UPDATE EXCLUSIVE and runs alongside reads and writes. On PostgreSQL 12 and later, SET NOT NULL can then use the validated CHECK and skip its own full scan; on 11 and earlier it always scans, so the CHECK is the migration on those versions.
- Build the index with CREATE INDEX CONCURRENTLY, which avoids ACCESS EXCLUSIVE at the cost of two table passes, cannot run inside a transaction block, and on failure leaves an INVALID index that must be dropped and rebuilt rather than reused.
- Cut over behind a flag: run the new and old expressions side by side for one reporting cycle and compare totals per day, since the cutoff rule will legitimately move entries near 17:00 across the boundary. Only once they reconcile do you retire the posted_at::date expression index, and you keep posted_at as the ordering key rather than repurposing it.
Worked solution 40 min
- Write the migration as numbered SQL statements, annotating each with its lock mode and expected duration, and set lock_timeout plus a retry around every ALTER.
- Rehearse on a 10-million-row copy under a concurrent insert load, measuring per-batch duration, WAL generated and replica lag.
- Kill the backfill at a random point, restart it, and confirm it resumes from the high-water mark.
- Add the CHECK as NOT VALID, VALIDATE it while inserts continue, then SET NOT NULL and build the index CONCURRENTLY.
- Run old and new date expressions side by side for one cycle and diff the daily totals before dropping the old expression index.
Follow-up
- Reporting now wants ledger_entry partitioned by business_date. Why can this not be another ALTER, and what is the migration instead?
- Nightly totals move for the days around the cutoff change. How do you tell a correct restatement from a backfill bug?
- The backfill is halfway done when a replica falls 20 minutes behind. What do you throttle, and what do you refuse to throttle?
Make a charge endpoint safe under concurrent duplicate retries
idempotency_key holds id, scope, key, request_fingerprint (SHA-256 over the canonicalised body), status (in_progress, completed, failed), response_status, response_body, locked_at, completed_at, expires_at, created_at. Fifty identical create-payment requests carrying the same scope and key reach four application instances inside the same 20 ms. Give the DDL constraint and the exact statements the handler runs so that exactly one payment_intent is created and all fifty callers receive the same response body. State what you return when that key arrives with a different fingerprint, and what an arrival after expires_at means.
Approach
- Put the concurrency control in the schema: UNIQUE (scope, key). A SELECT-then-INSERT cannot work because both transactions can read nothing before either commits, so the check passes twice and the constraint then surfaces as an error on a payment that succeeded.
- Claim the key with INSERT ... ON CONFLICT (scope, key) DO NOTHING RETURNING id. A conflict returns zero rows rather than the existing row, so branch on rowcount: the winner proceeds, the loser reads the stored row.
- Keep that path on READ COMMITTED deliberately. The loser's follow-up SELECT takes a fresh statement snapshot and therefore sees the winner's committed row; under REPEATABLE READ the transaction snapshot predates that commit, the row stays invisible and the loser concludes the key does not exist.
- Split the work across two transactions because the processor call cannot sit inside one: commit the in_progress row with locked_at first so losers can see a claim, perform the effect, then write payment_intent plus status=completed with response_status and response_body in a single second transaction.
- Handle the crash window explicitly: a row stuck in_progress past its lease is an unknown outcome, not a failure, so the reaper queries the processor for that key before deciding. A loser that sees in_progress returns 409 and retries rather than repeating the effect.
- Compare request_fingerprint before replaying anything. Same key with a different body is 409, never the cached response, because replaying confirms a payment the caller did not request; and set expires_at beyond the client's and the processor's maximum retry horizon, since a replay after it is a genuinely new request.
Follow-up
- The handler dies after the processor call and before the local commit. What does the next retry with that key observe, and how does the system converge on exactly one charge?
- Does the downstream processor honour an idempotency key of its own? Who mints it, and what breaks if a fresh one is generated per attempt?
- How do you purge rows past expires_at without the delete contending with the insert path?
Design a high-throughput early direct deposit system that ingests clea…
Design a high-throughput early direct deposit system that ingests clearinghouse files, validates funds, and credits accounts within milliseconds.
Approach
- State the consistency you need, and where you are willing to be stale.
- Fix the scope first: who calls this, how often, and what they do when it fails.
- Name the read and write paths separately; they rarely have the same bottleneck.
Follow-up
- How does this behave when that dependency is down for an hour?
- What would you drop to keep the system up under load?
Design a credit card transaction processing platform capable of real-t…
Design a credit card transaction processing platform capable of real-time balance validation, ledger updating, and fraud check orchestration.
Approach
- Fix the scope first: who calls this, how often, and what they do when it fails.
- State the consistency you need, and where you are willing to be stale.
- Name the failure you are designing for, then the recovery path.
Follow-up
- What would you drop to keep the system up under load?
- What breaks first when traffic grows ten times?
Capture endpoint that survives concurrent duplicate retries
POST /payments/{intent_id}/capture carries an Idempotency-Key header and an amount. Callers retry on timeout, and two retries can land concurrently on different instances. You have idempotency_key(id, scope, key, UNIQUE(scope,key), request_fingerprint bytea, status in_progress|completed|failed, response_status, response_body jsonb, locked_at, completed_at, expires_at). The processor capture takes 200 to 2,000 ms. Design the path so exactly one capture reaches the processor, every duplicate receives the identical response, and a crash between the processor call and your commit converges. Name the single statement that is the concurrency control.
Approach
- The concurrency control is INSERT INTO idempotency_key (...) VALUES (...) ON CONFLICT (scope, key) DO NOTHING RETURNING id, and it must commit before the processor call. A returned row means you own the effect; zero rows means you lost and must read the winner's result. A SELECT-then-INSERT check cannot substitute: two requests both read 'absent' and both proceed, and the window is exactly the concurrency you are defending against.
- Commit the in_progress row in its own short transaction. A concurrent duplicate insert blocks on an in-flight conflicting insert until that transaction ends, so wrapping the 2 s processor call in the same transaction turns every duplicate into a 2 s lock wait and a retry storm into pool exhaustion.
- Define loser behaviour per status rather than uniformly: completed replays response_status and response_body unchanged; in_progress returns 409 with Retry-After and performs nothing; failed splits by cause, since a terminal processor decline should replay but a transport failure should let the key be retried. Getting this wrong in the safe direction (replay a decline) is better than returning an error for a capture that succeeded.
- Fingerprint the canonicalised body with SHA-256 and compare on every hit. Same key with a different amount is a client bug and must return 409 or 422, never the cached response, because returning the cached body silently captures the old amount and looks successful.
- Pass the same key to the processor so deduplication holds end to end, and generate it once at the originating caller. A key regenerated per attempt leaves every line of idempotency code in place while disabling the mechanism entirely.
- Converge after a crash by querying the authoritative side rather than guessing: a reaper picks up rows in_progress past locked_at plus a bound, asks the processor for that key or the intent's processor_reference, and completes the row from the answer. Set expires_at longer than the caller's full retry schedule and document that a replay after expiry is a new request.
Worked solution 30 min
- Implement the endpoint with the ON CONFLICT DO NOTHING insert committed before the processor call, and a processor stub that counts calls and sleeps 1,500 ms.
- Drive 50 concurrent identical requests through two application instances.
- Repeat with the same key and the amount changed by one minor unit.
- Kill the instance between the stub's response and the local commit, restart, and run the reaper.
Follow-up
- The processor does not honour idempotency keys. What is the end-to-end design now, and what can you no longer promise?
- Two merchants send the same key value. What makes that safe?
- You keep keys for 24 hours at 3,000 requests/s. Size the table and the index, and say what expires them.
Authorisation p99 tripled overnight with no deploy
Payment orchestration serves about 3,000 authorisations per second against a 150 ms p99 budget. Since 02:00, p99 is 460 ms and rising about 8 ms per hour, while p50 is unchanged at 11 ms. There was no deploy, no traffic change and no processor degradation. The hot path does one INSERT into idempotency_key, which has UNIQUE (scope, key), then two UPDATEs on that row: locked_at, then status and response_body. A nightly reconciliation job started at 01:50 and is still running. Give the ordered diagnostic checklist and the cause.
Approach
- Read the shape first. p50 flat with p99 rising and no deploy is a resource or data-volume effect, not a code path, because a code change moves the median too. A tail that climbs monotonically at fixed workload means something monotonically grows.
- Ask what started at 01:50. In PostgreSQL an open transaction holds back the xmin horizon cluster-wide, so autovacuum can reclaim no dead tuple newer than that snapshot. Confirm with pg_stat_activity (state, now() - xact_start, backend_xmin) and with pg_stat_all_tables (n_dead_tup, last_autovacuum) for idempotency_key.
- Connect it to the write pattern. Three writes per key produce up to two dead tuples each, so at 3,000 rps the table sheds roughly 6,000 dead tuples per second. Heap-only tuple updates would keep those out of the index, but only when no indexed column changes and the page has room, and appending response_body grows the tuple enough to force a new page. So the unique index on (scope, key) grows too.
- Explain why only the tail suffers. A larger index means more pages per lookup and a rising fraction of them missing shared_buffers; the median request still hits cache while the tail pays physical I/O. This is exactly the p50-flat, p99-rising signature, and it is worth stating before acting.
- Verify before fixing rather than after. n_dead_tup in the millions and rising, last_autovacuum stale since about 01:50, and pg_relation_size on the unique index measured twice fifteen minutes apart showing growth at constant workload. Index bloat cannot be inferred from row count alone; use the size series or pgstattuple.
- Fix in two moves and prevent separately. End or chunk the long transaction so the reconciliation job commits per batch instead of holding one snapshot over 50M lines, then let autovacuum catch up or run REINDEX CONCURRENTLY. Add a transaction-age alert and a statement timeout on the reporting role. Collapsing the two UPDATEs into one and lowering fillfactor halves dead-tuple production, but that is an optimisation, not the cause.
Follow-up
- The reconciliation job legitimately needs a consistent view of 50M lines. How do you give it one without pinning the xmin horizon?
- Why did p50 not move at all?
- You also run an expires_at cleanup job that DELETEs old idempotency keys. During this incident, does running it help or hurt, and what design avoids the question entirely?
For a candidate senior enough that the loop turns on design and judgement rather than on whether the coding round gets finished. Five days build one system properly and then stress it; coding gets a single maintenance day, on the assumption that the risk at this level is an unexamined tradeoff rather than a missed algorithm.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Numbers before diagrams
- Build your own reference card of the figures you will re-derive all week: bytes for a realistic record, requests per second implied by a given daily active count, and the storage that a year at a given write rate produces. Derive each one rather than copying it, because the derivation is what survives a follow-up.
- Turn one product statement into capacity requirements. From ten million daily users at four writes and forty reads each, state the peak-to-average factor you are assuming and why, then produce peak write QPS, peak read QPS and a year of storage.
- Write the two numbers whose order of magnitude changes the design, the read-to-write ratio and the working-set size against memory per node, and state the threshold at which each one flips your answer.
Deliverable: A one-page numbers card and one worked capacity estimate with every assumption written down.
Practice prompt ↗Practice prompt ↗Worked solution ↗02One system, from requirements to schema
- Spend the first ten minutes producing only functional requirements, non-functional targets with numbers attached, a p99 latency, a durability expectation, a consistency requirement, and an explicit out-of-scope list.
- Define the interface before the boxes: the three or four endpoints, their parameters, what each returns, and which of them are idempotent.
- Write the data model, then write the single access pattern that justifies it, and state what the schema would have to become if the dominant access pattern were the other one.
Deliverable: One design carried to endpoint-and-schema depth, with non-functional targets expressed as numbers and a written out-of-scope list.
Practice prompt ↗Practice prompt ↗03The consistency you are actually buying
- Write out what a client sees under asynchronous replication when its write commits on the leader and its next read is served by a lagging follower, then write the two fixes, pinning that session's reads to the leader for a bounded window or carrying a version token the replica must reach, and the cost of each.
- Work the quorum arithmetic on paper for N of three with W and R of two, and separate what R + W > N does guarantee, that any read set intersects any write set, from what it does not: on its own it is not linearizability, and a sloppy quorum that accepts writes on nodes outside the preference list breaks even the intersection.
- Take two storage choices with different defaults, a single-leader relational store committing synchronously and a quorum-replicated store that converges eventually, and write the specific product behaviour that would be wrong under each, rather than a general statement about which is stronger.
Deliverable: A page separating what quorum overlap guarantees from what it does not, with one concrete product misbehaviour attached to each gap.
Practice prompt ↗Practice prompt ↗04Failure is the design
- For one write path, work through the case where the client times out after the server has already committed, then design the idempotency key: who generates it, how long it is retained, and what the duplicate request returns.
- Express the retry policy as parameters rather than as a word: maximum attempts, base delay, backoff factor, jitter, and which error classes are retried at all. Then state why retrying a non-idempotent write without a key is a correctness bug and not merely waste.
- Compute the fan-out effect on tail latency. If a request waits on ten backends and each independently exceeds its p99 one percent of the time, the chance at least one is slow is 1 - 0.99^10, about ten percent. Then write why independence is the optimistic assumption and what correlates them in practice.
- Name the backpressure mechanism for one queue or one dependency in the design, a bounded queue with shedding or a concurrency limit, and write what the caller is told when it engages.
Deliverable: One write path with an idempotency design, a parameterised retry policy, and a written tail-latency calculation with its assumption named.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Scaling the hot path
- Choose cache-aside or write-through for one read path and write the staleness window each produces, then name the invalidation event and what the system does when that event is lost.
- Design against the stampede: either coalesce requests so only one recomputes a missing key, or refresh early with jittered expiry, and write why identical TTLs on keys populated in the same moment produce a synchronised expiry and a thundering herd.
- Shard one table by a key you choose, then answer the two questions that break the choice: which queries now require a scatter-gather, and what happens to the distribution when one tenant is a hundred times larger than the median.
- Write the cost of adding a node under plain modulo placement, where nearly every key moves, against consistent hashing, where roughly one key in n+1 moves, and state what virtual nodes are for.
Deliverable: A caching and sharding decision for one path, each with its failure mode and its rebalancing cost written beside it.
Practice prompt ↗Practice prompt ↗06Keep the coding hand in, at the bar that applies to you
- Solve one medium problem in thirty minutes, then spend twenty more making it production-shaped: named invariants, validation at the boundary, and errors that distinguish a caller mistake from an internal fault.
- Write the tests you would require of a colleague's version of that function: one for empty input, one for the boundary, and one for the case the implementation is most likely to get wrong.
- Read a piece of your own code from six months ago and write the change you would ask for, phrased as you would actually phrase it in review.
Deliverable: One problem hardened to review standard, with its test list and one written review comment.
Practice prompt ↗Practice prompt ↗07Defend it while being interrupted
- Run a forty-five-minute design mock with an interviewer briefed to change a requirement halfway, a tenfold traffic increase or a new strict consistency requirement, and to push on one number you estimated.
- Rehearse the two sentences a senior loop is listening for: naming the tradeoff you are choosing against and why, and saying what you would measure to learn that the choice was wrong.
- Prepare the design you regret: a real decision, the constraint that produced it, what it cost, and what you changed afterwards.
Deliverable: Mock notes recording how the design changed under the new requirement, plus a written account of one regretted decision.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Nobody is scoring your stamina at three in the morning. What carries weight is which signal told you something was wrong, what you measured before touching anything, what you rolled back versus what you fixed forward, and why you picked one. 'We restarted it and it went away' is a story about not knowing.
How do you approach mentoring junior software engineers and fostering …
How do you approach mentoring junior software engineers and fostering engineering excellence across cross-functional teams?
Approach
- Close with what you would do differently, concretely.
- Name the disagreement and how you resolved it with evidence.
- State the situation in two sentences and spend the rest on the reasoning.
Follow-up
- How did you know your change caused the improvement?
- What did you decide not to do, and why?
Describe a time when you had an architectural disagreement with your m…
Describe a time when you had an architectural disagreement with your manager or team lead. How did you handle the discussion, and what was the outcome?
Approach
- Name the disagreement and how you resolved it with evidence.
- Close with what you would do differently, concretely.
- Give the blast radius: what could have broken, and what you measured.
Follow-up
- What would you do differently if you ran that again?
- What did you decide not to do, and why?
Unblock an engineer on double-posted interest accrual
An engineer two years into their career has a nightly accrual job that double-posts interest for some accounts whenever the batch is partially re-run after a failure. They have spent two days on it and are now rewriting the batch runner. You have 30 minutes. Describe how you have unblocked someone without taking the keyboard: the question you asked first, how you chose between handing over the answer and handing over the method, what you left behind so the next person does not get stuck here, and how you knew they were unblocked rather than deferring to you.
Approach
- The probe is whether you grow people or absorb their work. Open with the diagnostic question rather than the solution: ask what identifies one unit of work, because the answer reveals immediately that accrual is keyed by (account_id, accrual_date) and that the job has no uniqueness on it.
- Redirect from the runner to the write. The rewrite is aimed at never re-running, which is unachievable; the property needed is that re-running posts nothing new, enforced by a unique index on (account_id, accrual_date) or by passing the same idempotency key to the ledger posting operation so the second attempt is a no-op rather than a second transaction.
- Choose deliberately between answer and method and say why. Two days in and blocked on the wrong layer is usually the moment to hand over the framing (restartable at account granularity, idempotent per unit) and let them write the code, because the lesson is the framing and the code is the easy part.
- Leave an artefact, not a conversation: a test that re-runs one account twice and asserts one posting, plus two lines in the runbook stating that per-account work must be idempotent because the batch is always partially re-run.
- Check that they are unblocked by asking them to predict the failure that the fix does not cover, such as a mid-run rate change producing two different correct amounts for the same key. If they can find the next edge themselves, they own it; if they ask you to confirm each step, they are deferring and you have hidden the block rather than removed it.
- Say what you deliberately did not do. Not fixing it yourself before the standup is the whole exercise, and a strong answer names the pressure it resisted.
Follow-up
- The unique index rejects the re-run, but the first run posted the wrong amount. How should the job behave now?
- How do you tell whether you taught them or just unblocked them, a month later?
- The same engineer is blocked again next week on a similar problem. What does that tell you about your first intervention?
- 01
How do you approach mentoring junior software engineers and fostering engineering excellence across cross-functional teams?
- 02
Describe a time when you had an architectural disagreement with your manager or team lead. How did you handle the discussion, and what was the outcome?
- 03
An engineer two years into their career has a nightly accrual job that double-posts interest for some accounts whenever the batch is partially re-run after a failure. They have spent two days on it and are now rewriting the batch runner. You have 30 minutes. Describe how you have unblocked someone without taking the keyboard: the question you asked first, how you chose between handing over the answer and handing over the method, what you left behind so the next person does not get stuck here, and how you knew they were unblocked rather than deferring to you.
Is this an official Chime interview guide?
No. It is PracHub's own research and practice material for the Software Engineer role at Chime. Rounds and questions reflect what candidates have reported, not a process Chime has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗What is the overall difficulty level of the software engineering interviews at Chime?
The technical interviews at Chime are generally rated at a moderate to challenging difficulty level. Rather than asking abstract competitive programming or trick questions, interviewers focus primarily on practical coding problems, data structure manipulations, clean object-oriented design, and real-world system architecture challenges.
PracHub interview research ↗What technical stack is most commonly used by backend engineering teams at Chime?
Chime’s primary backend tech stack heavily utilizes Ruby on Rails for application logic and APIs, with an increasing transition toward Go for high-performance microservices and core infrastructure on the Financial Platform team. PostgreSQL, Redis, and AWS cloud infrastructure form the core of the data tier.
PracHub interview research ↗How are behavioral rounds evaluated at Chime?
Behavioral interviews at Chime assess your humility, self-awareness, cross-functional collaboration, and ownership. Interviewers look for candidate stories that highlight personal accountability, continuous learning from technical failures, and how you handle technical disagreements constructively without ego.
PracHub interview research ↗How fast is the decision timeline following the final onsite loop?
Chime generally strives to move quickly following the virtual onsite loop. Candidates typically receive feedback or next steps from their recruiter within three to five business days after the final interview rounds are completed.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-24 - 02PracHub Software Engineer practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-24 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-24