As a Software Engineer at Credit Karma, you build the core technology platform that helps over 120 million members make financial progress. Your engineering solutions directly drive products across credit monitoring, personalized financial recommendation engines, tax preparation, loans, and credit card matching algorithms. Working at the intersection of high-volume data processing and consumer-facing web and mobile applications, you will solve complex architectural challenges that handle massive real-time transaction traffic with low latency and high availability.
In this role, you collaborate within cross-functional Scrum teams consisting of product managers, data scientists, site reliability engineers, and product designers. You will be responsible for designing resilient microservices, building intuitive user interfaces, optimizing backend data pipelines, and establishing scalable API contracts. Engineering teams at Credit Karma prioritize clean object-oriented architecture, comprehensive automated testing, and long-term system extensibility over quick temporary patches.
Succeeding as a requires balancing rigorous computer science fundamentals with strong business empathy and communication. Whether you are scaling partner decisioning systems or building self-service financial tools, your code directly impacts consumer financial health on a national scale.
Initial Screening
reportedHalf of this call is the part candidates treat as small talk: start date, notice period, work authorisation and its timing, location and time zone, on-call, and the number. Those are what kill offers late, after several engineers have each spent a day. Surfacing a hard constraint now costs you nothing and occasionally buys you something, since a loop compressed to fit a competing deadline can usually only be arranged if it is asked for early. The common failure is deflecting the compensation question twice, then discovering at offer stage that the band never reached your number.
What to demonstrate
- Whether your hard constraints are compatible with the role before a loop gets booked: earliest start, notice period, what authorisation you hold and when it needs action, days on site, willingness to carry a pager
- Whether you give a compensation range with something behind it, such as current total compensation or a competing timeline, rather than leaving the band untested
- Whether your stated timeline is real, since a competing deadline raised now is something scheduling can sometimes work around and the same deadline raised at offer stage usually is not
How to prepare
- Write each constraint down in one line before the call and state them as facts rather than negotiating them live under a question you were not expecting
- Set your range from two or three current data points for that level and location, and name the structure you are quoting in, so the number is comparable to the one they are holding
- If another process is running, say where it stands and by when, and ask directly whether this loop can be scheduled inside that window
Technical Phone Screen
reportedBefore anything technical happens, someone has to decide which rung of the ladder your loop is calibrated to, and that decision sets the bar for every round after it. It comes from how you describe scope, not from your title, because titles do not convert cleanly between companies. The weak version of the answer is team size and years. The strong version names the largest change you shipped where nobody reviewed the design, what would have broken if you had been wrong, and what you were paged for. Get the level said out loud on this call, because the range and the loop both follow from it.
What to demonstrate
- Whether the scope in your own account maps onto a level the team actually has an opening at, so a mismatch ends the process cheaply rather than after four interviewers have spent a day
- Whether your title needs re-mapping: the same word describes very different amounts of independent decision-making at a twenty-person company and a ten-thousand-person one
- Whether your compensation expectation can be filled at that level in the structure the role pays in, which is why the number gets asked for before any engineer is scheduled
How to prepare
- Write down two changes from the last two years: the largest one you designed with nobody reviewing the design, and the largest one where someone more senior did. Lead with the first when scope comes up, and be ready to say which parts of the second were yours
- Ask which level the loop is calibrated to and what changes at the level above it, then plan your weeks from that answer rather than from the posting
- Settle a total-compensation range beforehand with the split named, base against bonus against equity and its vesting period, so a question about numbers gets a number instead of the word market
Onsite Evaluation
reportedNobody in the room with you decides this. Interviewers typically write their rounds up separately, often before seeing anyone else's, and the outcome is settled later from those write-ups. A split panel gets resolved by whichever note carries specific evidence, so what you want out of each room is one concrete thing that person could write down: a bug you caught yourself, a trade-off you named, a decision you owned. The rest is arithmetic. The project you describe in a behavioural conversation is often the same system you sketched an hour earlier, and the two accounts have to agree.
What to demonstrate
- Whether the scale, team size and timeline you attach to a project hold steady when that project resurfaces in a different round
- Whether each interviewer leaves with a specific thing to cite rather than a general impression of competence
- Whether a trade-off you defended in one round survives a challenge in another, instead of being quietly swapped for the answer the new interviewer seemed to want
- Whether a question you have already answered earlier in the day gets the same answer at the same depth, without visible impatience
How to prepare
- Write a one-page sheet per project fixing the figures you will quote — request volume, data size, team size, elapsed time, what broke — and say them aloud from the sheet until they come out identical every time
- For each round on the schedule, decide in advance the one sentence you want in that person's notes, then check in a mock that you said it outright instead of leaving it to be inferred
- Have someone ask you the same project question twice, an hour apart, and diff the two answers for numbers that moved or a trade-off that reversed
PracHub editorial advice for the preparation topics above.
Assuming the default isolation level enforces the invariant you wrote down
PostgreSQL defaults to READ COMMITTED, where every statement takes a fresh snapshot, so a read-modify-write on a balance loses updates under concurrency. Its REPEATABLE READ is snapshot isolation, which blocks that particular anomaly by aborting the loser with SQLSTATE 40001 but still permits write skew across two different rows; only SERIALIZABLE closes that, and both levels therefore require a bounded retry loop on 40001 that many implementations simply never write. MySQL's InnoDB REPEATABLE READ behaves differently again — it does not abort on a conflicting write, so the identical application code silently changes behaviour when the engine changes. Two-sided transfers add a second failure mode on top: without a deterministic lock ordering, such as always locking account ids in ascending order, concurrent opposing transfers deadlock (SQLSTATE 40P01).
Treating money as a decimal with two places
ISO 4217 exponents are 0 for currencies such as JPY and KRW, 2 for most, and 3 for BHD, KWD, JOD, OMR and TND, so a hard-coded multiply-by-100 is off by a factor of 100 or 10 depending on the currency, in opposite directions. Floating point is worse: IEEE 754 binary64 cannot represent 0.1 exactly, so repeated accrual accumulates drift that appears as a handful of minor units in the daily reconciliation and then gets 'fixed' by widening the match tolerance, which is how a genuine break becomes invisible. The only forms that survive a reconciliation are integer minor units with the exponent carried alongside the currency code, or a fixed-scale decimal type with exactly one documented rounding point.
Hardcoding to the sample inputs
Solve the stated problem rather than the two examples; special-casing a literal to make a sample pass is obvious immediately and reads as either a misunderstanding or an attempt to fake progress. If you genuinely cannot generalise yet, say which part is a stub and what would replace it.
Assuming fixed-width integer arithmetic cannot overflow
In languages with fixed-width integers, including C, C++, Java, Go and Rust, computing a midpoint as (lo + hi) / 2 overflows once the sum passes the type's maximum, so write lo + (hi - lo) / 2 instead. Say which language you are in: arbitrary-precision integers, as in Python or Ruby, remove this specific hazard and none of the others.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Write a function to parse, group, and summarize user transaction data …
Write a function to parse, group, and summarize user transaction data using hash maps or custom data structures.
Approach
- Walk one small example through your approach before writing the whole thing.
- State the target complexity and say which constraint rules the naive version out.
- Choose the data structure from the access pattern, not from familiarity.
Follow-up
- How does this change if the input no longer fits in memory?
- What is the worst case, and how likely is it on real data?
Solve a custom array manipulation or interval merging problem, explain…
Solve a custom array manipulation or interval merging problem, explaining your logic step-by-step during a live pair-programming session.
Approach
- Choose the data structure from the access pattern, not from familiarity.
- Walk one small example through your approach before writing the whole thing.
- State the target complexity and say which constraint rules the naive version out.
Follow-up
- How does this change if the input no longer fits in memory?
- Which test case would catch an off-by-one here?
Given a graph structure representing user connection networks or finan…
Given a graph structure representing user connection networks or financial reporting pipelines, implement a traversal method to find optimal or shortest execution paths.
Approach
- Walk one small example through your approach before writing the whole thing.
- Restate the input: its shape, its size, and what is guaranteed about it.
- Name the brute-force solution and its complexity before improving on it.
Follow-up
- Which test case would catch an off-by-one here?
- What is the worst case, and how likely is it on real data?
Solve a robot grid navigation problem with obstacles, implementing dyn…
Solve a robot grid navigation problem with obstacles, implementing dynamic programming or depth-first search while discussing time and space complexity optimizations.
Approach
- Walk one small example through your approach before writing the whole thing.
- State the target complexity and say which constraint rules the naive version out.
- Name the brute-force solution and its complexity before improving on it.
Follow-up
- How does this change if the input no longer fits in memory?
- What is the worst case, and how likely is it on real data?
Answer as-of balance queries over an append-only entry log
Given 400 million ledger_entry rows (entry_id, account_id, direction, amount_minor, currency, business_date) and 2 million queries of (account_id, currency, as_of_date) asking for the balance at the end of that business date, produce every answer. The obvious solution — per query, sum that account's entries with business_date <= as_of_date — is correct. Say precisely why it will not finish, then give one that will, with time and space complexity. Corrections are posted as new entries carrying their own business_date.
Approach
- Cost the naive version in numbers before rejecting it. Spread uniformly over 20 million accounts, each query touches about 20 rows behind a per-account index and 2 million queries is 4e7 row touches — perfectly fine. The problem is skew: one pooled clearing or merchant settlement account holding 3e7 entries, taking 10% of the queries, is 6e12 row touches. Name the skew; 'n is large' is not the reason.
- The structural fact that buys a cheap answer: entries are append-only and never updated, so a prefix sum over an account's entries ordered by
(business_date, entry_id)is stable — nothing behind position i can change. No mutable-balance design offers that, and it is why the storage is worth paying for. - Offline sweep, when all queries are known up front: externally sort entries by
(account_id, currency, business_date, entry_id)and queries by(account_id, currency, as_of_date), then merge-walk both with a running sum, emitting each query's answer as the sweep passes its date. O((n + q) log(n + q)) dominated by the sort, O(1) beyond sort buffers, one sequential pass over each input instead of 2 million random seeks. - Online alternative: materialise end-of-day snapshots — one row per
(account_id, currency, business_date)that had activity, holding the cumulative total. A query becomes one index seek for the latest snapshot at or beforeas_of_date, O(log n) per query, over far fewer rows than n. Use snapshots when queries arrive singly and the sweep when they arrive as a batch. - Corrections are the subtlety: an entry posted today but dated back changes historical answers, so every snapshot for that account from that date forward is stale. Either keep a Fenwick tree over dates per account (O(log D) update and prefix query) or recompute that account's snapshots from the corrected date onward. Then be precise about what reproducibility means — yesterday's statement is reproducible as of a stated snapshot time, not identical forever.
- Bound the resources: int64 sums throughout, no float; 400 million rows at roughly 48 bytes of the columns you actually need is about 19 GB, so the sort is external and its fan-out is chosen from the sort buffer, not from the row count.
Worked solution 40 min
- Compute both costs explicitly: the uniform case at about 4e7 row touches, and the skewed case at about 6e12. Showing that arithmetic is the answer to 'why'.
- Implement the offline sweep on a 10-million-row, 50,000-query fixture, merging on
(account_id, currency, business_date, entry_id). - Implement the naive version as the reference answer and assert both agree on every fixture query.
- Add a correction entry dated 30 days back, re-run, and assert that exactly the queries with
as_of_dateon or after that date move, all by the same signed amount. - Measure rows touched and wall time for each at 10 million rows, then extrapolate to 400 million and state the assumption that makes the extrapolation valid — sequential I/O, no random seeks.
Follow-up
- One account holds 30% of all entries. What does the external sort do with it, and what would you do for that one key instead?
- Queries now arrive online at 500 per second. Which design survives, and what does keeping the other one warm cost?
- A correction lands with a
business_date90 days back. Which snapshots are now wrong, and how does a reader find out?
Decide whether balance is derived or materialised, then hold the floor
Authorisation needs an account's available balance inside an 80 ms budget at 3,000 requests per second; that account already has 200 million ledger_entry rows. A product rule says the balance may never fall below the account's negative overdraft limit. Decide whether the balance is summed from entries or held in a materialised account_balance(account_id, balance_minor, floor_minor, currency, version) row, and justify the choice from the read pattern. Then give the write path that holds the floor, naming the isolation level, the anomaly a weaker level permits, and the SQLSTATE you retry.
Approach
- Size the derived read before arguing about it: summing 200 million rows is not an 80 ms operation under any index, since even a covering index on (account_id, entry_id) still reads work proportional to the rows. The authorisation read pattern forces one materialised row fetched by primary key; the statement read pattern, which is low-rate and historical, stays derived from entries. That is the whole justification for the denormalisation, and its price is a write on that row per posting.
- Put the floor where it can be a single-row constraint: CHECK (balance_minor >= floor_minor) on account_balance, updated in the same transaction as the entries. As a predicate over a set of entry rows it cannot be a CHECK at all, which is the reason the materialised row earns its keep twice.
- Write the mutation as one statement: UPDATE account_balance SET balance_minor = balance_minor - $1, version = version + 1 WHERE account_id = $2 AND balance_minor - $1 >= floor_minor. Zero rows updated means refused. This is safe even under READ COMMITTED, because the UPDATE re-evaluates its predicate against the locked, post-update version of the row.
- Name the anomaly in the shape that is not safe: SELECT the balance, compute a new value in application code, then UPDATE to that constant. Every statement under READ COMMITTED takes a fresh snapshot, so two concurrent 60-unit withdrawals against 100 both read 100 and both write 40. The row then claims 40 while 120 has actually left, so the true position is -20, below a floor of 0, and the row it is checked against cannot show it. PostgreSQL REPEATABLE READ is snapshot isolation and aborts the loser with SQLSTATE 40001; SERIALIZABLE additionally closes write skew across two rows. Both require a bounded retry with backoff, and InnoDB REPEATABLE READ does not abort at all, so identical code changes behaviour on a different engine.
- For a two-account transfer, acquire the rows in a deterministic order such as ascending account_id; without it, opposing concurrent transfers deadlock and the database kills one with SQLSTATE 40P01. That is a retry, not a correctness failure, but it is a retry somebody has to write.
- Finish on the ceiling: the hot row admits one committed write per lock hold, so at a 2 ms hold it caps near 500 per second regardless of cores. Sharding into N sub-rows multiplies throughput and immediately makes the floor check cross-row again, which then needs the shard sum under SERIALIZABLE or a per-shard reserved allowance.
Worked solution 35 min
- Seed one account with balance_minor = 100 and floor_minor = 0, then run two concurrent 60-unit withdrawals under READ COMMITTED using SELECT-then-UPDATE and record the final balance.
- Replace it with the single-statement conditional UPDATE and re-run the same race, asserting on rows affected rather than on an exception.
- Re-run under SERIALIZABLE with the read-modify-write shape, count the 40001 aborts, and add a retry loop with a fixed attempt cap and jittered backoff.
- Measure committed writes per second against that single row, then write the drift query: sum signed entries per account and compare against balance_minor.
Follow-up
- Shard the balance into eight sub-rows. Write exactly what the floor check now does, and what it costs per authorisation.
- Someone proposes an AFTER INSERT trigger on ledger_entry to maintain the balance. What does that change about ordering, about batch posting, and about failure handling?
- How do you detect that the materialised row has drifted from the entries, how often do you run it, and on which replica?
Explain why the outbox relay stopped using its partial index
outbox_event holds event_id, aggregate_type, aggregate_id, aggregate_version, event_type, payload jsonb, published_at, attempts, last_error, created_at, with index ix_unpub ON outbox_event (created_at) WHERE published_at IS NULL. The relay runs SELECT ... WHERE published_at IS NULL ORDER BY created_at LIMIT 500 FOR UPDATE SKIP LOCKED, then marks each row by setting published_at. Unpublished rows hold steady near 400, but the query has gone from 3 ms to 900 ms. Explain what EXPLAIN (ANALYZE, BUFFERS) will show, why it happens, and the fix.
Approach
- Read the plan for the gap between rows returned and work done: an index scan on ix_unpub returning 500 rows while touching tens of thousands of buffers is the signature. Rows Removed by Filter and the buffer counts name it; wall-clock alone does not, because a warm cache hides it.
- Explain the mechanism: marking a row published is an UPDATE, which writes a new tuple version. The new version fails the index predicate and leaves ix_unpub, but the dead old version's index entry stays until vacuum removes it, so the scan walks dead entries and discards them. PostgreSQL can hint an entry LP_DEAD once a scan has proved it dead, which cheapens repeat visits, but the index pages themselves still have to be read and are not reclaimed.
- Ask why vacuum is not reclaiming. Anything holding the xmin horizon back prevents removal: a long-running query, an idle-in-transaction session, an abandoned prepared transaction, or an inactive replication slot. Check the oldest xact_start in pg_stat_activity, pg_replication_slots, pg_prepared_xacts, and n_dead_tup with last_autovacuum in pg_stat_all_tables.
- Fix in order of leverage: delete or archive published rows instead of leaving them in place, so a queue table stays a queue; keep the transaction horizon short and alert on it; then tune autovacuum on this one table with an aggressive scale factor rather than changing the global setting.
- Rule out the other failure with the same symptom: a partial index is usable only when the planner can prove the query predicate implies the index predicate, so rewriting the filter as coalesce(published_at, 'epoch') = 'epoch' or wrapping the column in a function disqualifies the index entirely and produces a sequential scan instead of a bloated index scan.
- Verify by re-running EXPLAIN (ANALYZE, BUFFERS) after the horizon is released and a VACUUM completes, comparing shared buffer reads rather than elapsed time, and confirm the relay keeps per-destination ordering after the change.
Follow-up
- SKIP LOCKED means two relay workers never block each other. What else does it change about ordering guarantees for a single destination?
- You archive published rows to a second table. What does that do to the relay's crash recovery and to duplicate delivery?
- The relay batches 500 rows and publishes them, then marks them. Where exactly can it crash, and what does the consumer see?
Design a high-throughput offer decisioning pipeline that evaluates thi…
Design a high-throughput offer decisioning pipeline that evaluates third-party financial partner offers for active members simultaneously.
Approach
- Name the read and write paths separately; they rarely have the same bottleneck.
- Choose a partition key and say what query it makes expensive.
- State the consistency you need, and where you are willing to be stale.
Follow-up
- What breaks first when traffic grows ten times?
- How does this behave when that dependency is down for an hour?
Design a distributed rate-limiting microservice to safeguard internal …
Design a distributed rate-limiting microservice to safeguard internal payment and credit reporting endpoints from sudden traffic spikes.
Approach
- Name the failure you are designing for, then the recovery path.
- Choose a partition key and say what query it makes expensive.
- State the consistency you need, and where you are willing to be stale.
Follow-up
- How does this behave when that dependency is down for an hour?
- What would you drop to keep the system up under load?
Walk through how you apply SOLID design principles, modularity, and de…
Walk through how you apply SOLID design principles, modularity, and dependency injection to build maintainable, easily testable backend services.
Approach
- Clarify what is being asked and what a complete answer contains.
- Work from the requirement backwards to the design.
- State your assumptions explicitly before working the problem.
Follow-up
- What assumption would you test first?
- How would you know your answer was wrong?
Expose unknown outcomes and rate limit a retrying caller
At 3,000 authorisation requests per second, a processor slows past your callers' 2 s timeout and every merchant's client library retries three times. A caller cannot distinguish a lost response from a lost request. Design two things: the API surface that lets a caller resolve an unknown outcome without re-attempting the payment, and the rate limiting that sheds the resulting load without making outcomes unknowable. State the limiter's algorithm and scope, the status and headers returned on rejection, and which requests must never be shed.
Approach
- Order the writes so an unknown outcome always leaves evidence. Persist the intent and the idempotency key as in_progress and commit, then call the processor, passing your own client reference so the attempt is queryable by an identifier you chose. Calling first and persisting after is the one ordering that can produce a charge with no local record at all.
- Make the unknown state a first-class, readable resource: GET /payments/{id} and a lookup by idempotency key, both returning status processing with a hint. The caller's recovery then becomes a read rather than a second write. Retrying the write with the same key is also safe, but a read is cheaper, cannot be confused with a new attempt, and is what you want a panicking integration to reach for.
- Converge without the caller. A sweeper picks up in_progress rows past their lease, queries the processor by the stored reference, and resolves the row, so a caller that never returns still ends with a correct record and the reconciliation does not open a break.
- Scope the limiter to merchant plus operation and use a token bucket sized to the merchant's committed rate with a burst that covers one client's retry fan-out. A global limiter cannot tell whose retries are amplifying and will shed the innocent; a fixed window lets through twice the limit across a boundary.
- Reject with 429, Retry-After and remaining-quota headers, and make the rejection cheap and early: no idempotency row, no processor call, no database write. That is what makes a shed request unambiguously 'never happened' and therefore safe to retry. Exempt the status read from the write bucket, or give it its own generous one, because a limited caller that cannot read state can never learn the outcome of the write it already made.
- Cap amplification at the source and put it in the client contract: exponential backoff with full jitter, plus a retry budget capping retries at roughly a tenth of recent successful requests rather than a fixed count per call. Three fixed retries at 3,000 rps is 12,000 rps arriving precisely when the system is already saturated, which is how a slow dependency becomes an outage that outlives the slowdown.
Worked solution 45 min
- Write the ordering as a sequence diagram and mark the single crash point that can produce a charge with no local record; confirm your ordering eliminates it.
- Implement the status lookup by idempotency key and assert it returns processing, not 404, for a row that is in_progress.
- Put the limiter in front of everything, including authentication, and assert with a counter that a shed request touches no table and makes no outbound call.
- Load-test at 3,000 rps with an injected 3 s processor delay and a client doing three fixed retries, record the observed arrival rate at the origin, then re-run with backoff plus jitter and a 10% retry budget and compare.
- Run the sweeper against 1,000 abandoned in_progress rows and assert every one resolves to a terminal state with exactly one charge each.
Follow-up
- The processor's own status query starts timing out too. What does the sweeper do, and how does it avoid becoming the next amplifier?
- A merchant claims your 429 cost them a sale. What in the design lets you prove the request never reached the processor?
- One merchant is 90% of your traffic. What does per-merchant bucketing do for everyone else, and what does it fail to protect?
Authorisation p99 tripled overnight with no deploy
Payment orchestration serves about 3,000 authorisations per second against a 150 ms p99 budget. Since 02:00, p99 is 460 ms and rising about 8 ms per hour, while p50 is unchanged at 11 ms. There was no deploy, no traffic change and no processor degradation. The hot path does one INSERT into idempotency_key, which has UNIQUE (scope, key), then two UPDATEs on that row: locked_at, then status and response_body. A nightly reconciliation job started at 01:50 and is still running. Give the ordered diagnostic checklist and the cause.
Approach
- Read the shape first. p50 flat with p99 rising and no deploy is a resource or data-volume effect, not a code path, because a code change moves the median too. A tail that climbs monotonically at fixed workload means something monotonically grows.
- Ask what started at 01:50. In PostgreSQL an open transaction holds back the xmin horizon cluster-wide, so autovacuum can reclaim no dead tuple newer than that snapshot. Confirm with pg_stat_activity (state, now() - xact_start, backend_xmin) and with pg_stat_all_tables (n_dead_tup, last_autovacuum) for idempotency_key.
- Connect it to the write pattern. Three writes per key produce up to two dead tuples each, so at 3,000 rps the table sheds roughly 6,000 dead tuples per second. Heap-only tuple updates would keep those out of the index, but only when no indexed column changes and the page has room, and appending response_body grows the tuple enough to force a new page. So the unique index on (scope, key) grows too.
- Explain why only the tail suffers. A larger index means more pages per lookup and a rising fraction of them missing shared_buffers; the median request still hits cache while the tail pays physical I/O. This is exactly the p50-flat, p99-rising signature, and it is worth stating before acting.
- Verify before fixing rather than after. n_dead_tup in the millions and rising, last_autovacuum stale since about 01:50, and pg_relation_size on the unique index measured twice fifteen minutes apart showing growth at constant workload. Index bloat cannot be inferred from row count alone; use the size series or pgstattuple.
- Fix in two moves and prevent separately. End or chunk the long transaction so the reconciliation job commits per batch instead of holding one snapshot over 50M lines, then let autovacuum catch up or run REINDEX CONCURRENTLY. Add a transaction-age alert and a statement timeout on the reporting role. Collapsing the two UPDATEs into one and lowering fillfactor halves dead-tuple production, but that is an optimisation, not the cause.
Follow-up
- The reconciliation job legitimately needs a consistent view of 50M lines. How do you give it one without pinning the xmin horizon?
- Why did p50 not move at all?
- You also run an expires_at cleanup job that DELETEs old idempotency keys. During this incident, does running it help or hurt, and what design avoids the question entirely?
For a candidate senior enough that the loop turns on design and judgement rather than on whether the coding round gets finished. Five days build one system properly and then stress it; coding gets a single maintenance day, on the assumption that the risk at this level is an unexamined tradeoff rather than a missed algorithm.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Numbers before diagrams
- Build your own reference card of the figures you will re-derive all week: bytes for a realistic record, requests per second implied by a given daily active count, and the storage that a year at a given write rate produces. Derive each one rather than copying it, because the derivation is what survives a follow-up.
- Turn one product statement into capacity requirements. From ten million daily users at four writes and forty reads each, state the peak-to-average factor you are assuming and why, then produce peak write QPS, peak read QPS and a year of storage.
- Write the two numbers whose order of magnitude changes the design, the read-to-write ratio and the working-set size against memory per node, and state the threshold at which each one flips your answer.
Deliverable: A one-page numbers card and one worked capacity estimate with every assumption written down.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02One system, from requirements to schema
- Spend the first ten minutes producing only functional requirements, non-functional targets with numbers attached, a p99 latency, a durability expectation, a consistency requirement, and an explicit out-of-scope list.
- Define the interface before the boxes: the three or four endpoints, their parameters, what each returns, and which of them are idempotent.
- Write the data model, then write the single access pattern that justifies it, and state what the schema would have to become if the dominant access pattern were the other one.
Deliverable: One design carried to endpoint-and-schema depth, with non-functional targets expressed as numbers and a written out-of-scope list.
Practice prompt ↗Practice prompt ↗03The consistency you are actually buying
- Write out what a client sees under asynchronous replication when its write commits on the leader and its next read is served by a lagging follower, then write the two fixes, pinning that session's reads to the leader for a bounded window or carrying a version token the replica must reach, and the cost of each.
- Work the quorum arithmetic on paper for N of three with W and R of two, and separate what R + W > N does guarantee, that any read set intersects any write set, from what it does not: on its own it is not linearizability, and a sloppy quorum that accepts writes on nodes outside the preference list breaks even the intersection.
- Take two storage choices with different defaults, a single-leader relational store committing synchronously and a quorum-replicated store that converges eventually, and write the specific product behaviour that would be wrong under each, rather than a general statement about which is stronger.
Deliverable: A page separating what quorum overlap guarantees from what it does not, with one concrete product misbehaviour attached to each gap.
Practice prompt ↗Practice prompt ↗04Failure is the design
- For one write path, work through the case where the client times out after the server has already committed, then design the idempotency key: who generates it, how long it is retained, and what the duplicate request returns.
- Express the retry policy as parameters rather than as a word: maximum attempts, base delay, backoff factor, jitter, and which error classes are retried at all. Then state why retrying a non-idempotent write without a key is a correctness bug and not merely waste.
- Compute the fan-out effect on tail latency. If a request waits on ten backends and each independently exceeds its p99 one percent of the time, the chance at least one is slow is 1 - 0.99^10, about ten percent. Then write why independence is the optimistic assumption and what correlates them in practice.
- Name the backpressure mechanism for one queue or one dependency in the design, a bounded queue with shedding or a concurrency limit, and write what the caller is told when it engages.
Deliverable: One write path with an idempotency design, a parameterised retry policy, and a written tail-latency calculation with its assumption named.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Scaling the hot path
- Choose cache-aside or write-through for one read path and write the staleness window each produces, then name the invalidation event and what the system does when that event is lost.
- Design against the stampede: either coalesce requests so only one recomputes a missing key, or refresh early with jittered expiry, and write why identical TTLs on keys populated in the same moment produce a synchronised expiry and a thundering herd.
- Shard one table by a key you choose, then answer the two questions that break the choice: which queries now require a scatter-gather, and what happens to the distribution when one tenant is a hundred times larger than the median.
- Write the cost of adding a node under plain modulo placement, where nearly every key moves, against consistent hashing, where roughly one key in n+1 moves, and state what virtual nodes are for.
Deliverable: A caching and sharding decision for one path, each with its failure mode and its rebalancing cost written beside it.
Practice prompt ↗Practice prompt ↗06Keep the coding hand in, at the bar that applies to you
- Solve one medium problem in thirty minutes, then spend twenty more making it production-shaped: named invariants, validation at the boundary, and errors that distinguish a caller mistake from an internal fault.
- Write the tests you would require of a colleague's version of that function: one for empty input, one for the boundary, and one for the case the implementation is most likely to get wrong.
- Read a piece of your own code from six months ago and write the change you would ask for, phrased as you would actually phrase it in review.
Deliverable: One problem hardened to review standard, with its test list and one written review comment.
Practice prompt ↗Practice prompt ↗07Defend it while being interrupted
- Run a forty-five-minute design mock with an interviewer briefed to change a requirement halfway, a tenfold traffic increase or a new strict consistency requirement, and to push on one number you estimated.
- Rehearse the two sentences a senior loop is listening for: naming the tradeoff you are choosing against and why, and saying what you would measure to learn that the choice was wrong.
- Prepare the design you regret: a real decision, the constraint that produced it, what it cost, and what you changed afterwards.
Deliverable: Mock notes recording how the design changed under the new requirement, plus a written account of one regretted decision.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Counting review comments or mentees proves nothing. The useful version is a specific change you approved with a reservation you stated, or one you blocked and the delay that cost. Say which standard you were holding and why it was worth the friction. A mentoring story needs the thing the other person can now do without you.
How would you explain the concept and structure of an API to a non-tec…
How would you explain the concept and structure of an API to a non-technical stakeholder using real-world analogies, without using foreign technical jargon?
Approach
- Close with what you would do differently, concretely.
- Name the disagreement and how you resolved it with evidence.
- Pick a story where you made the decision, not one where you watched it.
Follow-up
- What would you do differently if you ran that again?
- How did you know your change caused the improvement?
Describe a time when you received constructive feedback on code extens…
Describe a time when you received constructive feedback on code extensibility or system performance, and how you integrated that input into your future engineering work.
Approach
- State the situation in two sentences and spend the rest on the reasoning.
- Give the blast radius: what could have broken, and what you measured.
- Close with what you would do differently, concretely.
Follow-up
- What did you decide not to do, and why?
- How did you know your change caused the improvement?
Force an implicit timeout behaviour into an explicit decision
The risk decision service has an 80 ms p99 budget inside a roughly 2 s caller timeout. Today, when its feature store is unavailable, the timeout handler returns approve. Nobody chose that; it is what the code does. You need a real decision: fail open, fail closed, or refer, potentially differing by amount band. Describe a time you turned an accidental behaviour into an owned decision. State who had to be in the room, the data you brought, what you did when nobody wanted to own it, and where the decision was recorded so it outlived you.
Approach
- The probe is whether you can drive a cross-functional decision rather than escalating and waiting. Lead with the framing that makes it undeniable: this is already a product decision, it is currently being made by an exception handler, and the only question is whether anyone reviews it.
- Bring the two losses side by side instead of arguing a principle. Fail open costs expected fraud loss on approved-but-should-have-declined volume during the outage; fail closed costs declined good payments, which is lost revenue plus customer harm and a support queue; refer costs manual review capacity, which is a headcount number and saturates within minutes at 3,000 decisions per second. Give each as a rate per minute of outage using real volume.
- Propose the banded answer as the default, because the two losses cross over at an amount: below some threshold the expected fraud loss is smaller than the expected decline loss, above it the reverse, and the crossover is computable from observed fraud rate by band. That converts a values argument into an arithmetic one.
- Name the attendees by the decision they own, not by title: whoever carries fraud loss, whoever carries approval rate, and whoever staffs manual review. Three people who can each say yes is a decision; eight people who can each say no is a meeting.
- Say what you did when ownership was contested. A strong answer has a forcing function: propose a default in writing with a review date and state that it ships unless someone objects, which converts inaction into consent rather than into another meeting.
- Record it where the code can find it: the decision, its date, its owner, the amount thresholds, and a test asserting the fallback behaviour, so the next engineer reading the timeout handler learns it was chosen. A wiki page nobody links from the code is the generic answer.
Follow-up
- The feature store is degraded rather than down and the model is scoring on stale features. Is that the same decision?
- How do you stop the banded thresholds from silently rotting as fraud patterns shift?
- Nobody objects to your written default, and six months later there is an outage and a loss. Who owns it?
- 01
How would you explain the concept and structure of an API to a non-technical stakeholder using real-world analogies, without using foreign technical jargon?
- 02
Describe a time when you received constructive feedback on code extensibility or system performance, and how you integrated that input into your future engineering work.
- 03
The risk decision service has an 80 ms p99 budget inside a roughly 2 s caller timeout. Today, when its feature store is unavailable, the timeout handler returns approve. Nobody chose that; it is what the code does. You need a real decision: fail open, fail closed, or refer, potentially differing by amount band. Describe a time you turned an accidental behaviour into an owned decision. State who had to be in the room, the data you brought, what you did when nobody wanted to own it, and where the decision was recorded so it outlived you.
Is this an official Credit Karma interview guide?
No. It is PracHub's own research and practice material for the Software Engineer role at Credit Karma. Rounds and questions reflect what candidates have reported, not a process Credit Karma has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗What is the primary coding language used during Credit Karma engineering interviews?
You are generally free to use whatever object-oriented or mainstream programming language you are most comfortable with during coding and pair-programming sessions, including Java, Python, C++, JavaScript, or Go. The evaluation focuses on your underlying problem-solving logic, OOP clean code practices, and communication rather than syntax memorization.
PracHub interview research ↗How difficult are the technical coding questions compared to standard industry platforms?
The algorithmic coding problems at Credit Karma are generally rated as average in difficulty. Rather than asking hyper-abstract dynamic programming puzzles, interviewers prefer practical, domain-adjacent problems (such as grid navigation, array/string parsing, or class structure implementation) that test real-world software engineering skills.
PracHub interview research ↗What differentiates successful candidates in the System Design round?
Successful candidates distinguish themselves by driving an interactive discussion rather than giving a static lecture. They proactively clarify performance requirements and throughput constraints, justify database choices, design clean API contracts, and address failure points, caching mechanisms, and scalability trade-offs clearly.
PracHub interview research ↗Is pair programming a significant part of the onsite interview process?
Yes, several coding rounds are conducted as interactive pair-programming sessions. Interviewers evaluate how well you collaborate, how clearly you talk through your thought process while typing, and how receptively you incorporate feedback or hints provided during the exercise.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-24 - 02PracHub Software Engineer practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-24 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-24