Software Engineers at Figma work on a collaborative design tool that runs in the browser. The work can include the core canvas engine, real-time commenting, and infrastructure for AI-driven design tools. That work shows up in the interviews. Reported coding questions ask you to model layers with properties, apply and undo changes, commit changes in batches and keep edit history small in memory. Reported design questions ask for multi-user editing with real-time comments, component templates and instances, and a real-time trending system. Two more are separate reported questions: an asynchronous job scheduler that handles immediate tasks and delayed timers, and handling worker crashes in a distributed task system.
The role involves state management, hierarchical document models and distributed systems. In practice, you should be able to build a data structure yourself rather than reach for a library, explain its memory cost, and change it when the scale grows (one reported question asks how to optimize memory when storing property changes for 10K+ layers).
Candidates report four stages, moving from implementation detail in the coding screens to broader architecture in system design, followed by behavioral interviews and a final onsite that may include several rounds with team members. Neither the reported questions nor this guide map a specific question to a specific stage. Prepare each category well enough that it would hold up in any of them, and confirm the current format with your recruiter.
Technical Screens
reportedCandidates describe the technical screens as initial coding rounds focused on specific implementation details. The reported coding questions are not tied to a particular stage, and they fall into two groups. One is document and layer management: a class that stores layers with properties, an update and a get-by-ID function, apply/undo/redo, and a commit_batch function under memory limits. The other is algorithms: topological sort over folder access, longest common prefix for file auto-complete, sorting canvas documents left to right while respecting parent-child relationships, and validating brackets or turning a nested string into a nested list. Expect follow-up questions on memory optimisation and edge cases, so treat your first working version as the start of the discussion, not the end.
What to demonstrate
- Whether you can build the data structure yourself, such as a history stack of deltas, a trie or an in-degree map, instead of relying on a library to hide the logic
- How you break down an ambiguous prompt: clarifying requirements and scope and stating trade-offs before you write code
- Whether your solution survives a follow-up that raises scale or tightens memory, and whether you can state the Big-O effect of the change
- Whether you trace test cases, including empty input, a single layer and an undo with nothing to undo, before you call the code finished
How to prepare
- Implement a layer store with init, update_property and get_by_id, then add undo and redo. Keep only the changed property's old and new value per edit, not a copy of the whole layer.
- Extend it with commit_batch so that several edits undo as one history entry, and write the rule that a new edit clears the redo stack
- Solve the reported algorithm set in plain code without helper libraries: Kahn's topological sort, a trie for longest common prefix, and a stack for bracket validity and nested-list parsing
- For canvas ordering, clarify the ordering rule before coding, then sort each parent's children by x and emit them with a pre-order DFS so every subtree stays under its parent; a single flat sort on (x, tree position) mixes children of different parents and breaks the hierarchy
System Design Discussions
reportedCandidates describe this stage as broader architectural discussion to assess design skills. Reported design questions, not tied to a particular stage, include multi-user editing of images and documents with real-time commenting, a real-time trending system for hot files or search queries, representing components efficiently as templates and instances, an asynchronous job scheduler that handles immediate tasks and delayed timers, and handling worker crashes in a distributed task system. Expect to address conflicts between users editing at the same time, propagating updates in real time, read/write throughput, optimistic UI updates and caching for hot files. Answers that name one standard architecture without saying what it gives up are the ones to avoid.
What to demonstrate
- Whether you say how concurrent edits to the same object resolve, and what a client sees while its optimistic update is still pending
- Whether you separate the real-time write path from the read path, and say where caching hot files helps and where it serves stale data
- Whether your design handles failure explicitly, including a worker dying mid-task, a client reconnecting after missed updates, and a delayed job firing late
- Whether trade-offs are stated against concrete constraints such as memory limits, concurrency on one file and persistence
How to prepare
- For collaborative editing and comments, write down the unit of conflict (document, layer or property), the ordering authority, how a reconnecting client catches up, and what happens to a comment whose anchored node is deleted
- For templates and instances, model an instance as a reference to its main component plus a sparse set of overrides, then walk through what a change to the main component does to instances that override that property
- For the trending system, choose a time window and a decay function, then explain how per-shard top-k counts get merged and how one extremely hot key is kept from overloading a single partition
- For the job scheduler and worker crashes, cover leases with expiry, re-queueing on expiry, idempotent handlers or a fencing token that rejects a stale worker's late write, and a dead-letter path after repeated failure
Behavioral Interviews
reportedCandidates describe the behavioral interviews as assessing cultural fit and first-principles thinking. Several behavioral items in the question bank are project discussions: explain what you built, why it mattered, the trade-offs, risk and scope decisions you made, and what you learned. Other bank items cover common behavioral and hiring-manager prompts, behavioral stories, a leadership deep-dive and adapting how you communicate. First-principles thinking means you can show how you reached a decision from the constraints in front of you rather than following a pattern you had seen before.
What to demonstrate
- Whether you can explain a project's technical decisions from the constraints that drove them, not only from the result
- How clearly you separate your own contribution from the team's, and how you talk about scope you cut and risks you accepted
- Whether you can adjust depth and vocabulary to the person asking, as the bank's communication prompt asks
How to prepare
- Pick two projects and write each as a list of decisions, with the constraint behind each one, the alternative you rejected, and what you would change now
- Rehearse the project deep-dive at three depths: a one-minute summary, the architecture, and a single hard bug or trade-off in full detail
- Prepare one story where a hint or piece of feedback changed your approach mid-course, and treat the interview itself like a pair-programming session: think out loud and be open to hints
Final Onsite Rounds
reportedThe final onsite is reported as a set of interviews that may include several rounds with team members. Candidates do not report a round-by-round breakdown, so do not prepare on an assumed agenda. Prepare every category above to the same standard: the layer model and undo/redo, the tree and graph algorithms, the real-time and asynchronous design questions, and your project stories. Different interviewers may probe the same material from new angles, so your answers need to stay consistent from one conversation to the next.
What to demonstrate
- Whether your coding holds up, including test cases and follow-ups on memory and edge cases
- Whether your design reasoning stays consistent when someone new challenges an assumption you defended earlier
- Whether your project stories match in detail across interviewers and show how you would work with the team you are meeting
How to prepare
- Run one back-to-back mock covering a layer-model coding problem, a collaborative-editing design and a project deep-dive, with a different person asking each if you can
- Keep a one-page sheet of your design defaults (conflict resolution, caching, retry and idempotency) so your answers stay consistent across rounds
- Ask your recruiter which categories the onsite covers and who is in each conversation, and prepare questions about the team's current technical problems
6 candidate reports. Individual accounts describe a particular role and hiring cycle.
Figma Software Engineer Interview Experience — Breezed Through the Legendary Layer Question, Blanked on the Easy One
View report detailsFigma Data Scientist Interview Experience — A Low-Energy Hiring Manager Round Worth Avoiding
Sharing my interview experience with Figma DS for the HM round — overall it was really bad. This is the only company where I finished the interview and immediately wanted to warn everyone else away from it. They started by asking about my background and introducing the role. The HM seemed really low energy, like they were just going through the motions. They asked about a project and my work back…
Read full experienceFigma Software Engineer Interview Experience — Memory-Efficient Undo/Redo on a Document Layer Problem
View report detailsFigma AI Team Software Engineer Interview Experience — Rejected 2 Days After a 5-Round Virtual Onsite
After the recruiter call, the first round was HM + coding. The VO (final onsite) was 5 rounds — coding + 2 system design rounds + behavioral + a project deep dive. I hadn't seen the AI-flavored system design round anywhere else before. The VO was spread across 3 days, and I got rejected 2 business days after that. The HM round asked about past projects, not much technical. The first coding round…
Read full experienceFigma Machine Learning Engineer Interview Experience — Long-Winded Onsite, Two ML Design Rounds, and References Before an Offer
View report detailsPracHub editorial advice for the preparation topics above.
Storing a full copy of the layer (or the whole document) on every edit in an undo/redo question
A reported question asks how to optimize memory when storing property changes for 10K+ layers, and per-edit snapshots do not scale to that. For each edit, store only the change: layer id, property, old value, new value. Undo writes the old values back in reverse order, and redo writes the new values forward. A commit_batch becomes a single history entry holding a list of these changes, which undoes as one unit. Say what you would cap (history depth, or merging repeated edits to the same property inside one batch) before you are asked.
Leaving the redo stack intact after a new edit, or undoing only part of a batch
Write down the history invariants before you code. A new edit clears the redo stack. Undo moves an entry from the undo stack to the redo stack. A batch is applied and reversed as a whole. Then trace the sequence edit, edit, undo, new edit, redo out loud: redo should do nothing. Also cover undo with an empty history and an update to a layer id that does not exist. Bugs that surface during initial testing are a common way to lose points, so trace these cases before you say you are done.
Designing multi-user editing and real-time comments as a message bus with no conflict or reconnect story
Name the unit of conflict and the ordering authority. A common default is a server-assigned order with last-writer-wins per property, and you should say what that loses when two users change the same property. Explain how optimistic local updates are reconciled when the server's order differs, how a client that was offline catches up without replaying everything, and what a comment anchored to a deleted or moved layer points to afterwards.
Answering worker crashes with retries alone, so a slow worker and its retry both complete the task
Give each task a lease with an expiry and a heartbeat, and re-queue it when the lease lapses. Accept that this is at-least-once delivery and make it safe: an idempotency key on the side effect, or a fencing token so the store rejects a write from a worker whose lease has already passed to someone else. Add bounded retries with backoff and a dead-letter queue. For delayed timers, say what happens when the scheduler restarts with jobs that became due while it was down.
Writing the folder-access topological sort without deciding what a cycle or an unreachable node means
Before coding, state the direction of the edges (folder to child, or team to folder) and what "highest level of access" resolves to when two paths grant different levels. With Kahn's algorithm, compare the count of processed nodes to the total and say what you return when they differ. Test with a single node, a diamond where two paths reach one folder, and a cycle.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Given a list of files and an input string, return the longest common p…
Given a list of files and an input string, return the longest common prefix (auto-complete).
Approach
- State the target complexity and say which constraint rules the naive version out.
- Name the brute-force solution and its complexity before improving on it.
- Walk one small example through your approach before writing the whole thing.
Follow-up
- How does this change if the input no longer fits in memory?
- What is the worst case, and how likely is it on real data?
Determine if a string with brackets is valid or convert a nested strin…
Determine if a string with brackets is valid or convert a nested string into a nested list.
Approach
- Restate the input: its shape, its size, and what is guaranteed about it.
- Choose the data structure from the access pattern, not from familiarity.
- State the target complexity and say which constraint rules the naive version out.
Follow-up
- Which test case would catch an off-by-one here?
- How does this change if the input no longer fits in memory?
Sort documents on a canvas from left to right, including hierarchical …
Sort documents on a canvas from left to right, including hierarchical (parent-child) relationships.
Approach
- Restate the input: its shape, its size, and what is guaranteed about it.
- State the target complexity and say which constraint rules the naive version out.
- Choose the data structure from the access pattern, not from familiarity.
Follow-up
- How does this change if the input no longer fits in memory?
- Which test case would catch an off-by-one here?
Parse and verify a timestamped multi-signature webhook header
An inbound webhook carries a signature header of at most 1 KiB shaped t=<unix seconds>,v1=<64 hex chars>, with up to five v1 values during secret rotation and possibly unknown scheme keys. You hold the raw request body bytes and the currently active signing secrets. Write the parser and the verifier: accept when any active secret reproduces a signature and the timestamp is within a five-minute tolerance in either direction, reject otherwise. Single left-to-right pass over the header, no regular expression. State what is inside the MAC and why.
Approach
- Parse in one scan: split on
,, then on the first=only, since a value may itself contain=under a future scheme. Accepttexactly once and treat a secondtas a reject rather than last-wins. Push everyv1onto a short list and ignore any other key, so av2can be introduced later without breaking this verifier. - Say what is signed: HMAC-SHA256 over the exact byte string
<t>.<raw body bytes>, yielding 32 bytes or 64 hex characters. The timestamp sits inside the MAC because otherwise an attacker replays yesterday's body with its still-valid signature and only has to edit the header timestamp. - Hash the bytes as received. Verifying against a re-serialised JSON body is the usual defect: key order, whitespace and number formatting all change the bytes while the parsed objects compare equal, so signatures fail for honest senders and the popular 'fix' is to stop checking.
- Compare in constant time over fixed-length digests. Decode the hex to 32 bytes, accumulate
acc |= a[i] ^ b[i]across the whole length, and testacc == 0at the end. Evaluate every candidate without an early exit; at five candidates that is five HMACs over the body, linear in body size and negligible beside the network. - Apply the tolerance as a two-sided bound, rejecting when
|now - t| > 300seconds. A sender whose clock runs ahead of yours is an ordinary case, and an unbounded future timestamp is a free replay window. - Complexity: O(L) over the header producing k candidates, plus k HMACs at O(|body|) each. Space is O(k) beyond the body itself. Do the cheap rejections, including the tolerance check, before any cryptography runs.
Worked solution 15 min
- Write the grammar on one line before coding:
header := field (',' field)*,field := key '=' value, split on the first=only. - Implement the parser to return
{t: int, v1: [hex, ...]}, rejecting a missingt, a duplicatet, anyv1that is not 64 hex characters, and a header over 1 KiB, all before any cryptography runs. - Implement the verifier: for each active secret compute
HMAC-SHA256(secret, f'{t}.'.encode() + raw_body), compare it in constant time against each parsedv1, and OR the results with no early exit. - Test with a valid signature; the same body with
tmoved 400 seconds into the past; the same body witht400 seconds into the future; a header carrying an unknownv2=alongside a validv1; and a body re-serialised with different JSON key order.
Follow-up
- The body is 40 MB. What changes about where you verify, and what can you do before the whole body has arrived?
- A customer reports that signatures fail for exactly the requests whose body contains a non-ASCII character. What is your first hypothesis?
- How do you rotate the signing secret with no failed deliveries, and how long do both secrets stay live?
Enforce a concurrent-run quota that survives simultaneous requests
A plan allows at most 20 concurrently running rows in job_run per tenant. The table holds run_id, tenant_id, workspace_id, status (queued, leased, running, succeeded, failed, timed_out, cancelled, lost), lease_token, leased_until, started_at and finished_at. Today the service runs select count(*) from job_run where tenant_id = $1 and status = 'running', compares the result to 20, then inserts. Under load a tenant exceeds the cap by exactly the number of concurrent requests. Name the anomaly, say which isolation levels do and do not prevent it, and give a version that holds, as SQL.
Approach
- Name it: write skew. Each transaction reads a predicate (the count of running rows), neither modifies what the other read, and both then insert rows that jointly violate an invariant no single row expresses. Read committed permits it. So does repeatable read, because snapshot isolation's first-updater-wins check fires only on conflicting row updates, and these are inserts touching disjoint rows.
- Enumerate the fixes with their real costs. SERIALIZABLE works: PostgreSQL's SSI tracks the predicate read and aborts one transaction with SQLSTATE 40001, which obliges the caller to retry and makes the abort rate rise with contention on a hot tenant. Folding the predicate into the write as
insert ... select ... where (select count(*) ...) < 20narrows the race to the statement's snapshot but does not close it under read committed. - Give the version that holds at read committed: serialise on a row both transactions must touch.
update tenant_concurrency set running = running + 1 where tenant_id = $1 and running < 20 returning runningupdates zero rows when the cap is reached, and zero rows is the rejection. This works because at read committed a blocked UPDATE re-evaluates its WHERE clause against the newly committed row; at repeatable read the same statement raises a serialisation error instead, so the isolation level changes the calling contract. - State the cost you just bought. That row is now a per-tenant serialisation point, so admission throughput for the tenant is bounded by one divided by the lock hold time; at a 2 ms hold that is roughly 500 admissions/second. Keep the critical section to the single UPDATE, with no network call or scheduling decision inside the transaction, and decrement in the same transaction that writes the terminal status.
- Close the leak the status enum implies: a run can end as
lost, so a crashed worker otherwise consumes a slot forever. Reconcile on a schedule againststatus = 'running' and leased_until < now(), and treat the counter as a fast path overjob_run, which stays the system of record.
Follow-up
- Write the retry loop for the SERIALIZABLE version. What does the caller see when it keeps aborting, and what bounds the retries?
- Two regions each keep a counter. What is the effective cap, and what does admission do when the counter store is unreachable?
- The cap changes mid-flight on a plan upgrade. Do running jobs get killed, and what does the counter row look like during the change?
Paginate a tenant's delivery export without skipping rows
A customer exports webhook_delivery: delivery_id (bigint identity), subscription_id, tenant_id, event_id, status, attempt_count, next_attempt_at, created_at, delivered_at, updated_at. The endpoint runs select ... where tenant_id = $1 order by created_at desc limit 100 offset $2, and customers report rows missing from exports taken while new deliveries are being inserted. Write the replacement query and the index that supports it, paging a tenant's deliveries newest first at constant cost per page. State why updated_at cannot be the cursor column.
Approach
- Name the defect precisely. OFFSET is a position in a result set that is recomputed on every request, so a row inserted ahead of the window shifts everything back by one and the next page starts after a row the client never received. Nothing errors and no identifier gap appears, so the loss is silent.
- Replace the position with a value predicate over a stable, unique, indexed ordering:
where tenant_id = $1 and (created_at, delivery_id) < ($2, $3) order by created_at desc, delivery_id desc limit 100. The row comparison is load-bearing: created_at alone is not unique, so ties straddling a page boundary are dropped or repeated, which is the same bug in a smaller window. - Index
(tenant_id, created_at, delivery_id). PostgreSQL scans a btree in either direction, so an all-DESC ORDER BY is served by an ASC index read backwards and no DESC modifiers are needed; they only matter when the ORDER BY mixes directions. Confirm the plan has no Sort node above the index scan, or the LIMIT stops being an early exit. - Price both forms: keyset is one index descent plus 100 adjacent leaf entries per page, constant regardless of depth, while OFFSET still produces and discards every skipped row, so page N costs time proportional to N times the page size and a deep page on a large table goes from milliseconds to seconds.
- Rule out updated_at as the cursor from the precondition, not from taste: a cursor column must never change value for a row already paged past. updated_at moves on every delivery attempt, so a row the client already emitted re-enters a later page and is exported twice. created_at and delivery_id are immutable, which is the whole qualification.
Worked solution 20 min
- Load about 50k deliveries for one tenant, then walk them with the OFFSET query while a writer inserts 10 rows/second, collecting every returned delivery_id.
- Compare the distinct ids collected against the set of ids that existed when the walk started, and record the shortfall.
- Repeat the walk with the keyset query and confirm every pre-existing id is returned exactly once.
- Run
explain (analyze, buffers)on page 1 and page 500 of each form and compare shared buffer hits.
Follow-up
- The client wants a snapshot as of one instant rather than a live tail. Compare a repeatable-read transaction held open, an added
created_at <= $snapshotbound, and a materialised export table. - A retention job deletes deliveries older than 90 days. What does a client mid-walk see, and does keyset pagination help at all?
- The customer wants to resume an export from yesterday's last cursor. What must be true of the cursor for that to be safe?
How would you design a real-time trending system for hot files or sear…
How would you design a real-time trending system for hot files or search queries?
Approach
- Name the read and write paths separately; they rarely have the same bottleneck.
- Name the failure you are designing for, then the recovery path.
- Choose a partition key and say what query it makes expensive.
Follow-up
- What breaks first when traffic grows ten times?
- How does this behave when that dependency is down for an hour?
Discuss how you would handle worker crashes in a distributed task syst…
Discuss how you would handle worker crashes in a distributed task system.
Approach
- State the consistency you need, and where you are willing to be stale.
- Choose a partition key and say what query it makes expensive.
- Fix the scope first: who calls this, how often, and what they do when it fails.
Follow-up
- What breaks first when traffic grows ten times?
- How does this behave when that dependency is down for an hour?
How would you efficiently represent components using templates and ins…
How would you efficiently represent components using templates and instances?
Approach
- Choose a partition key and say what query it makes expensive.
- Name the read and write paths separately; they rarely have the same bottleneck.
- State the consistency you need, and where you are willing to be stale.
Follow-up
- How does this behave when that dependency is down for an hour?
- What would you drop to keep the system up under load?
Implement a class to manage layers with properties, including methods …
Implement a class to manage layers with properties, including methods for applying changes and undoing those changes.
Approach
- State your assumptions explicitly before working the problem.
- Work from the requirement backwards to the design.
- Clarify what is being asked and what a complete answer contains.
Follow-up
- What assumption would you test first?
- How would you know your answer was wrong?
Authorisation cache with a bounded revocation window
The edge gateway serves about 30k requests/second from roughly 120 pods across three regions and may add no more than 10 ms at p99. Each request presents an API key that must resolve to an authorisation context: tenant, workspace, scopes, entitlements and credential version. The control plane that owns those rows takes tens of writes/second. A revoked credential must stop authorising within a bound you state as a number. Design the cache - what is keyed, what invalidates it, how many tiers - and specify what the gateway does for the duration of a control-plane outage.
Approach
- Fix the entry shape before the topology: key on SHA-256 of the presented secret, value is the resolved context plus the principal's auth_version and a fetched_at. Cache negative lookups too, with a much shorter TTL and a bounded-size structure, because otherwise every sprayed invalid key is a control-plane round trip, and unbounded negative entries let a sprayer evict live ones.
- Compute the control-plane read load before choosing a TTL, and notice the multiplier is pods, not regions, when the cache is in-process: distinct_active_keys x pods / TTL. At 50,000 active credentials, 120 pods and a 60 s TTL that is 100,000 reads/second against a single-writer primary with read replicas, which is not serviceable - so the design needs two tiers, a per-region shared cache in front of the control plane with the in-process cache held to a few seconds.
- State the bound as the sum of the tiers, not as a hope: with a 10 s in-process TTL over a 60 s regional TTL, worst-case staleness absent any invalidation message is 70 s, because an in-process entry can be filled from a regional entry that was itself about to expire. Publish-subscribe invalidation on every credential and entitlement mutation makes the typical case sub-second, but it is lossy under partition, so the TTL is the only enforced bound and both tiers must subscribe.
- Make auth_version propagate through the same path: a password reset or sign-out-everywhere bumps the principal and revokes its keys with no hook of its own, so the invalidation publisher has to expand principal -> credentials and publish per key, or the cache keeps serving keys whose auth_version no longer matches.
- Decide the partition behaviour in advance and write it as two rules: on a cache hit past TTL, serve from the stale entry up to a grace ceiling (say 10 minutes); on a cache miss, refuse, because authorising something never seen converts a control-plane outage into an authorisation bypass. Worst-case revoked-key lifetime during an outage is then TTL + grace, about 11 minutes, and that number is the price of not turning a control-plane outage into a total data-plane outage.
- Protect the refill path: per-key single-flight so a mass invalidation or a cold pod does not stampede the control plane, TTL jitter so entries created together do not expire together, and a small separately replicated deny-list for compromised keys that is consulted on the hot path and survives control-plane loss.
Worked solution 35 min
- Write the cache entry shape - key, value fields, and which of those fields a request actually reads on the hot path - and mark which field makes a password reset propagate.
- Compute control-plane reads/second for TTLs of 10 s, 60 s and 300 s using distinct_keys x cache_instances / TTL, once with cache_instances = 3 regions and once with cache_instances = 120 pods, and note which of the two the in-process design actually implies.
- Enumerate the four states a revocation can be in - published and received, published and dropped, control plane unreachable, pod started after the publish - and write which entry serves the next request in each.
- Write the outage policy as two rules (hit past TTL within grace: serve; miss: refuse) and compute worst-case revoked-key lifetime as the sum of both tier TTLs plus the grace.
Follow-up
- A key is found in a public repository and must stop working in seconds, not minutes. What changes, and what does it cost on the request path?
- One region is partitioned from the control plane while the control plane itself is healthy. What do that region's pods do, and how do you distinguish this from a control-plane outage?
- How would you measure the actual revocation bound in production rather than asserting it from the configuration?
Gateway p99 spikes on a five-minute cadence
edge-gateway caches each credential-to-authorisation-context decision for five minutes. p99 sits at 6 ms except for a spike to 900 ms roughly every five minutes, worst in the region with the most pods, and control-plane CPU and read latency rise in step with it. The error rate stays near zero. Customers are told a revoked credential stops authorising within 60 seconds. Give the ordered checklist that identifies the mechanism, and a fix that removes the spike without weakening the 60-second bound.
Approach
- Test periodicity before anything else: take the spike timestamps modulo the TTL in seconds. A tight cluster at a fixed offset means expiry phase, while traffic-driven spikes scatter.
- Overlay pod start times. Entries filled at first request inherit the phase of the pod that filled them, so a cohort of pods deployed together expires together and the amplitude should track cohort size rather than tenant count.
- Separate a herd from a capacity shortfall by measuring control-plane requests per second during a spike against baseline. A stampede shows a step of roughly (pods x hot keys) for one interval with hit rate collapsing to near zero, not a gradual climb that would indicate the dependency is simply undersized.
- Apply three independent controls: randomise each key's TTL by a factor drawn uniformly from something like 0.8 to 1.0 so cohorts de-phase; coalesce concurrent misses per key per pod so exactly one refresh is in flight; and serve the stale value while that refresh runs so a miss costs the stale read rather than the dependency's queue.
- Bound staleness against the published contract rather than against comfort: serve-stale is admissible only up to the 60-second revocation bound, so the TTL floor and the stale window together must stay inside it, and the published invalidation must delete the entry rather than schedule a refresh.
- Decide in advance what a miss does when the control plane is unreachable, because that is now the only uncached path: failing closed converts a dependency outage into a total outage, while extending stale service past the bound breaks the revocation promise. Pick one and configure it explicitly.
Follow-up
- Publish-subscribe invalidation is lossy under a partition. Given that, what actually enforces the 60-second bound, and what number would you put in the contract if asked to defend it?
- One tenant's key is hot enough that a single pod's coalesced refresh still matters. What changes?
- Would a shared cache tier in front of the control plane help or shift the problem, and what new failure does it add?
For someone fluent in a dynamic language who has shipped real work but has never had to say what the runtime is doing underneath. The week is built on measuring and deliberately breaking things, because the questions that expose this background are the ones where the interviewer asks why a second time.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Layer model: init, update, get by ID
- Implement the reported trio from scratch: an initialization function, a property update function and retrieval by ID. Decide and write down what an update to an unknown id or an unknown property does.
- Add apply and undo to the same class, recording only the changed property's old and new value per edit (reported question: the layer class with apply and undo)
- Before running anything, write test cases for an empty store, one layer, repeated updates to one property, and undo with no history
Deliverable: A working layer class with apply/undo and a written list of the test cases it passes, including edge cases.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Undo/redo, commit_batch and memory
- Add redo and the rule that a new edit clears the redo stack, then trace edit, edit, undo, new edit, redo aloud
- Implement commit_batch so a group of edits undoes as one entry. Then solve the bank variant that groups dependency-ordered edits into batches with Kahn's algorithm and a maximum batch size.
- Answer the reported memory question for 10K+ layers: compare snapshots with deltas, estimate bytes per history entry for each, and say what you would cap
Deliverable: An undo/redo/batch implementation plus a short written comparison of snapshot and delta storage with a size estimate.
Practice prompt ↗Practice prompt ↗03Trees, graphs and parsing from the reported algorithm set
- Implement a topological sort for the highest level of folder access a team member has, and handle cycles and two paths that grant different access levels
- Solve longest common prefix for file auto-complete twice, by scanning and with a trie, and state when the trie pays off
- Sort canvas documents left to right while respecting parent-child relationships: sort each parent's children by x and walk the tree pre-order so every subtree stays under its parent. Then validate bracket strings and convert a nested string into a nested list with a stack.
- If time remains, apply the same stack technique to two related bank titles: validating a balanced INDENT and DEDENT token stream, and converting leveled rich-text tokens into nested lists
Deliverable: Four solutions with their complexity written out, and for each the one test case most likely to catch a bug.
Practice prompt ↗Practice prompt ↗04Design: collaborative editing, comments, templates and instances
- Design multi-user editing of documents with real-time commenting: unit of conflict, ordering, optimistic updates, reconnect and catch-up, and comment anchoring when a layer is deleted
- Design components as templates and instances, with each instance stored as a reference plus sparse overrides, and trace a change to the main component through instances that override it
- Present one of the two designs aloud and write down each point where you named a technology without naming the trade-off
Deliverable: Two one-page designs, each with an explicit conflict or propagation rule and a list of trade-offs.
Practice prompt ↗Practice prompt ↗Worked solution ↗05Design: trending, async scheduling and worker crashes
- Design the real-time trending system for hot files or search queries: window, decay, per-shard top-k with a merge step, and protection against one hot key
- Design the asynchronous job scheduler for immediate tasks and delayed timers, then take the separate worker-crash question: leases, re-queueing, idempotency or fencing, and a dead-letter path
- For caching hot files, state a staleness bound as a number: pick a cache TTL, add the invalidation delay, and say what a reader sees during that window right after another user edits the file
Deliverable: Two designs, each with a failure section that says what a client sees when a component is down, plus a stated staleness bound for the hot-file cache.
Practice prompt ↗Practice prompt ↗06Behavioral: project deep-dive and first-principles stories
- Write two project stories as decision lists: the constraint behind each decision, the alternative you rejected and what you would change now
- Rehearse the bank prompts "Discuss your projects" and "Discuss a project you have worked on" at three depths: a one-minute summary, the architecture, and one hard trade-off
- Prepare one story about adapting how you communicate to a listener, and one about a hint or piece of feedback that changed your approach
Deliverable: Two decision-list stories and one recording of yourself giving each at all three depths.
Practice prompt ↗Practice prompt ↗07Mock onsite across every category
- Run a back-to-back mock: one layer-model or undo/redo coding problem, one algorithm from day 3, one design from days 4-5, and one project deep-dive
- In the coding parts, say your test cases before you run the code and explain each optimisation as a change in Big-O
- Review the notes from the mock, redo the weakest problem from scratch, and write your one-page sheet of design defaults for the onsite
Deliverable: Mock notes listing each stumble, a redone solution for the weakest problem, and a one-page sheet of design defaults.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Candidates describe the behavioral interviews as assessing cultural fit and first-principles thinking, and several behavioral items in the question bank are project discussions. Prepare stories that show how you reasoned from the constraints to each decision: what you built, why it mattered, what you cut or deferred, what risk you accepted and what you learned. Tell each story in the first person singular so your own part is clear, and be ready to go deeper on any decision you mention.
Implement a topological sort to find the highest level of folder acces…
Implement a topological sort to find the highest level of folder access for a team member.
Approach
- State the situation in two sentences and spend the rest on the reasoning.
- Pick a story where you made the decision, not one where you watched it.
- Close with what you would do differently, concretely.
Follow-up
- What did you decide not to do, and why?
- What would you do differently if you ran that again?
Unblock an engineer on a job run that finished twice
An engineer two weeks into the team brings you a job_run row showing status succeeded with an exit_code written by a worker declared dead ten minutes earlier; the retry attempt also shows succeeded. They have spent a day adding logging and are no closer. You have twenty minutes and you do not want to take the keyboard. Describe how you unblock someone: the question you ask first, what you let them find themselves, the concept you name and when, and how you check the next day that they own the fix rather than having watched you produce it.
Approach
- Ask what they expect rather than what they see: which statement set status to succeeded, and what did it check before writing? That question points directly at the update's WHERE clause, which is where the answer lives, and it costs them nothing to answer, so it does not read as a test.
- Let them build the timeline themselves from the row: queued_at, started_at, leased_until, finished_at and worker_id, on both the original run and the retry. Two different worker_ids with a lease expiry between them tells the whole story, and they will see it before you say it.
- Name the concept once the evidence has earned it. A lease bounds time; it does not prevent a write. The store has to reject a stale writer, which means the update carries a fencing token the row compares — update job_run set status = 'succeeded' where run_id = $1 and lease_token = $2 and status = 'running' — and a long garbage-collection pause or a brief partition is enough to produce what they are looking at.
- Point at the second, less obvious half and let them decide it: 'lost' exists in the status enum precisely so a run whose worker vanished is not recorded as failed, because failed asserts an outcome nobody observed and the system then bills and retries on that assertion. Ask them what these two rows should have said.
- Leave them with the next step rather than the patch — a test that kills the first worker after the sandbox exits and before the row is written — and say when you are available again, so the offer is real rather than polite.
- Check ownership the next day by what they produced, not by asking if it went well: a test that reproduces the window proves they understood it; a test that only asserts the new WHERE clause proves they copied it. Ask them to explain it to a third person and listen for whether the explanation is theirs.
Follow-up
- They propose a longer lease instead of a token. What do you say, and what breaks when legitimate runs last thirty minutes?
- How can you tell whether your explanation landed or they simply deferred to you?
- The same engineer hits a variant of this next month. What did you fail to teach the first time?
Own the incident where invoices undercounted metered usage
A metering consumer acknowledged each batch before committing the fold into usage_rollup_hourly. A rolling deploy restarted consumers mid-batch for two hours; roughly 1.4M usage_event rows were acknowledged and never folded, and 61 invoices sealed against the resulting rollups before anyone noticed. Take the owner's role. Describe an incident of comparable blast radius you owned: how it surfaced, the query that sized the loss, what you stopped first, and how the money was corrected. Give a wall-clock timeline and one thing you got wrong while it was still live.
Approach
- Open with the invariant that broke and the direction of the error, because they determine everything else: acknowledging before committing makes the consumer at-most-once, so this loses events rather than duplicating them, and loss raises no error anywhere. A listener who hears 'we lost revenue silently' knows immediately why detection took two hours.
- Size it with a stated reconciliation rather than an adjective: sum(quantity) from usage_event grouped by (tenant_id, sku, hour of occurred_at) over the window, against usage_rollup_hourly.quantity_sum on the same keys, filtered to environment='production' because staging and sandbox are metered but not billed. Then bisect by hour and tenant until single cells explain the gap. Say how long that ran and whether a replica could serve it while the incident was live.
- Separate mitigation from fix and say which came first. Mitigation is holding the sealing job, because a sealed row is frozen by design and every minute of sealing converts a recoverable rollup into an invoice correction. The fix is moving the acknowledgement after the commit, which re-introduces duplicates that the dedup check on (tenant_id, idempotency_key) must now absorb.
- State the correction path in the domain's own terms: sealed periods are never edited, so each affected tenant gets an adjustment line on the next invoice with kind='adjustment' and voided_by_line_id pointing at the line it reverses, priced against the same rate tier and carrying the watermark it priced against. That is four separate numbers — tenants affected, minor units, the cycle the adjustment lands in, and when customers were told.
- Close on one prevention control with its cost, not five: a per-hour reconciliation comparing raw sum to rollup sum that pages above a threshold. Name the threshold and the false-page rate you accepted, because a detector nobody will keep staffed is not prevention.
- Name a mistake you made inside the response window — the wrong first hypothesis, a mitigation that made it worse — rather than a design mistake from six months earlier. That is the part candidates rehearse away and interviewers weight heavily.
Follow-up
- Your fix moves the acknowledgement after the commit. What breaks now, and what absorbs it?
- One undercharged tenant has since churned. Do you bill them, and who decides?
- How would you have caught this in ten minutes instead of two hours, and what would that detector cost you in pages per week?
- 01
Walk through a project you built: what problem it solved, why it mattered, and the main trade-off you made.
- 02
Describe a technical decision you derived from constraints rather than from a familiar pattern. What were the constraints, and what did you reject?
- 03
Tell me about a time you cut or deferred scope to ship. How did you decide what to drop, and what risk did you accept?
- 04
Describe a time you changed how you explained something technical because the listener needed a different level of detail.
- 05
Tell me about a hint or piece of feedback that changed your approach partway through a problem.
- 06
Lead me through a project where you had to align other engineers on a design. What was the disagreement, and how was it settled?
Is this an official Figma interview guide?
No. It is PracHub's own research and practice material for the Software Engineer role at Figma. The rounds and questions reflect what candidates have reported, not a process Figma has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗What topics come up in the Figma Software Engineer interview?
The topics candidates report most often are the layered document model, undo/redo, commit batching (batch_apply / commit_batch), real-time updates for collaboration and comments, and graph algorithms, especially topological sorting. Reported design questions also cover a real-time trending system and components as templates and instances, plus two separate questions on asynchronous work: a job scheduler that handles immediate tasks and delayed timers, and handling worker crashes in a distributed task system.
PracHub Software Engineer practice ↗How hard are the coding questions?
The problems are often set in realistic situations such as managing layers, and the follow-ups push on memory use and edge cases. Plan to optimise and extend your first working solution, not only to reach it.
PracHub interview research ↗Do I need a background in design tools?
No. You do not need design-tool experience to prepare. Spend the time on the problems of collaborative software (latency, state synchronization and performance) instead. Practise the collaborative-editing and undo/redo questions until you can explain those problems in concrete terms.
PracHub interview research ↗How should I prepare for the system design stage?
Focus on trade-offs rather than a standard architecture. For each reported design question (multi-user editing with comments, trending, templates and instances, the async scheduler, worker crashes), explain how the system handles memory limits, many users on one file, persistence, and failures such as a crashed worker or a client reconnecting.
PracHub interview research ↗Which questions should I practise first?
Start with the layer model: a class with init, update and get-by-ID, then apply/undo/redo, then commit_batch with memory limits. It is a foundational topic, and it leads directly into the reported memory and batching questions. Next, do the graph and tree problems (topological sort for folder access, canvas ordering, bracket parsing), then the collaborative-editing design.
PracHub Software Engineer practice ↗How long does the hiring process take?
Estimates vary. One puts it at about 3-5 weeks across four stages, and another at 2-4 weeks from the first screen to a final decision. Treat both as rough and ask your recruiter for the schedule once the first screen is booked.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-24 - 02PracHub Software Engineer practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-24 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-24