The 60-second brief
Cloudflare runs edge data centers in hundreds of cities worldwide and handles millions of HTTP requests per second. Its engineering work covers edge compute services (including Cloudflare Workers), content delivery, DNS infrastructure, distributed security systems such as DDoS mitigation, and high-throughput data pipelines such as log aggregation. Software Engineers write production code mainly in Go, Rust, C++ or Python, and TypeScript/JavaScript also appears among the core languages. The role includes operational ownership: monitoring production metrics, investigating failures and taking part in on-call.
Cloudflare candidates report 5 rounds · ≈ 4-6 weeks. The stages below are what candidates describe, not a published process.
What a Software Engineer does at Cloudflare
The role involves operational ownership as well as feature work. That includes watching production metrics, investigating unexpected behaviour and edge failures, joining on-call rotations and improving resilience through post-mortems. Engineers work with product management, systems performance, SRE and information security teams. They write unit and integration tests and review each other's code.
For preparation, this means depth in three areas. First, networking and operating-system fundamentals: DNS, TCP and UDP, TLS, HTTP status codes, threads versus processes, and IPC. Second, practical coding: circular buffers, LRU caches, log parsing and rule-based rate limiting. Third, design problems that span many locations, such as a global rate limiter or encrypted log collection from edge data centers. Spread your preparation across all three instead of spending it only on algorithms.
The Cloudflare interview process
Cloudflare candidates report 5 rounds · ≈ 4-6 weeks. The stages below are what candidates describe, not a published process.
Initial Screening Call
Candidates report that the first call is with a recruiter or hiring manager and covers your background, your core technical skills and how you fit the team's work. Describe each project so that someone outside engineering could repeat it accurately: what was breaking, what you changed and what happened after. Avoid internal system names. This call is also the time to ask about format. Find out whether the technical screen is an online assessment or live pair-programming, which language you may use, and whether you read input from standard input.
What it evaluates
- Whether your background maps onto the team's work in terms a recruiter can repeat accurately
- Whether each project has a shape (problem, change, result) and is not just a list of technologies
- Whether you can say which part of a team project was yours
How to prepare
- Rewrite each headline project as two sentences with no internal system names. Say them to someone outside engineering and have them repeat the sentences back
- Prepare a specific answer to the reported question about why you want Cloudflare and which technical domain or product area interests you
- Ask about the screen's format, language choice and input handling, and write the answers down
Technical Screening
Candidates describe this stage as either an online assessment or a live pair-programming session. It covers algorithms and data structures, practical backend tasks, or networking concepts. Reported coding questions (not tied to a specific round) are practical: a thread-safe circular queue, an LRU cache, counting requests per domain from a log file, a rule-based isRateLimited function, and a sliding-window or k-Sum problem. Candidates also warn that some screens run in HackerRank-style environments where you parse standard input yourself. That costs time if you have not practised it.
What it evaluates
- Whether you can write working code for a practical data structure or parsing task and handle its edge cases
- Whether you can parse input and format output correctly without losing problem-solving time
- Whether you can answer a networking question, such as what happens between typing a URL and seeing the page, with the steps in the right order
How to prepare
- Keep a template for reading stdin in your chosen language, and practise writing it from memory before the screen
- Implement a circular buffer, an LRU cache and a per-domain log counter, with tests for their edge cases
- Rehearse the URL-to-page lifecycle and the TCP versus UDP trade-offs out loud
Full Interview Loop
Candidates describe the full loop as practical live coding, system design, debugging or pair-programming exercises, and conversations with product managers or team leads. Reported design questions (not tied to a specific round) include a global rate-limiting gateway, encrypted log collection from edge data centers under an SLA, a distributed key-value store across 100 edge databases, a high-availability HTTPS load balancer and a distributed hit counter. Candidates also report that some technical loops include AI-assisted coding or debugging tasks. Prepare each strand separately.
What it evaluates
- Whether your designs say what stays local to each location, what is synchronised, and what happens when a location loses its link
- Whether you debug by forming hypotheses from logs, requests and code and testing them, rather than guessing
- Whether you can explain a technical constraint to a product manager and reach a trade-off
- Whether code you write live handles invalid input and errors, not just the happy path
How to prepare
- Sketch the rate-limiting gateway and the encrypted log pipeline, and write one failure scenario for each
- Practise diagnosing a service you broke on purpose, saying each hypothesis before you test it
- If an AI-assisted task is possible, practise prompting an assistant, reading its diff and checking it with a test
- Prepare one story about resolving a conflict between product requirements and a technical constraint
Behavioral Assessment
Candidates call this stage the 'Orange Cloud' behavioral assessment and say it covers communication, ownership, collaboration and decisions under ambiguity. Reported behavioral questions include the hardest project you owned from architecture to launch, a severe production bug or outage, working with product managers when technical constraints conflict with requirements, and an urgent decision made with incomplete information. Use the STAR structure. Keep each answer on what you decided yourself and why, based on what you knew at the time.
What it evaluates
- Whether an outage story names the signal, the investigation, the fix and what stopped it from happening again
- Whether the reasoning behind a decision was available when you made it, not only after the outcome was known
- Whether you can say which part of a team effort was yours
How to prepare
- Write STAR answers for the four reported prompts, and cut any 'we' sentence that does not say what you did
- For the outage story, write down the metric or log that alerted you and what you measured before changing anything
- Rehearse the ambiguous-decision story up to the moment you decided, then have someone ask what you would do next
Final Conversational Call
Candidates report that some teams add a final conversational call with an executive or senior director before an offer is finalised, and other teams do not. No questions are reported for this stage. Prepare as you would for a conversation with a senior leader about your work and your interest in the team. Have a clear account of your most significant project, a specific answer about which technical domain at Cloudflare interests you and why, and questions of your own. The project facts you give here should match what you said in earlier rounds.
What it evaluates
- Whether the scale, team size and timeline you give for a project stay the same across rounds
- Whether your interest in the team's work is specific and not generic
How to prepare
- Write a one-page sheet per project that fixes the figures you will quote, and say them out loud until they come out the same every time
- Prepare three questions about the team's work that public material does not answer
- Ask your recruiter whether your team includes this call
Questions Cloudflare candidates report
These are the prompts PracHub has collected for this role and company. They are a record of what candidates describe being asked, so treat them as the shape of the conversation rather than a script Cloudflare works from.
Data structures and algorithms
Ask which resource the problem actually takes away, because it is not always time. If the input does not fit in memory, the set you were going to build is gone and the bound that matters is space. If the data arrives once and cannot be revisited, anything needing a second pass is out whatever its complexity, and a constraint banning modification of the input removes the in-place trick that would otherwise be free. Naming the scarce resource first cuts the candidates down to a handful, and it stops you optimising a dimension the problem was never short of.
- Implement a thread-safe circular queue data structure with comprehensive unit tests and edge case handling.
- Implement a memory-layered LRU Cache data structure, ensuring optimal time complexity for read and write operations.
- Given a log file containing request records for various domain names, output the total request count per domain and handle edge cases cleanly.
- Write a function
isRateLimitedthat evaluates incoming HTTP requests against configurable rule sets based on criteria like client IP, HTTP method, request path, user ID, and tenant ID.
System design and architecture
A design is a set of promises about what happens when something is slow, down, or delivered twice, and that is the part usually left unsaid. For each dependency, say what the system does when it is unavailable: serve stale data, reject the write, or accept it and settle later. For each relaxation of consistency, describe the anomaly a user would actually notice, in the product's own words, rather than naming a consistency model and moving on. The separator is volunteering the failure behaviour as part of the design instead of producing it only when pushed.
- How would you architect a global rate-limiting gateway capable of filtering millions of requests per second across distributed locations?
- Walk me through the step-by-step lifecycle of what happens when a user types a URL into a browser and presses enter, covering DNS resolution, TCP handshake, TLS negotiation, and HTTP request routing.
Behavioural and engineering judgement
A reversal is the easiest of these to tell badly, because the honest version sounds like a mistake and the flattering version sounds like nothing happened. Give the trigger: the specific observation that made continuing worse than stopping, dated to when it arrived rather than when it became undeniable. Then give what stopping cost, whether that was work discarded, a migration left half-finished, or a team told the dependency they had built against was moving. The answer that stands out names a stopping condition set before it was needed, or says plainly that none had been set.
- Describe a situation where you had to make an urgent technical decision under high ambiguity or with incomplete information.
- Why do you want to work at Cloudflare, and which specific technical domain or product area interests you most?
Practice exercises with worked solutions
The prompts above tell you what gets asked. These are PracHub exercises written to drill the same skills, with the reasoning spelled out so you can check your own answer instead of guessing whether it was good enough.
Migrate a live partitioned event table without blocking ingest
usage_event is range-partitioned daily on ingested_at, holds roughly 250M rows per day across 400 live partitions, and is written at 10-40k rows/second. Two changes are required: quantity must move from double precision to numeric(20,6), and a new environment column must become NOT NULL with a default of 'production'. Ingest cannot stop. Give the ordered plan, naming for each step the lock it takes, what that lock blocks, and roughly how long it is held. Identify the one step that cannot be rolled back cleanly once traffic depends on it.
How a strong candidate works it
- Classify the two changes before planning anything. Adding a column with a non-volatile default has been metadata-only since PostgreSQL 11, so it is cheap. Changing double precision to numeric is not binary-coercible, so
alter column ... typerewrites every partition under ACCESS EXCLUSIVE and rebuilds its indexes; on this volume that is hours of blocked ingest and is simply not an option, which is why the plan is expand-and-contract rather than one statement. - Expand: add
quantity_numeric numeric(20,6)andenvironmentwith its default on the parent. Both are catalogue-only but both take a brief ACCESS EXCLUSIVE that cascades to partitions, so run each withlock_timeoutset to a second or two and retry on failure. A queued ACCESS EXCLUSIVE request blocks every reader behind it, which is how a metadata-only change turns into an outage. - Dual-write: deploy producer code that populates both columns on every insert, and leave it running before anything reads the new column. This is the step that cannot be reverted cleanly. Once readers depend on quantity_numeric, reverting the writer leaves rows with a null there, and the gap is only discoverable by re-reading the old column, which the readers have stopped doing.
- Backfill older partitions in batches keyed on the primary key, oldest first, committing every few thousand rows with a pause between batches, and skipping the partition still receiving writes until it rotates. Each batch is an ordinary UPDATE taking row locks only. The cost is bloat and WAL rather than blocking, so watch dead tuples and let autovacuum keep pace instead of wrapping 400 partitions in one transaction.
- Make NOT NULL cheap with the three-step form:
add constraint ... check (environment is not null) not valid(brief ACCESS EXCLUSIVE, no scan), thenvalidate constraint(SHARE UPDATE EXCLUSIVE, scans while reads and writes continue), thenset not null, which from PostgreSQL 12 uses the validated check and skips its own full scan. Do this per partition, then on the parent. - Switch and contract: move reads to the new column behind a flag, verify over a full period that both columns agree on freshly written rows, drop the old column (metadata-only), and only then remove the dual-write. Any index on the new column goes on with CREATE INDEX CONCURRENTLY per partition, since CIC is not supported on a partitioned parent: create the parent index with ONLY, build each child concurrently, then ALTER INDEX ... ATTACH PARTITION until the parent index becomes valid.
Expected result (45 min): An ordered expand, dual-write, backfill, switch, contract plan with the lock named per step: metadata-only ADD COLUMN under a brief ACCESS EXCLUSIVE taken with lock_timeout; batched UPDATE backfill under row locks; CHECK NOT VALID, then VALIDATE under SHARE UPDATE EXCLUSIVE, then SET NOT NULL; CREATE INDEX CONCURRENTLY per partition with ATTACH PARTITION; and the dual-write-to-read switch identified as the step that cannot be reverted cleanly.
Check yourself
- No step in the plan holds ACCESS EXCLUSIVE for longer than the configured lock_timeout.
- The writer's error rate stays at zero through the rehearsal apart from the deliberate lock timeout.
- After VALIDATE, SET NOT NULL completes in milliseconds on a table where the naive form scanned every row.
- Freshly written rows agree on both columns for a full verification period before the old column is dropped.
Follow-ups
- A CREATE INDEX CONCURRENTLY fails halfway through the partition list. What state is the table in, how do you detect it, and what do you run?
- The producer computes quantity itself. What happens to a request already in flight when the dual-write deploy lands, and does it matter?
- Give two queries that prove the backfill is complete: one cheap enough to run every minute, one authoritative.
Decide which facts an invoice line copies instead of joining
invoice_line_item already denormalises tenant_id, which is reachable through invoice_id, and stores amount_minor even though quantity times unit_price_micros would recompute it. A reviewer asks you to normalise both away, and separately asks whether the tenant's legal name and billing address should be copied onto the invoice header. Decide each case. For every field you keep denormalised, name the read pattern or the invariant that justifies it, the anomaly the copy can develop, and the mechanism that prevents that anomaly here.
How a strong candidate works it
- Split the question into two kinds of copy, because they fail differently. A copy of a currently mutable fact is a cache: it drifts and needs invalidation. A copy of a fact frozen at write time is not a cache at all, it is the record of what happened, and normalising it away destroys information the source no longer holds.
- Keep tenant_id on the line. It costs 8 bytes, it leads every index on the table so no read is ever accidentally cross-tenant, and it turns a wrong join into an empty result rather than another tenant's money. Prevent the drift structurally: a unique constraint on invoice (invoice_id, tenant_id) plus a composite foreign key from the line on (invoice_id, tenant_id) makes a mismatched pair impossible, so the database enforces agreement instead of a code review.
- Keep amount_minor. Rounding must happen exactly once, at a named site, with a stated mode (half-even here). If readers recompute from quantity and unit_price_micros, every reader owns a rounding decision, and half-up and half-even diverge systematically across thousands of lines rather than cancelling out. A check constraint can bound the stored value but deliberately cannot re-derive it.
- Copy the legal name and billing address onto the invoice header, written once and never updated. The statement must show what was true when it was sealed, and the tenant record will change afterwards. This is a snapshot for the same reason
source_rollup_watermarkis stored per line: without it, nobody can reconstruct what the customer was told. - Name the read pattern that pays for all of it. Rendering, dispute response and export are per-tenant, per-period reads over thousands of lines that would otherwise join back to slowly changing dimensions that no longer hold the historical value. The write side is a once-per-period batch, so the extra columns cost nothing that matters.
- Concede the case where the reviewer is right: a mutable operational attribute such as the tenant's current plan name has no business on a line. If a report wants it, join. If a statement needs the plan as of the period, that is another snapshot and it belongs on the header with the rest.
Follow-ups
- Write the composite foreign key and the unique constraint it requires on the parent. What does it cost on every line insert, and what does it do to a bulk load?
- A tenant is renamed after being invoiced. Which rows change, and what does the customer see on last quarter's PDF?
- Where does currency live, and what breaks if a tenant's billing currency changes between two periods?
Seal an hour under late data with bounded memory
Metering ingest reads 256 partitions at 10,000 to 40,000 events/second. Events carry occurred_at and ingested_at, and during a producer replay the gap between them is hours. Seal each UTC hour once no more than 50 parts per million of that hour's eventual quantity can still arrive, using memory that does not grow with the size of the replay. Define the watermark, the lateness parameter and how you measure it, the structure holding open hours, and the write that performs the seal. State what an idle partition does to your watermark.
How a strong candidate works it
- Two clocks, two jobs. Bucket by
occurred_at, because that is the hour the customer is billed for, and advance the watermark oningested_at, because that is what the fold has consumed and whatsource_max_ingested_atrecords. Conflating them is what makes late data invisible. - The global watermark is the min over partitions of each partition's committed
ingested_at, not the max: the fold is trustworthy only as far as the slowest partition. The consequence is that one idle partition pins the watermark forever and nothing seals, so an idle partition must promote its watermark to wall clock after a stated idle timeout, and that timeout becomes a correctness parameter, because a partition that is slow rather than idle gets sealed past. - Choose the lateness L from the measured distribution of
ingested_at - occurred_at, weighted by quantity rather than by event count. The target is 50 ppm of the hour's quantity, and a replay is rare in events while carrying disproportionate mass, so an event-weighted quantile picks an L that is comfortably wrong at exactly the moment it matters. - Measure that quantile in bounded memory. A Greenwald-Khanna summary gives epsilon-approximate quantiles in O((1/epsilon) log(epsilon n)) space; a t-digest costs more per merge but has relative error that tightens at the tails, which is the half of the distribution you are reading at p99.99. Keep a separate summary per tenant class, because one tenant's batch importer is not the population.
- Hold open hours in a min-heap keyed by
hour_start. When the watermark advances, pop every hour withhour_end + L < Wand seal it: O(log H_open) per advance and O(1) amortised per event to touch its bucket. Memory is open hours multiplied by distinct(tenant, workspace, sku)keys, so cap the number of simultaneously open hours and spill the oldest intousage_rollup_hourlyasstatus='open'with arevisionbump. While an hour is open the row is upsertable, so the store is your overflow. - The seal itself is a conditional write:
update ... set status='sealed', sealed_at=now() where status='open' returning .... Two sealers race on every restart, and the loser must see zero rows and stop rather than write a second value. After the seal, an event for that hour is not an upsert but an adjustment, andsource_max_ingested_atis what proves it arrived afterwards.
Expected result (40 min): The quantity-weighted p99.99 is hours larger than the event-weighted one. Choosing L from the event-weighted number lets roughly the tail's 3% of quantity land after the seal, 600 times the 50 ppm target. With the min watermark and no idle timeout, the stalled partition blocks all sealing; with a 60-second timeout, the stall is sealed past and its events arrive late.
Check yourself
- Post-seal quantity as a fraction of the hour's total is at or below 50 ppm at the chosen L, measured rather than assumed.
- The second sealer's update reports zero rows affected and writes nothing.
- Heap size equals the number of open hours and does not grow with the replay's event count.
Follow-ups
- A replay starts during the sealing window for a period you are about to close. What do you do, and what is the customer-visible consequence of each option?
- Your measured quantity-weighted p99.99 lateness is six hours and the invoice must be issued at 02:00 UTC on the first. How do you reconcile those two numbers?
- How would you detect that L has drifted before it costs you an hour's quantity?
Webhook fan-out with per-endpoint isolation and backoff
One domain event fans out to every matching subscription, producing a webhook_delivery row per (subscription_id, event_id, redelivery_seq). Peak unique event rate is 20k/second; attempts run five to ten times that once fan-out and retries are counted. One customer endpoint has returned 503 for six hours and its backlog holds days of events; every other customer must be unaffected. Design the delivery system: how a worker claims work, the backoff schedule, the per-endpoint circuit breaker, the queue partitioning, and whether you offer ordering per subscription. State the delivery guarantee in one sentence.
How a strong candidate works it
- State the guarantee first, because it determines the rest: at-least-once with a stable event_id, and the consumer documented as responsible for idempotency. Exactly-once over HTTP is not deliverable - the 200 can be lost after the customer has already committed - so any design that promises it is either lying or is really offering at-most-once.
- Partition work per subscription rather than into one global pool, with a concurrency cap per subscription. With a shared pool, the endpoint that has been dead for six hours consumes workers on retries that will fail, and every other customer's delivery latency rises: head-of-line blocking across tenants is the exact failure being designed against here.
- Claim by compare-and-set with a fencing token: UPDATE webhook_delivery SET status = 'in_flight', lease_token = $new, leased_until = now() + interval '60 seconds' WHERE delivery_id = $1 AND status IN ('pending','failed_retryable') AND (leased_until IS NULL OR leased_until < now()), and make the terminal write carry AND lease_token = $new so a paused worker's late write is rejected rather than overwriting a newer attempt. Find due work through the partial index on next_attempt_at WHERE status IN ('pending','failed_retryable'), so the scan is proportional to live rows rather than to the terminal rows that outnumber them by orders of magnitude.
- Use full jitter: sleep uniformly in [0, min(cap, base x 2^(attempt-1))]. Plain exponential backoff hands a recovering endpoint its entire backlog as one synchronised herd and knocks it over again; full jitter de-correlates it. Then check the schedule actually spans the retention you promise - with base 1 s and a 3,600 s cap, twenty attempts have an expected total elapsed time of only about 4.6 hours, so a twenty-four-hour promise needs roughly fifty-nine attempts or a larger cap.
- Trip a circuit per endpoint on consecutive failures or a failure ratio over a rolling window: stop dispatching, push next_attempt_at out or mark new deliveries dropped_circuit_open, and half-open with exactly one probe rather than a batch. Bound the backlog explicitly with a per-subscription cap or retention, and decide in advance whether a recovered endpoint receives six hours of events at full rate or a pointer telling it to fetch what it missed.
- Offer ordering only as an opt-in mode of one in-flight attempt per subscription, and price it honestly: with parallel attempts a retried event overtakes a newer one, so ordering requires serialisation, and serialisation means one slow endpoint blocks its own queue entirely. That converts a shared problem into that customer's own problem, which is the right place for it, but it is still a real cost.
Expected result (35 min): At-least-once with a stable event_id; per-subscription lanes with a concurrency cap; lease-and-fence claiming; full jitter whose schedule matches the retention promised - with base 1 s and a one-hour cap, twenty attempts covers only about 4.6 hours of expected elapsed time, so twenty-four hours needs roughly fifty-nine attempts or a larger cap; a per-endpoint breaker with a single half-open probe; ordering offered only as opt-in single-in-flight with head-of-line blocking named as its price.
Check yourself
- Sum the expected sleeps in your schedule and compare against the retention window you promised the customer. If the two numbers disagree, one of them is wrong and it is usually the promise.
- Execute the terminal write with a stale lease_token and confirm it affects zero rows. If it succeeds, the lease was doing the work alone and the token is decorative.
- The dead endpoint's backlog must not be visible in any queue another subscription's workers poll. If one index scan serves all subscriptions, the isolation is nominal.
Follow-ups
- The endpoint recovers. Does it receive six hours of events at full rate, and what does that do to it?
- Trace the exact code path by which an event belonging to one tenant could be signed and sent to another tenant's endpoint.
- A customer insists they never received an event your row marks delivered. What evidence do you have, and what does payload_digest let you prove?
Webhook workers leak until OOM and drop in-flight deliveries
webhook-delivery workers grow from 400 MB to a 2 GB limit over about 36 hours, are OOM-killed, restart, and repeat. Each restart abandons in-flight attempts, so webhook_delivery rows sit in in_flight until their leases expire and the backlog spikes. The live set measured after a forced full collection also grows. The fleet serves tens of thousands of subscriptions, several thousand of which have been failing for weeks. Give an ordered checklist, the measurement separating retention from fragmentation, and the fix.
How a strong candidate works it
- Separate the two failure shapes with one measurement: track resident set size against the live set after a forced full collection. A live set that climbs monotonically is retention; a flat live set under a rising RSS is fragmentation, off-heap or native allocation, or an allocator that never returns pages. The stated symptom puts this in the first category, which rules out allocator tuning as a fix.
- Characterise the curve rather than the total. Growth linear in uptime implies an unbounded structure keyed by something that keeps arriving; step growth implies buffering a large object. Correlate the slope against event rate and separately against the count of distinct subscriptions seen, because those two diverge and only one of them will fit.
- Diff two heap snapshots an hour apart by retained size grouped by dominant root, not by allocation count, which is dominated by short-lived objects and will point at the wrong thing.
- Expect a per-subscription map with no eviction: circuit-breaker or backoff state created on first failure and never removed, so the retained set grows with endpoints that have ever failed, and the several thousand permanently dead endpoints hold theirs forever.
- Fix in two places. Bound the in-memory structure with a size-capped LRU or a TTL keyed on last use, and move state that must survive a restart onto the subscription or webhook_delivery row, since the worker holding it in memory is exactly why a restart loses it.
- Repair the second-order damage separately, because it will outlive the leak: workers claim by compare-and-set with leased_until, so a bounded lease returns in_flight rows to pending on a known schedule, and a graceful shutdown releases leases instead of waiting them out.
Follow-ups
- The backlog spike after a restart is itself a thundering herd against customer endpoints. What stops the recovery from becoming a second incident?
- Suppose the live set had been flat while RSS still climbed. Name two causes and the measurement that separates them.
- How would you size the LRU, and what does a miss on an evicted circuit-breaker entry cost a customer whose endpoint is down?
Reverse a webhook ordering decision after measuring its cost
You argued for strict per-subscription ordering in webhook-delivery, which means one in-flight attempt per subscription. It shipped. Three months later a single unresponsive endpoint holds one subscription's queue at a six-hour backlog, and two customers report events arriving out of order anyway once their own retries are counted. Describe a decision you reversed: what you originally optimised for, the measurement that changed your mind, what the reversal cost in engineering time and customer change, and how you told the people who had already built on the original guarantee.
How a strong candidate works it
- State the original decision as a trade you made knowingly. Ordering across a network requires a single in-flight attempt per subscription, and its price is head-of-line blocking whenever one endpoint is slow. 'We priced it wrong' is a much stronger opening than 'we did not realise', and it is usually the true one.
- Bring the measurement that flipped it, not the anecdote: backlog age at the ninety-ninth percentile per subscription, the share of subscriptions where one slow endpoint gated an otherwise healthy queue, and the delivery throughput lost to serialisation. A reversal justified by complaints is indistinguishable from a reversal justified by fatigue.
- Name what you learned about the guarantee itself, which is the engineering content of this story. At-least-once delivery means a retried event already arrives after newer ones and the consumer already must be idempotent, so a guarantee the customer has to defend against anyway was never worth what it cost to provide.
- Describe the migration, because reversing a published contract is the hard half and the part candidates skip. Parallel attempts behind a per-subscription flag, a monotonically increasing sequence number added to the envelope so order-sensitive consumers can sort or discard, documentation that states at-least-once and unordered in those words, and a deprecation measured in quarters because the client is a pinned SDK inside a build pipeline you cannot see or redeploy.
- Give the cost in the two currencies that matter: engineer-weeks, and how many customers had to change code. Then say who you told before it shipped rather than in a changelog afterwards, and which large customer you left on the old behaviour and for how long.
- Close with the signal you now weight differently, stated as something you would do earlier next time: measuring the blocking cost on the slowest decile of endpoints before committing to the guarantee, rather than after a customer noticed.
Follow-ups
- A customer insists they need ordering. What do you offer them that is not global serialisation?
- How did you choose the deprecation window given that you cannot see or redeploy the clients?
- What would have to be true for you to reverse back?
Where candidates lose points
Losing time in the technical screen to reading input
Candidates warn that some technical screens run in HackerRank-style environments where you read standard input and print the output yourself. Before the screen, write a small template in your language that reads all of stdin, splits lines, parses integers and strings, and prints in the exact format required. Practise it on the reported log question (count requests per domain), so that handling malformed lines, blank lines and mixed-case hostnames is already routine. Ask your recruiter which environment the screen uses.
Writing a data structure that only works on the happy path
Each reported data-structure question (circular queue, ring buffer, LRU cache) has a known edge-case trap. In a circular buffer, head == tail can mean empty or full unless you track a count or leave one slot unused, so say which one you chose. In an LRU cache, updating an existing key must move it to the front without changing the size, and eviction must remove the key from both the map and the list. For the thread-safe queue, say what blocks or fails when the queue is full or empty. Then write tests for capacity one, wrap-around, and several producers and consumers running at once.
Writing isRateLimited before deciding how the rules behave
The reported prompt matches requests on client IP, method, path, user ID and tenant ID. Before you code, settle three things. Can several rules match one request, and is the request limited if any one of them is exhausted? Is each counter keyed by the rule plus the matched values? Is the window fixed or sliding, and what happens at its boundary? Then state how memory grows (one counter per distinct key) and how idle keys expire. One global counter cannot express the prompt's per-rule limits, so design the per-rule key first.
Designing a global rate limiter or log pipeline as if every location shared one synchronous database
The reported design prompts span many locations: a rate-limiting gateway across distributed locations, log collection from edge data centers into central storage with encryption, and a key-value store across 100 edge databases. Say early what is counted or stored locally, what is synchronised, and how stale a remote count can be. Say what a location does when it loses its link to the core. For example, it buffers to disk up to a limit, and you state what gets dropped past that limit. For the log question, say where encryption happens, who holds the keys, and how the SLA is measured end to end.
Answering the URL-to-page or 429-versus-500 question at textbook level only
Networking fundamentals come up across the reported questions. When you walk through a request, include the parts candidates skip: resolver caching and TTLs, where a CDN or reverse proxy appears in the DNS answer, the TLS handshake and certificate check, and which hop terminates TLS. For status codes, know the troubleshooting path. A 429 means a limiter refused the request, so check which rule fired and whose quota ran out. A 500 means the server failed, so check the logs and recent deploys. A 504 from a proxy means the origin did not answer in time. Practise saying each one in order, out loud, without notes.
Breadth then depth, for a generalist facing a mixed loop
Four days sample coding, design, fundamentals and the practical rounds at deliberately shallow depth, which is enough to surface the topics you did not know were in scope. That map, rather than a guess made on day one, decides where the last three days go.
Day 1 — Input parsing and the log question
- Build a stdin template in the language you will use in the interview. It should read all input, split lines, parse fields and print in an exact format. Write it from memory twice.
- Solve the reported log question: count requests per domain. Handle blank and malformed lines, mixed-case hosts, trailing dots and ports. State the O(n) time and O(d) memory, where d is the number of distinct domains.
- Prepare the follow-up where the file no longer fits in memory. Explain streaming the file, and if there are too many domains to hold in memory, hashing them into partitions on disk and counting each partition.
Deliverable: A stdin template you can write from memory, and a tested domain counter with its edge cases listed.
Day 2 — Circular buffers, queues and the LRU cache
- Implement a fixed-capacity circular buffer with O(1) enqueue, dequeue and peek. Choose between a count field and an unused slot to tell full from empty, and say why.
- Make the buffer thread-safe with one mutex and two condition variables (not full, not empty). Write tests for capacity one, wrap-around, and several producers and consumers.
- Implement an LRU cache with a hash map and a doubly linked list, with O(1) get and put. Test updating an existing key and evicting at capacity. Then sketch the 'memory-layered' variant: a small hot tier in front of a larger tier, and what moves between them.
Deliverable: Three implementations with tests. Each has a one-line note on its concurrency or eviction rule.
Day 3 — Rule-based rate limiting in code
- Write isRateLimited against rules on client IP, method, path, user ID and tenant ID. Before coding, decide how multiple matching rules behave and which window type to use, and say both out loud.
- Implement a fixed window, then a sliding window (a request log or a two-bucket approximation). Compare the memory each key needs and how each behaves at the window boundary.
- Add expiry for idle keys, and a test that sends a burst straddling a window boundary.
- Warm up with one sliding-window or k-Sum frequency problem from the reported coding category, and state its complexity.
Deliverable: A rate limiter with tests for multi-rule matching and boundary bursts, plus a written comparison of the two window types.
Day 4 — Networking and OS fundamentals
- Walk through what happens from typing a URL to seeing the page, out loud: DNS resolution through the hierarchy and caches, the TCP handshake, TLS negotiation, the HTTP request, and where a CDN or reverse proxy steps in. Record it, and redo the step you rushed.
- Write short answers to four questions: the trade-offs between TCP and UDP, forward versus reverse proxies and host-header rewriting, troubleshooting a 429 versus a 500, and isolating a 504 to the network, DNS or the application.
- Cover the OS topics in the question bank: threads versus processes, interprocess communication, and the Linux commands you would use to inspect sockets, processes and logs.
Deliverable: A recorded URL walkthrough, plus a one-page sheet of short networking and OS answers.
Day 5 — Edge system design
- Design a global rate-limiting gateway across distributed locations. Cover local counters, how often they sync, how much overshoot you allow, and what a location does when it is cut off from the others.
- Design encrypted log collection from edge data centers to central storage under an SLA. Cover buffering at the edge, encryption in transit, who owns the keys, backpressure, and how you measure the SLA.
- Work through the webhook fan-out design exercise in this guide for per-tenant isolation, jittered backoff and circuit breakers. Note which ideas carry over to the rate limiter and the log pipeline.
Deliverable: Two design sketches. Each lists what is local and what is global, and walks through one failure scenario.
Day 6 — Debugging and production incidents
- Take a small service you own and break it on purpose, for example with a wrong config, an exhausted connection pool or a slow dependency. Practise diagnosing it in order from logs, requests and code, saying each hypothesis out loud.
- Work the bank question on an API service stuck under load. Then do this guide's debugging drill on workers leaking memory until OOM, and name the measurement that tells retention apart from fragmentation.
- In case your loop includes AI-assisted debugging, practise with an assistant: write a precise prompt, read the suggested diff before accepting it, and confirm the fix with a test.
Deliverable: Two narrated debugging walkthroughs, and a checklist you follow before accepting any change.
Day 7 — Orange Cloud behavioral prep and a full mock
- Write STAR answers for the reported behavioral prompts: the hardest project you owned, a severe production bug or outage, a conflict between product requirements and technical constraints, and an urgent decision under ambiguity.
- Write your answer to why you want Cloudflare. Name one specific technical domain or product area and what you have actually read or built with it.
- Write a one-page sheet per project with the figures you will quote. Then run one coding round and one behavioral round back to back, and check that the figures did not change between them.
- Prepare questions for a possible final call with an executive or senior director.
Deliverable: Four STAR answers, a specific motivation answer, and project sheets that held up in a mock.
Frequently asked questions
Is this an official Cloudflare interview guide?
No. This is PracHub's own research and practice material for the Software Engineer role at Cloudflare. The rounds and questions come from what candidates have reported, not from a process Cloudflare has published, and they change over time. Confirm the current format and scope with your recruiter.
How difficult are the technical interviews at Cloudflare compared to other major tech companies?
This guide cannot rate difficulty reliably. What the reported questions do show is the mix. Coding questions are practical: a thread-safe circular queue, an LRU cache, counting requests per domain from a log, and a rule-based rate limiter, alongside sliding-window or k-Sum problems. Networking fundamentals come up throughout, and design prompts span many edge locations. Prepare for that mix instead of algorithm puzzles alone.
Can I use any programming language I prefer during coding assessments?
Candidate reports say most general technical rounds let you choose your language. Some systems teams may prefer a language from their stack, such as Go, Rust or C++, and tell you in advance. Ask your recruiter in the first call, then practise the stdin template and your data-structure implementations in that language.
What is the typical timeframe from initial recruiter contact to an offer?
Candidate reports describe five stages over roughly four to six weeks. Some accounts stretch to eight weeks, depending on team availability, region and scheduling. Staying in touch with your recruiter or coordinator helps you keep track of where you are.
How deep does the System Design interview go?
Reported design prompts include a global rate-limiting gateway, encrypted log collection from edge data centers under an SLA, a key-value store across 100 edge databases, a high-availability HTTPS load balancer, and a distributed hit counter. They all involve many locations. Be ready to say what stays local and what is synchronised, which consistency you accept and where, and what a location does when it loses its link to the core. Also cover how you would measure an SLA end to end.
Do I need to handle input parsing myself in the technical screen?
Possibly. Candidates report that some screens run in HackerRank-style environments that expect you to read standard input and format the output yourself. Keep a stdin template you can write from memory, and practise it on the reported log-counting question so that parsing does not eat into your problem-solving time.
Will the loop include AI-assisted coding or debugging?
Candidates report that some technical loops include AI-assisted coding or debugging tasks, but not all. If yours might, practise writing precise prompts, reading every change before you accept it, spotting logic errors in generated code, and confirming each fix with a test. Ask your recruiter whether your loop includes this format.
How much networking do I need to know?
The reported questions cover the URL-to-page lifecycle (DNS, the TCP handshake, TLS, HTTP routing), how DNS resolution works and where proxies and CDNs fit into it, TCP versus UDP, and troubleshooting a 429 versus a 500. Other reported topics include forward versus reverse proxies, isolating a 504, DNS record types and TTLs, and HTTP/2 and HTTP/3. Practise explaining each one out loud and in order.
Keep practising
- Software Engineer question bank — cross-company practice.
- Interview preparation framework
- All interview guides