New Relic · Software Engineer
Updated · 2026-09-24

New Relic Software Engineer
Interview Guide

THE 60-SECOND BRIEF

A Software Engineer at New Relic is responsible for building and maintaining the core systems that power one of the world's leading observability platforms. New Relic ingest massive volumes of telemetry data—metrics, events, logs, and traces (MELT)—from thousands of enterprise customers in real time. As an engineer here, you will design and implement highly available, low-latency, and fault-tolerant data pipelines capable of handling millions of data points per second. Your work directly impacts how developers and operations teams worldwide monitor, troubleshoot, and optimize their digital systems.

The loop does not sample the job evenly, and arguing about that in the room costs you. Daily work is mostly incremental change inside code someone else wrote, while the loop samples narrow slices of it; prepare for the slices and save the realism argument for your questions at the end.

New Relic candidates report 4 rounds · ≈ 3-5 weeks. The stages below are what candidates describe, not a published process.

Build at-least-once pipelines with explicit deduplication horizonsBound blast radius with per-tenant concurrency limitsEvolve APIs without breaking pinned SDK clients

40 min read

Practice 16 Software Engineer prompts
4Candidate experiences ↗Read their reports
16Practice promptsAcross five skill areas
3With worked solutionsIncluded in the practice prompts

A Software Engineer at New Relic is responsible for building and maintaining the core systems that power one of the world's leading observability platforms. New Relic ingest massive volumes of telemetry data—metrics, events, logs, and traces (MELT)—from thousands of enterprise customers in real time. As an engineer here, you will design and implement highly available, low-latency, and fault-tolerant data pipelines capable of handling millions of data points per second. Your work directly impacts how developers and operations teams worldwide monitor, troubleshoot, and optimize their digital systems.

The engineering organization at New Relic operates at an extraordinary scale, which introduces unique technical challenges. You will work on distributed systems, high-throughput streaming platforms, and complex data storage layers. Whether you are optimizing a Kafka ingestion pipeline, developing real-time visualization features in React, or scaling a backend service in Java, your contributions will focus on performance, reliability, and developer experience.

Successful candidates at New Relic are not just strong coders; they are pragmatic problem solvers who care deeply about system architecture, code quality, and operational excellence. The company values a culture of continuous learning, collaboration, and transparency. You will be expected to make data-driven decisions, take ownership of your services, and work closely with product managers and site reliability engineers to deliver high-impact features.

01

Recruiter Phone Screen

reported

The title covers product work, platform work, infrastructure, mobile and frontend, and those are different jobs with different loops behind them. A screening call is the cheapest place to find out which one the seat is, and asking reads as experienced rather than fussy. The questions that separate them: what the team is on call for, what the last three projects were, and whether any round happens inside an existing repository instead of a blank file. Then say which of that you have done and which you have not. Claiming the whole posting is the fastest way to be found out one round later.

What to demonstrate

  • Whether you can locate your experience inside one flavour of the role honestly instead of claiming the entire requirements list
  • Whether you name what you have not done, which an experienced screener reads as a level signal and can plan the loop around
  • Whether what you want next matches what the seat is: someone who wants greenfield work landing on a team that mostly operates an existing system is a hire that leaves within the year

How to prepare

  • Mark every line of the posting as done, adjacent or new, and write one sentence for each adjacent line naming the closest thing you actually built
  • Split your last two years into rough percentages across feature work, operating and debugging live systems, and design or review, so a question about scope gets numbers rather than adjectives
  • Bring three questions that discriminate between seats: what the team is paged for, how much of the work is changing existing code versus standing up something new, and what shipped in the last quarter
PracHub interview research ↗
02

Hiring Manager Screening

reported

Part of what is being decided is what your manager's week looks like once you are on the team: how they find out that something you shipped is broken, and whether they hear it from you or from a customer. The material that moves this round is therefore not the launch, it is the week after. Describe how you knew the change was working, which number you watched and for how long, and what you would have reverted to. A project that ends at the word shipped leaves the manager guessing about the part they care about most.

What to demonstrate

  • How a change of yours was verified in production, and whether that check existed before the deploy or was assembled afterwards once something looked wrong
  • Whether the rollback story is specific, including the case where a revert is not enough: once a migration has dropped a column or rewritten data, going back is its own change with its own risk
  • Whether you can describe an incident you caused, how it was found, and how much time passed between it starting and anyone noticing
  • Whether bad news in your account travels early and from you, or consistently arrives via somebody else

How to prepare

  • For your last significant change, write down what you watched after the deploy, for how long, and the number that would have made you revert. If nothing was watched, say that plainly instead of inventing a dashboard.
  • Work out the rollback answer for a change that touched stored data, and be able to say what made it reversible or what you would have had to do instead. That distinction is a level signal on its own.
  • Write the two-minute version of an incident you owned: what broke, how it surfaced, what you did in the first ten minutes, and the change that stopped it recurring. Get the timeline straight, because this is the story most likely to be interrupted with questions.
PracHub interview research ↗
03

Technical Assessment

reported

Most of the time lost in this format is not lost to thinking. It goes to a standard-library call you half-remember, an off-by-one in a loop bound, and a debugging loop that mutates code at random until something passes. When output is wrong, stop re-reading the whole function: take the smallest input that reproduces it and walk the state through by hand, printing intermediates if the environment allows. Guessing at a fix without a failing case you understand is how a five-minute bug becomes twenty, and the clock does not pause while you do it.

What to demonstrate

  • Whether you reach the right structure without a detour, and can write it from memory rather than only recall that one exists
  • Whether overflow is considered where the language has fixed-width integers, since a signed 32-bit value stops at 2,147,483,647 and then wraps in Java, is undefined behaviour in C++, and does not arise in Python, whose integers grow instead
  • Whether recursion depth is treated as a constraint on large inputs, given that CPython's default limit is 1000 frames and a deep recursion can exhaust the stack in any language where an iterative version would not
  • Whether a failing case is isolated and explained before any edit is made to the code

How to prepare

  • From an empty file and with no references open, implement the pieces you lean on most: a heap push and pop, an iterative DFS with an explicit stack, and a binary search whose midpoint is written lo + (hi - lo) / 2, which avoids the overflow that (lo + hi) / 2 can hit in a fixed-width integer type
  • Time yourself on the ten library calls you look up most, such as sorting with a custom comparator, splitting and joining strings, and finding the next key at or above a value in an ordered map, until the lookup is gone
  • Take a solution you know is broken and, before touching it, write one sentence naming the input, the expected value and the actual value. Repeat until you do it without deciding to.
PracHub interview research ↗
04

Virtual Onsite Loop

reported

Nobody in the room with you decides this. Interviewers typically write their rounds up separately, often before seeing anyone else's, and the outcome is settled later from those write-ups. A split panel gets resolved by whichever note carries specific evidence, so what you want out of each room is one concrete thing that person could write down: a bug you caught yourself, a trade-off you named, a decision you owned. The rest is arithmetic. The project you describe in a behavioural conversation is often the same system you sketched an hour earlier, and the two accounts have to agree.

What to demonstrate

  • Whether the scale, team size and timeline you attach to a project hold steady when that project resurfaces in a different round
  • Whether each interviewer leaves with a specific thing to cite rather than a general impression of competence
  • Whether a trade-off you defended in one round survives a challenge in another, instead of being quietly swapped for the answer the new interviewer seemed to want
  • Whether a question you have already answered earlier in the day gets the same answer at the same depth, without visible impatience

How to prepare

  • Write a one-page sheet per project fixing the figures you will quote — request volume, data size, team size, elapsed time, what broke — and say them aloud from the sheet until they come out identical every time
  • For each round on the schedule, decide in advance the one sentence you want in that person's notes, then check in a mock that you said it outright instead of leaving it to be inferred
  • Have someone ask you the same project question twice, an hour apart, and diff the two answers for numbers that moved or a trade-off that reversed
PracHub interview research ↗

4 candidate reports. Individual accounts describe a particular role and hiring cycle.

Backend Engineer

New Relic Backend Engineer Interview Experience — String Compression, Longest Increasing Path, and a Kafka Trace-Grouping Design

Onsite

Question 1: Implement a method that does basic string compression by counting consecutive repeated characters. If the compressed string is not shorter than the original string, return the original string. Discuss time complexity and space complexity Discuss edge cases Discuss code optimization Question 2: Longest Increasing Path in a Matrix. You can only move in four directions (up, down, left, r…

Read full experience
Software Engineer

New Relic Software Engineer Interview Experience: good rounds followed by silence

Technical ScreenOutcome: ghosted

The early part moved quickly. I had technical rounds that went well in a single day, then a hiring-manager conversation a few days later. That conversation felt positive, and communication during it was easy and on track. What threw me off was everything afterward. HR went quiet, and when I messaged, I got nothing back. After about a week of silence, I wanted feedback, even negative feedback, bec…

Read full experience
Software Engineer

New Relic Software Engineer interview experience: SDE-2 Codility DSA

Technical ScreenOutcome: rejected

I interviewed for an SDE-2 track and was rejected after the first technical round. It was a DSA-focused coding session on Codility, centered on data structures and algorithms. The interviewer gave a brief introduction and moved straight to the coding questions. The prompts used real-life scenarios rather than only abstract problems. I did not move past that initial technical checkpoint, and it fe…

Read full experience
Software Engineer

Software Engineer interview at New Relic: delayed Portland process

Outcome: rejected

After the recruiter call, I waited almost two weeks before the hiring manager scheduled the Portland interview. On that call, the manager seemed disengaged and did little to build momentum, so the conversation felt flat and uninterested in my experience. Another couple of weeks passed with no update. The recruiter eventually said the search had narrowed and they were pausing candidates, followed…

Read full experience

PracHub editorial advice for the preparation topics above.

01

Paginating a growing table with limit and offset

Two unrelated defects share the idiom. Correctness: rows inserted or deleted between page requests shift the window, so a consumer walking an export skips rows and sees others twice, which for a customer-facing sync is silent data loss rather than an error anyone notices. Cost: the database still produces and discards the skipped rows, so page N costs time proportional to N times the page size and a deep page on a large table degrades from milliseconds to seconds. Keyset pagination over a stable, unique, indexed ordering -- where (created_at, id) < ($1, $2) order by created_at desc, id desc limit $3 -- is constant-cost per page and immune to shifting, on the precondition that the cursor columns never change value for a row, which disqualifies updated_at as a cursor.

02

Holding money in a floating-point type, or rounding it more than once

Binary floating point cannot represent 0.01 or 0.1 exactly, so sums drift and two code paths that should agree disagree by cents nobody can trace back. The fix is integer minor units or an exact decimal type end to end, with sub-cent rates expressed as scaled integers such as micro-units, because a per-request price genuinely is smaller than a cent. The second half of the trap is rounding position: rounding each line and then summing gives a different total from summing and rounding once, and half-up and half-even diverge systematically across many lines, so rounding must happen at one named place and every downstream reader must carry the rounded value rather than recompute it from quantity and rate.

03

Quoting amortised or average cost as if it were a worst-case guarantee

Appending to a dynamic array is amortised O(1), but the append that triggers a resize copies every element, and hash lookup is constant only while the hash spreads the actual keys. Say which guarantee you are offering when the caller cares about the latency of one call rather than the total over many.

04

Assuming fixed-width integer arithmetic cannot overflow

In languages with fixed-width integers, including C, C++, Java, Go and Rust, computing a midpoint as (lo + hi) / 2 overflows once the sum passes the type's maximum, so write lo + (hi - lo) / 2 instead. Say which language you are in: arbitrary-precision integers, as in Python or Ruby, remove this specific hazard and none of the others.

Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.

13 technical prompts3 include a worked solution

Implement both Bubble Sort and Merge Sort from scratch, and explain th…

medium
data structures and algorithms

Implement both Bubble Sort and Merge Sort from scratch, and explain their respective time and space complexities.

Approach
  1. Name the brute-force solution and its complexity before improving on it.
  2. Choose the data structure from the access pattern, not from familiarity.
  3. Walk one small example through your approach before writing the whole thing.
Follow-up
  • How does this change if the input no longer fits in memory?
  • Which test case would catch an off-by-one here?

Find the minimum subsequence in an array that meets a specific target …

medium
data structures and algorithms

Find the minimum subsequence in an array that meets a specific target condition.

Approach
  1. State the target complexity and say which constraint rules the naive version out.
  2. Name the brute-force solution and its complexity before improving on it.
  3. Walk one small example through your approach before writing the whole thing.
Follow-up
  • What is the worst case, and how likely is it on real data?
  • Which test case would catch an off-by-one here?

Check if a given string is a palindrome, optimizing for space complexi…

medium
data structures and algorithms

Check if a given string is a palindrome, optimizing for space complexity.

Approach
  1. Name the brute-force solution and its complexity before improving on it.
  2. Restate the input: its shape, its size, and what is guaranteed about it.
  3. Choose the data structure from the access pattern, not from familiarity.
Follow-up
  • What is the worst case, and how likely is it on real data?
  • Which test case would catch an off-by-one here?

Implement a thread-safe stack or queue and explain how you handle conc…

medium
data structures and algorithms

Implement a thread-safe stack or queue and explain how you handle concurrent access.

Approach
  1. Choose the data structure from the access pattern, not from familiarity.
  2. Name the brute-force solution and its complexity before improving on it.
  3. State the target complexity and say which constraint rules the naive version out.
Follow-up
  • Which test case would catch an off-by-one here?
  • What is the worst case, and how likely is it on real data?

What are the SOLID design patterns, and how would you apply them to re…

medium
languages, concurrency and fundamentals

What are the SOLID design patterns, and how would you apply them to refactor a tightly coupled codebase?

Approach
  1. Reach for the cheapest primitive that closes the race, not the broadest lock.
  2. Name what is shared across threads and what owns each piece of state.
  3. Distinguish a value from a reference to it, and say which one you handed out.
Follow-up
  • Where could this allocate more than you expect?
  • What happens if two callers reach this at the same time?

Explain the internal implementation of a HashMap in your primary progr…

medium
languages, concurrency and fundamentals

Explain the internal implementation of a HashMap in your primary programming language and discuss how collisions are handled.

Approach
  1. Name what is shared across threads and what owns each piece of state.
  2. Distinguish a value from a reference to it, and say which one you handed out.
  3. Identify the window where an invariant is briefly untrue.
Follow-up
  • What happens if two callers reach this at the same time?
  • Where could this allocate more than you expect?

Locate a billing reconciliation gap without rescanning ninety million events

hardWorked solution
reconciliationdimensional-bisectionwatermarkshypothesis-testing

A tenant's sealed invoice total is 0.4% below the sum of its raw usage_event rows for the period. That tenant has 90 million events over 30 days in a table partitioned daily on ingested_at, and its rollups carry source_max_ingested_at, revision and sealed_at. Recomputing all 30 days from raw is correct, and you are not going to do it. Give the procedure that locates the divergent (workspace, sku, hour) cell, the cost of each probe, and the one query you run before any of it.

Approach
  1. Run the free query first. Sum raw quantity for the period restricted to ingested_at <= source_max_ingested_at of the sealed rollups, and compare that against the unrestricted sum. The rollup stores the watermark precisely so this can be answered without a scan. If the whole 0.4% sits above the watermark, nothing is broken: it is late data, it becomes an adjustment line, and the investigation ends in one query.
  2. Only if the gap survives that test do you bisect, and you bisect by dimension rather than by rows. Compare 30 per-day totals, then inside the offending day compare the 6 SKUs, then the workspaces, then the 24 hours. That is roughly 30 + 6 + W + 24 grouped probes, each an indexed range scan over one daily partition for one tenant, against O(N) per attempt for the naive re-fold.
  3. Quantify why naive is not merely slow but unusable mid-incident: at a generous 200,000 rows/second sequential, 90 million rows is about 7.5 minutes per attempt, you will want ten attempts, and every one competes for I/O on the same partitions live ingest is writing. The diagnostic worsens the backlog it is diagnosing.
  4. Before fetching each comparison, state what it would look like under each hypothesis. Two adjacent hours off by equal and opposite amounts is occurred_at versus ingested_at bucketing. A whole day offset by exactly N hours is a timezone applied at the wrong layer. A gap confined to one SKU in one workspace is an environment filter. The same (tenant_id, idempotency_key) present in two ingested_day partitions is the dedup horizon losing a retry that crossed midnight.
  5. Make the next bisection cheap by storing the aggregate you keep recomputing. A per-(tenant_id, ingested_day) count and quantity checksum turns step two from thirty probes into one read, and it is the same number the reconciliation job already produces.
  6. Whatever you find, the sealed period does not change value. The correction is an adjustment line pointing at the line it reverses, carrying its own source_rollup_watermark, because the original invoice is the evidence of what the customer was charged.
Worked solution 35 min
  1. Reproduce the shape locally: generate 2 million events for one synthetic tenant over 5 days, fold them, then inject three defects, namely 0.2% of events bucketed by ingested_at, a handful of duplicate idempotency_key values whose retries cross midnight, and a block of events ingested after the seal.
  2. Write the watermark-bounded query first and record how much of the gap it explains on its own.
  3. Bisect by day, then SKU, then hour, recording the probe count and the rows each probe touches.
  4. For each located cell, write down the predicted signature before querying it, then check whether the data matches the prediction.
  5. Time a full re-fold of the 2 million rows and extrapolate to 90 million.
EXPECTED RESULTThe watermark-bounded query accounts for the post-seal block entirely and removes it from the investigation. The `ingested_at` bucketing appears as adjacent hours off by equal and opposite amounts. The midnight-crossing duplicates appear as one `(tenant_id, idempotency_key)` present in two `ingested_day` partitions. Bisection reaches each cell in fewer than 70 probes.
Follow-up
  • The gap is 0.4% in one direction on one day and 0.4% the other way the next day. What does that shape rule in, and what does it rule out?
  • How do you distinguish a duplicate from a restatement, given revision and recomputed_at on the rollup?
  • Ingest is still running while you investigate. What makes your two numbers comparable at all?

Day one measures instead of guessing, under a fixed rubric, and the remaining hours are allocated in proportion to the gaps before any studying begins. The allocation is deliberately not renegotiated midweek, because the area that feels worst on day three is usually the one that is moving.

Small steps. Visible outcomes.0 / 7 completed
ONE WEEK · YOUR PACE

Prepare, practise & reflect

One practical outcome each day. Spend longer where you need it.

0 / 7 done
01Diagnostic, scored before you study anything
  • Sit a 110-minute diagnostic in four blocks: forty-five minutes on two coding problems, twenty-five on one design prompt taken to interface and data model, twenty of short-answer fundamentals, and twenty delivering two behavioural answers aloud.
  • Score each block from 0 to 3 on a fixed rubric where 3 is correct and fluent, 2 is correct but slow or prompted, 1 is partially correct and 0 is stuck, grading the artifact rather than how the attempt felt.
  • Allocate days two to five in proportion to 3 minus each block's score, write the allocation down, and commit to leaving it alone.

Deliverable: A scored rubric and a fixed hour allocation for the rest of the week.

Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗
02Largest gap: find the boundary rather than the subject
  • Split the weakest area into named sub-skills and rate each separately. For coding those are restating the problem, choosing the structure, stating the invariant, turning the invariant into loop bounds, handling empty and single-element input, and accounting for complexity out loud.
  • Attempt three items positioned just above where the rating drops off, and for each write the first move you failed to make.
  • Re-attempt one of them from blank four hours later with nothing open.

Deliverable: A sub-skill map with the two blocking sub-skills circled.

Practice prompt ↗Practice prompt ↗Practice prompt ↗
03Drill the blocking sub-skill by repeating the shape
  • Do eight short repetitions of the same shape rather than eight different problems, so what gets practised is the pattern and not the puzzle.
  • State the rule you now hold in one sentence, then test it against a case built to break it, a sliding window over an array containing negative values, or a cache-aside read path whose invalidation message is dropped.
  • Have someone else read your one-sentence rule and find the precondition you left out.

Deliverable: One rule statement with its preconditions attached and one counterexample that would have caught the incomplete version.

Practice prompt ↗Practice prompt ↗
04Second gap, plus maintenance on the strongest area
  • Run the same sub-skill decomposition on the second-largest gap in half the time.
  • Spend twenty-five timed minutes on the block you scored highest, choosing the hardest item you can still finish rather than a warm-up.
  • Write whether each area fails you on recall, on setup, or on execution, and set the fix accordingly: repetition for recall, a written checklist for setup, timed work for execution.

Deliverable: A second sub-skill map plus a one-line failure diagnosis for each area.

Practice prompt ↗Practice prompt ↗Worked solution ↗
05The gap that is not a skill
  • Record one technical and one behavioural answer, then count two things in the playback: seconds before your first clarifying question, and sentences you began without knowing where they would end.
  • Practise saying that you do not know, followed by how you would find out, without letting it soften into a guess, and practise stating a complexity or an estimate before being asked for it.
  • Redeliver one answer under a hard ninety-second cap, which forces structure ahead of detail.

Deliverable: Two recordings with a counted reduction in time-to-first-question.

Practice prompt ↗Practice prompt ↗
06Retest under day-one conditions
  • Sit the same 110-minute structure with new prompts of comparable difficulty and score it on the identical rubric.
  • For any block that did not move, change the method rather than adding hours: a block stuck at 1 usually means the practice was too varied, not too short.
  • Write down which single block you would still lose the offer on.

Deliverable: A second scored rubric placed beside the first, with one named remaining risk.

Practice prompt ↗Practice prompt ↗
07Full loop under interview conditions
  • Run a sixty-minute mock over the two blocks that moved least, with an interviewer briefed to interrupt and change direction mid-answer.
  • Write the recovery script for going blank: restate the question, state your assumption, name the first thing you would check.
  • Say every rule from the week aloud without reading it, and cut any you cannot state in a single sentence, since a rule you have to reconstruct mid-answer will not survive an interruption.

Deliverable: A one-page card holding the recovery script and only the rules you could state from memory.

Practice prompt ↗Practice prompt ↗Worked solution ↗

Expand any day for tasks and deliverables. Your progress is saved on this device.

Counting review comments or mentees proves nothing. The useful version is a specific change you approved with a reservation you stated, or one you blocked and the delay that cost. Say which standard you were holding and why it was worth the friction. A mentoring story needs the thing the other person can now do without you.

Tell me about a time when you had to learn a complex new technology un…

medium
behavioural and engineering judgement

Tell me about a time when you had to learn a complex new technology under a tight deadline to deliver a project.

Approach
  1. State the situation in two sentences and spend the rest on the reasoning.
  2. Pick a story where you made the decision, not one where you watched it.
  3. Give the blast radius: what could have broken, and what you measured.
Follow-up
  • How did you know your change caused the improvement?
  • What would you do differently if you ran that again?

Disclose a cross-tenant webhook delivery to affected customers

medium
cross-tenant leakdisclosureblast radiusauthorisation checks

An enqueue path took the subscription from one lookup and the payload from another. For nineteen minutes, webhook_delivery rows were created whose tenant_id did not match the subscription's tenant, and eleven payloads were signed and sent to four endpoints belonging to other customers. You hold payload_digest, delivery timestamps and response codes. Describe how you handle a disclosure of this kind: what the records prove, what they cannot prove, what you say before you know everything, the one code change that closes it, and which parts you personally drove.

Approach
  1. Bound the population before saying anything externally. The affected set is deliveries in the window where the event's tenant and the subscription's tenant differ; the ones that actually left are those with delivered_at set and a 2xx in last_response_code. Attempted and delivered are two different counts and a disclosure has to use the right one in the right sentence.
  2. Separate what the records prove from what they do not, and say both halves rather than the flattering one. They prove which payloads were signed, where they went, and — through payload_digest — exactly which bytes. They do not prove what the receiving system did with them, and they do not bound the window more precisely than your deploy timestamps do.
  3. Communicate on the facts you hold, with the scope stated as an upper bound: 'at most eleven payloads, four recipient endpoints, these fields, this window' is more useful and more honest than waiting a day for certainty. The field list matters more than the event count, because a customer cannot assess exposure from 'an event'.
  4. Name the code change precisely, because this class never originates in the delivery worker. Compare the event's tenant against the subscription's tenant at enqueue and again immediately before the payload is signed, and make the second comparison drop the delivery rather than log a warning. Say why one check is insufficient: the enqueue check protects against the bug you know about, the pre-signing check protects the boundary itself.
  5. Run the history question in parallel and say so: a query over historical deliveries for the same mismatch tells you whether this was nineteen minutes or a year, and you would rather find the second case yourself than have a customer find it after your disclosure.
  6. Split the response into workstreams with owners — recipients asked to delete, affected customers notified, the check landed with a test, history swept — and say which you personally drove and which you handed off. Claiming all four is not credible and claiming none is not ownership.
Follow-up
  • The historical sweep finds two more instances from last year. What changes in what you have already told people?
  • Who approves the wording, and what do you do when you are asked to soften the scope?
  • A customer asks you to prove a redelivery contained the same bytes as the original. What do you show them?

Argue against failing open when the control plane is unreachable

hard
revocationcache stalenessfail-open vs fail-closedinfluence

The gateway caches credential-to-context decisions with a sixty-second TTL. A design proposal says that when control-plane reads fail, pods should keep serving from expired entries indefinitely so a control-plane outage never becomes a product outage. You believe that converts every revocation into an unbounded one. Describe a design you argued against while it was still a live proposal: what you measured or modelled to make the case, what you conceded, who decided, and what happened afterwards. Say what would have changed your mind before the decision, not after it.

Approach
  1. Reframe it from a values argument into a bounded-staleness argument. Both sides already accept the cache; the disagreement is only about the ceiling on how long a revoked credential keeps authorising. Put a number on the table — serve stale for up to fifteen minutes, then fail closed — and make the other side argue against a number rather than against a principle.
  2. Bring arithmetic rather than adjectives: the rate of revocations with revoked_reason in ('suspected_leak','auth_version_bump'), the observed distribution of control-plane unavailability, and the product of the two, which is expected requests served by revoked credentials per outage-hour. At 30k requests/second the unbounded version is not a subtle exposure and the number says so.
  3. Concede the strong half of the opposing case first, because that is what buys you the room: failing closed turns one service's outage into a total outage across three regions, and a control plane doing tens of writes per second is not engineered to the gateway's availability target. A proposal you have not steelmanned reads as reflex.
  4. Propose the asymmetry that usually resolves this: stale entitlements cost bounded money (a quota fifteen minutes out of date over-serves by a computable amount), while a stale revocation costs unbounded access. Split the cached decision by what it authorises, give the two halves different staleness ceilings, and let the entitlement half fail open while the revocation half fails closed.
  5. State the propagation dependency plainly, since it is the part that is missed: validity is also derived from the principal's auth_version, so password reset and sign-out-everywhere flow through this same cache. A design that bounds staleness for explicit revocation and not for auth_version bumps has only fixed half of it.
  6. Say who decided, and what you did afterwards in either outcome: write the decision down with its number and a review date, and instrument the exposure you were worried about so the next round of the argument is settled by data instead of by seniority.
Follow-up
  • Publish-subscribe invalidation is lossy under a partition, and a TTL is the only hard bound. What TTL do you pick, and what does it cost you at 30k requests/second?
  • The key was revoked because it was found in a public repository. Does your answer change, and where does that urgency live in the design?
  • You lost the argument and six weeks later the failure you predicted happens. What do you say in the review, and what do you not say?
  • 01

    Tell me about a time when you had to learn a complex new technology under a tight deadline to deliver a project.

  • 02

    An enqueue path took the subscription from one lookup and the payload from another. For nineteen minutes, webhook_delivery rows were created whose tenant_id did not match the subscription's tenant, and eleven payloads were signed and sent to four endpoints belonging to other customers. You hold payload_digest, delivery timestamps and response codes. Describe how you handle a disclosure of this kind: what the records prove, what they cannot prove, what you say before you know everything, the one code change that closes it, and which parts you personally drove.

  • 03

    The gateway caches credential-to-context decisions with a sixty-second TTL. A design proposal says that when control-plane reads fail, pods should keep serving from expired entries indefinitely so a control-plane outage never becomes a product outage. You believe that converts every revocation into an unbounded one. Describe a design you argued against while it was still a live proposal: what you measured or modelled to make the case, what you conceded, who decided, and what happened afterwards. Say what would have changed your mind before the decision, not after it.

PracHub interview preparation framework ↗
Is this an official New Relic interview guide?

No. It is PracHub's own research and practice material for the Software Engineer role at New Relic. Rounds and questions reflect what candidates have reported, not a process New Relic has published, and they change over time. Confirm the current format and scope with your recruiter.

PracHub interview research ↗
How technical is the interview process compared to other tech companies?

The process is highly technical but pragmatic. New Relic prioritizes real-world engineering skills—such as writing concurrent code, designing distributed systems, and testing—over obscure algorithmic brainteasers. The take-home assignment is rigorous and serves as a major evaluation point.

PracHub interview research ↗
Can I complete the take-home challenge in any language?

This depends on the specific team and role you are interviewing for. While some generalist roles allow any language, many backend teams require the challenge to be completed in Java or Go, while frontend roles require JavaScript/TypeScript. Clarify this expectation with your recruiter early in the process.

PracHub interview research ↗
What is the company culture like for engineers?

Engineers at New Relic highly praise the culture of collaboration, transparency, and psychological safety. The team values constructive feedback, mentoring, and continuous learning. However, because the systems operate at massive scale, there is a strong emphasis on operational excellence and ownership.

PracHub interview research ↗
How long does the entire interview process take?

On average, the process takes between three to five weeks. Delays can occasionally occur during the review of the take-home assignment, as active engineers conduct these reviews alongside their daily responsibilities.

PracHub interview research ↗
Sources & methodology 3 sources ↗

Official role evidence, timestamped platform data and clearly labeled preparation advice.