A Software Engineer at New Relic is responsible for building and maintaining the core systems that power one of the world's leading observability platforms. New Relic ingest massive volumes of telemetry data—metrics, events, logs, and traces (MELT)—from thousands of enterprise customers in real time. As an engineer here, you will design and implement highly available, low-latency, and fault-tolerant data pipelines capable of handling millions of data points per second. Your work directly impacts how developers and operations teams worldwide monitor, troubleshoot, and optimize their digital systems.
The engineering organization at New Relic operates at an extraordinary scale, which introduces unique technical challenges. You will work on distributed systems, high-throughput streaming platforms, and complex data storage layers. Whether you are optimizing a Kafka ingestion pipeline, developing real-time visualization features in React, or scaling a backend service in Java, your contributions will focus on performance, reliability, and developer experience.
Successful candidates at New Relic are not just strong coders; they are pragmatic problem solvers who care deeply about system architecture, code quality, and operational excellence. The company values a culture of continuous learning, collaboration, and transparency. You will be expected to make data-driven decisions, take ownership of your services, and work closely with product managers and site reliability engineers to deliver high-impact features.
Recruiter Phone Screen
reportedThe title covers product work, platform work, infrastructure, mobile and frontend, and those are different jobs with different loops behind them. A screening call is the cheapest place to find out which one the seat is, and asking reads as experienced rather than fussy. The questions that separate them: what the team is on call for, what the last three projects were, and whether any round happens inside an existing repository instead of a blank file. Then say which of that you have done and which you have not. Claiming the whole posting is the fastest way to be found out one round later.
What to demonstrate
- Whether you can locate your experience inside one flavour of the role honestly instead of claiming the entire requirements list
- Whether you name what you have not done, which an experienced screener reads as a level signal and can plan the loop around
- Whether what you want next matches what the seat is: someone who wants greenfield work landing on a team that mostly operates an existing system is a hire that leaves within the year
How to prepare
- Mark every line of the posting as done, adjacent or new, and write one sentence for each adjacent line naming the closest thing you actually built
- Split your last two years into rough percentages across feature work, operating and debugging live systems, and design or review, so a question about scope gets numbers rather than adjectives
- Bring three questions that discriminate between seats: what the team is paged for, how much of the work is changing existing code versus standing up something new, and what shipped in the last quarter
Hiring Manager Screening
reportedPart of what is being decided is what your manager's week looks like once you are on the team: how they find out that something you shipped is broken, and whether they hear it from you or from a customer. The material that moves this round is therefore not the launch, it is the week after. Describe how you knew the change was working, which number you watched and for how long, and what you would have reverted to. A project that ends at the word shipped leaves the manager guessing about the part they care about most.
What to demonstrate
- How a change of yours was verified in production, and whether that check existed before the deploy or was assembled afterwards once something looked wrong
- Whether the rollback story is specific, including the case where a revert is not enough: once a migration has dropped a column or rewritten data, going back is its own change with its own risk
- Whether you can describe an incident you caused, how it was found, and how much time passed between it starting and anyone noticing
- Whether bad news in your account travels early and from you, or consistently arrives via somebody else
How to prepare
- For your last significant change, write down what you watched after the deploy, for how long, and the number that would have made you revert. If nothing was watched, say that plainly instead of inventing a dashboard.
- Work out the rollback answer for a change that touched stored data, and be able to say what made it reversible or what you would have had to do instead. That distinction is a level signal on its own.
- Write the two-minute version of an incident you owned: what broke, how it surfaced, what you did in the first ten minutes, and the change that stopped it recurring. Get the timeline straight, because this is the story most likely to be interrupted with questions.
Technical Assessment
reportedMost of the time lost in this format is not lost to thinking. It goes to a standard-library call you half-remember, an off-by-one in a loop bound, and a debugging loop that mutates code at random until something passes. When output is wrong, stop re-reading the whole function: take the smallest input that reproduces it and walk the state through by hand, printing intermediates if the environment allows. Guessing at a fix without a failing case you understand is how a five-minute bug becomes twenty, and the clock does not pause while you do it.
What to demonstrate
- Whether you reach the right structure without a detour, and can write it from memory rather than only recall that one exists
- Whether overflow is considered where the language has fixed-width integers, since a signed 32-bit value stops at 2,147,483,647 and then wraps in Java, is undefined behaviour in C++, and does not arise in Python, whose integers grow instead
- Whether recursion depth is treated as a constraint on large inputs, given that CPython's default limit is 1000 frames and a deep recursion can exhaust the stack in any language where an iterative version would not
- Whether a failing case is isolated and explained before any edit is made to the code
How to prepare
- From an empty file and with no references open, implement the pieces you lean on most: a heap push and pop, an iterative DFS with an explicit stack, and a binary search whose midpoint is written lo + (hi - lo) / 2, which avoids the overflow that (lo + hi) / 2 can hit in a fixed-width integer type
- Time yourself on the ten library calls you look up most, such as sorting with a custom comparator, splitting and joining strings, and finding the next key at or above a value in an ordered map, until the lookup is gone
- Take a solution you know is broken and, before touching it, write one sentence naming the input, the expected value and the actual value. Repeat until you do it without deciding to.
Virtual Onsite Loop
reportedNobody in the room with you decides this. Interviewers typically write their rounds up separately, often before seeing anyone else's, and the outcome is settled later from those write-ups. A split panel gets resolved by whichever note carries specific evidence, so what you want out of each room is one concrete thing that person could write down: a bug you caught yourself, a trade-off you named, a decision you owned. The rest is arithmetic. The project you describe in a behavioural conversation is often the same system you sketched an hour earlier, and the two accounts have to agree.
What to demonstrate
- Whether the scale, team size and timeline you attach to a project hold steady when that project resurfaces in a different round
- Whether each interviewer leaves with a specific thing to cite rather than a general impression of competence
- Whether a trade-off you defended in one round survives a challenge in another, instead of being quietly swapped for the answer the new interviewer seemed to want
- Whether a question you have already answered earlier in the day gets the same answer at the same depth, without visible impatience
How to prepare
- Write a one-page sheet per project fixing the figures you will quote — request volume, data size, team size, elapsed time, what broke — and say them aloud from the sheet until they come out identical every time
- For each round on the schedule, decide in advance the one sentence you want in that person's notes, then check in a mock that you said it outright instead of leaving it to be inferred
- Have someone ask you the same project question twice, an hour apart, and diff the two answers for numbers that moved or a trade-off that reversed
4 candidate reports. Individual accounts describe a particular role and hiring cycle.
New Relic Backend Engineer Interview Experience — String Compression, Longest Increasing Path, and a Kafka Trace-Grouping Design
Question 1: Implement a method that does basic string compression by counting consecutive repeated characters. If the compressed string is not shorter than the original string, return the original string. Discuss time complexity and space complexity Discuss edge cases Discuss code optimization Question 2: Longest Increasing Path in a Matrix. You can only move in four directions (up, down, left, r…
Read full experienceNew Relic Software Engineer Interview Experience: good rounds followed by silence
The early part moved quickly. I had technical rounds that went well in a single day, then a hiring-manager conversation a few days later. That conversation felt positive, and communication during it was easy and on track. What threw me off was everything afterward. HR went quiet, and when I messaged, I got nothing back. After about a week of silence, I wanted feedback, even negative feedback, bec…
Read full experienceNew Relic Software Engineer interview experience: SDE-2 Codility DSA
I interviewed for an SDE-2 track and was rejected after the first technical round. It was a DSA-focused coding session on Codility, centered on data structures and algorithms. The interviewer gave a brief introduction and moved straight to the coding questions. The prompts used real-life scenarios rather than only abstract problems. I did not move past that initial technical checkpoint, and it fe…
Read full experienceSoftware Engineer interview at New Relic: delayed Portland process
After the recruiter call, I waited almost two weeks before the hiring manager scheduled the Portland interview. On that call, the manager seemed disengaged and did little to build momentum, so the conversation felt flat and uninterested in my experience. Another couple of weeks passed with no update. The recruiter eventually said the search had narrowed and they were pausing candidates, followed…
Read full experiencePracHub editorial advice for the preparation topics above.
Paginating a growing table with limit and offset
Two unrelated defects share the idiom. Correctness: rows inserted or deleted between page requests shift the window, so a consumer walking an export skips rows and sees others twice, which for a customer-facing sync is silent data loss rather than an error anyone notices. Cost: the database still produces and discards the skipped rows, so page N costs time proportional to N times the page size and a deep page on a large table degrades from milliseconds to seconds. Keyset pagination over a stable, unique, indexed ordering -- where (created_at, id) < ($1, $2) order by created_at desc, id desc limit $3 -- is constant-cost per page and immune to shifting, on the precondition that the cursor columns never change value for a row, which disqualifies updated_at as a cursor.
Holding money in a floating-point type, or rounding it more than once
Binary floating point cannot represent 0.01 or 0.1 exactly, so sums drift and two code paths that should agree disagree by cents nobody can trace back. The fix is integer minor units or an exact decimal type end to end, with sub-cent rates expressed as scaled integers such as micro-units, because a per-request price genuinely is smaller than a cent. The second half of the trap is rounding position: rounding each line and then summing gives a different total from summing and rounding once, and half-up and half-even diverge systematically across many lines, so rounding must happen at one named place and every downstream reader must carry the rounded value rather than recompute it from quantity and rate.
Quoting amortised or average cost as if it were a worst-case guarantee
Appending to a dynamic array is amortised O(1), but the append that triggers a resize copies every element, and hash lookup is constant only while the hash spreads the actual keys. Say which guarantee you are offering when the caller cares about the latency of one call rather than the total over many.
Assuming fixed-width integer arithmetic cannot overflow
In languages with fixed-width integers, including C, C++, Java, Go and Rust, computing a midpoint as (lo + hi) / 2 overflows once the sum passes the type's maximum, so write lo + (hi - lo) / 2 instead. Say which language you are in: arbitrary-precision integers, as in Python or Ruby, remove this specific hazard and none of the others.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Implement both Bubble Sort and Merge Sort from scratch, and explain th…
Implement both Bubble Sort and Merge Sort from scratch, and explain their respective time and space complexities.
Approach
- Name the brute-force solution and its complexity before improving on it.
- Choose the data structure from the access pattern, not from familiarity.
- Walk one small example through your approach before writing the whole thing.
Follow-up
- How does this change if the input no longer fits in memory?
- Which test case would catch an off-by-one here?
Find the minimum subsequence in an array that meets a specific target …
Find the minimum subsequence in an array that meets a specific target condition.
Approach
- State the target complexity and say which constraint rules the naive version out.
- Name the brute-force solution and its complexity before improving on it.
- Walk one small example through your approach before writing the whole thing.
Follow-up
- What is the worst case, and how likely is it on real data?
- Which test case would catch an off-by-one here?
Check if a given string is a palindrome, optimizing for space complexi…
Check if a given string is a palindrome, optimizing for space complexity.
Approach
- Name the brute-force solution and its complexity before improving on it.
- Restate the input: its shape, its size, and what is guaranteed about it.
- Choose the data structure from the access pattern, not from familiarity.
Follow-up
- What is the worst case, and how likely is it on real data?
- Which test case would catch an off-by-one here?
Implement a thread-safe stack or queue and explain how you handle conc…
Implement a thread-safe stack or queue and explain how you handle concurrent access.
Approach
- Choose the data structure from the access pattern, not from familiarity.
- Name the brute-force solution and its complexity before improving on it.
- State the target complexity and say which constraint rules the naive version out.
Follow-up
- Which test case would catch an off-by-one here?
- What is the worst case, and how likely is it on real data?
What are the SOLID design patterns, and how would you apply them to re…
What are the SOLID design patterns, and how would you apply them to refactor a tightly coupled codebase?
Approach
- Reach for the cheapest primitive that closes the race, not the broadest lock.
- Name what is shared across threads and what owns each piece of state.
- Distinguish a value from a reference to it, and say which one you handed out.
Follow-up
- Where could this allocate more than you expect?
- What happens if two callers reach this at the same time?
Explain the internal implementation of a HashMap in your primary progr…
Explain the internal implementation of a HashMap in your primary programming language and discuss how collisions are handled.
Approach
- Name what is shared across threads and what owns each piece of state.
- Distinguish a value from a reference to it, and say which one you handed out.
- Identify the window where an invariant is briefly untrue.
Follow-up
- What happens if two callers reach this at the same time?
- Where could this allocate more than you expect?
Locate a billing reconciliation gap without rescanning ninety million events
A tenant's sealed invoice total is 0.4% below the sum of its raw usage_event rows for the period. That tenant has 90 million events over 30 days in a table partitioned daily on ingested_at, and its rollups carry source_max_ingested_at, revision and sealed_at. Recomputing all 30 days from raw is correct, and you are not going to do it. Give the procedure that locates the divergent (workspace, sku, hour) cell, the cost of each probe, and the one query you run before any of it.
Approach
- Run the free query first. Sum raw quantity for the period restricted to
ingested_at <= source_max_ingested_atof the sealed rollups, and compare that against the unrestricted sum. The rollup stores the watermark precisely so this can be answered without a scan. If the whole 0.4% sits above the watermark, nothing is broken: it is late data, it becomes an adjustment line, and the investigation ends in one query. - Only if the gap survives that test do you bisect, and you bisect by dimension rather than by rows. Compare 30 per-day totals, then inside the offending day compare the 6 SKUs, then the workspaces, then the 24 hours. That is roughly 30 + 6 + W + 24 grouped probes, each an indexed range scan over one daily partition for one tenant, against O(N) per attempt for the naive re-fold.
- Quantify why naive is not merely slow but unusable mid-incident: at a generous 200,000 rows/second sequential, 90 million rows is about 7.5 minutes per attempt, you will want ten attempts, and every one competes for I/O on the same partitions live ingest is writing. The diagnostic worsens the backlog it is diagnosing.
- Before fetching each comparison, state what it would look like under each hypothesis. Two adjacent hours off by equal and opposite amounts is
occurred_atversusingested_atbucketing. A whole day offset by exactly N hours is a timezone applied at the wrong layer. A gap confined to one SKU in one workspace is an environment filter. The same(tenant_id, idempotency_key)present in twoingested_daypartitions is the dedup horizon losing a retry that crossed midnight. - Make the next bisection cheap by storing the aggregate you keep recomputing. A per-
(tenant_id, ingested_day)count and quantity checksum turns step two from thirty probes into one read, and it is the same number the reconciliation job already produces. - Whatever you find, the sealed period does not change value. The correction is an adjustment line pointing at the line it reverses, carrying its own
source_rollup_watermark, because the original invoice is the evidence of what the customer was charged.
Worked solution 35 min
- Reproduce the shape locally: generate 2 million events for one synthetic tenant over 5 days, fold them, then inject three defects, namely 0.2% of events bucketed by
ingested_at, a handful of duplicateidempotency_keyvalues whose retries cross midnight, and a block of events ingested after the seal. - Write the watermark-bounded query first and record how much of the gap it explains on its own.
- Bisect by day, then SKU, then hour, recording the probe count and the rows each probe touches.
- For each located cell, write down the predicted signature before querying it, then check whether the data matches the prediction.
- Time a full re-fold of the 2 million rows and extrapolate to 90 million.
Follow-up
- The gap is 0.4% in one direction on one day and 0.4% the other way the next day. What does that shape rule in, and what does it rule out?
- How do you distinguish a duplicate from a restatement, given
revisionandrecomputed_aton the rollup? - Ingest is still running while you investigate. What makes your two numbers comparable at all?
Explain why the metering dashboard scans every daily partition
usage_event is range-partitioned daily on ingested_at and holds tenant_id, workspace_id, environment, sku, quantity numeric(20,6), occurred_at and ingested_at. The only relevant index is on (occurred_at). A dashboard runs select sku, sum(quantity) from usage_event where tenant_id = $1 and date_trunc('hour', occurred_at) >= $2 and environment = 'production' group by sku, and EXPLAIN shows a sequential scan of every partition. Give each distinct reason, rewrite the predicate so an index can serve it, propose the index, and state the write cost its column order adds.
Approach
- Separate the three causes rather than blaming one. First,
date_trunc('hour', occurred_at)wraps the column, so the predicate is not sargable against a btree on the bare column. Second, pruning keys off ingested_at while the query constrains occurred_at, so no partition can be excluded. Third, even made sargable, (occurred_at) is not tenant-leading, so for one tenant among thousands the scan reads the whole time range and discards nearly all of it. - Rewrite the bound carefully, because the obvious rewrite is only conditionally equivalent.
date_trunc('hour', x) >= $2equalsx >= $2only when $2 is already hour-aligned; for an arbitrary $2 it meansx >= date_trunc('hour', $2) + interval '1 hour'. Normalise the parameter in the caller and leave the column bare. - Restore pruning with a second, redundant predicate on the partition key:
ingested_at >= $2 - interval '<late-data horizon>'. State both sides of it. It prunes to a handful of partitions, and it silently omits any event whose ingest lagged past that horizon, which is precisely what a producer replay produces. Either document the horizon as a stated bound, or partition on occurred_at and move the problem into the dedup window instead. - Propose
(tenant_id, occurred_at) include (sku, quantity)per partition. A partial indexwhere environment = 'production'mostly saves size rather than selectivity, since production dominates the three environments; take it if non-production is a meaningful share and skip it otherwise. - Price the write path honestly. At roughly 250M rows/day each extra index is another insert plus WAL per row, and a tenant-leading key scatters inserts across one hot leaf per active tenant instead of appending to a single rightmost leaf, so page dirtying and random I/O both rise. An INCLUDE payload widens every leaf entry and enlarges the index accordingly.
- Add the index-only-scan caveat before someone reports it as a regression: on a freshly appended table the visibility map is not yet set for recent pages, so the INCLUDE columns still cost heap fetches until autovacuum has been through, and the newest hour is exactly the data the dashboard reads.
Worked solution 30 min
- Build 30 daily partitions with skewed tenants, one holding about 40% of the rows, then ANALYZE.
- Run
explain (analyze, buffers)on the original query and record how many partitions were scanned and the rows removed by filter. - Apply the rewritten predicate and the index, re-run, and confirm the plan lists only the partitions inside the ingested_at bound.
- Re-run with $2 set to a non-hour-aligned timestamp and confirm the rewritten and original predicates return identical rows.
- Insert an event with ingested_at six hours past occurred_at and check whether the pruning predicate excludes it.
Follow-up
- CREATE INDEX CONCURRENTLY is not supported on a partitioned parent. Give the sequence that gets this index onto 400 existing partitions without blocking ingest.
- One tenant holds 200 times the median row count and the dashboard still times out for them with the index in place. What changes?
- Should this read hit
usage_rollup_hourlyinstead? State what that costs in freshness and what the watermark lets you promise.
Migrate a live partitioned event table without blocking ingest
usage_event is range-partitioned daily on ingested_at, holds roughly 250M rows per day across 400 live partitions, and is written at 10-40k rows/second. Two changes are required: quantity must move from double precision to numeric(20,6), and a new environment column must become NOT NULL with a default of 'production'. Ingest cannot stop. Give the ordered plan, naming for each step the lock it takes, what that lock blocks, and roughly how long it is held. Identify the one step that cannot be rolled back cleanly once traffic depends on it.
Approach
- Classify the two changes before planning anything. Adding a column with a non-volatile default has been metadata-only since PostgreSQL 11, so it is cheap. Changing double precision to numeric is not binary-coercible, so
alter column ... typerewrites every partition under ACCESS EXCLUSIVE and rebuilds its indexes; on this volume that is hours of blocked ingest and is simply not an option, which is why the plan is expand-and-contract rather than one statement. - Expand: add
quantity_numeric numeric(20,6)andenvironmentwith its default on the parent. Both are catalogue-only but both take a brief ACCESS EXCLUSIVE that cascades to partitions, so run each withlock_timeoutset to a second or two and retry on failure. A queued ACCESS EXCLUSIVE request blocks every reader behind it, which is how a metadata-only change turns into an outage. - Dual-write: deploy producer code that populates both columns on every insert, and leave it running before anything reads the new column. This is the step that cannot be reverted cleanly. Once readers depend on quantity_numeric, reverting the writer leaves rows with a null there, and the gap is only discoverable by re-reading the old column, which the readers have stopped doing.
- Backfill older partitions in batches keyed on the primary key, oldest first, committing every few thousand rows with a pause between batches, and skipping the partition still receiving writes until it rotates. Each batch is an ordinary UPDATE taking row locks only. The cost is bloat and WAL rather than blocking, so watch dead tuples and let autovacuum keep pace instead of wrapping 400 partitions in one transaction.
- Make NOT NULL cheap with the three-step form:
add constraint ... check (environment is not null) not valid(brief ACCESS EXCLUSIVE, no scan), thenvalidate constraint(SHARE UPDATE EXCLUSIVE, scans while reads and writes continue), thenset not null, which from PostgreSQL 12 uses the validated check and skips its own full scan. Do this per partition, then on the parent. - Switch and contract: move reads to the new column behind a flag, verify over a full period that both columns agree on freshly written rows, drop the old column (metadata-only), and only then remove the dual-write. Any index on the new column goes on with CREATE INDEX CONCURRENTLY per partition, since CIC is not supported on a partitioned parent: create the parent index with ONLY, build each child concurrently, then ALTER INDEX ... ATTACH PARTITION until the parent index becomes valid.
Follow-up
- A CREATE INDEX CONCURRENTLY fails halfway through the partition list. What state is the table in, how do you detect it, and what do you run?
- The producer computes quantity itself. What happens to a request already in flight when the dual-write deploy lands, and does it matter?
- Give two queries that prove the backfill is complete: one cheap enough to run every minute, one authoritative.
Design a real-time monitoring system capable of tracking server status…
Design a real-time monitoring system capable of tracking server status across multiple enterprise clients (e.g., handling high availability, message queues, and caching).
Approach
- Name the read and write paths separately; they rarely have the same bottleneck.
- Name the failure you are designing for, then the recovery path.
- Fix the scope first: who calls this, how often, and what they do when it fails.
Follow-up
- What breaks first when traffic grows ten times?
- What would you drop to keep the system up under load?
Design a high-throughput TCP socket server that can handle millions of…
Design a high-throughput TCP socket server that can handle millions of concurrent connections and write incoming data to a file.
Approach
- Choose a partition key and say what query it makes expensive.
- Name the failure you are designing for, then the recovery path.
- State the consistency you need, and where you are willing to be stale.
Follow-up
- What would you drop to keep the system up under load?
- How does this behave when that dependency is down for an hour?
Sealing a billing period against late-arriving usage
billing seals a tenant's period once the metering watermark passes the period end, prices the sealed rollups against the plan (included allowance, tier boundaries, negotiated discount) and writes invoice_line_item rows: quantity numeric(28,6), unit_price_micros bigint, amount_minor bigint, currency char(3), source_rollup_watermark timestamptz. The job is re-run after failures and two workers may attempt the same tenant. Events legitimately arrive with occurred_at inside the period and ingested_at after it. Specify the seal transition, the idempotency key for every write including the payment-processor call, the rounding position and mode, and the fate of a post-seal event.
Approach
- Make sealing a conditional write rather than a read followed by a write: UPDATE usage_rollup_hourly SET status = 'sealed', sealed_at = now() WHERE tenant_id = $1 AND hour_start >= $2 AND hour_start < $3 AND status = 'open' RETURNING rollup_id. Two concurrent sealers cannot both win because the loser's predicate no longer matches, and the loser learns it lost from an empty result rather than from a lock timeout.
- Gate the seal on the watermark, not on the clock - the period closes when the fold is trustworthy past period_end - but cap the wait, because an unsealed period blocks the whole billing run. Seal anyway after a stated maximum (six hours past period end is a reasonable default) and let whatever arrives afterwards become an adjustment. The trade-off is named in one line: adjustments are cheap, a missed billing cycle is not.
- Make every write converge under re-run. Line items are idempotent on (invoice_id, sku, rate_tier) with ON CONFLICT DO UPDATE permitted only while the invoice is draft; the payment-processor call carries (tenant_id, billing_period_start) as its idempotency key, because a timeout there is an unknown outcome, not a failed one, and a plain retry is how a customer gets charged twice.
- Keep the arithmetic exact and round exactly once. quantity stays numeric, the rate is an integer unit_price_micros in millionths of a minor unit because per-request prices are genuinely below a cent, and amount_minor = round_half_even(quantity x unit_price_micros / 1,000,000) is computed once per line and stored. Tiering walks the boundaries in order emitting one line per (sku, rate_tier) with tier 0 as the included allowance at price zero; the invoice total is the sum of stored amount_minor values and is never recomputed downstream from quantity and rate.
- Handle the post-seal event as a forward-only correction. The sealed rollup is frozen, so the difference between the sealed value and the restated value becomes a new invoice_line_item with kind = 'adjustment' and voided_by_line_id pointing at the line it reverses, carrying its own source_rollup_watermark. Nothing is edited in place, because the original line is the only evidence of what the customer was actually charged and is precisely what a dispute, a refund or an audit asks to see.
Worked solution 35 min
- Write the seal as a single conditional UPDATE ... WHERE status = 'open' RETURNING and state exactly what the losing worker sees.
- Price one line by hand: quantity 4,318,904.250000 at unit_price_micros 1200 gives 5,182.6851 minor units and 5,183 after a single half-even rounding. Then do quantity 2,500,000.000000 at unit_price_micros 1, which lands exactly on 2.5 and rounds to 2 under half-even but 3 under half-up.
- Take four hundred synthetic lines with fractional minor units and compute the total two ways - sum of per-line rounded amounts, and a single rounding of the summed exact amounts - then record the gap.
- Trace one event with occurred_at inside the period and ingested_at two days after the seal all the way to the row that eventually reflects it.
Follow-up
- Two workers start the same tenant's seal a millisecond apart and the winner crashes after sealing but before writing any line. What does the second worker observe, and is the resulting invoice correct?
- Show the divergence between rounding each line and rounding the total once over four hundred lines, and say which direction it goes.
- A customer disputes a charge from two quarters ago. Which rows answer it, and which single design decision made that answer possible?
Webhook workers leak until OOM and drop in-flight deliveries
webhook-delivery workers grow from 400 MB to a 2 GB limit over about 36 hours, are OOM-killed, restart, and repeat. Each restart abandons in-flight attempts, so webhook_delivery rows sit in in_flight until their leases expire and the backlog spikes. The live set measured after a forced full collection also grows. The fleet serves tens of thousands of subscriptions, several thousand of which have been failing for weeks. Give an ordered checklist, the measurement separating retention from fragmentation, and the fix.
Approach
- Separate the two failure shapes with one measurement: track resident set size against the live set after a forced full collection. A live set that climbs monotonically is retention; a flat live set under a rising RSS is fragmentation, off-heap or native allocation, or an allocator that never returns pages. The stated symptom puts this in the first category, which rules out allocator tuning as a fix.
- Characterise the curve rather than the total. Growth linear in uptime implies an unbounded structure keyed by something that keeps arriving; step growth implies buffering a large object. Correlate the slope against event rate and separately against the count of distinct subscriptions seen, because those two diverge and only one of them will fit.
- Diff two heap snapshots an hour apart by retained size grouped by dominant root, not by allocation count, which is dominated by short-lived objects and will point at the wrong thing.
- Expect a per-subscription map with no eviction: circuit-breaker or backoff state created on first failure and never removed, so the retained set grows with endpoints that have ever failed, and the several thousand permanently dead endpoints hold theirs forever.
- Fix in two places. Bound the in-memory structure with a size-capped LRU or a TTL keyed on last use, and move state that must survive a restart onto the subscription or webhook_delivery row, since the worker holding it in memory is exactly why a restart loses it.
- Repair the second-order damage separately, because it will outlive the leak: workers claim by compare-and-set with leased_until, so a bounded lease returns in_flight rows to pending on a known schedule, and a graceful shutdown releases leases instead of waiting them out.
Follow-up
- The backlog spike after a restart is itself a thundering herd against customer endpoints. What stops the recovery from becoming a second incident?
- Suppose the live set had been flat while RSS still climbed. Name two causes and the measurement that separates them.
- How would you size the LRU, and what does a miss on an evicted circuit-breaker entry cost a customer whose endpoint is down?
Day one measures instead of guessing, under a fixed rubric, and the remaining hours are allocated in proportion to the gaps before any studying begins. The allocation is deliberately not renegotiated midweek, because the area that feels worst on day three is usually the one that is moving.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Diagnostic, scored before you study anything
- Sit a 110-minute diagnostic in four blocks: forty-five minutes on two coding problems, twenty-five on one design prompt taken to interface and data model, twenty of short-answer fundamentals, and twenty delivering two behavioural answers aloud.
- Score each block from 0 to 3 on a fixed rubric where 3 is correct and fluent, 2 is correct but slow or prompted, 1 is partially correct and 0 is stuck, grading the artifact rather than how the attempt felt.
- Allocate days two to five in proportion to 3 minus each block's score, write the allocation down, and commit to leaving it alone.
Deliverable: A scored rubric and a fixed hour allocation for the rest of the week.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗02Largest gap: find the boundary rather than the subject
- Split the weakest area into named sub-skills and rate each separately. For coding those are restating the problem, choosing the structure, stating the invariant, turning the invariant into loop bounds, handling empty and single-element input, and accounting for complexity out loud.
- Attempt three items positioned just above where the rating drops off, and for each write the first move you failed to make.
- Re-attempt one of them from blank four hours later with nothing open.
Deliverable: A sub-skill map with the two blocking sub-skills circled.
Practice prompt ↗Practice prompt ↗Practice prompt ↗03Drill the blocking sub-skill by repeating the shape
- Do eight short repetitions of the same shape rather than eight different problems, so what gets practised is the pattern and not the puzzle.
- State the rule you now hold in one sentence, then test it against a case built to break it, a sliding window over an array containing negative values, or a cache-aside read path whose invalidation message is dropped.
- Have someone else read your one-sentence rule and find the precondition you left out.
Deliverable: One rule statement with its preconditions attached and one counterexample that would have caught the incomplete version.
Practice prompt ↗Practice prompt ↗04Second gap, plus maintenance on the strongest area
- Run the same sub-skill decomposition on the second-largest gap in half the time.
- Spend twenty-five timed minutes on the block you scored highest, choosing the hardest item you can still finish rather than a warm-up.
- Write whether each area fails you on recall, on setup, or on execution, and set the fix accordingly: repetition for recall, a written checklist for setup, timed work for execution.
Deliverable: A second sub-skill map plus a one-line failure diagnosis for each area.
Practice prompt ↗Practice prompt ↗Worked solution ↗05The gap that is not a skill
- Record one technical and one behavioural answer, then count two things in the playback: seconds before your first clarifying question, and sentences you began without knowing where they would end.
- Practise saying that you do not know, followed by how you would find out, without letting it soften into a guess, and practise stating a complexity or an estimate before being asked for it.
- Redeliver one answer under a hard ninety-second cap, which forces structure ahead of detail.
Deliverable: Two recordings with a counted reduction in time-to-first-question.
Practice prompt ↗Practice prompt ↗06Retest under day-one conditions
- Sit the same 110-minute structure with new prompts of comparable difficulty and score it on the identical rubric.
- For any block that did not move, change the method rather than adding hours: a block stuck at 1 usually means the practice was too varied, not too short.
- Write down which single block you would still lose the offer on.
Deliverable: A second scored rubric placed beside the first, with one named remaining risk.
Practice prompt ↗Practice prompt ↗07Full loop under interview conditions
- Run a sixty-minute mock over the two blocks that moved least, with an interviewer briefed to interrupt and change direction mid-answer.
- Write the recovery script for going blank: restate the question, state your assumption, name the first thing you would check.
- Say every rule from the week aloud without reading it, and cut any you cannot state in a single sentence, since a rule you have to reconstruct mid-answer will not survive an interruption.
Deliverable: A one-page card holding the recovery script and only the rules you could state from memory.
Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Counting review comments or mentees proves nothing. The useful version is a specific change you approved with a reservation you stated, or one you blocked and the delay that cost. Say which standard you were holding and why it was worth the friction. A mentoring story needs the thing the other person can now do without you.
Tell me about a time when you had to learn a complex new technology un…
Tell me about a time when you had to learn a complex new technology under a tight deadline to deliver a project.
Approach
- State the situation in two sentences and spend the rest on the reasoning.
- Pick a story where you made the decision, not one where you watched it.
- Give the blast radius: what could have broken, and what you measured.
Follow-up
- How did you know your change caused the improvement?
- What would you do differently if you ran that again?
Disclose a cross-tenant webhook delivery to affected customers
An enqueue path took the subscription from one lookup and the payload from another. For nineteen minutes, webhook_delivery rows were created whose tenant_id did not match the subscription's tenant, and eleven payloads were signed and sent to four endpoints belonging to other customers. You hold payload_digest, delivery timestamps and response codes. Describe how you handle a disclosure of this kind: what the records prove, what they cannot prove, what you say before you know everything, the one code change that closes it, and which parts you personally drove.
Approach
- Bound the population before saying anything externally. The affected set is deliveries in the window where the event's tenant and the subscription's tenant differ; the ones that actually left are those with delivered_at set and a 2xx in last_response_code. Attempted and delivered are two different counts and a disclosure has to use the right one in the right sentence.
- Separate what the records prove from what they do not, and say both halves rather than the flattering one. They prove which payloads were signed, where they went, and — through payload_digest — exactly which bytes. They do not prove what the receiving system did with them, and they do not bound the window more precisely than your deploy timestamps do.
- Communicate on the facts you hold, with the scope stated as an upper bound: 'at most eleven payloads, four recipient endpoints, these fields, this window' is more useful and more honest than waiting a day for certainty. The field list matters more than the event count, because a customer cannot assess exposure from 'an event'.
- Name the code change precisely, because this class never originates in the delivery worker. Compare the event's tenant against the subscription's tenant at enqueue and again immediately before the payload is signed, and make the second comparison drop the delivery rather than log a warning. Say why one check is insufficient: the enqueue check protects against the bug you know about, the pre-signing check protects the boundary itself.
- Run the history question in parallel and say so: a query over historical deliveries for the same mismatch tells you whether this was nineteen minutes or a year, and you would rather find the second case yourself than have a customer find it after your disclosure.
- Split the response into workstreams with owners — recipients asked to delete, affected customers notified, the check landed with a test, history swept — and say which you personally drove and which you handed off. Claiming all four is not credible and claiming none is not ownership.
Follow-up
- The historical sweep finds two more instances from last year. What changes in what you have already told people?
- Who approves the wording, and what do you do when you are asked to soften the scope?
- A customer asks you to prove a redelivery contained the same bytes as the original. What do you show them?
Argue against failing open when the control plane is unreachable
The gateway caches credential-to-context decisions with a sixty-second TTL. A design proposal says that when control-plane reads fail, pods should keep serving from expired entries indefinitely so a control-plane outage never becomes a product outage. You believe that converts every revocation into an unbounded one. Describe a design you argued against while it was still a live proposal: what you measured or modelled to make the case, what you conceded, who decided, and what happened afterwards. Say what would have changed your mind before the decision, not after it.
Approach
- Reframe it from a values argument into a bounded-staleness argument. Both sides already accept the cache; the disagreement is only about the ceiling on how long a revoked credential keeps authorising. Put a number on the table — serve stale for up to fifteen minutes, then fail closed — and make the other side argue against a number rather than against a principle.
- Bring arithmetic rather than adjectives: the rate of revocations with revoked_reason in ('suspected_leak','auth_version_bump'), the observed distribution of control-plane unavailability, and the product of the two, which is expected requests served by revoked credentials per outage-hour. At 30k requests/second the unbounded version is not a subtle exposure and the number says so.
- Concede the strong half of the opposing case first, because that is what buys you the room: failing closed turns one service's outage into a total outage across three regions, and a control plane doing tens of writes per second is not engineered to the gateway's availability target. A proposal you have not steelmanned reads as reflex.
- Propose the asymmetry that usually resolves this: stale entitlements cost bounded money (a quota fifteen minutes out of date over-serves by a computable amount), while a stale revocation costs unbounded access. Split the cached decision by what it authorises, give the two halves different staleness ceilings, and let the entitlement half fail open while the revocation half fails closed.
- State the propagation dependency plainly, since it is the part that is missed: validity is also derived from the principal's auth_version, so password reset and sign-out-everywhere flow through this same cache. A design that bounds staleness for explicit revocation and not for auth_version bumps has only fixed half of it.
- Say who decided, and what you did afterwards in either outcome: write the decision down with its number and a review date, and instrument the exposure you were worried about so the next round of the argument is settled by data instead of by seniority.
Follow-up
- Publish-subscribe invalidation is lossy under a partition, and a TTL is the only hard bound. What TTL do you pick, and what does it cost you at 30k requests/second?
- The key was revoked because it was found in a public repository. Does your answer change, and where does that urgency live in the design?
- You lost the argument and six weeks later the failure you predicted happens. What do you say in the review, and what do you not say?
- 01
Tell me about a time when you had to learn a complex new technology under a tight deadline to deliver a project.
- 02
An enqueue path took the subscription from one lookup and the payload from another. For nineteen minutes, webhook_delivery rows were created whose tenant_id did not match the subscription's tenant, and eleven payloads were signed and sent to four endpoints belonging to other customers. You hold payload_digest, delivery timestamps and response codes. Describe how you handle a disclosure of this kind: what the records prove, what they cannot prove, what you say before you know everything, the one code change that closes it, and which parts you personally drove.
- 03
The gateway caches credential-to-context decisions with a sixty-second TTL. A design proposal says that when control-plane reads fail, pods should keep serving from expired entries indefinitely so a control-plane outage never becomes a product outage. You believe that converts every revocation into an unbounded one. Describe a design you argued against while it was still a live proposal: what you measured or modelled to make the case, what you conceded, who decided, and what happened afterwards. Say what would have changed your mind before the decision, not after it.
Is this an official New Relic interview guide?
No. It is PracHub's own research and practice material for the Software Engineer role at New Relic. Rounds and questions reflect what candidates have reported, not a process New Relic has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗How technical is the interview process compared to other tech companies?
The process is highly technical but pragmatic. New Relic prioritizes real-world engineering skills—such as writing concurrent code, designing distributed systems, and testing—over obscure algorithmic brainteasers. The take-home assignment is rigorous and serves as a major evaluation point.
PracHub interview research ↗Can I complete the take-home challenge in any language?
This depends on the specific team and role you are interviewing for. While some generalist roles allow any language, many backend teams require the challenge to be completed in Java or Go, while frontend roles require JavaScript/TypeScript. Clarify this expectation with your recruiter early in the process.
PracHub interview research ↗What is the company culture like for engineers?
Engineers at New Relic highly praise the culture of collaboration, transparency, and psychological safety. The team values constructive feedback, mentoring, and continuous learning. However, because the systems operate at massive scale, there is a strong emphasis on operational excellence and ownership.
PracHub interview research ↗How long does the entire interview process take?
On average, the process takes between three to five weeks. Delays can occasionally occur during the review of the take-home assignment, as active engineers conduct these reviews alongside their daily responsibilities.
PracHub interview research ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-24 - 02PracHub Software Engineer practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-24 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-24