Candidate reports describe the Axon Software Engineer role as spanning hardware, cloud infrastructure, real-time communications and AI. The products named in those reports include Evidence.com, a cloud digital evidence management platform, as well as body-worn camera ingestion pipelines, real-time officer safety systems, dispatch routing and AI transcription services. Depending on the team, the role can involve backend services, device-facing APIs or frontend work for these workflows.
The reported coding questions are standard problems with domain framing. They include simplifying a GPS path by dropping collinear middle points, returning an officer's location at the nearest preceding timestamp, BFS shortest path on a city grid with walls, O(1) body camera check-in and check-out, and resolving device access through nested user groups. Under the story, each one tests binary search, geometry with integer arithmetic, graph traversal or hash-map state. The quickest way to prepare is to see which one it is and name the access pattern before you write any code.
The reported design questions include a video evidence upload and auto-transcription pipeline, a device logging system that must absorb reconnect surges at shift change, a physical evidence storage backend with API, schema and permission model, and a real-time voice and dispatch pipeline. The topics listed for this role are system design, hash tables, handling large blobs, time and space complexity analysis, and chain of custody. For each design, prepare large binary uploads, async processing through queues, peak-load shaping, access control and an audit trail.
Reports also say that live coding often happens in a plain shared editor without autocomplete, so practise writing complete code by hand and tracing it on a small input. The drills in this guide on API keys, pagination, job graphs, billing, webhooks and worker memory are editorial practice for the same skills. They are not reported Axon questions.
Initial HR Conversation
reportedReports describe this first stage as either a conversation with HR or an asynchronous technical screening assessment, so find out which one it is before you prepare for it. For a conversation, have a short account of your background and a real reason for applying ready. A reported behavioral prompt pairs 'why Axon' with the ethics of public safety technology, so it helps to have thought about both. For an asynchronous assessment, reports say screens at this stage emphasize fundamental data structures, array manipulation and basic object-oriented design.
What to demonstrate
- Whether you can summarise your background and one recent project clearly and specifically, without a long preamble
- If the stage is a conversation, whether your reason for applying is specific enough to hold up when the ethics of public safety technology come up
- If the stage is an asynchronous assessment, whether your code is correct on fundamental data-structure and array problems with no interviewer to ask
How to prepare
- Ask the recruiter in writing whether this stage is an HR conversation or an asynchronous assessment, and for an assessment, which language and environment it uses
- Write and say aloud a short background summary that ends on one project you owned, with the decision you made and its result
- If an assessment is likely, practise array, hash-map and small class-design problems under a self-imposed cap in the environment you will use
Technical Screen
reportedReports describe the Technical Screen as a Zoom or phone call with an engineering manager or senior engineer. It covers your technical background and then a live coding or system design problem. Many live screens reportedly use a simple shared text editor without autocomplete or compilation. Prepare for both formats. For coding, write a correct plain version first, state its complexity, then improve it. For design, open with clarifying questions about scale, users and failure before you draw anything.
What to demonstrate
- Whether your account of past technical work holds up under follow-up questions about what you built and why
- Whether live code is correct on edge cases you check without being asked: empty input, a single element, duplicates, no valid answer
- Whether the complexity you state matches the code you actually wrote
- If the problem is a design question, whether you pin down requirements and scale before you pick components
How to prepare
- Practise the reported coding prompts (nearest preceding timestamp, GPS path simplification, grid BFS, body camera check-in/check-out, nested group access) in a plain editor, and hand-trace each one before running it
- Prepare a short walkthrough of one system you worked on: what it did, the hardest decision in it, and what you would change
- Keep a clarifying-question checklist for design (who calls it, peak load, data size, consistency needs, failure behaviour) and use it for the first moves of any design prompt
Final Interview Loop
reportedReports describe the Final Interview Loop as 4 to 5 individual sessions that include live coding and system design. Reports say it can also cover low-level object-oriented design, such as check-in/check-out asset tracking or game-state logic like Minesweeper, and that senior loops lean toward distributed systems design. The reported design questions involve large video uploads, reconnect surges, chain of custody and real-time audio. Reports also say some design prompts are deliberately broad, so scoping the problem is part of the answer.
What to demonstrate
- Whether coding solutions on domain-framed problems are correct and use the data structure the access pattern calls for
- Whether object-oriented designs keep state transitions explicit and reject invalid ones, such as checking out a camera that is already out
- Whether system designs handle peak load, large binary data, access control and an auditable record of who touched what
- Whether you narrow a broad prompt with clarifying questions and state your assumptions before designing
How to prepare
- Outline each reported design prompt (video evidence and transcription, shift-change device logging, physical evidence storage, real-time voice dispatch) with requirements, data model, write path, read path and one failure you designed for
- Build one small stateful class end to end, such as the body camera check-in/check-out system with explicit states and error returns, and explain the complexity of each operation
- Rehearse coding, design and behavioral sessions back to back so your structure still holds in the later ones
Axon Culture Evaluation
reportedReports describe the Axon Culture Evaluation as a dedicated behavioral interview focused on cultural fit, which candidates report is framed around Axon Cultural Excellence (ACE). The reported prompts cover a production mistake and the safeguards you added afterwards, raising quality standards under deadline pushback, a design disagreement with a senior engineer or architect, evaluating the long-term cost of new technologies or AI tools, and why Axon along with the ethical and social questions around public safety technology. Prepare a specific story for each prompt rather than one general story you adapt.
What to demonstrate
- Whether a story about a production mistake gives the impact, how you fixed it and the concrete safeguard you added afterwards
- Whether a disagreement was settled with evidence such as a benchmark, prototype or written comparison, and how you kept the working relationship intact
- Whether your answer on why Axon addresses the ethics half of the question with a specific position and a real example
- Whether your own decisions are clearly separated from the team's work
How to prepare
- Write one Situation, Task, Action, Result outline per reported prompt, with the decision you personally made and a result you can back up
- For the ethics prompt, choose one concrete consideration, such as who may view sensitive recordings or how access is audited, and one work example where you raised a data-handling or ethical concern
- Say each story aloud and cut setup that comes before your decision, so follow-up questions have something specific to probe
6 candidate reports. Individual accounts describe a particular role and hiring cycle.
Axon Software Engineer Interview Experience: a take home system design presentation
The process began with a recruiter conversation about logistics and my background. That was followed by a hiring manager discussion focused on my previous work experience, and that part went very well. I then completed a take home system design deliverable. I submitted a written document and presented the design to two team members. Their follow ups went into deeper detail, so it was not simply a…
Read full experienceAxon Software Engineer Interview Experience — A Rolling-Window Officer Status Lookup System Design Question
I applied for this position online on a whim and didn't expect to actually land an interview. The company works in public safety. The question was to design a system for looking up an officer's status record near a given point in time. The system continuously receives data from police officers' devices. For each officer, while they're active, their device generates a record every 5 seconds. Each…
Read full experienceAxon Software Engineer Interview Experience — Three Onsite Rounds With a Body Cam Checkout System Design
View report detailsAxon Software Engineer Interview Experience — Rejected After a Behavioral-Heavy Hiring Manager Screen
I applied on a whim through LinkedIn Jobs. The first round was an HM screen, mostly chit-chat and behavioral questions. It wasn't the hiring manager for the role I actually applied to, but the manager of a larger group — maybe our personalities just didn't click, because I got rejected. Here are the questions I remember: Why did you choose Axon? Walk me through your career path — what's driven yo…
Read full experienceAxon Senior Software Engineer Interview Experience — Camera Footage Upload Design, Advancing to a Five-Round Onsite
The well-known stun gun company Axon, phone screen. I applied for Senior Engineer I. HR told me the phone screen would already need to happen in San Diego, and only after that would they decide which direction to move me in. The interviewer was a senior engineer, with a newly joined staff engineer shadowing. The question was pretty much the same as other write-ups on this forum — it was about upl…
Read full experiencePracHub editorial advice for the preparation topics above.
Scanning linearly, or returning the closest point in either direction, on the reported timestamped location lookup
The reported prompt asks for the officer's location at the nearest preceding timestamp when there is no exact match. Keep a per-officer list of (timestamp, location) sorted by time and answer with a binary search for the last timestamp at or before the query (bisect_right minus one). Say what a query earlier than the first record returns. State O(log n) per query, and say what inserts cost if points arrive out of order: O(n) into a sorted array, O(log n) in a balanced tree. A full scan, or returning the next point after the query, answers a different question.
Testing collinearity with floating-point slopes in the GPS path simplification
Comparing slopes divides by zero on vertical segments and misjudges floats that are almost equal. Use the cross product instead: a, b and c are collinear when (b.x - a.x)(c.y - a.y) - (b.y - a.y)(c.x - a.x) is zero. That stays exact for integer input. With real coordinates, use a tolerance and say so. Keep the kept points on a stack, and before pushing each new point, pop while the last two kept points and the new one are collinear. That way each new triple is checked against kept points, not original neighbours. Decide aloud how repeated identical points are handled, and trace a straight run of four or more points before calling the solution done.
Sending multi-gigabyte body-camera video through API servers and leaving chain of custody until the end of the design
In the reported video evidence and transcription design, the upload path and the custody record are the core of the answer. Have devices upload in resumable chunks straight to object storage with short-lived signed URLs, so the API handles only metadata and authorisation. Compute a content hash on the device, verify it when the upload completes, and write an append-only audit record for every upload, view and export. Trigger transcription from a completion event on a queue so a failed job retries without a re-upload, and make the job idempotent on the evidence id. Say which data must be strongly consistent (the custody log, permissions) and which may lag (transcripts, search indexes).
Sizing the reported shift-change device logging system for average traffic
This prompt is about the spike when many units reconnect to docks at once, so size for that peak and then shape it. Accept batches into a durable queue and acknowledge only after the write is durable. Have devices keep their logs until acknowledged and retry with exponential backoff and jitter. Return a retry-after signal when the ingest tier is saturated. Give each batch an id so re-sent batches are deduplicated rather than double-counted. Keep the ingest path separate from the audit query path so a surge does not slow searches.
Answering 'why Axon' with enthusiasm alone and skipping the ethics half of the question
The reported prompt pairs your motivation with how you handle the ethical and social questions around public safety technology, and the question bank also lists 'How did you handle an ethical issue?'. Prepare a specific answer to both halves. Name one concrete consideration you have thought about, such as access control over sensitive recordings or auditing who viewed what, and say what you would do if you had concerns about a feature. Back it with one real work example where you raised a data-handling or ethical concern, and say what happened. General support for the mission does not answer the second half.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Body Camera Check-In/Check-Out System: Design a data management system…
Body Camera Check-In/Check-Out System: Design a data management system for checking body cameras in and out of a precinct station. Ensure both check-in and check-out operations execute in O(1) time complexity while maintaining asset availability states.
Approach
- State the target complexity and say which constraint rules the naive version out.
- Restate the input: its shape, its size, and what is guaranteed about it.
- Choose the data structure from the access pattern, not from familiarity.
Follow-up
- What is the worst case, and how likely is it on real data?
- Which test case would catch an off-by-one here?
Order a job dependency graph and find its critical path
A workspace defines up to 50,000 jobs with up to 200,000 dependency edges and an estimated duration_seconds per job. Given the edge list, reject the graph if it contains a cycle and name one cycle's nodes; otherwise return a valid execution order, the earliest possible completion time with unlimited workers, and the set of jobs whose slack is zero. Then say which single job to shorten in order to cut the completion time, and by exactly how much. State the complexity of each part.
Approach
- Kahn's algorithm for the order: compute indegrees, seed a queue with zero-indegree nodes, emit and decrement. O(V + E), which at 50,000 and 200,000 is milliseconds. If fewer than V nodes are emitted, the graph contains a cycle.
- Kahn detects a cycle but cannot name one. The nodes left with indegree above zero contain every cycle, so run one DFS restricted to that residual subgraph with three-colour marking and report the stack slice from the grey node the back edge points at. That is the difference between a usable error message and 'dependency cycle detected'.
- Earliest completion with unlimited workers is the longest path, which is NP-hard on a general graph and linear on a DAG. State the precondition, then relax in topological order:
earliest_finish[v] = duration[v] + max(earliest_finish[u] for u in preds(v)), taking the max over an empty predecessor set as zero. The makespan T is the maximum over all nodes. O(V + E). - Second pass in reverse topological order for
latest_finish, thenslack[v] = latest_finish[v] - earliest_finish[v]. Zero-slack nodes form the critical path, and there can be several disjoint critical paths, so return the set rather than one chain.slack[v] = 0is exactly the statement that some longest path runs through v; equivalently, the longest path through v has lengthT - slack[v]. - The speed-up bound is the point of the question, and the obvious form of it is wrong. Shortening a zero-slack job v by d, with 0 <= d <= duration[v], cuts the makespan by
min(d, T - L_avoid(v)), whereL_avoid(v)is the longest path in the graph with v deleted: the longest path that avoids v, not the second-longest path overall. The two coincide only when the runner-up path misses v. Counterexample: A of 10 s feeds both B of 5 s and C of 4 s, so T = 15 s and the second-longest path is 14 s, yet shortening A by 10 s leaves a makespan of 5 s. The realised gain is the full 10 s, because both paths ran through A and shrank together, whilemin(10, 15 - 14)predicts 1 s. The reason is structural: shortening v reduces every path through v by d and leaves every other path alone, so the new makespan ismax(T - d, L_avoid(v)). - Compute
L_avoid(v)the direct way: delete v and re-run the same forward relaxation, O(V + E) per candidate. The cheaper equivalent skips the deletion, sinceL_avoid(v)only ever matters through that max: setduration[v] := 0, recompute the makespan asT0(v) = max(T - duration[v], L_avoid(v)), and the gain ismin(d, T - T0(v)), which is identical for every d <= duration[v]. Only zero-slack jobs are candidates, because shortening a job with positive slack changes the completion time not at all. One relaxation is milliseconds at this size, so ranking a critical set in the hundreds costs O(k(V + E)) and is worth doing exactly; a critical set in the tens of thousands is not, and there you evaluate a shortlist, longest jobs first, and say that the answer is the best of that shortlist rather than the optimum.
Worked solution 30 min
- Build four fixtures. A: 12 jobs, two branches of 100 s and 95 s that share no job. B: fixture A plus one back edge. C: two disjoint paths tied at 100 s. D: the shared-prefix case, one job of 10 s feeding a 5 s job and a 4 s job, so the longest path is 15 s and the runner-up is 14 s.
- Run Kahn; on fixture B confirm it emits fewer than V nodes, then run the residual-subgraph DFS and print the actual cycle.
- Compute
earliest_finishforward andlatest_finishbackward, and list the zero-slack set for each fixture. - For each zero-slack job v, recompute the makespan with
duration[v] := 0to getT0(v), and record both the correct boundT - T0(v)and the wrong one,T - second_longest_path, side by side. - Apply the shortening for real (20 s off the critical branch of A, 10 s off the shared prefix of D) and diff the recomputed makespan against each prediction.
Follow-up
- Only m workers are available. What happens to your answer, and what can you still promise about the schedule you produce?
- Edges arrive incrementally as the customer edits the pipeline. How do you detect a cycle at insert time without re-running Kahn over 250,000 elements?
- Durations are estimates. How would you express completion time as a distribution, and what breaks about the critical path once you do?
Locate a billing reconciliation gap without rescanning ninety million events
A tenant's sealed invoice total is 0.4% below the sum of its raw usage_event rows for the period. That tenant has 90 million events over 30 days in a table partitioned daily on ingested_at, and its rollups carry source_max_ingested_at, revision and sealed_at. Recomputing all 30 days from raw is correct, and you are not going to do it. Give the procedure that locates the divergent (workspace, sku, hour) cell, the cost of each probe, and the one query you run before any of it.
Approach
- Run the free query first. Sum raw quantity for the period restricted to
ingested_at <= source_max_ingested_atof the sealed rollups, and compare that against the unrestricted sum. The rollup stores the watermark precisely so this can be answered without a scan. If the whole 0.4% sits above the watermark, nothing is broken: it is late data, it becomes an adjustment line, and the investigation ends in one query. - Only if the gap survives that test do you bisect, and you bisect by dimension rather than by rows. Compare 30 per-day totals, then inside the offending day compare the 6 SKUs, then the workspaces, then the 24 hours. That is roughly 30 + 6 + W + 24 grouped probes, each an indexed range scan over one daily partition for one tenant, against O(N) per attempt for the naive re-fold.
- Quantify why naive is not merely slow but unusable mid-incident: at a generous 200,000 rows/second sequential, 90 million rows is about 7.5 minutes per attempt, you will want ten attempts, and every one competes for I/O on the same partitions live ingest is writing. The diagnostic worsens the backlog it is diagnosing.
- Before fetching each comparison, state what it would look like under each hypothesis. Two adjacent hours off by equal and opposite amounts is
occurred_atversusingested_atbucketing. A whole day offset by exactly N hours is a timezone applied at the wrong layer. A gap confined to one SKU in one workspace is an environment filter. The same(tenant_id, idempotency_key)present in twoingested_daypartitions is the dedup horizon losing a retry that crossed midnight. - Make the next bisection cheap by storing the aggregate you keep recomputing. A per-
(tenant_id, ingested_day)count and quantity checksum turns step two from thirty probes into one read, and it is the same number the reconciliation job already produces. - Whatever you find, the sealed period does not change value. The correction is an adjustment line pointing at the line it reverses, carrying its own
source_rollup_watermark, because the original invoice is the evidence of what the customer was charged.
Follow-up
- The gap is 0.4% in one direction on one day and 0.4% the other way the next day. What does that shape rule in, and what does it rule out?
- How do you distinguish a duplicate from a restatement, given
revisionandrecomputed_aton the rollup? - Ingest is still running while you investigate. What makes your two numbers comparable at all?
Model credential revocation so history survives the delete
tenant_api_key stores key_id, tenant_id, workspace_id, name, key_prefix, secret_hash, scopes text[], status (active, revoked, expired, compromised), auth_version, created_at, expires_at, last_used_at, revoked_at, revoked_reason. Rotation inserts a new row and revocation never deletes, because an incident review asks which credential served a request last quarter. Write the constraints that enforce: a label is unique only among a tenant's live keys, revoked_at and status can never disagree, and scopes is never empty. Then write the authentication lookup predicate, and name one column in this table that must stay out of it.
Approach
- Reach for a partial unique index rather than a plain UNIQUE:
create unique index on tenant_api_key (tenant_id, name) where revoked_at is null. Any number of revoked rows may share a label, the live namespace stays unique per tenant, and the revoked majority is not in the index at all, so it stays small on a table that only grows. - Tie the nullable timestamp to the enum so the two cannot drift:
check ((revoked_at is not null) = (status in ('revoked','compromised')))andcheck ((revoked_at is null) = (revoked_reason is null)). A revocation that records no reason is the one an incident review cannot use. - Write the emptiness check as
check (cardinality(scopes) > 0), notarray_length(scopes, 1) > 0. array_length returns NULL for an empty array, a CHECK constraint passes when its expression is NULL, so the array_length version accepts exactly the value it was written to reject. - Make the lookup a single index probe with every liveness condition inside it:
where secret_hash = $1 and revoked_at is null and (expires_at is null or expires_at > now()) and auth_version = $2, backed by a unique index on secret_hash. Nothing is filtered in application code, so there is no path that forgets a clause. - Keep last_used_at out of that predicate. It is written asynchronously and is allowed to lag by a minute, so it is a usage signal; feeding it into an authorisation decision makes the decision depend on a write that may be late, batched away or lost.
- Flag the modelling smell while you are here:
expiredis derivable fromexpires_at < now(), so storing it as a status obliges a job to keep it true and guarantees the column is wrong between the expiry instant and that job's next run. Derive it in the predicate; keep the stored status for states that are decisions rather than clock readings.
Worked solution 20 min
- Create the table with all three constraints on a scratch database and insert two revoked rows sharing (tenant_id, name); the partial index should accept both.
- Insert a second live row with that same name and confirm the violation names the partial index.
- Run
update tenant_api_key set revoked_at = now()leaving status = 'active' and confirm the CHECK rejects it; then tryinsert ... scopes = '{}'against both the cardinality and the array_length forms and note that only one rejects it. - Run
explain (analyze, buffers)on the lookup predicate for a live key and confirm an index scan on secret_hash with rows removed by filter equal to zero.
Follow-up
- Rotation issues a replacement while the old key stays live for a 30-day overlap. What does the uniqueness rule become, and what does the UI show to tell two same-named keys apart?
- A password reset bumps the principal's auth_version. No row in this table changed. How does the next request fail, and what query counts how many keys that bump just killed?
- A key turns up in a public repository. Which columns let you find it, and what do you write to the row?
Paginate a tenant's delivery export without skipping rows
A customer exports webhook_delivery: delivery_id (bigint identity), subscription_id, tenant_id, event_id, status, attempt_count, next_attempt_at, created_at, delivered_at, updated_at. The endpoint runs select ... where tenant_id = $1 order by created_at desc limit 100 offset $2, and customers report rows missing from exports taken while new deliveries are being inserted. Write the replacement query and the index that supports it, paging a tenant's deliveries newest first at constant cost per page. State why updated_at cannot be the cursor column.
Approach
- Name the defect precisely. OFFSET is a position in a result set that is recomputed on every request, so a row inserted ahead of the window shifts everything back by one and the next page starts after a row the client never received. Nothing errors and no identifier gap appears, so the loss is silent.
- Replace the position with a value predicate over a stable, unique, indexed ordering:
where tenant_id = $1 and (created_at, delivery_id) < ($2, $3) order by created_at desc, delivery_id desc limit 100. The row comparison is load-bearing: created_at alone is not unique, so ties straddling a page boundary are dropped or repeated, which is the same bug in a smaller window. - Index
(tenant_id, created_at, delivery_id). PostgreSQL scans a btree in either direction, so an all-DESC ORDER BY is served by an ASC index read backwards and no DESC modifiers are needed; they only matter when the ORDER BY mixes directions. Confirm the plan has no Sort node above the index scan, or the LIMIT stops being an early exit. - Price both forms: keyset is one index descent plus 100 adjacent leaf entries per page, constant regardless of depth, while OFFSET still produces and discards every skipped row, so page N costs time proportional to N times the page size and a deep page on a large table goes from milliseconds to seconds.
- Rule out updated_at as the cursor from the precondition, not from taste: a cursor column must never change value for a row already paged past. updated_at moves on every delivery attempt, so a row the client already emitted re-enters a later page and is exported twice. created_at and delivery_id are immutable, which is the whole qualification.
Follow-up
- The client wants a snapshot as of one instant rather than a live tail. Compare a repeatable-read transaction held open, an added
created_at <= $snapshotbound, and a materialised export table. - A retention job deletes deliveries older than 90 days. What does a client mid-walk see, and does keyset pagination help at all?
- The customer wants to resume an export from yesterday's last cursor. What must be true of the cursor for that to be safe?
Digital Video Evidence & Auto-Transcription Pipeline: Architect an end…
Digital Video Evidence & Auto-Transcription Pipeline: Architect an end-to-end cloud platform that accepts multi-gigabyte video uploads from thousands of body cameras simultaneously, stores them securely for legal chain of custody, and automatically triggers an asynchronous background job to generate text transcripts.
Approach
- Name the failure you are designing for, then the recovery path.
- Name the read and write paths separately; they rarely have the same bottleneck.
- State the consistency you need, and where you are willing to be stale.
Follow-up
- What would you drop to keep the system up under load?
- How does this behave when that dependency is down for an hour?
Real-Time Voice & Dispatch System: Design a low-latency voice processi…
Real-Time Voice & Dispatch System: Design a low-latency voice processing pipeline capable of streaming live audio from officer radios, extracting key urgency keywords, and notifying dispatchers in real time.
Approach
- State the consistency you need, and where you are willing to be stale.
- Fix the scope first: who calls this, how often, and what they do when it fails.
- Name the failure you are designing for, then the recovery path.
Follow-up
- What would you drop to keep the system up under load?
- What breaks first when traffic grows ten times?
High-Surge Device Logging System: Design a centralized log auditing sy…
High-Surge Device Logging System: Design a centralized log auditing system capable of receiving log batches from offline devices. Address how the system handles massive traffic spikes when thousands of police units reconnect to network docks at shift changes (e.g., 6:00 PM).
Approach
- Name the failure you are designing for, then the recovery path.
- Name the read and write paths separately; they rarely have the same bottleneck.
- State the consistency you need, and where you are willing to be stale.
Follow-up
- How does this behave when that dependency is down for an hour?
- What breaks first when traffic grows ten times?
Specify webhook signature verification a customer can implement
The webhook-delivery service signs each payload before POSTing it to a customer endpoint. Write the signature specification a customer implements in their own language: the header format, exactly which bytes are signed, the algorithm, how replay is bounded, and how a signing secret rotates without a delivery gap. Then write the verification steps the customer performs, in order, including what they compare and what they return on failure. Constraint: most customers reach for their web framework's parsed JSON body by default. Deliverable: the spec section plus reference pseudocode.
Approach
- Sign the concatenation of the timestamp and the raw body,
t + "." + body, and emit a header of the formt=<unix seconds>,v1=<hex>. The timestamp has to be inside the MAC, or an attacker re-stamps a captured body and the tolerance window buys nothing. - Require the raw request bytes. A framework that parses JSON and re-serialises it changes key order, whitespace and number formatting, so the spec must tell the customer to capture the body before the parser runs and give the middleware note for each common framework.
- Use HMAC-SHA256, not sha256(secret || body): SHA-256 is a Merkle-Damgard construction, so the naive form admits length extension. Require a constant-time comparison as well, since a short-circuiting byte compare leaks the expected prefix under repeated probing.
- Bound replay in two layers: reject when |now - t| exceeds a stated tolerance such as 300 seconds, then deduplicate on the event identifier header. The tolerance is what makes the customer's dedup store finite rather than unbounded.
- Rotate by allowing two live secrets and emitting both signatures in one header (
v1=<old>,v1=<new>); the customer accepts if any candidate matches, so neither side needs an instantaneous cutover. A failed verification returns 400 and the body is not processed.
Worked solution 15 min
- Write the header grammar and one real example line with a plausible timestamp and hex digest.
- Write the signed string construction explicitly as a byte concatenation, and add the sentence telling the customer where in their framework to obtain the raw body.
- Write the five verification steps in order: extract t and candidates, check the tolerance, recompute the HMAC over t + '.' + raw body, compare in constant time against each candidate, then deduplicate on the event identifier.
- Add the rotation paragraph: two active secrets, both signatures sent, overlap window stated in the dashboard.
- State the failure response and the fact that the payload is not processed, plus what the sender does with that 400.
Follow-up
- A customer's verification passes locally and fails in production behind a proxy that re-encodes the response body. Where do you look first?
- Why sign with a per-endpoint secret rather than the tenant's API key?
Webhook workers leak until OOM and drop in-flight deliveries
webhook-delivery workers grow from 400 MB to a 2 GB limit over about 36 hours, are OOM-killed, restart, and repeat. Each restart abandons in-flight attempts, so webhook_delivery rows sit in in_flight until their leases expire and the backlog spikes. The live set measured after a forced full collection also grows. The fleet serves tens of thousands of subscriptions, several thousand of which have been failing for weeks. Give an ordered checklist, the measurement separating retention from fragmentation, and the fix.
Approach
- Separate the two failure shapes with one measurement: track resident set size against the live set after a forced full collection. A live set that climbs monotonically is retention; a flat live set under a rising RSS is fragmentation, off-heap or native allocation, or an allocator that never returns pages. The stated symptom puts this in the first category, which rules out allocator tuning as a fix.
- Characterise the curve rather than the total. Growth linear in uptime implies an unbounded structure keyed by something that keeps arriving; step growth implies buffering a large object. Correlate the slope against event rate and separately against the count of distinct subscriptions seen, because those two diverge and only one of them will fit.
- Diff two heap snapshots an hour apart by retained size grouped by dominant root, not by allocation count, which is dominated by short-lived objects and will point at the wrong thing.
- Expect a per-subscription map with no eviction: circuit-breaker or backoff state created on first failure and never removed, so the retained set grows with endpoints that have ever failed, and the several thousand permanently dead endpoints hold theirs forever.
- Fix in two places. Bound the in-memory structure with a size-capped LRU or a TTL keyed on last use, and move state that must survive a restart onto the subscription or webhook_delivery row, since the worker holding it in memory is exactly why a restart loses it.
- Repair the second-order damage separately, because it will outlive the leak: workers claim by compare-and-set with leased_until, so a bounded lease returns in_flight rows to pending on a known schedule, and a graceful shutdown releases leases instead of waiting them out.
Follow-up
- The backlog spike after a restart is itself a thundering herd against customer endpoints. What stops the recovery from becoming a second incident?
- Suppose the live set had been flat while RSS still climbed. Name two causes and the measurement that separates them.
- How would you size the LRU, and what does a miss on an evicted circuit-breaker entry cost a customer whose endpoint is down?
Four days sample coding, design, fundamentals and the practical rounds at deliberately shallow depth, which is enough to surface the topics you did not know were in scope. That map, rather than a guess made on day one, decides where the last three days go.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Pin down the first two rounds and warm up on fundamentals
- Ask the recruiter whether the first stage is an HR conversation or an asynchronous technical assessment, and whether the Technical Screen will be live coding or a system design problem, in which language and editor. Write the answers down, because they decide how to weight days 2 to 5.
- Write a short background summary that ends on one project you owned. You will use it in the Initial HR Conversation and again at the start of the Technical Screen.
- Warm up in a plain text editor with no autocomplete on two questions from the reported bank: the Tic-Tac-Toe winner check, and an LRU cache with O(1) get and set using a hash map and a doubly linked list. Hand-trace each one before running it.
Deliverable: Written answers about the format of the first two stages, a background summary you can say aloud, and two hand-traced solutions.
Practice prompt ↗Practice prompt ↗Worked solution ↗02Coding: timestamped locations and GPS paths
- Implement the nearest-preceding-timestamp lookup with per-officer sorted lists and a binary search for the last time at or before the query. Cover a query before the first record, exact matches and duplicate timestamps.
- Implement the GPS recorder that drops middle points when three consecutive kept points are collinear, using the cross product and a stack so every newly formed triple is checked.
- For both, write down time and space complexity, and what changes if points arrive out of order.
- Trace one input by hand for each solution before running it, the way you would have to in a plain shared editor.
Deliverable: Two working implementations with edge-case tests and a written complexity statement for each.
Practice prompt ↗Practice prompt ↗03Coding: graphs, access resolution and asset state
- Solve BFS shortest path on a grid with walls, from an officer to a target. Mark cells visited when you enqueue them and return an agreed sentinel when the target cannot be reached.
- Solve nested group access: run a BFS from the officer through group memberships to devices, with a visited set so cyclic group nesting terminates.
- Build the body camera check-in/check-out system with O(1) operations, using a map from camera id to state and holder plus a set of available ids. Decide what a double check-out or an unknown id returns.
- For more graph depth, work the existing exercise 'Order a job dependency graph and find its critical path' and run its fixtures. Optionally, try currency conversion queries from the bank as a weighted graph traversal.
Deliverable: Three solutions with stated complexity and error behaviour, plus the critical-path exercise with its fixtures passing.
Practice prompt ↗Practice prompt ↗04System design: video evidence and chain of custody
- Design the reported video evidence and auto-transcription pipeline end to end. Ask clarifying questions first, then cover chunked resumable upload to object storage, the metadata API, the custody audit log and an async transcription job off a queue.
- Write a short table of which data must be strongly consistent and which may lag, and describe what users see while the transcription service is down.
- Work the existing exercise 'Specify webhook signature verification a customer can implement' to practise HMAC, integrity and replay vocabulary you can reuse for tamper evidence.
- Sketch the bank topic 'Design a tamper-evident video chain-of-custody audit system' at the level of data model and verification step.
Deliverable: One full design with a diagram, a consistency table, and a list of failure modes with the recovery for each.
Practice prompt ↗Practice prompt ↗Worked solution ↗05System design: reconnect surges, live audio and permission models
- Design the shift-change device logging system: size for the peak, use a durable queue, acknowledge after the durable write, have devices retry with jitter, and deduplicate by batch id.
- Design the real-time voice and dispatch pipeline at the level of streaming ingest, keyword extraction and dispatcher notification, and say where latency is spent at each hop.
- Design the physical evidence storage backend with its API, schema, authentication and permission model. Then work the existing SQL exercise 'Model credential revocation so history survives the delete' to practise constraints and audit history.
Deliverable: Three design outlines, each stating its peak-load assumption and one failure it was designed to survive.
Practice prompt ↗Practice prompt ↗06Axon Culture Evaluation: stories for the reported prompts
- Write one story for each reported prompt: a production mistake and the safeguards you added, pushing for quality under deadline pushback, a design disagreement with a senior engineer, and weighing the long-term cost of a new technology or AI tool.
- Prepare the 'why Axon' answer with both halves: your reason for applying, and a concrete position on the ethical and social questions around public safety technology, backed by one real example of handling an ethical issue.
- In each story, replace collective phrasing with what you personally decided and did, and add a result you can back up.
- Say each story aloud and cut the setup that comes before your decision.
Deliverable: Six story outlines, each with your decision, its result and the follow-up question you expect.
Practice prompt ↗Practice prompt ↗07Rehearse the Technical Screen and the Final Interview Loop
- Run a mock Technical Screen: a background walkthrough, then one reported coding prompt in a plain editor. Talk through the approach and hand-trace the code before saying you are done.
- Run back-to-back sessions for the final loop: one coding problem, one object-oriented design (the check-in/check-out system with explicit states), one system design from day 4 or 5, and one behavioral story.
- Note where your structure slipped in the later sessions, and rewrite the opening you use for that kind of round.
Deliverable: Mock notes for each session, plus a one-page card with your clarifying-question checklist, complexity reminders and story list.
Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Reports describe the Axon Culture Evaluation as a dedicated behavioral interview on cultural fit, and reports connect the behavioral prompts to Axon Cultural Excellence (ACE). The prompts below are the reported ones plus one from the question bank. For each, prepare a story where the decision was yours. Say what you knew when you made it, what you did, and a result you can back up. The ethics prompt needs a concrete position and a real example, not general support for the mission.
Why do you want to work at Axon, and how do you navigate the ethical a…
Why do you want to work at Axon, and how do you navigate the ethical and social considerations associated with public safety technology?
Approach
- Name the disagreement and how you resolved it with evidence.
- Pick a story where you made the decision, not one where you watched it.
- State the situation in two sentences and spend the rest on the reasoning.
Follow-up
- What would you do differently if you ran that again?
- What did you decide not to do, and why?
Describe a situation where you had a fundamental disagreement with a s…
Describe a situation where you had a fundamental disagreement with a senior engineer or architect regarding system design. How did you present your arguments and arrive at a resolution?
Approach
- Close with what you would do differently, concretely.
- State the situation in two sentences and spend the rest on the reasoning.
- Pick a story where you made the decision, not one where you watched it.
Follow-up
- What would you do differently if you ran that again?
- How did you know your change caused the improvement?
Ship metered billing with a named deduplication horizon
Metered billing must be on in three weeks. usage_event is partitioned daily, so its unique index must include the partition key and deduplicates only within a day: a producer retry that crosses midnight, or a replay run a week later, gets through. A cross-partition dedup store is two weeks you do not have. Describe shipping with debt you named in advance: what you shipped, what you wrote down, the detector you added, the trigger and date for paying it off, and what you would have refused to ship under the same pressure.
Approach
- Show you can separate the two kinds of debt, because that distinction is what the question actually probes. Debt that costs engineering time later is shippable on a deadline. Debt that silently corrupts a number a customer gets charged for is not shippable unless the corruption is detectable, and detectability is the whole negotiation.
- Make the exposure narrow and measured rather than gestural. The hole is duplicates whose occurrences straddle a UTC day boundary, plus any replay older than partition retention. Measure it before arguing about it: how often an idempotency_key recurs at all, and the distribution of the gap between first and last occurrence. If the ninety-ninth percentile of that gap is four minutes, the residual risk is a small band around midnight and you can say so numerically.
- Add the detector before the feature, not after. A nightly job counting keys that appear in more than one partition is one grouped scan over recent partitions, and it converts a silent overcount into a page. State what it costs to run and what it fires on.
- Buy the cheap half of the real fix immediately: extend partition retention so the dedup horizon exceeds the producer's maximum retry window plus the longest replay you intend to support. That reframes retention as a correctness parameter rather than a storage cost, which is the sentence you need on record before someone optimises the bill.
- Make repayment mechanical instead of aspirational: a dated entry with a named owner, plus a threshold that pulls the date forward — first detector hit above N events, or first customer dispute. Debt with a trigger gets paid; debt with only a date does not.
- Answer the second half honestly by naming what you would refuse under identical pressure: the sealing path, because a sealed row is frozen and a wrong number there stops being a bug and becomes an adjustment line, a dispute and an audit question.
Follow-up
- The detector fires on forty duplicate events for one tenant, and two of their invoices have already sealed. What happens next?
- Whom did you tell that the billing numbers had a known hole, and in what words?
- Finance asks you to cut storage by shortening partition retention. What do you say, and to whom?
- 01
Tell me about a time you made a significant technical mistake on a production system. What was the impact, how did you resolve it, and what operational safeguards did you implement afterward?
- 02
How do you approach driving engineering excellence and higher quality standards when encountering pushback or resistance from teammates or product managers under tight deadlines?
- 03
Describe a situation where you had a fundamental disagreement with a senior engineer or architect regarding system design. How did you present your arguments and arrive at a resolution?
- 04
How do you evaluate the long-term operational and maintenance impacts of introducing new technologies or AI tools into an existing software stack?
- 05
Why do you want to work at Axon, and how do you navigate the ethical and social considerations associated with public safety technology?
- 06
How did you handle an ethical issue?
Is this an official Axon interview guide?
No. It is PracHub's own research and practice material for the Software Engineer role at Axon. Rounds and questions reflect what candidates have reported, not a process Axon has published, and they change over time. Confirm the current format and scope with your recruiter.
PracHub interview research ↗What happens in the Initial HR Conversation?
Reports describe this stage as either a conversation with HR or an asynchronous technical screening assessment, so ask your recruiter which one you will get. For a conversation, prepare a short background summary and a specific reason for applying. For an assessment, reports say screens at this stage emphasize fundamental data structures, array manipulation and basic object-oriented design.
PracHub Software Engineer practice ↗What kind of coding questions are reported for Axon Software Engineer candidates?
Mostly standard problems with domain framing. Reported ones include simplifying a GPS path by removing collinear middle points, returning an officer's location at the nearest preceding timestamp, BFS shortest path on a city grid with walls, O(1) body camera check-in and check-out, and resolving device access through nested user groups. The question bank also lists an LRU cache, a Tic-Tac-Toe winner check, currency conversion over a weighted graph, and recursive dynamic programming on strings. Prepare hash maps, binary search, BFS and DFS, and small stateful classes, and practise stating the complexity of each.
PracHub interview research ↗How much system design should I expect?
Reports say system design carries more weight from SDE II upward, usually as two separate architectural rounds, and that senior loops lean toward distributed systems design. Reported design prompts include a video evidence upload and auto-transcription pipeline, a device logging system that absorbs reconnect surges at shift change, a physical evidence storage backend with API, schema and permission model, and a real-time voice and dispatch pipeline. For each, prepare storage choices, queues for async work, peak-load handling, access control and an audit trail.
PracHub interview research ↗What does the Axon Culture Evaluation cover?
Reports describe it as a dedicated behavioral interview on cultural fit, which candidates report is framed around Axon Cultural Excellence (ACE). Reported prompts cover a production mistake and the safeguards you added afterwards, raising quality standards under deadline pushback, a design disagreement with a senior engineer, the long-term cost of adopting new technologies or AI tools, and why Axon along with the ethics of public safety technology. Use a Situation, Task, Action, Result structure, make your own decisions explicit, and only quote numbers you can back up.
PracHub interview research ↗Will I be coding in a full IDE?
Reports say many live coding screens use a simple shared text editor without autocomplete or compilation. Practise writing complete code without IDE help, trace it by hand on a small input, and talk through your test cases before saying you are done.
PracHub interview research ↗Are all the practice questions in this guide reported Axon questions?
No. The practice items on body camera check-in and check-out, the video evidence and transcription pipeline, real-time voice dispatch and shift-change device logging come from candidate reports, as do the two behavioral prompts on why Axon and on a design disagreement with a senior engineer. The drills on API key revocation, export pagination, job dependency graphs, billing reconciliation, webhook signatures, worker memory leaks and shipping with named technical debt are editorial practice. They exercise skills the reported questions rely on, such as schema constraints, graph algorithms, integrity checks and debugging method.
PracHub Software Engineer practice ↗Sources & methodology 3 sources ↗
Official role evidence, timestamped platform data and clearly labeled preparation advice.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
platform · Accessed 2026-09-24 - 02PracHub Software Engineer practice ↗
Cross-company practice questions for this role.
platform · Accessed 2026-09-24 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
platform · Accessed 2026-09-24