The reported coding questions are small, self-contained Python problems built around data a network product handles: request logs, a blacklist of URLs, a stream too large for memory, and a topology graph. Each one has a short correct answer and several ways to get it subtly wrong, such as a window that lets a double burst through, a run counter that forgets the last run in the file, or a Trie that matches a prefix when the question needs an exact host. Practise writing these as complete, runnable code and testing them aloud.
The ML side is described as three areas: theory and AI safety (evaluation metrics such as precision/recall versus ROC-AUC, adversarial robustness, and the mechanics of Transformers, CNNs and RNNs), engineering in Python with PyTorch or TensorFlow, and architecture for end-to-end ML pipelines (serving at scale, data drift, feature stores, low-latency inference). Because the role is tied to traffic and threat detection, practise answering with class imbalance, false-positive budgets and per-request latency in mind.
Senior and Staff candidates are described as needing several years of applied ML or software experience and the ability to explain trade-offs to technical and non-technical people. Prepare real projects you can walk through from data exploration to deployment and monitoring, with the numbers you moved and what you would do differently.
Preparation focus
editorialCandidates describe a path that starts with a recruiter screen, moves through technical assessments and ends in onsite loops, where they meet peers, cross-functional partners and engineering leadership. Early conversations are reported as high level and later ones as algorithmic and architectural. The stage order is reported, but round names, counts and per-round content are not, so this step covers every area: Python coding (rate limiter, failed-login run in a log file, chunked generator, Trie, shortest path), ML theory and AI safety, ML engineering in PyTorch or TensorFlow, and ML pipeline design. Start with the coding set, because it is the most concrete, then build ML answers around network and security examples.
What to demonstrate
- Python fluency on data-processing problems: streaming input, a Trie, a rate limiter and graph traversal
- ML theory and AI safety: evaluation metrics, adversarial robustness, and Transformer, CNN and RNN mechanics
- ML engineering in PyTorch or TensorFlow and end-to-end pipeline design: training and evaluation code, serving, drift, feature stores and low-latency inference
- Clear explanation of why an approach was chosen, including trade-offs
How to prepare
- Implement the rate limiter, the failed-login run finder, the chunking generator and the Trie from a blank file, with tests for empty input and boundaries
- Prepare a one-page note on precision, recall, PR-AUC and ROC-AUC for an imbalanced attack-versus-benign classifier, including how you would set a threshold
- Write a PyTorch training and evaluation loop for an imbalanced traffic classifier with a class-weighted loss, then export it and time its inference path at batch size 1
- Outline an ML serving pipeline for traffic features with a baseline model, a latency budget, drift monitoring and a rollback path, and rehearse three STAR project stories that each end with a measured result
PracHub editorial advice for the preparation topics above.
A rate limiter that uses a fixed window counter and ignores the boundary burst
State which algorithm you chose and why. A fixed window lets up to twice the limit through around a window boundary. A token bucket refills lazily from elapsed time (tokens = min(capacity, tokens + elapsed * rate)) and handles bursts deliberately; a sliding-window log is exact but stores one timestamp per request. Use a monotonic clock, key state per client, guard it with a lock if threads are involved, and say how idle keys get evicted.
Log parsing that reads the whole file or drops the final run
Iterate the file line by line, or with a generator that yields chunks, so memory stays constant. For the longest run of failed logins, update the best length when a success breaks the run and once more after the loop ends, otherwise a file ending in failures returns the wrong answer. Ask whether contiguous means across the whole file or per user or IP, and handle malformed lines on purpose.
A Trie answer that never settles what a match means
Ask whether the blacklist is matched as an exact URL, a prefix or a host. Normalise case, scheme and fragments before inserting, decide whether to store host labels in reverse order, and keep a terminal flag per node so that a stored string is not confused with a prefix of another. Mention memory (dict children versus arrays, compressing single-child chains) and when a plain set or Bloom filter is enough.
Quoting ROC-AUC alone for rare-attack traffic
When positives are rare, ROC-AUC can look strong while precision at the operating point is poor. Report precision, recall and PR-AUC, pick the threshold from a stated false-positive budget, and explain what an alert costs. Add how you would evaluate after drift, for example on a recent time-split validation set rather than a random one.
Jumping to a deep model in a serving design without a baseline or latency budget
Clarify traffic volume, per-request latency, labels and who consumes the output first. Start with a simple baseline, then say what would justify a larger model, and cover feature consistency between training and serving, batching or quantisation, drift monitoring, shadow deployment and rollback.
Choose a category, try a prompt, then open its approach, worked solution or follow-up when you need it.
Implement a custom rate limiter for an API endpoint.
Implement a custom rate limiter for an API endpoint.
Approach
- Pin down the contract first: what the limit is keyed on (API key, user or client IP), the limit and window, whether bursts are allowed, single process or many servers, and what a rejected call gets back (HTTP 429 with a Retry-After header).
- Default to a token bucket per key: store tokens and last_refill, and on each call refill lazily with tokens = min(capacity, tokens + (now - last_refill) * rate), then allow and decrement if tokens >= 1. That is O(1) time and O(1) state per key, and it permits bursts up to capacity. Within one process, read time.monotonic(), not wall-clock time, so a clock adjustment cannot refill or freeze buckets.
- Compare the alternatives by their failure mode. A fixed-window counter is cheapest but lets up to twice the limit through around a window boundary. A sliding-window log (a deque of timestamps, evicting ones older than the window) is exact but costs O(limit) memory per key. A sliding-window counter weights the previous window's count to approximate the log in O(1).
- Make the check-and-decrement atomic: a lock per key (or one lock) in a threaded server; across several servers, keep state in a shared store such as Redis and do the read-modify-write in a single Lua script, or use INCR plus EXPIRE for a fixed window. Monotonic readings have an arbitrary origin and cannot be compared across hosts, so take now from the store inside the script (Redis TIME), or clock skew between servers corrupts refill.
- Cover the edges in tests: the first request for a new key, a burst exactly at capacity, a request that costs more than one token, and idle keys, which need a TTL or LRU eviction so per-IP state cannot grow without bound.
Follow-up
- How would you keep the limit correct when the endpoint runs on ten servers behind a load balancer, and what happens if the shared store is unreachable: fail open or fail closed?
- Show the fixed-window boundary burst with concrete timestamps, then explain how your chosen algorithm avoids it.
- A flood of spoofed or rotating client IPs creates millions of one-request buckets. How do you bound memory?
Given a log file of network requests, write a script to parse the file…
Given a log file of network requests, write a script to parse the file and return the longest contiguous sequence of failed login attempts.
Approach
- Clarify the format and the meaning of contiguous before coding: which field marks a login attempt (path, event type) and which marks failure (status 401/403 or a result field), whether contiguous means consecutive lines in file order or consecutive attempts per user or IP, and whether non-login lines break a run or are skipped.
- Stream the file in one pass with for line in f, so memory stays O(1) however large the log is. Parse each line with str.split or a regex compiled once outside the loop, and skip and count malformed lines instead of crashing.
- Keep the current run (start line or timestamp, length) and the best run. A failed attempt extends the current run; a successful attempt resets it. Return the run itself (start, end, count), not just its length, and state the tie rule (the first longest run wins).
- Compare the current run against the best once more after the loop ends: forgetting that final check is the classic bug when the longest run sits at the end of the file. An empty file or a file with no failures returns an empty result, not an exception.
- For the per-source variant, keep a dict from user or IP to its current run; memory becomes O(distinct sources). If lines can arrive out of order, an in-memory sort breaks the O(1)-memory claim: a log too large for memory needs an external sort by timestamp, or a bounded reorder buffer (a min-heap over a time window) if lateness is limited. Otherwise state the ordering you assume.
Follow-up
- Attempts more than N seconds apart should not count as one run. How does the state change?
- The logs are split across many files or arrive as a live stream. How do you merge runs that straddle a file boundary?
- How would you test the parser against malformed lines, duplicate timestamps and a run that ends on the last line?
Write a Python generator that processes a massive dataset in chunks to…
Write a Python generator that processes a massive dataset in chunks to prevent memory overflow.
Approach
- Write a generator that yields fixed-size batches: open the file in a with block, then loop chunk = list(itertools.islice(f, chunk_size)); stop when chunk is empty; otherwise yield chunk. Peak memory is O(chunk_size) regardless of file size, and the final chunk may be shorter.
- For binary or unstructured data, use iter(lambda: f.read(n), b'') and carry the bytes after the last newline into the next chunk so a record split across a chunk boundary is never parsed in two halves.
- Compose stages as generators (read, parse, filter) and consume them with an incremental aggregate (running sum, counts, min/max). Building a list of all results at the end defeats the point and is the most common mistake.
- Explain cleanup: if the consumer stops early, calling close() on the generator, or garbage-collecting it, raises GeneratorExit at the paused yield, so the with block exits and the file closes. For CSV input, pandas.read_csv(path, chunksize=n) already returns an iterator of DataFrames.
- Name the trade-off: small chunks cost Python-level overhead per chunk, large ones raise peak memory. Vectorise work inside each chunk with NumPy or pandas and choose a size that keeps one chunk plus its intermediates well under available memory.
Follow-up
- How would you parallelise processing across chunks, and what do you lose if results must stay in input order?
- Sums and counts combine across chunks; how would you compute an exact median or a distinct count, and when would an approximate sketch (t-digest, HyperLogLog) be acceptable?
- What happens to the open file handle if the caller breaks out of the loop after the first chunk?
Implement a Trie data structure to efficiently store and search a larg…
Implement a Trie data structure to efficiently store and search a large blacklist of malicious URLs.
Approach
- Settle the match semantics first: exact URL, or prefix blocking where listing a domain or path blocks everything beneath it. Normalise every URL the same way at insert and lookup: lowercase scheme and host, drop default ports and fragments, decode percent-encoding consistently, and convert international hosts to punycode.
- Node = dict of children plus an is_terminal flag. Insert and lookup are O(L) in the length of the key, independent of how many entries the list holds. For prefix blocking, walk the query and return True as soon as you reach a terminal node.
- Tokenise by URL component instead of by character: reversed host labels (com, example, evil) followed by path segments. Subdomains then share their parent's prefix, so blocking a domain blocks its subdomains, and a character-level prefix such as evil.com can no longer match a different host like evil.com.attacker.net.
- Discuss memory: a dict per node is heavy in Python, so compress single-child chains into a radix (Patricia) trie. If only exact matches are needed, a hash set gives O(1) average lookup. A Bloom filter in front can reject most clean URLs cheaply; it has false positives but no false negatives, so positives still need an exact check.
- Test the edges: an empty key, duplicate inserts, deleting an entry (unset the flag and prune nodes left with no children and no terminal flag), a URL that is a prefix of a listed one, and trailing-slash variants.
Follow-up
- The list grows to hundreds of millions of entries and no longer fits in one process's memory. What changes?
- How do you apply blacklist updates while lookups keep serving traffic, for example by building a new trie and swapping a reference atomically?
- How would you support wildcard entries such as *.example.com, or path patterns, without falling back to a linear scan?
Archive a resource graph without breaking live references or recursing
Resources reference other resources within a tenant; for the largest tenant the reference table holds up to 2,000,000 nodes and 8,000,000 edges. Archiving a resource must archive everything reachable from it that nothing outside the set still references, refuse when a live external referrer exists, and terminate when references form cycles, which they legitimately do. Produce the archive order and the refusal list, targeting O(V+E). Say what stops the traversal crossing a tenant boundary, and why recursion is the wrong control structure at this size.
Approach
- Load the subgraph with the tenant predicate on both endpoints of the edge, not only on the side you started from. Scoping the left table alone is the classic cross-tenant leak: one mis-entered edge then pulls another tenant's resources into the traversal and, worse, into the archive.
- Traverse iteratively with an explicit stack. A 2,000,000-node graph can hold a chain deep enough to exhaust a native stack in the low tens of thousands of frames, and that failure is a process crash rather than an error you can return.
- Treat cycles as data rather than corruption: compute strongly connected components with Tarjan in O(V+E) using its own explicit stack, then condense. The condensation is a DAG, so a topological order over it gives the archive order, and every member of a component archives in one transaction because no order within a cycle is valid.
- Decide refusals with reverse edges. A candidate is archivable only if every in-edge originates inside the candidate set, so build the transpose or count in-degrees restricted to the visited set, and emit each blocked resource with the id of the external referrer, which is the only part of the answer an operator can act on.
- Store the graph as CSR rather than a map of lists: an offsets array of V+1 8-byte entries plus E 8-byte targets is about 80 MB at this size, where boxed adjacency lists cost several times that and lose cache locality on every hop.
- Run Kahn over the condensation for the order in O(V+E). If the emitted count is short of the component count the condensation step itself is wrong, since a condensation cannot contain a cycle, which makes the check free.
Worked solution 30 min
- Write the edge-loading query with the tenant predicate on both endpoints and state what it does with a cross-tenant edge.
- Implement iterative Tarjan with an explicit stack and confirm on a three-node cycle that it emits one component of size three.
- Build the transpose restricted to the visited set and mark every node with an in-edge from outside it as refused, carrying the referrer id.
- Run Kahn over the condensation and verify the emitted order against the referrer-before-referenced rule.
- Size the CSR arrays for 2,000,000 nodes and 8,000,000 edges and compare against a boxed adjacency map.
Follow-up
- The graph is read in one query and the archive writes a minute later. What can change in between, and how do you make the write safe?
- The candidate set is 400,000 resources. Is that one transaction, and if not, what does a half-finished archive look like to a reader?
- An edge points at a resource in another tenant. Is that a refusal, an error, or an alert?
Keep soft-deleted accounts from blocking re-registration
app_user holds user_id, tenant_id, email CITEXT, password_hash (NULL for SSO principals), email_verified_at, auth_version, status ('invited','active','suspended','deactivated'), created_at, updated_at, deleted_at. Two live accounts for one address inside a tenant must be impossible, but an address freed by a soft delete must be reusable, and the same tenant may delete and re-register it repeatedly. Write the uniqueness DDL for PostgreSQL 16, then the equivalent for MySQL 8 where partial indexes do not exist, and say what each permits once three deleted rows already hold that address.
Approach
- Start from what is actually unique: not (tenant_id, email), but (tenant_id, email) among live rows. PostgreSQL says that directly — CREATE UNIQUE INDEX app_user_live_email ON app_user (tenant_id, email) WHERE deleted_at IS NULL. A full constraint over the same two columns burns the address permanently the first time someone deletes an account.
- Keep case-insensitivity in the type or the index, never in the application: CITEXT as given, or UNIQUE (tenant_id, lower(email)) as an expression index where the extension is unavailable. A case-sensitive unique column is exactly how two accounts for one human appear.
- For MySQL 8 the predicate has to move inside the key: add a discriminator column that is a constant 0 while the row is live and is set to user_id on delete, with UNIQUE (tenant_id, email, deleted_marker). Live rows share the constant and still collide; deleted rows differ from each other and stop colliding.
- State the NULL variant and its dependency: leaving the marker NULL for deleted rows also works, because a unique index treats NULLs as distinct — true in MySQL, and true in PostgreSQL only under the default NULLS DISTINCT, which PostgreSQL 15 lets you reverse. Check the polarity against the three existing deleted rows: constant-on-live is what preserves the collision you want, and reversing it silently admits duplicate live accounts.
- Say what a soft delete must do besides setting deleted_at: increment auth_version so existing tokens stop validating, leave resource.owner_user_id and resource_revision.actor_user_id intact, and accept that the address is retained — erasure is a different requirement answered by scrubbing the column, not by a DELETE that would break those references.
Follow-up
- A deleted account re-registers with the same address the next day. Do the old resource rows follow the new user_id, and how does the API keep the two principals apart?
- How do you honour an erasure request while resource_revision.actor_user_id still references this table?
- What changes if a user may hold membership in two tenants?
Denormalise tenant onto revisions and backfill it live
resource_revision (revision_id, resource_id, version, actor_user_id, change_kind, patch, request_id, created_at) has 400M rows and no tenant column; tenant_id lives only on resource. Two reads need it: a tenant-scoped audit feed ordered by created_at DESC, and an offboarding purge. Both join back to resource today. Justify adding tenant_id to resource_revision against those two reads, name the anomaly the copy introduces and the constraint that prevents it, then give the ordered migration for a live table taking 1.2k writes/second — the lock each step takes, how the backfill is batched, and where each step stops being reversible. PostgreSQL 16.
Approach
- Justify from the access path rather than from taste. Without the column, the audit feed either scans resource_revision by created_at and discards other tenants' rows, or resolves the tenant's resource_ids first and probes with them — both proportional to the tenant's whole history rather than to one page. With (tenant_id, created_at DESC, revision_id DESC) it is a seek that stops at 50 rows, and the purge becomes a ranged delete instead of a join.
- Name the cost exactly: a second copy of a fact can disagree with the first. Make the disagreement unwritable rather than documented — add UNIQUE (resource_id, tenant_id) on resource so it can serve as a foreign-key target, then FOREIGN KEY (resource_id, tenant_id) REFERENCES resource (resource_id, tenant_id) on the revision table. A revision can then only ever carry its parent's tenant.
- Step one, expand: ALTER TABLE resource_revision ADD COLUMN tenant_id BIGINT NULL, with no default, so it is a catalogue change and no rewrite. It still needs ACCESS EXCLUSIVE for an instant, and that instant queues behind the longest open transaction on the table while every later query queues behind it — set lock_timeout to 2s and retry rather than wait.
- Step two, dual-write: deploy the writer that populates tenant_id on every new revision while reads still use the join. Reversible by redeploying the previous build, because nothing reads the column yet.
- Step three, backfill: batch by primary key rather than by created_at so the cursor is dense and resumable — UPDATE resource_revision rr SET tenant_id = r.tenant_id FROM resource r WHERE r.resource_id = rr.resource_id AND rr.revision_id > $1 AND rr.revision_id <= $1 + 5000 AND rr.tenant_id IS NULL — committing per batch and persisting the cursor. Throttle on replica replay lag and on dead-tuple count, since each batch writes 5,000 new row versions. Run the backfill before the index exists so those updates can stay HOT.
- Step four, index then enforce then contract: CREATE INDEX CONCURRENTLY (cannot run inside a transaction block, scans the table twice, waits on open transactions, and leaves an INVALID index to drop concurrently if it fails); ADD CONSTRAINT ... CHECK (tenant_id IS NOT NULL) NOT VALID, then VALIDATE CONSTRAINT, which takes only SHARE UPDATE EXCLUSIVE, after which SET NOT NULL uses the validated check instead of re-scanning on PostgreSQL 12 and later. Only then move the audit reads onto the column and, in a later deploy, delete the join path.
Worked solution 40 min
- Write the five steps as separate scripts and state, for each, the lock mode it acquires and the deploy it pairs with.
- On a 20M-row copy, run the ADD COLUMN while a 30-second transaction holds a lock on the table, and record how long unrelated queries queue behind it.
- Run the batched backfill at 5,000 rows, kill it mid-run, restart from the persisted cursor, and confirm no row is processed twice and none is skipped.
- Build the index concurrently under concurrent write load, then add the CHECK ... NOT VALID, VALIDATE it and SET NOT NULL, timing each.
- Compare the audit-feed plan before and after: join-and-filter versus an index seek with no Sort.
Follow-up
- The backfill is half finished and a rollback is required. What state is the table in, and what does the previous build do with a half-populated column?
- How do you verify the backfill actually finished, given rows are still being inserted while it runs?
- A resource must now be movable between tenants. What does that do to the composite foreign key and to the revisions already written?
Publish rate-limit and deadline semantics the edge actually enforces
The edge API serves about 3k requests/second steady and 9k at peak against a 400 ms p99 budget, with an explicit bounded concurrency limit per instance. Limits exist per principal and per tenant. Callers are a partner integration running nightly bulk loads and a browser app. Specify the counting algorithm and window, which limit a request is charged against, the headers a well-behaved client reads, the status and body when a limit is hit, how that differs from the response when an instance is shedding load, and what each caller does with each.
Approach
- Choose the counter and name its failure mode. Fixed windows admit nearly twice the limit across a boundary - a full burst at the end of one window and another at the start of the next. A token bucket states sustained rate and burst separately, which is exactly what a nightly bulk load needs. A sliding-window counter is more faithful and costs more state per key. State the choice and the burst it permits.
- Charge each request against both keys and reject on the stricter. The tenant limit protects the shared primary, which absorbs roughly 1.2k writes/second in total; the per-principal limit stops one credential inside a tenant from consuming that tenant's whole allowance. The tenant is the fairness unit for the same reason it is the leading column of every index.
- Advertise limit, remaining and reset for the binding key on every response, not only on rejections, so a client can pace before it is refused. Pick one naming scheme - the RateLimit-* draft fields or an X-prefixed set - document the units, and never change them afterwards.
- Separate two rejections that look identical to a naive client. 429 means this caller exceeded its own share and Retry-After is a real schedule it should obey. 503 means the instance is at its concurrency bound and shedding, which is a statement about the server; a fleet-wide 503 retried on a fixed delay resynchronises every client into one stampede, so full jitter is mandatory there and the delay is the client's guess, not ours.
- Make shedding cheap and early - before the token is verified against the database, before any downstream call - because a rejection that costs as much as the work relieves nothing. Drop requests whose client deadline has already elapsed rather than serving them; the caller has stopped listening and the work is pure cost.
- Write the caller behaviours down: the bulk loader paces against
remainingand treats a 429 as a defect in its own pacing; the browser surfaces the wait and must never retry a 429 inside a render loop, which turns one limited user into a self-inflicted flood.
Worked solution 20 min
- Write the bucket parameters for both keys: sustained rate, burst size, and the refill interval, with the arithmetic that ties them to the 3k/9k figures.
- Draft the three response headers and one example 429 body carrying a code, the limit that bound, and Retry-After.
- Write the 429-versus-503 decision as a two-line rule an on-call engineer can apply to a log line.
- State where in the request pipeline the rejection happens and which work it skips.
Follow-up
- One tenant stays under its limit and still degrades everyone else during a backfill. What changes - the limiter, the worker concurrency caps, or both?
- How are counters kept correct across 20 to 40 stateless instances, and what does your answer cost per request?
Rebuild the search projection while it serves nine thousand queries
The read-model service answers about 9k queries/second at a 120 ms p99 from a projection built off the event log, normally under 2 seconds behind. A mapping change forces a full rebuild from resource_revision, during which apply lag rises to minutes. Writes continue at 1.2k/second and events at 4k/second. Design the rebuild: how the new index is populated and cut over, how position is tracked per log partition, what the API returns alongside results so a client can tell a stale answer from a current one, and the criterion for cutting over.
Approach
- Build into a second index and swap an alias rather than mutating the live one. The rebuild is then reversible by pointing the alias back, so a bad mapping costs a wasted rebuild instead of an outage. The price is peak storage for two full copies and double apply load during catch-up, and both numbers should be stated up front rather than discovered when disk fills.
- Track position per log partition, not globally. Order is guaranteed only within an aggregate's partition, so progress is a vector, and the only number safe to publish is taken from the least advanced partition - the slowest one is what bounds completeness. Publishing the most advanced partition's position declares the projection current while another partition sits twenty minutes behind.
- Make apply idempotent so the backfill and the live tail can overlap without a freeze. Each document records the aggregate_version it reflects, and any event at or below that version is discarded. This is why the event carries the full fact and not a delta: a delta cannot be discarded safely, and a consumer that calls back to read current state applies a state newer than the event it is processing, which is how a projection ends up with changes applied out of order.
- Turn lag into a contract instead of a surprise. Return the watermark with every result set, and return the version produced by a write so the client can compare the two. A client that wrote version 7 and receives results at a watermark older than its own commit can show that its change is still landing, rather than rendering the previous value as current. Blocking the read until the projection catches up would convert a staleness problem into an availability problem at 9k queries/second, and choosing not to do that is the trade.
- Define the cutover numerically. Writes are untouched by the rebuild - they commit to the primary and land in the outbox - so the only coupling is apply throughput. If catch-up applies slower than the 4k events/second arriving, it never converges. The cutover criterion is that lag is measurably decreasing and below a stated threshold, not that the backfill loop reached the end of its range.
Follow-up
- The rebuilt index disagrees with the primary tables for 300 documents. Which is authoritative, and how do you decide without freezing writes?
- The rebuild doubles load on the log and pushes the projection p99 from 120 ms to 400 ms. What do you throttle, and which signal sets how much?
- Clients start polling until the watermark passes their write. What does that do at 9k queries/second, and what do you offer instead?
Every query on one table stalls for forty seconds mid-deploy
During a release on PostgreSQL, every query touching resource times out for about 40 seconds and then recovers with no intervention. The release ran one migration, ALTER TABLE resource ADD COLUMN archived_reason TEXT, and the migration log shows it completing in 6 ms. Unrelated tables showed no change in error rate. Explain how a 6 ms statement caused a 40-second stall, give the ordered checks you would run on a live system to confirm it, and give the migration procedure that prevents a repeat.
Approach
- Separate the statement's duration from the lock's duration. ADD COLUMN with no default is a catalogue-only change and genuinely runs in milliseconds, but it requires ACCESS EXCLUSIVE, and it cannot acquire that until every transaction already touching the table has finished.
- Account for the queueing, which is the part that surprises people. A lock request that is waiting blocks later requests for conflicting modes behind it rather than letting them overtake, so one long-open transaction holds the DDL and the DDL holds all the traffic. The stall length is set by the longest open transaction, not by the size of the change.
- Confirm on a live system in this order: pg_stat_activity for that table ordered by xact_start, looking for the oldest transaction and specifically for state = idle in transaction; then pg_locks where granted = false to find the waiter; then join them on pid to name blocker and blocked. pg_blocking_pids() does that join for you and is the fastest single call.
- Prevent rather than merely time it better. Set lock_timeout to a second or two on the migration session so the DDL abandons the queue after a bounded wait and is retried, instead of holding it for as long as the oldest transaction lives. Be exact about what that buys: queries arriving during the wait still queue behind the pending ACCESS EXCLUSIVE request, so each attempt costs them up to one lock_timeout of added latency. The outage goes from 40 seconds to about one second per attempt, not to zero. Also run migrations away from deploy-time peaks, and put a statement timeout and an idle-in-transaction timeout on the analytics role that opens the long transactions.
- Know the lock each change takes, since the mitigation differs by change. A column with a non-volatile default is a metadata-only change from PostgreSQL 11 and still needs the brief ACCESS EXCLUSIVE; an index needs CREATE INDEX CONCURRENTLY, which cannot run inside a transaction block and leaves an INVALID index to drop if it fails; a check or foreign key is added NOT VALID and then VALIDATE CONSTRAINT as a separate statement under a weaker lock.
Follow-up
- The same release also wants NOT NULL on that column. What is the sequence that gets there without a long lock?
- Your lock_timeout retry fails ten times in a row because the analytics transaction is always open. What do you change?
- How does this differ on MySQL with InnoDB online DDL, and what is the equivalent of the waiting-lock queue there?
The plan starts with the reported Python coding questions, moves to ML theory, AI safety and serving design with PyTorch code on the metrics and serving days, and ends with a timed session and behavioral stories. Practice drills from PracHub's own library are attached where they match the day.
Prepare, practise & reflect
One practical outcome each day. Spend longer where you need it.
0 / 7 done01Rate limiter from scratch
- Implement a token-bucket limiter class with an allow(key, now) method, lazy refill from elapsed time and per-key state; inject the clock so tests need no sleeping.
- Write the fixed window, sliding-window log and token bucket side by side and list for each: memory per key, burst behaviour at a window boundary and accuracy.
- Add a lock and an eviction rule for idle keys, then note what moves to a shared store such as Redis if several processes enforce the limit, including where the clock comes from.
- Work through the PracHub practice drill on publishing rate-limit semantics that the edge enforces.
Deliverable: A working limiter with tests for burst, refill and idle eviction, a three-algorithm comparison table, and notes on the shared-store variant.
Practice prompt ↗Practice prompt ↗Worked solution ↗02Log parsing and streaming generators
- Write the failed-login run finder as a single pass over an iterator, returning both the length and the matching lines, and update the best run after the loop for files that end in failures.
- Extend it to per-user or per-IP runs with a dict of current runs, and decide how to treat malformed lines.
- Write a Python generator that yields fixed-size chunks with itertools.islice, including a final short chunk, and chain it with a parse step so no stage holds the file in memory.
- Test empty files, a file without a trailing newline, one-line files and chunk size larger than the file.
Deliverable: A tested run finder (global and per-key versions) and a chunking generator pipeline, plus a test list covering the four edge cases.
Practice prompt ↗Practice prompt ↗03Trie and shortest path
- Implement a Trie with insert, exact search and starts-with, using dict children and a terminal flag; add URL normalisation and decide between full-string and host-label matching.
- Compare the Trie with a set and a Bloom filter for a large blacklist on memory, prefix queries and false positives.
- Write BFS with a deque and a parent map for an unweighted network graph and rebuild the path; then write Dijkstra with heapq for weighted links such as latency.
- Cover unreachable nodes, cycles and a source equal to the target; practise the PracHub drill on archiving a resource graph without recursion for iterative traversal habits.
Deliverable: A tested Trie with a written matching policy, BFS and Dijkstra implementations with path reconstruction, and a one-paragraph choice between the Trie, a set and a Bloom filter.
Practice prompt ↗Practice prompt ↗04Evaluation metrics and adversarial robustness in PyTorch
- Take a classifier with 1 attack in 1,000 flows. Compute precision, recall and F1 from a confusion matrix at two thresholds, explain why ROC-AUC and PR-AUC can disagree, and pick an operating threshold from a false-positive budget.
- In PyTorch, train a small attack-versus-benign classifier with a class-weighted loss (BCEWithLogitsLoss with pos_weight, or CrossEntropyLoss with weight), evaluate under model.eval() and torch.no_grad(), and report precision, recall and PR-AUC on a time-based split rather than a random one.
- Write FGSM on the attack flows only, since evasion means making attacks look benign: model.eval(), x_adv = x.clone().detach().requires_grad_(True), loss.backward(), then x_adv = (x_adv + eps * x_adv.grad.sign()).clamp(lo, hi).detach(); measure how far recall drops as eps grows.
- Note that feature-space changes to ports, flags or counts are often not realisable traffic, so a realistic evasion test perturbs only features an attacker controls. List evasion, poisoning and model-extraction attacks, with adversarial training, input validation and monitoring as defences.
Deliverable: A metrics sheet with the worked confusion matrices and threshold choice, a PyTorch training, evaluation and FGSM script with a recall-versus-epsilon table for attack flows, and a table of attack types with one defence each.
Worked solution ↗05Architectures and LLM safety evaluation
- Write from memory: scaled dot-product attention softmax(QK^T / sqrt(d))V, its O(n^2) cost in sequence length, a CNN receptive field, and why plain RNNs suffer vanishing gradients while LSTMs mitigate it.
- For a sequence of network events, argue for and against a Transformer, a CNN and an RNN on latency, data needs and long-range dependencies.
- Outline an automated evaluation framework for a generative model: a labelled test set, scripted red-team prompts, automated and human grading, regression gates before release, and live output monitoring with guardrails.
- Name two failure modes of using another model as the grader and how you would check it.
Deliverable: An architecture cheat sheet with the attention equation and trade-off table, and a one-page outline of an LLM evaluation and guardrail pipeline.
06ML pipeline and serving design
- Design a real-time threat detection pipeline. Begin with clarifying questions on traffic volume, latency target, label source and who consumes alerts, then specify a baseline (rules or a tree-based model) and what would justify a neural model.
- Apply torch.ao.quantization.quantize_dynamic(model, {nn.Linear}, dtype=torch.qint8) to the eager day-4 model, then trace it (for ONNX, export the float model and quantise the file with onnxruntime.quantization.quantize_dynamic); measure p50 and p99 latency at batch size 1 and small batches against the original, checking recall.
- Cover feature stores and training-serving consistency, drift checks (population stability index or a KS test on key features), shadow deployment, canary rollout and rollback.
- Say where the model runs relative to the traffic path, what happens when it is slow or unavailable, and where batching, distillation or caching would help.
Deliverable: A one-page design with assumed numbers, the baseline and upgrade path, a latency and recall table for the original and quantised model, a monitoring list and a fallback behaviour.
07Timed mock and behavioral stories
- Re-solve the rate limiter and the Trie under a timer, speaking your assumptions, complexity and tests as you go.
- Answer one ML theory question and one design question out loud, ending each with the trade-off you chose.
- Write three STAR stories from your real work: a model shipped to production, a trade-off explained to a non-technical stakeholder, and a result you measured with numbers.
- Practise the PracHub behavioral prompts on shipping under a deadline and on estimating unfamiliar work.
Deliverable: A timed-run log for both coding problems, three STAR stories with measured results, and a list of the weakest points to review.
Practice prompt ↗Practice prompt ↗Practice prompt ↗Practice prompt ↗Worked solution ↗Expand any day for tasks and deliverables. Your progress is saved on this device.
Candidates are advised to use the STAR method with a measured result, and the role is described as involving threat researchers, core software engineers and product managers, plus explaining trade-offs to technical and non-technical stakeholders. Prepare stories about taking a model from prototype to production, working with domain experts, choosing a simple baseline over a complex model, and admitting what you did not know.
Ship under a deadline and bound the debt you chose
You have four days to ship a tenant-facing listing endpoint. The version you would defend uses keyset pagination over (tenant_id, status, updated_at DESC, resource_id DESC); the version you can finish uses LIMIT/OFFSET with no matching index. Describe a deadline call you actually made of this shape: what you shipped, what you knowingly deferred, how you bounded the damage with a mechanism rather than an intention, and the specific numeric condition that would force the follow-up. Name who you told and where you wrote it down.
Approach
- Name the deferred failure precisely instead of calling it slow. OFFSET n makes the database produce and discard n rows, so cost grows with page depth; without an index matching the sort, every matching row is read and sorted before the limit applies; and rows inserted between two page fetches shift across the boundary so items are skipped or repeated with nothing in the response to signal it.
- Bound the blast radius with something mechanical rather than a promise: cap maximum page depth, cap page size, restrict the endpoint to one internal caller, or keep it behind a flag. State which failure each cap removes and which it leaves standing.
- Attach a number to the trigger and wire it to an alarm: the first tenant crossing N resources, or the endpoint's p99 crossing its share of the 400 ms budget, so the debt announces itself instead of waiting to be remembered.
- Write it where the next engineer looks, which is the code and the ticket, not a chat message: what was deferred, why, the cap, and the trigger.
- Report what actually happened in your real example, including the case where the trigger never fired and the debt was correctly never repaid.
Follow-up
- At what page depth does the offset version breach your latency budget, given your page size and row counts?
- What breaks first when you switch to keyset pagination later, and what does a client holding an old page token see?
- Who would have overruled you if you had asked for two more days, and did you ask?
Tell callers you do not own that their integration breaks
A field in a write endpoint's response must change shape. You own the endpoint; you do not own the four internal callers or the outbound webhook consumers who read it. Describe a deprecation you were responsible for: what you shipped first, how you established who was actually reading the field, the window you gave and what set its length, what you did about the consumer who never moved, and how you decided removal was safe. Name the signal you used, not the announcement you sent.
Approach
- Establish the reader set empirically rather than from a wiki of owners: per-field usage counters keyed by principal, or access logs attributed to a consumer. State the blind spot of whichever you pick, since a consumer that reads the field only on a monthly job will not appear in a week of logs.
- Ship additive first. Populate the new field alongside the old one so no reader is forced to move, which is also what keeps a rolling deploy safe, because old and new instances answer the same requests at the same time and a rollback must still find the old shape present.
- Set the window from the slowest legitimate consumer's release cadence, not from your calendar, and decide separately what to do for a consumer with no release process at all, such as an external webhook endpoint you can only email.
- Convert silence into evidence before you rely on it: a short, low-traffic removal window that makes a still-dependent consumer fail visibly and loudly while you are watching, rather than at three in the morning after you have moved on.
- State the removal criterion as a measurement with a duration attached, such as observed reads at zero across a full billing cycle, and keep the change reversible for one release after removal.
Follow-up
- How would you detect a consumer that reads the field only during a monthly export?
- One caller refuses to move and has a commercial relationship behind it. What changes in your plan and what does not?
- After removal, what makes the change irreversible, and how long before you cross that line?
Estimate work you have never done and defend the range
You are asked to estimate a change you have never attempted: add a column to a 100-million-row table, populate it, move reads across, and drop the old shape. Give a range with the assumptions that generate it, including batch size, the signal your backfill throttles on, and wall-clock hours, and name the three unknowns that would move the number most. Then describe a real estimate you gave under comparable ignorance: how you expressed its uncertainty, what you committed to, and how wrong you turned out to be.
Approach
- Decompose into independently deployable steps before estimating anything: add the column nullable, write both shapes, backfill in batches, verify, move reads, stop writing the old shape, drop it. That is four deploys spread over days, and the calendar estimate is dominated by them rather than by the loop's runtime.
- Do the arithmetic aloud for the part that has arithmetic in it: batch size times number of batches times per-batch duration, at a write rate the primary can absorb alongside roughly 1.2k writes per second of production traffic. The loop is throttled by replication lag and lock waits, not by how fast it can issue statements.
- Price the schema step by its lock rather than its statement duration. In PostgreSQL an ALTER TABLE taking ACCESS EXCLUSIVE waits for every open transaction on that table while later queries queue behind it, so a millisecond change issued during a thirty-second analytics query stalls that table for thirty seconds. Adding a nullable column with a non-volatile default avoids a rewrite from version 11; a new index wants CREATE INDEX CONCURRENTLY, which cannot run inside a transaction block and leaves an invalid index behind if it fails.
- Express the answer as a range whose endpoints each trace to a stated assumption, then name the cheapest experiment that collapses it, which is almost always running one real batch against the real table and multiplying.
- Commit to a checkpoint rather than a completion date: the day you report a measured number from that first batch. That is a promise you can keep under uncertainty, and it is what the asker actually needs in order to plan.
Follow-up
- How do you verify the backfill genuinely finished, given rows written by production traffic while it ran?
- Where does the backfill resume from after a worker is killed mid-batch, and what makes that resume point trustworthy?
- Your first batch comes back ten times slower than assumed. What do you tell the person waiting on the estimate, and when?
- 01
Describe a model you took from prototype to production and what you monitored after release.
- 02
Tell me about a time you explained an accuracy versus latency trade-off to someone non-technical.
- 03
Describe a project where you started with a simple baseline and what made you add complexity, or not.
- 04
Tell me about working with domain experts, such as security researchers, to turn their knowledge into model features.
- 05
Describe a time you were asked about a method you did not know and how you handled it.
- 06
Give an example of a result you quantified, how you measured it and what the baseline was.
Is this an official A10 Networks interview guide?
No. It is PracHub's own research and practice material for the Machine Learning Engineer role at A10 Networks. The questions and areas reflect what candidates have reported, not a process the company has published, and they can change. Confirm the current format with your recruiter.
PracHub interview research ↗Which coding questions do candidates report?
Five: implement a custom rate limiter for an API endpoint; parse a log file of network requests and return the longest contiguous sequence of failed login attempts; write a Python generator that processes a massive dataset in chunks to prevent memory overflow; implement a Trie to store and search a large blacklist of malicious URLs; and solve a graph traversal problem such as the shortest path in a network topology. Practise them as complete, runnable Python.
PracHub Machine Learning Engineer practice ↗How long should I prepare?
Candidates describe spending roughly two to three weeks on core algorithms, coding in a shared editor and refreshing ML fundamentals. The seven-day plan here compresses that: three days on coding, two on ML theory and AI safety, one on serving design and one on a timed mock plus stories. Stretch the coding days if Python data structures are not yet automatic.
PracHub Machine Learning Engineer practice ↗What ML topics should I cover?
Candidates describe evaluation metrics (precision/recall versus ROC-AUC) and custom evaluation frameworks for generative models, adversarial robustness, and the mechanics of Transformers, CNNs and RNNs. On the engineering side they describe serving at scale, data drift, feature stores and low-latency inference. Python with PyTorch or TensorFlow is the stack the role lists.
PracHub Machine Learning Engineer practice ↗How should I prepare for AI safety questions?
Candidates are advised to favour practical implementation over abstract frameworks: how you would build automated guardrails, a red-teaming pipeline and live monitoring of model outputs. Prepare an outline covering a test set, scripted adversarial prompts, grading, release gates and production monitoring, and be ready to discuss bias, hallucination and adversarial vulnerability.
PracHub Machine Learning Engineer practice ↗Does the loop change for Senior and Staff levels?
Candidates report that the role targets Senior to Staff levels with several years of ML, data science or software experience, and that senior onsite loops lean on architecture and cross-functional influence. Prepare design answers that state trade-offs and a project you led end to end. Round counts and order are not reported, so ask your recruiter.
PracHub Machine Learning Engineer practice ↗Sources & methodology 3 sources ↗
No official company page is cited. Rounds and questions come from candidate reports and PracHub editorial material; each source shows the date it was read.
- 01PracHub interview research ↗
PracHub editorial research into this company and role, maintained with this guide. Candidate-reported, not an employer publication.
PracHub page · Accessed 2026-09-30 - 02PracHub Machine Learning Engineer practice ↗
Cross-company practice questions for this role.
PracHub page · Accessed 2026-09-30 - 03PracHub interview preparation framework ↗
The framework the preparation plan follows.
PracHub page · Accessed 2026-09-30