An interview at NinjaTrader, a small brokerage.
Server Log Analyzer and Anomaly Detector
Problem environment: an IDE with server.log (sample logs), analyzer_stub.py (skeleton code), and test.py (pytest unit tests). You have to implement the functions below.
Input data spec (JSON Schema): the input is multi-line text, and each line is a JSON string containing these 4 required fields:
- timestamp: a UTC timestamp in ISO-8601 format (e.g. "2026-09-24T07:30:15Z")
- level: log level ("INFO", "WARN", "ERROR")
- latency_ms: request response time (numeric)
- status_code: HTTP status code (integer, e.g. 200, 404, 500)
Task 1: Log parsing and data cleaning (parse_logs)
- Function signature:
parse_logs(raw: str) -> list[dict] - Requirements:
- Parse the incoming log text line by line, filter out empty lines, and parse each line into a JSON dict.
- Validate that every record contains all 4 required fields: ('timestamp', 'level', 'latency_ms', 'status_code').
- Output stream restriction: if you hit a line with invalid JSON or a missing required field, you must skip that line, and the warning must be written to standard error (sys.stderr) (e.g.
print(..., file=sys.stderr)). Writing it directly to standard output (stdout) is strictly forbidden. - Return a list made up of all the valid records.
Task 2: Overall summary statistics (print_summary)
- Function signature:
print_summary(entries: list[dict]) -> None - Requirements: if there are no records, just exit; otherwise compute and print to standard output (stdout) a report covering these dimensions:
- Total count: the total number of valid log entries.
- Log level distribution: the count of INFO, WARN, and ERROR entries and each one's percentage (1 decimal place).
- Latency metrics:
- Average latency (Avg Latency, 1 decimal place)
- P95 latency (95th percentile, taken by index after sorting)
- Max latency (Max Latency)
- Status code distribution: show how often each status code appears, sorted by HTTP status code in ascending order.
- Output format:
=== Log Summary ===
Total entries: 100
Log levels:
INFO: 80 (80.0%)
WARN: 15 (15.0%)
ERROR: 5 (5.0%)
Latency (ms):
Avg: 112.4 | P95: 245 | Max: 410
Status codes:
200: 80 | 400: 10 | 404: 5 | 500: 5
Task 3: Time-bucket window anomaly detection (detect_anomalies)
- Function signature:
detect_anomalies(entries: list[dict]) -> None - Requirements: group the valid log entries into 1-minute time buckets by timestamp, then go through each bucket and evaluate the anomaly metrics:
- Error rate spike alert: fires when ERROR logs make up more than 10% of the current bucket:
- Format:
[YYYY-MM-DDTHH:MM:SSZ] ERROR rate spike: X.X% (threshold: 10%)
- Format:
- Latency spike alert: fires when the average latency of the requests in the current bucket exceeds 200ms:
- Format:
[YYYY-MM-DDTHH:MM:SSZ] Latency spike: XXXms avg (threshold: 200ms)
- Format:
- If there are no anomalies, output
"\nNo anomalies detected.".
- Error rate spike alert: fires when ERROR logs make up more than 10% of the current bucket:
- Output format:
=== Anomalies Detected ===
[2026-09-24T07:31:00Z] ERROR rate spike: 25.0% (threshold: 10%)
[2026-09-24T07:35:00Z] Latency spike: 284ms avg (threshold: 200ms)
Task 4: Scaling to large data (asked verbally)
- Question: "If the log file grows to tens of GB and can't be loaded into memory all at once, how would you optimize the current code structure?"
- What it tests: streaming line-by-line parsing with Python generators (yield), and keeping memory usage under control with a sliding window built on a double-ended queue (collections.deque).
Discussion
Loading comments…