Server Log Analyzer: Parse JSON Lines, Summarize Latency, Flag Per-Minute Anomalies

Read the full interview experience this question came from →

Quick Overview

A practical Python exercise to finish a server log analyzer: parse JSON-lines logs while sending warnings for bad lines to stderr, print a summary with level shares, average, P95 and max latency and status-code counts, and flag per-minute error-rate and latency spikes. A verbal follow-up asks how to handle log files too large for memory.

Server Log Analyzer: Parse JSON Lines, Summarize Latency, Flag Per-Minute Anomalies

Company: Ninjatrader

Role: Software Engineer

Category: Software Engineering Fundamentals

Difficulty: hard

Interview Round: Technical Screen

Complete a small server-log analyzer in Python. The online IDE provides three files: `server.log` (a sample log), `analyzer_stub.py` (skeleton code with the function stubs) and `test.py` (pytest unit tests). Implement the functions in Parts 1 to 3 so that the unit tests pass, then answer the verbal question in Part 4. The log input is multi-line text. Each line is one JSON object with four required fields: | Field | Type | Meaning | |---|---|---| | `timestamp` | string | UTC timestamp in ISO-8601 format, for example `"2026-09-24T07:30:15Z"` | | `level` | string | Log level: `"INFO"`, `"WARN"` or `"ERROR"` | | `latency_ms` | number | Request response time in milliseconds | | `status_code` | integer | HTTP status code, for example `200`, `404` or `500` | ### Constraints and Clarifications - Standard output (stdout) is reserved for the reports printed in Parts 2 and 3. Any warning about a skipped line must go to standard error, for example `print(..., file=sys.stderr)`; printing it to stdout is not allowed. - The report formats in Parts 2 and 3 are fixed. The samples show the content and order of every line; exact whitespace such as indentation or blank lines follows the provided tests. - Percentages and the average latency in the summary are printed with one decimal place. - Part 4 is discussed verbally; no code is required for it. ### Clarifying Questions - Are log lines guaranteed to be in timestamp order? This matters for bucketing in Part 3 and for streaming in Part 4. - Should `parse_logs` validate field values (a known level, a numeric latency, an integer status code, a parseable timestamp), or only check that the four fields are present? - Can a record contain extra fields, and should they be preserved? - Is a required field that is present with the value `null` considered missing? ### Part 1 — Parse and clean the logs ```python def parse_logs(raw: str) -> list[dict]: ``` 1. Process the text line by line, skip empty lines, and parse each remaining line as JSON into a dictionary. 2. Check that every record contains all four required fields: `timestamp`, `level`, `latency_ms` and `status_code`. 3. If a line is not valid JSON or is missing a required field, skip that line and write a warning to `sys.stderr`. The warning must never go to stdout. 4. Return the list of all valid records. ```hint One gate per line Send every line through a single check that ends in exactly one of two ways: a valid record, or one warning on stderr. Then no bad line can reach stdout or stop the loop. ``` #### Clarifying Questions for this Part - Does a line that contains only whitespace count as empty? - Should the warning include the line number and the reason the line was rejected? - Is a line that is valid JSON but not an object, such as `[1, 2]` or `42`, treated as invalid? #### What This Part Should Cover - Robust per-line handling of malformed JSON, non-object JSON, missing fields and blank lines. - Strict stream discipline: warnings only on stderr. - Valid records returned in input order. ### Part 2 — Summary report ```python def print_summary(entries: list[dict]) -> None: ``` If `entries` is empty, return without printing anything. Otherwise print a report to stdout with: 1. **Total entries**: the number of valid records. 2. **Log level distribution**: the count of `INFO`, `WARN` and `ERROR` records, and each count's percentage of the total, to one decimal place. 3. **Latency metrics**: the average latency (one decimal place), the P95 latency (the 95th percentile, taken by sorting the latencies and selecting the value at the corresponding index), and the maximum latency. 4. **Status code distribution**: how many times each HTTP status code occurs, listed in ascending numeric order of the code. Sample output: ```text === Log Summary === Total entries: 100 Log levels: INFO: 80 (80.0%) WARN: 15 (15.0%) ERROR: 5 (5.0%) Latency (ms): Avg: 112.4 | P95: 245 | Max: 410 Status codes: 200: 80 | 400: 10 | 404: 5 | 500: 5 ``` ```hint Pin the percentile index Write down which index "the 95th percentile after sorting" means for a list of n values, and check it for n = 1, n = 20 and n = 100 before coding it. ``` #### Clarifying Questions for this Part - Which index defines P95 for n sorted values: `int(0.95 * n)`, or the nearest-rank index `ceil(0.95 * n) - 1`? - Should a level with no records still be printed as `0 (0.0%)`? - How should P95 and Max be printed when latencies are not whole numbers? - How should a record whose level is not one of the three known levels be counted? #### What This Part Should Cover - Correct aggregates and formatting: one-decimal rounding and percentages of the total. - A precisely defined P95 index that can never go out of range. - Deterministic ordering: a fixed level order and ascending status codes. ### Part 3 — Per-minute anomaly detection ```python def detect_anomalies(entries: list[dict]) -> None: ``` Assign each valid record to a 1-minute time bucket based on its timestamp, then evaluate every bucket: 1. **Error rate spike**: when `ERROR` records make up more than 10% of the bucket's records, print `[YYYY-MM-DDTHH:MM:SSZ] ERROR rate spike: X.X% (threshold: 10%)` 2. **Latency spike**: when the average latency of the bucket's records is more than 200 ms, print `[YYYY-MM-DDTHH:MM:SSZ] Latency spike: XXXms avg (threshold: 200ms)` 3. If no bucket has an anomaly, print `"\nNo anomalies detected."`, that is, a newline followed by the message. Sample output: ```text === Anomalies Detected === [2026-09-24T07:31:00Z] ERROR rate spike: 25.0% (threshold: 10%) [2026-09-24T07:35:00Z] Latency spike: 284ms avg (threshold: 200ms) ``` In the sample, the bracketed time is the start of the bucket's minute. ```hint Choose the bucket key Derive one key per record that is identical for every timestamp in the same UTC minute and that sorts chronologically; the printed label can be produced from the same key. ``` #### Clarifying Questions for this Part - Is the `=== Anomalies Detected ===` header printed even when there are no anomalies, before the "No anomalies detected." line? - How is the average latency turned into the whole number in the latency line: rounded or truncated? - Are buckets reported in chronological order, and if one bucket triggers both alerts, which line comes first? - Can timestamps carry fractional seconds or an offset other than `Z`? #### What This Part Should Cover - Correct bucketing by UTC minute and chronological iteration over buckets. - Strict thresholds ("more than", not "at least") and exact line formats. - The no-anomaly path and empty input. ### Part 4 — Scaling to very large log files (verbal) "If the log file grows to tens of gigabytes and can no longer be loaded into memory at once, how would you restructure the current code?" ```hint Audit every full-list step Go through the three functions and mark each place that needs all records at once. For each one, decide what state is really needed to produce the same output. ``` #### What This Part Should Cover - Processing the input as a stream, so memory does not grow with the file size. - Which statistics can be kept incrementally, and what to do about the percentile, which needs more than a running total. - Bounded state for the per-minute buckets, and the assumption about timestamp order that makes it work. ### What a Strong Answer Covers - A clean separation of parsing, aggregation and formatting, so each piece can be tested against the provided pytest file. - Exact adherence to the output formats and to the stdout and stderr split. - Explicit decisions on every ambiguous rule (percentile index, rounding, zero-count levels, bucket order), stated before coding. - Edge cases: empty input, all lines invalid, a single record, and a bucket at exactly 10% errors or exactly 200 ms average latency. - Time and memory complexity, and a credible streaming design for Part 4. ### Follow-up Questions - How would you compute an exact P95 over a file that does not fit in memory, and when is an approximation acceptable? - If log lines can arrive a little out of order, how would you decide that a minute bucket is complete and can be reported? - How would you make the 10% and 200 ms thresholds configurable without breaking the output format the tests expect? - How would you split the analysis across several processes or machines and merge the partial results?

Overview: A practical Python exercise to finish a server log analyzer: parse JSON-lines logs while sending warnings for bad lines to stderr, print a summary with level shares, average, P95 and max latency and status-code counts, and flag per-minute error-rate and latency spikes. A verbal follow-up asks how to handle log files too large for memory.

Read the full Ninjatrader Software Engineer interview experience this question came from

|Home/Software Engineering Fundamentals/Ninjatrader
Ninjatrader logo
Ninjatrader
Sep 9, 2026
hardSoftware EngineerTechnical ScreenSoftware Engineering Fundamentals
0
0

Complete a small server-log analyzer in Python. The online IDE provides three files: server.log (a sample log), analyzer_stub.py (skeleton code with the function stubs) and test.py (pytest unit tests). Implement the functions in Parts 1 to 3 so that the unit tests pass, then answer the verbal question in Part 4.

The log input is multi-line text. Each line is one JSON object with four required fields:

FieldTypeMeaning
timestampstringUTC timestamp in ISO-8601 format, for example "2026-09-24T07:30:15Z"
levelstringLog level: "INFO", "WARN" or "ERROR"
latency_msnumberRequest response time in milliseconds
status_codeintegerHTTP status code, for example 200, 404 or 500

Constraints and Clarifications

  • Standard output (stdout) is reserved for the reports printed in Parts 2 and 3. Any warning about a skipped line must go to standard error, for example print(..., file=sys.stderr) ; printing it to stdout is not allowed.
  • The report formats in Parts 2 and 3 are fixed. The samples show the content and order of every line; exact whitespace such as indentation or blank lines follows the provided tests.
  • Percentages and the average latency in the summary are printed with one decimal place.
  • Part 4 is discussed verbally; no code is required for it.

Clarifying Questions Guidance

  • Are log lines guaranteed to be in timestamp order? This matters for bucketing in Part 3 and for streaming in Part 4.
  • Should parse_logs validate field values (a known level, a numeric latency, an integer status code, a parseable timestamp), or only check that the four fields are present?
  • Can a record contain extra fields, and should they be preserved?
  • Is a required field that is present with the value null considered missing?

Part 1 — Parse and clean the logs

def parse_logs(raw: str) -> list[dict]:
  1. Process the text line by line, skip empty lines, and parse each remaining line as JSON into a dictionary.
  2. Check that every record contains all four required fields: timestamp , level , latency_ms and status_code .
  3. If a line is not valid JSON or is missing a required field, skip that line and write a warning to sys.stderr . The warning must never go to stdout.
  4. Return the list of all valid records.

Clarifying Questions for this Part Guidance

  • Does a line that contains only whitespace count as empty?
  • Should the warning include the line number and the reason the line was rejected?
  • Is a line that is valid JSON but not an object, such as [1, 2] or 42 , treated as invalid?

What This Part Should Cover Guidance

  • Robust per-line handling of malformed JSON, non-object JSON, missing fields and blank lines.
  • Strict stream discipline: warnings only on stderr.
  • Valid records returned in input order.

Part 2 — Summary report

def print_summary(entries: list[dict]) -> None:

If entries is empty, return without printing anything. Otherwise print a report to stdout with:

  1. Total entries : the number of valid records.
  2. Log level distribution : the count of INFO , WARN and ERROR records, and each count's percentage of the total, to one decimal place.
  3. Latency metrics : the average latency (one decimal place), the P95 latency (the 95th percentile, taken by sorting the latencies and selecting the value at the corresponding index), and the maximum latency.
  4. Status code distribution : how many times each HTTP status code occurs, listed in ascending numeric order of the code.

Sample output:

=== Log Summary ===
Total entries: 100
Log levels:
INFO: 80 (80.0%)
WARN: 15 (15.0%)
ERROR: 5 (5.0%)
Latency (ms):
Avg: 112.4 | P95: 245 | Max: 410
Status codes:
200: 80 | 400: 10 | 404: 5 | 500: 5

Clarifying Questions for this Part Guidance

  • Which index defines P95 for n sorted values: int(0.95 * n) , or the nearest-rank index ceil(0.95 * n) - 1 ?
  • Should a level with no records still be printed as 0 (0.0%) ?
  • How should P95 and Max be printed when latencies are not whole numbers?
  • How should a record whose level is not one of the three known levels be counted?

What This Part Should Cover Guidance

  • Correct aggregates and formatting: one-decimal rounding and percentages of the total.
  • A precisely defined P95 index that can never go out of range.
  • Deterministic ordering: a fixed level order and ascending status codes.

Part 3 — Per-minute anomaly detection

def detect_anomalies(entries: list[dict]) -> None:

Assign each valid record to a 1-minute time bucket based on its timestamp, then evaluate every bucket:

  1. Error rate spike : when ERROR records make up more than 10% of the bucket's records, print [YYYY-MM-DDTHH:MM:SSZ] ERROR rate spike: X.X% (threshold: 10%)
  2. Latency spike : when the average latency of the bucket's records is more than 200 ms, print [YYYY-MM-DDTHH:MM:SSZ] Latency spike: XXXms avg (threshold: 200ms)
  3. If no bucket has an anomaly, print "\nNo anomalies detected." , that is, a newline followed by the message.

Sample output:

=== Anomalies Detected ===
[2026-09-24T07:31:00Z] ERROR rate spike: 25.0% (threshold: 10%)
[2026-09-24T07:35:00Z] Latency spike: 284ms avg (threshold: 200ms)

In the sample, the bracketed time is the start of the bucket's minute.

Clarifying Questions for this Part Guidance

  • Is the === Anomalies Detected === header printed even when there are no anomalies, before the "No anomalies detected." line?
  • How is the average latency turned into the whole number in the latency line: rounded or truncated?
  • Are buckets reported in chronological order, and if one bucket triggers both alerts, which line comes first?
  • Can timestamps carry fractional seconds or an offset other than Z ?

What This Part Should Cover Guidance

  • Correct bucketing by UTC minute and chronological iteration over buckets.
  • Strict thresholds ("more than", not "at least") and exact line formats.
  • The no-anomaly path and empty input.

Part 4 — Scaling to very large log files (verbal)

"If the log file grows to tens of gigabytes and can no longer be loaded into memory at once, how would you restructure the current code?"

What This Part Should Cover Guidance

  • Processing the input as a stream, so memory does not grow with the file size.
  • Which statistics can be kept incrementally, and what to do about the percentile, which needs more than a running total.
  • Bounded state for the per-minute buckets, and the assumption about timestamp order that makes it work.

What a Strong Answer Covers Guidance

  • A clean separation of parsing, aggregation and formatting, so each piece can be tested against the provided pytest file.
  • Exact adherence to the output formats and to the stdout and stderr split.
  • Explicit decisions on every ambiguous rule (percentile index, rounding, zero-count levels, bucket order), stated before coding.
  • Edge cases: empty input, all lines invalid, a single record, and a bucket at exactly 10% errors or exactly 200 ms average latency.
  • Time and memory complexity, and a credible streaming design for Part 4.

Follow-up Questions Guidance

  • How would you compute an exact P95 over a file that does not fit in memory, and when is an approximation acceptable?
  • If log lines can arrive a little out of order, how would you decide that a minute bucket is complete and can be reported?
  • How would you make the 10% and 200 ms thresholds configurable without breaking the output format the tests expect?
  • How would you split the analysis across several processes or machines and merge the partial results?
Loading comments...