Interview conceptStatistics & Math

Robust Statistical Metrics

Asked of: Software Engineer

Last updated

What's being tested

This tests whether you can choose and implement a metric whose behavior matches the operational question, rather than treating “average” as a default. For a Software Engineer, the interviewer is probing your ability to reason about aggregation semantics, numerical stability, streaming or distributed computation, and edge cases such as outliers, empty samples, and skewed trading returns. Hudson River Trading cares because a seemingly small metric choice can change monitoring, debugging, alerting, and decisions about whether a trading system is behaving safely.

Core knowledge

  • Mean is xˉ=1n∑i=1nxi\bar{x}=\frac{1}{n}\sum_{i=1}^{n}x_i and uses every observation, making it useful for total expected profit per trade when extreme values are legitimate and economically important. It is highly sensitive to outliers.

  • Median is the 50th percentile: after sorting, it is the middle observation, or the average of the two middle observations for even nn. It is robust to a small number of arbitrarily large wins or losses.

  • Robustness means a metric changes relatively little when a small fraction of observations are contaminated. Median, trimmed mean, and winsorized mean are more robust than ordinary mean, but they answer different questions.

  • A trading-profit distribution is often heavy-tailed and skewed: many small outcomes may coexist with rare large gains or losses. The median describes a typical trade; the mean describes expected profit per trade and preserves tail contribution.

  • Define the unit of observation before aggregating. “Mean profit per fill,” “median profit per order,” and “daily total P&L” are different metrics; mixing fills, orders, and days can overweight high-activity periods.

  • For a bounded-memory streaming mean, use Welford’s algorithm for stable running statistics. Maintain (n,μ,M2)(n,\mu,M_2); update with δ=x−μ\delta=x-\mu, μ←μ+δ/n\mu\leftarrow\mu+\delta/n, and M2←M2+δ(x−μ)M_2\leftarrow M_2+\delta(x-\mu).

  • A median requires order information. For an in-memory batch, sorting costs O(nlog⁡n)O(n\log n) time and O(n)O(n) space; Quickselect finds an exact median in expected O(n)O(n) time, while two heaps support online updates in O(log⁡n)O(\log n) per value.

  • The two-heaps algorithm keeps a max-heap for the lower half and a min-heap for the upper half, maintaining size difference at most one and every lower-half value no greater than every upper-half value. Define behavior explicitly for even sample counts.

  • Distributed means merge efficiently using count and sum, or count, mean, and variance state. Exact distributed medians require retaining or merging ordered summaries; approximate quantile sketches such as t-digest or KLL trade accuracy and memory for scalability.

  • Do not silently discard or reinterpret outliers. First determine whether a large trade is a data error, a genuine tail event, or a risk signal; filtering it changes the metric’s meaning and should be visible in the definition.

  • Report complementary metrics when one statistic is insufficient: mean, median, sample count, standard deviation or MAD, selected percentiles, and total P&L. A median alone can hide profitable rare events; a mean alone can hide typical losses.

  • Handle implementation edge cases explicitly: empty input, one observation, negative profits, integer overflow, floating-point cancellation, NaN or infinity, duplicate values, and ties. For money, prefer fixed-point integer cents or a decimal type where the system’s precision requirements justify it.

Worked example: Choose Mean or Median for Trading Profit Metrics

For “Choose Mean or Median for Trading Profit Metrics,” I would first ask whether the metric is intended to represent typical trade experience, expected profit per trade, total economic contribution, or an alerting signal. I would also clarify the observation unit, time window, treatment of canceled or zero-profit trades, and whether extreme values are valid trades or data-quality errors. My answer would have four pillars: inspect the distribution, match the statistic to the business or operational question, define aggregation semantics, and specify a robust implementation. I would say median is usually better for describing a typical trade when P&L is highly skewed, while mean is necessary for expected value and reconciliation to total P&L because mean×n=total P&L\text{mean}\times n=\text{total P\&L}. The explicit tradeoff is that median is stable under rare tail events but can completely ignore their economic magnitude, whereas mean reflects those tails but can be dominated by one exceptional trade. For a production service, I would compute both, use Welford-style state for the mean, and use a batch sort, two heaps, or a quantile sketch for the median depending on scale and exactness requirements. I would close by saying that, with more time, I would validate the choice against historical distributions and add monitoring for sample count, tail percentiles, and changes in the fraction of missing or invalid observations.

A second angle

The same decision can arise when choosing a metric for a live trading-system dashboard rather than a one-time interview calculation. A dashboard showing median per-trade P&L may be stable and useful for detecting broad degradation, while mean P&L and total P&L are needed to detect a few large losses or gains. The framing shifts from “which single statistic is correct?” to “which small metric set gives operators enough information without encouraging a misleading interpretation.” An engineer should also consider windowing, update latency, reset behavior, and whether distributed workers can merge the metric without bias. In practice, exposing mean, median, count, and tail quantiles is often safer than forcing one statistic to serve every purpose.

Common pitfalls

Pitfall: “Always use the median because trading data has outliers” is too simplistic. Legitimate rare profits and losses are part of the economics, so the better answer distinguishes typical behavior from expected value and total P&L.

Choosing the right statistic but ignoring aggregation boundaries is an analytical mistake. Computing a mean over all fills may overweight high-volume strategies, while averaging daily means gives each day equal weight; state which population and weighting scheme the metric represents.

Pitfall: Saying “I would calculate the median” without discussing how is shallow for a Software Engineer. Mention exact versus approximate computation, memory limits, streaming updates, distributed merging, and behavior for empty or even-sized inputs.

Connections

An interviewer may pivot to percentiles, quantile sketches, histograms, numerical stability, or distributed aggregation. They may also ask how to detect data-quality anomalies without masking genuine trading-tail behavior.

Practice questions

Related concepts