Compute Top-K word frequencies under a path
Company: Box
Role: Software Engineer
Category: Coding & Algorithms
Difficulty: medium
Interview Round: Onsite
Given a filesystem path that may contain nested subdirectories and files, compute the top K most frequent words across all files. Describe an in-memory solution and its complexity. Follow-up: when the corpus is too large to fit in memory, propose scalable approaches (e.g., external sorting/partitioning, MapReduce-style sharding and merge) and an approximate heavy-hitters approach (e.g., Count–Min Sketch with Space-Saving), including accuracy/latency/storage trade-offs.
Quick Answer: Compute Top-K word frequencies under a path evaluates algorithm design, data structures, correctness, complexity, edge cases, and implementation details in a realistic interview setting. A strong answer states assumptions, handles edge cases, explains trade-offs, and shows how to validate the result clearly.
Solution
# Solution Alignment
The prompt asks for an implementation-level answer. The safest way to present it is to define the state, maintain clear invariants, then walk through complexity and tests.
## Problem Restatement
Given a filesystem path that may contain nested subdirectories and files, compute the top K most frequent words across all files. Describe an in-memory solution and its complexity. Follow-up: when the corpus is too large to fit in memory, propose scalable approaches (e.g., external sorting/partitioning, MapReduce-style sharding and merge) and an approximate heavy-hitters approach (e.g., Count–Min Sketch with Space-Saving), including accuracy/latency/storage trade-offs.
## Recommended Approach
For one-time top-K, use a size-K min-heap or quickselect plus sorting the selected K. For streaming windows, maintain counts in a hash map plus a heap with lazy deletion or bucketed frequency structures when updates must be near O(1). Define deterministic tie-breaking.
## Correctness
The implementation should maintain an invariant after each loop or operation that directly matches the problem statement. At termination, that invariant implies the returned value has considered every valid candidate exactly once, or has preserved the required data-structure state after every API call.
## Complexity
One-time heap: O(n log k) time and O(k) space. Quickselect: expected O(n) plus O(k log k) to order output. Streaming complexity depends on window eviction and tie-breaking.
## Edge Cases and Tests
k = 0, k > n, duplicate values, ties, negative values, stale heap entries, and deterministic output ordering.