Compute Top-K word frequencies under a path

Read the full interview experience this question came from →

Quick Overview

Compute Top-K word frequencies under a path evaluates algorithm design, data structures, correctness, complexity, edge cases, and implementation details in a realistic interview setting. A strong answer states assumptions, handles edge cases, explains trade-offs, and shows how to validate the result clearly.

Compute Top-K word frequencies under a path

Company: Box

Role: Software Engineer

Category: Coding & Algorithms

Difficulty: medium

Interview Round: Onsite

Given a filesystem path that may contain nested subdirectories and files, compute the top K most frequent words across all files. Describe an in-memory solution and its complexity. Follow-up: when the corpus is too large to fit in memory, propose scalable approaches (e.g., external sorting/partitioning, MapReduce-style sharding and merge) and an approximate heavy-hitters approach (e.g., Count–Min Sketch with Space-Saving), including accuracy/latency/storage trade-offs.

Overview: Compute Top-K word frequencies under a path evaluates algorithm design, data structures, correctness, complexity, edge cases, and implementation details in a realistic interview setting. A strong answer states assumptions, handles edge cases, explains trade-offs, and shows how to validate the result clearly.

Read the full Box Software Engineer interview experience this question came from

Solution

# Solution Alignment The prompt asks for an implementation-level answer. The safest way to present it is to define the state, maintain clear invariants, then walk through complexity and tests. ## Problem Restatement Given a filesystem path that may contain nested subdirectories and files, compute the top K most frequent words across all files. Describe an in-memory solution and its complexity. Follow-up: when the corpus is too large to fit in memory, propose scalable approaches (e.g., external sorting/partitioning, MapReduce-style sharding and merge) and an approximate heavy-hitters approach (e.g., Count–Min Sketch with Space-Saving), including accuracy/latency/storage trade-offs. ## Recommended Approach For one-time top-K, use a size-K min-heap or quickselect plus sorting the selected K. For streaming windows, maintain counts in a hash map plus a heap with lazy deletion or bucketed frequency structures when updates must be near O(1). Define deterministic tie-breaking. ## Correctness The implementation should maintain an invariant after each loop or operation that directly matches the problem statement. At termination, that invariant implies the returned value has considered every valid candidate exactly once, or has preserved the required data-structure state after every API call. ## Complexity One-time heap: O(n log k) time and O(k) space. Quickselect: expected O(n) plus O(k log k) to order output. Streaming complexity depends on window eviction and tie-breaking. ## Edge Cases and Tests k = 0, k > n, duplicate values, ties, negative values, stale heap entries, and deterministic output ordering.
|Home/Coding & Algorithms/Box
Box logo
Box
Aug 1, 2025
mediumSoftware EngineerOnsiteCoding & Algorithms
28
0

Compute Top-K word frequencies under a path

Given a filesystem path that may contain nested subdirectories and files, compute the top K most frequent words across all files. Describe an in-memory solution and its complexity. Follow-up: when the corpus is too large to fit in memory, propose scalable approaches (e.g., external sorting/partitioning, MapReduce-style sharding and merge) and an approximate heavy-hitters approach (e.g., Count–Min Sketch with Space-Saving), including accuracy/latency/storage trade-offs.

Clarifying Questions to Ask Guidance

  • Clarify input sizes, value ranges, mutability, return format, and tie-breaking.
  • State the target time and space complexity before coding.
  • Call out edge cases such as empty inputs, duplicates, invalid values, overflow, and boundary sizes.

What a Strong Answer Covers Guidance

  • A clear algorithm with the right data structures and enough pseudocode or code-level detail to implement it.
  • A correctness argument that explains why the algorithm covers all required cases.
  • Time and space complexity, plus at least one alternative approach when relevant.
  • Focused tests for normal cases, edge cases, and failure modes.

Follow-up Questions Guidance

  • How would the approach change if the input were streaming or too large for memory?
  • What invariants would you assert in production code?
  • Which tests would catch off-by-one, duplicate, or tie-breaking bugs?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...