Implement a high-throughput web crawler safely
Company: Amazon
Role: Data Scientist
Category: Coding & Algorithms
Difficulty: medium
Interview Round: Technical Screen
Overview: This question evaluates competency in concurrent systems, scalable data ingestion, and algorithmic design for a high-throughput web crawler, covering URL normalization, deduplication, per-host rate limiting, fault-tolerant checkpointing, and bounded-memory queueing.
Read the full Amazon Data Scientist interview experience this question came from