Design a Wiki Crawler That Keeps Every Stored Page at Most Five Days Stale
Company: Lyft
Role: Software Engineer
Category: System Design
Difficulty: medium
Interview Round: Onsite
Design a system that crawls all of the content of a wiki, stores it, and keeps the stored copy up to date within a bounded time: the stored version of any page must never be more than five days behind the live wiki.
```hint Freshness is a schedule
Turn the five-day bound into a rule for how often every page must be revisited, and from there into a minimum crawl rate.
```
```hint Not every page changes
Think about which signals tell you a page has changed, so the crawl budget goes where it is needed while every page still meets the bound.
```
### Constraints and Clarifications
- The five-day bound applies to every stored page, including pages that rarely change.
- The size of the wiki is not given. Ask for it, or state assumptions and size the design from them.
### Clarifying Questions
- Is this one wiki site or many wikis, and roughly how many pages are there and how large is a page?
- Does the bound mean every page is rechecked at least every five days, or that any edit appears in storage within five days of being made?
- Does the wiki offer an API, a recent-changes feed or periodic dumps, or only HTML pages?
- Should storage keep every version of a page, or only the latest?
- Should deleted or renamed pages be removed from storage, or kept with a marker?
- What rate limits or crawling rules does the wiki impose?
### What a Strong Answer Covers
- A capacity estimate that turns the five-day bound and the page count into a required fetch rate and storage size
- A crawl architecture: discovery, frontier and scheduler, fetchers, parsing and deduplication
- A scheduling policy that guarantees the bound for every page, including the order in which a backlog is served, while spending more effort on frequently changing pages
- Storage for page content, metadata and versions
- Politeness, failure handling, and monitoring of how stale the stored copy is
### Follow-up Questions
- The fetchers are down for two days. Which pages break the five-day bound, and how should the crawler catch up after they recover?
- How would you detect that a page has not meaningfully changed, so you do not store a new version?
- If the bound were tightened from five days to one hour, what would change?
- How do you handle pages that render differently on every fetch, for example because they show the current time?
Overview: A system design question about crawling and storing all the content of a wiki while keeping every stored page no more than five days stale. It tests capacity estimates, the crawl frontier and scheduling, change detection, storage of page versions, politeness limits, failure recovery, and monitoring of freshness.
Read the full Lyft Software Engineer interview experience this question came from