Design an in-memory database
System Design: In-Memory Database Engine (Single-Node Core, Scale-Out Aware)
Context
Design the core of an in-memory, single-node database engine intended to be the building block for a larger distributed system. Data is primarily memory-resident on each node; durability and scale-out are achieved via logging, snapshots, and replication/sharding across multiple such nodes.
Requirements
-
CRUD operations on records.
-
Efficient querying: point lookups and range scans; basic filtering and ordering.
-
Transaction isolation: define levels supported (at least read committed and snapshot isolation; discuss path to serializable).
-
Persistence and backup: crash recovery, snapshots/checkpoints, WAL, and restore.
-
Horizontal scalability: strategy for replication and sharding across multiple single-node instances.
Deliverable
Provide a design that covers:
-
Data model and API surface.
-
In-memory storage layout and indexing.
-
Query execution approach.
-
Concurrency control and transaction management.
-
Durability: WAL, checkpoints/snapshots, recovery.
-
Backup/restore strategy.
-
Horizontal scaling: replication and sharding.
-
Operational concerns: garbage collection, observability, and pitfalls.
Constraints & Assumptions
-
Preserve the scope, facts, inputs, and requested outputs from the prompt above.
-
If the prompt leaves a detail unspecified, state a reasonable assumption before relying on it.
-
Keep the answer interview-ready: concise enough to present, but concrete enough to implement or evaluate.
Clarifying Questions to Ask
-
Clarify users, core use cases, read/write patterns, scale, latency, availability, and data retention.
-
State explicit assumptions before making sizing or architecture decisions.
-
Prioritize the functional path first, then address reliability, security, observability, and rollout.
What a Strong Answer Covers
-
A scoped requirements summary with concrete non-goals and success metrics.
-
API, data model, architecture, consistency, capacity, and operations.
-
Reasoned trade-offs among simple and scalable designs, including bottlenecks and failure modes.
-
A validation, monitoring, migration, and launch plan appropriate for the risk level.
Follow-up Questions
-
What breaks first at 10x traffic or data volume?
-
How would you degrade gracefully during dependency failures?
-
What metrics and alerts would prove the design is healthy after launch?