Design an in-memory database

Quick Overview

This interview question evaluates requirements, scale assumptions, API/data design, architecture, trade-offs, failure modes, and rollout in a realistic interview setting. A strong answer for Design an in-memory database states assumptions, handles edge cases, explains trade-offs, and shows how to validate the result clearly.

Design an in-memory database

Company: OpenAI

Role: Software Engineer

Category: System Design

Difficulty: hard

Interview Round: Technical Screen

##### Question Design an in-memory database that supports CRUD operations, efficient querying, transaction isolation, persistence/backup strategy, and horizontal scalability within a single-node memory-resident architecture.

Overview: This interview question evaluates requirements, scale assumptions, API/data design, architecture, trade-offs, failure modes, and rollout in a realistic interview setting. A strong answer for Design an in-memory database states assumptions, handles edge cases, explains trade-offs, and shows how to validate the result clearly.

|Home/System Design/OpenAI
OpenAI logo
OpenAI
Aug 4, 2025
hardSoftware EngineerTechnical ScreenSystem Design
35
0

Design an in-memory database

System Design: In-Memory Database Engine (Single-Node Core, Scale-Out Aware)

Context

Design the core of an in-memory, single-node database engine intended to be the building block for a larger distributed system. Data is primarily memory-resident on each node; durability and scale-out are achieved via logging, snapshots, and replication/sharding across multiple such nodes.

Requirements

  • CRUD operations on records.
  • Efficient querying: point lookups and range scans; basic filtering and ordering.
  • Transaction isolation: define levels supported (at least read committed and snapshot isolation; discuss path to serializable).
  • Persistence and backup: crash recovery, snapshots/checkpoints, WAL, and restore.
  • Horizontal scalability: strategy for replication and sharding across multiple single-node instances.

Deliverable

Provide a design that covers:

  1. Data model and API surface.
  2. In-memory storage layout and indexing.
  3. Query execution approach.
  4. Concurrency control and transaction management.
  5. Durability: WAL, checkpoints/snapshots, recovery.
  6. Backup/restore strategy.
  7. Horizontal scaling: replication and sharding.
  8. Operational concerns: garbage collection, observability, and pitfalls.

Clarifying Questions to Ask Guidance

  • Clarify users, core use cases, read/write patterns, scale, latency, availability, and data retention.
  • State explicit assumptions before making sizing or architecture decisions.
  • Prioritize the functional path first, then address reliability, security, observability, and rollout.

What a Strong Answer Covers Guidance

  • A scoped requirements summary with concrete non-goals and success metrics.
  • API, data model, architecture, consistency, capacity, and operations.
  • Reasoned trade-offs among simple and scalable designs, including bottlenecks and failure modes.
  • A validation, monitoring, migration, and launch plan appropriate for the risk level.

Follow-up Questions Guidance

  • What breaks first at 10x traffic or data volume?
  • How would you degrade gracefully during dependency failures?
  • What metrics and alerts would prove the design is healthy after launch?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...