System Architecture Design: Core Principles

Quick Overview
Master system architecture design for technical interviews. Explore core principles, reliability patterns, and the trade-offs interviewers evaluate.
Good system architecture starts with a clear contract: what the system must do, which failures it must tolerate, and what trade-offs the team can operate. A larger diagram does not automatically represent a better design.
In an interview, make your decisions visible. Explain why a request is synchronous, who owns each piece of data, what a retry can do, and which evidence would justify more infrastructure. This guide uses an illustrative order-processing service to connect those decisions.
Start with the user journey and constraints
Before drawing services, describe one successful journey and one failure. For an order service, a customer submits an order, receives a stable identifier, and can check its status. If a payment provider times out, the customer needs an honest pending state and a safe way to recover.
Separate functional requirements from operating targets:
| Requirement | Question to resolve | Design consequence |
|---|---|---|
| Correctness | Can the same order be charged twice? | Deduplication and payment reconciliation |
| Latency | Must payment finish before the response? | Synchronous confirmation or an explicit pending state |
| Availability | What can users do during provider downtime? | Durable acceptance, status reads, and recovery |
| Scale | Which reads or writes dominate? | Capacity estimates and indexes for the hot path |
| Recovery | How much data loss and downtime are acceptable? | Backup, replication, and restore strategy |
Treat any numbers introduced in an interview as assumptions to validate, not facts about the company. A modest workload and a global, highly available payment system need different designs.
If you need a preparation sequence, use PracHub's complete system design interview guide before practicing a full timed answer.
Choose boundaries you can explain
Begin with a modular application unless the requirements justify separate services. An order module, a payment integration, and a notification worker can have clear responsibilities without independent network deployments.
A service boundary earns its cost when it supports a concrete need: independent scaling, separate ownership, a different release cadence, or useful failure containment. It also introduces contracts, network errors, deployment coordination, and monitoring obligations.
Ask what would force the boundary to change. If notifications become expensive but orders remain small, a worker pool may solve the problem without splitting the entire application into microservices. If two teams need independent control of stable domains, separate deployment can be valuable.
Interview prompt: Which module would you extract first, and what measurement would justify doing it? “Microservices scale better” is incomplete unless you identify the constrained workload and the coordination cost.
Design a durable write path
For the example service, a simple first design is an API, a relational database, and a worker. The API validates the request and records the order. The worker coordinates payment and later side effects.
One possible flow is:
- The client submits an order with an idempotency key.
- The API checks that the key belongs to the same caller and request payload.
- A database transaction writes the order and a durable work record.
- The API returns the order identifier and its current status.
- A worker processes the durable record and persists the outcome.
Writing the order and its work record together avoids a gap where the database commit succeeds but an in-memory task disappears. An outbox is one implementation of this idea; it still needs a dispatcher, retry behavior, and monitoring.
Do not claim that a database commit proves the payment happened. Model states such as pending, paid, and failed, and decide how an ambiguous provider timeout is reconciled. Consider access control on status reads as part of the design.
Work through payment processing with idempotency and reconciliation to practice explaining the uncertain-outcome case.
Make retries safe before making them automatic
A timeout tells the caller it did not receive a response in time. It does not prove that the receiver did nothing. Retrying a payment, reservation, or resource creation can therefore repeat an effect.
An idempotency design needs more than a header name. Explain:
- The scope of the key: caller, operation, and relevant business object.
- How the server detects a reused key with a different payload.
- How concurrent requests with the same key are serialized or rejected.
- Which result is stored and how a repeated request retrieves it.
- How retention and recovery work after the deduplication record expires.
Use database uniqueness or another atomic mechanism for the claim step. A separate “check, then insert” sequence can race. The consumer of a redelivered message also needs its own duplicate-handling strategy.
PracHub's article “Just make it idempotent” is half an answer explores the reasoning behind these follow-ups. Read it, then state the invariants for your own design before choosing a queue.
Contain dependency failures
Set an overall deadline for the user request and budget downstream calls within it. Retry only failures that are plausibly transient, and only when replay is safe. A bounded attempt count, backoff, and jitter help avoid adding more load to an impaired dependency.
A circuit breaker tracks failures and can temporarily stop calls to a struggling dependency. It serves a different purpose from a retry loop. It needs a recovery probe and a defined caller response when open; returning an empty success is not a meaningful fallback. Microsoft's circuit breaker pattern describes the states and operational considerations.
Keep resource limits explicit. If payment is slow, it should not consume every API worker or connection. Separate pools, bounded concurrency, queue limits, and backpressure are useful only when you explain what users experience at those limits.
Practice that reasoning with a rate limiter using per-user token quotas. Distinguish a rate limit from a concurrency limit, and consider whether the decision must be coordinated across instances.
Scale the part that is actually constrained
Trace the hot path before introducing a cache or a shard. For an order-status endpoint, an index on the authorized lookup key and a small response may be enough. For a notification backlog, adding API replicas does not solve the worker bottleneck.
| Change | Useful when | New responsibility |
|---|---|---|
| Read cache | Repeated reads tolerate bounded staleness | Invalidation, authorization-safe keys, and miss behavior |
| Additional workers | Independent jobs are waiting | Duplicate handling, ordering, and downstream limits |
| Read replicas | Read traffic dominates | Replica lag and read-after-write expectations |
| Partitioning | A single storage unit is constrained | Key choice, hot partitions, and cross-partition work |
Never cache a personalized response under a shared public key. Decide which fields may be stale and which require an authoritative read. In the order example, a cached product description and a payment status have different correctness requirements.
The next change should respond to evidence: queue age, connection wait time, p95 latency, lock contention, or a concentrated partition. Name the metric and the threshold you would investigate instead of promising unlimited scale.
Explain recovery as a procedure
Availability and recovery are related but different. Redundant application instances help with some failures; they do not restore accidentally deleted data. A replica can faithfully reproduce an unwanted write.
Define the recovery time objective, the tolerated interruption, and the recovery point objective, the tolerated data-loss window. Choose backup and replication arrangements that can meet those objectives, then describe how a restore is tested.
For the example order service, recovery also involves reconciling the database with the payment provider. A restored order record may be older than an external charge. An operational runbook must identify pending work, avoid repeating charges, and preserve an audit trail.
A useful interview follow-up is: “What happens if the database is restored to yesterday but the provider still has today's payments?” A concrete reconciliation plan demonstrates more understanding than adding a second region to the diagram.
Use a 45-minute interview deliberately
The following is a practice allocation, not a universal interviewer rubric:
| Time | Focus | Deliverable |
|---|---|---|
| First 5 minutes | Scope and constraints | One user journey, assumptions, and priorities |
| Next 10 minutes | API, data, and boundaries | A small design with clear ownership |
| Next 15 minutes | Main path and failure cases | Write flow, retries, and consistency decisions |
| Next 10 minutes | Scale and recovery | A bottleneck, mitigation, and operational cost |
| Final 5 minutes | Review the trade-off | Limits and the evidence that would change the design |
Start with one of PracHub's system design questions. Explain the design aloud, ask what fails next, and revise one decision. For a structured foundation, follow the Foundations of System Design learning track.
A strong closing answer is specific: “This design meets the stated workload and recovery assumptions. I would add separate worker capacity when queue age grows, and revisit the storage boundary when measured write contention becomes the constraint.” The value is in the reasoning connecting the requirement to the next change.
Comments (0)