Design a Supercharger Backend: Billing, Stall Telemetry and Nearest Available Charger
Company: Tesla
Role: Software Engineer
Category: System Design
Difficulty: hard
Interview Round: Technical Screen
Design the backend for a network of electric-vehicle fast-charging stations (Superchargers). Each site has several charging stalls. The design must meet three functional requirements:
- **FR1 — Billing:** measure the amount of energy delivered in each charging session and store it in the backend so the customer can be billed for it.
- **FR2 — Stall health telemetry:** collect health metrics and telemetry from every charging stall and keep them in a long-term data store.
- **FR3 — Nearest available charger:** on the car's screen or in the in-car application, the driver touches the map and sees the nearest available Superchargers together with their price.
Expect the interviewer to go deep on each requirement rather than accept one quick high-level diagram per requirement. For billing alone, they may want to discuss routing, secure communication, DNS, local buffers, and synchronous versus asynchronous calls. Manage the time so that all three requirements get designed.
### Clarifying Questions
- How many sites and stalls are in scope, in which regions, and how many sessions run at the same time? The prompt gives no scale.
- How is a session tied to a customer: does the vehicle identify itself when plugged in, or does the driver start the session from an app? Is a payment method always on file?
- How is price defined: per kWh, per minute, varying by site or by time of day? Who changes it, and how often?
- How reliable is connectivity at a site? Must a stall keep charging while the backend is unreachable?
- Which clients consume the data: billing and payments, field service, engineering analysis, the in-car application?
### Part 1 — Billing for energy delivered (FR1)
Design how a stall measures the energy delivered during a session, and how that measurement reaches the backend and becomes a billing record. Be ready to explain in detail:
- how a stall's requests are routed to the right backend service, including the role DNS plays;
- how communication between the stall and the backend is secured;
- what the stall or site buffers locally, and what happens when the backend is unreachable;
- which interactions are synchronous and which are asynchronous, and why.
```hint Assume the link drops mid-session
Decide what the stall must record locally, and in what form, so that a session that loses connectivity halfway through can still be billed correctly and exactly once.
```
#### Clarifying Questions for this Part
- Must the backend authorize a session before charging starts, or may a stall start charging while offline?
- Is the stall's own meter the source of truth for billed energy?
- Is the customer charged per session, or are sessions aggregated into periodic invoices?
#### What This Part Should Cover
- The session lifecycle and a data model for sessions, meter readings and billing records.
- The stall-to-backend path: device identity, transport security, DNS and routing to a backend region.
- Local buffering, idempotent upload and exactly-once billing across retries and outages.
- Where the synchronous and asynchronous boundaries sit, and why.
### Part 2 — Stall health metrics and telemetry (FR2)
Design how health metrics and telemetry get from every stall into long-term storage, and how they are read back.
```hint Let the reads choose the storage
Telemetry is written continuously and read mostly as a time range for one stall or as aggregates across many stalls. Let those access patterns, and how long data must be kept, decide the storage layout.
```
#### Clarifying Questions for this Part
- Which metrics are collected and how often? Are there discrete events, such as fault codes, as well as periodic samples?
- How long must raw data be kept, and is downsampled data acceptable for older periods?
- Who reads the data: real-time alerting, dashboards, or offline engineering analysis?
#### What This Part Should Cover
- An ingestion path that is separate from billing and absorbs bursts, for example when many sites reconnect at once.
- Storage tiers, retention and downsampling for long-term data.
- The main query patterns and how the layout serves each one, including alerting on faults.
### Part 3 — Nearest available Supercharger with price (FR3)
When the driver touches a point on the map in the car, return the nearest Supercharger sites that have an available stall, together with the current price at each.
```hint Two very different freshness needs
Where a site is almost never changes; whether its stalls are free changes constantly. Consider storing, indexing and caching those two kinds of data separately.
```
#### Clarifying Questions for this Part
- Is "nearest" measured in straight-line distance or in driving distance or time?
- How many results should be returned, and within what radius?
- How fresh must availability be, and what should the car show when data is stale?
#### What This Part Should Cover
- A geospatial lookup for the nearest sites.
- A pipeline that turns stall status into per-site availability, and how it is cached.
- Price lookup, and consistency between the price shown on the map and the price billed.
- The API, latency budget and behavior on a poor in-car connection.
### What a Strong Answer Covers
- A clear split of the three flows by their requirements: exactly-once correctness for billing, high-volume ingestion and retention for telemetry, low-latency reads with bounded staleness for search.
- Depth on each component when probed, not only boxes on a diagram: identity and transport security, DNS and routing, buffering, idempotency, storage choices.
- An explicit design for offline sites and partial failures in every flow.
- Observability and reconciliation that would reveal lost, duplicated or wrong billing data.
- Time management that reaches all three requirements.
### Follow-up Questions
- A site loses connectivity for several hours while cars keep charging. Walk through what happens to billing, telemetry and the availability shown on the map, before and after it reconnects.
- A firmware bug inflated energy readings on some stalls for a week. How would you detect it and correct the affected bills?
- How would you show predicted availability at the driver's arrival time instead of availability right now?
- How would you rotate stall credentials or roll out new firmware across the fleet without interrupting charging?
Overview: System design question on the backend for an electric-vehicle fast-charging network: billing each session by the energy delivered, storing stall health telemetry long term, and showing drivers the nearest available charger with its price. Probes secure device communication, DNS and routing, local buffering and sync versus async choices.
Read the full Tesla Software Engineer interview experience this question came from