Design an LLM Inference Serving System
Company: Baseten
Role: Software Engineer
Category: ML System Design
Difficulty: hard
Interview Round: Technical Screen
Overview: Design production inference serving for one or more large language models from admission through streamed token generation. Explore model lifecycle, accelerator-aware batching and scheduling, memory pressure, fairness, cancellation, autoscaling, crash recovery, and latency-throughput observability.