Debug Transformer and Add KV Cache
Company: OpenAI
Role: Machine Learning Engineer
Category: Machine Learning
Difficulty: medium
Interview Round: Onsite
Quick Answer: This question evaluates debugging and implementation skills for transformer-based autoregressive language models, focusing on attention mechanics, positional embeddings, causal masking, and integrating a key-value (KV) cache for incremental decoding.