Explain FlashAttention, KV cache, and RoPE
Company: TikTok
Role: Machine Learning Engineer
Category: Machine Learning
Difficulty: medium
Interview Round: Technical Screen
Quick Answer: This question evaluates understanding of transformer attention optimizations (FlashAttention), autoregressive decoding state management (KV cache), and positional encoding mechanisms (RoPE), focusing on competencies in memory and compute trade-offs, inference efficiency, and long-context behavior.