Design multi-GPU matrix multiplication
Company: Google
Role: Machine Learning Engineer
Category: ML System Design
Difficulty: hard
Interview Round: Technical Screen
Design and implement computing C = A × B across two GPUs when A and B must reside on both devices. Specify data partitioning (row/column/block tiling), communication primitives (e.g., all-reduce, all-gather, point-to-point), compute scheduling (tiled GEMM with overlap of compute and communication), memory layout and buffer reuse, numerical precision, synchronization, how you aggregate and return C, and discuss scalability and failure handling.
Overview: This question evaluates proficiency in multi-GPU parallelism and system-level ML engineering, covering data partitioning, inter-GPU communication primitives, compute scheduling and overlap, memory layout and buffer reuse, numerical precision trade-offs, synchronization, scalability, and failure handling.
Community answers
Answer by ankita10yadav10
Data partitioning: 1D row-wise split of C
Input Layout: A and B fully replicated on both GPUs
Compute: cuBLAS/cuBLASLt GEMM on each GPU
Communication: None during GEMM; cudaMemcpyPeerAsync or NCCL AllGather after computation if a full C is required
Overlap: Separate compute and communication streams with CUDA events
Precision: FP16/BF16 inputs, FP32 accumulation using Tensor Cores
Synchronization | CUDA streams, events, and final device/NCCL synchronization
Memory: Reuse cuBLAS workspaces, contiguous buffers, pointer offsets for submatrices
Scalability: Use 2D block partitioning (SUMMA) and hierarchical NCCL collectives for many GPUs
Failure Handling: Detect CUDA/NCCL errors, retry communications, recreate communicators, and degrade to fewer GPUs if needed