Optimizing BERT inference latency and throughput for a production NLP service

Read the full interview experience this question came from →

Quick Overview

An ML engineering question on speeding up a fine-tuned BERT encoder that has become too slow or too costly to serve in production. It tests profiling, model compression, runtime and serving optimizations such as batching and padding control, and how to verify that accuracy does not regress.

Optimizing BERT inference latency and throughput for a production NLP service

Company: eBay

Role: Applied Scientist

Category: Machine Learning

Difficulty: easy

Interview Round: Technical Screen

A fine-tuned BERT-style encoder serves predictions in production (for example, classifying or tagging short text). It is now too slow or too expensive to serve. Explain how you would optimize its inference, covering changes to the model, to how it is executed, and to how requests reach it, and how you would confirm that quality has not regressed. ```hint Profile first Find where the time actually goes: tokenization, the encoder layers, padding, or the serialization and network around the model. ``` ```hint Where compute grows Think about which property of the input makes the encoder more expensive, and how your traffic is distributed along it. ``` ### Constraints and Clarifications - No latency, throughput, hardware, or accuracy numbers were given; ask for them before choosing techniques. - Assume the model was fine-tuned from a standard pretrained BERT checkpoint and that you can retrain or distill it if needed. ### Clarifying Questions - Is the target p99 latency per request, throughput per unit of cost, or both? - Is serving on CPU or GPU, and can that change? - Is the task sequence-level (one label per text) or token-level (one label per token)? - What is the distribution of input lengths, and what maximum sequence length does the model use today? - How much accuracy loss, if any, is acceptable? - Is all traffic online and latency-sensitive, or can some of it be processed in offline batches? ### What a Strong Answer Covers - Profiling and a clearly stated target before any optimization - Model-level compression (distillation, quantization, pruning, shorter maximum length) and its accuracy trade-offs - Execution-level optimization matched to the hardware (exported and fused graphs, reduced precision, optimized runtimes) - Serving-level techniques: dynamic batching, length bucketing to cut padding, caching, concurrency and autoscaling - Validation: accuracy parity on held-out data, benchmarks on the real length distribution, and a guarded online rollout ### Follow-up Questions - After INT8 quantization, recall on one minority class drops noticeably. What do you try next? - Dynamic batching raised throughput but made p99 latency worse. How do you tune it? - A long tail of inputs exceeds the maximum sequence length. How do you handle them without slowing every request? - How would you decide between distilling to a smaller model on CPU and keeping the full model on GPUs?

Overview: An ML engineering question on speeding up a fine-tuned BERT encoder that has become too slow or too costly to serve in production. It tests profiling, model compression, runtime and serving optimizations such as batching and padding control, and how to verify that accuracy does not regress.

Read the full eBay Applied Scientist interview experience this question came from

|Home/Machine Learning/eBay
eBay logo
eBay
Sep 24, 2026
easyApplied ScientistTechnical ScreenMachine Learning
0
0

A fine-tuned BERT-style encoder serves predictions in production (for example, classifying or tagging short text). It is now too slow or too expensive to serve. Explain how you would optimize its inference, covering changes to the model, to how it is executed, and to how requests reach it, and how you would confirm that quality has not regressed.

Constraints and Clarifications

  • No latency, throughput, hardware, or accuracy numbers were given; ask for them before choosing techniques.
  • Assume the model was fine-tuned from a standard pretrained BERT checkpoint and that you can retrain or distill it if needed.

Clarifying Questions Guidance

  • Is the target p99 latency per request, throughput per unit of cost, or both?
  • Is serving on CPU or GPU, and can that change?
  • Is the task sequence-level (one label per text) or token-level (one label per token)?
  • What is the distribution of input lengths, and what maximum sequence length does the model use today?
  • How much accuracy loss, if any, is acceptable?
  • Is all traffic online and latency-sensitive, or can some of it be processed in offline batches?

What a Strong Answer Covers Guidance

  • Profiling and a clearly stated target before any optimization
  • Model-level compression (distillation, quantization, pruning, shorter maximum length) and its accuracy trade-offs
  • Execution-level optimization matched to the hardware (exported and fused graphs, reduced precision, optimized runtimes)
  • Serving-level techniques: dynamic batching, length bucketing to cut padding, caching, concurrency and autoscaling
  • Validation: accuracy parity on held-out data, benchmarks on the real length distribution, and a guarded online rollout

Follow-up Questions Guidance

  • After INT8 quantization, recall on one minority class drops noticeably. What do you try next?
  • Dynamic batching raised throughput but made p99 latency worse. How do you tune it?
  • A long tail of inputs exceeds the maximum sequence length. How do you handle them without slowing every request?
  • How would you decide between distilling to a smaller model on CPU and keeping the full model on GPUs?
Loading comments...