Design a Real-Time Question-Answering Chatbot: Training Data, RL Tasks and Rewards
Company: Meta
Role: Machine Learning Engineer
Category: ML System Design
Difficulty: medium
Interview Round: Technical Screen
Design a chatbot that answers users' questions in real time. Beyond the serving system, the discussion must cover how you would collect training data, which reinforcement-learning tasks you would set up to improve the model, and the weaknesses of your reward design and how you would address them.
### Clarifying Questions
- Does "real time" mean low-latency streamed responses, answers that reflect up-to-date information, or both?
- What kinds of questions dominate: open-domain factual questions, questions about a particular product or knowledge base, or open-ended conversation?
- What latency target and traffic scale should the design meet?
- Is a pretrained large language model available as a starting point?
- Which failure matters most: wrong answers, unsafe answers or slow answers?
### Part 1 — System design for real-time answers
Describe the end-to-end system that takes a user's question and returns an answer: its components, the request path, how fresh information reaches the answer when it is needed, and how latency is controlled.
```hint Budget the latency
List every step between the question arriving and the first token reaching the user, and decide which steps can run in parallel, be cached, or be skipped for easy questions.
```
#### What This Part Should Cover
- The request path from question to streamed answer, including any retrieval or tool use.
- A latency budget and the techniques that meet it.
- Safety checks, fallbacks and online metrics.
### Part 2 — Training data collection
Explain how you would collect and curate training data for the model behind the chatbot.
```hint Match data to each training stage
Supervised fine-tuning, preference learning and evaluation each need data of a different shape. Ask where each one comes from and how you would check its quality.
```
#### What This Part Should Cover
- Sources: human-written demonstrations, product logs and feedback, and synthetic data, with the privacy constraints on logs.
- Quality control, deduplication, and decontamination against evaluation sets.
- Coverage of hard cases, such as questions that need fresh information and questions that should be declined.
### Part 3 — RL tasks and reward design
Define the reinforcement-learning tasks you would train on and the reward for each. Then discuss the weaknesses of your reward design and how you would address them.
```hint Name the exploit
For each reward, describe the behavior that would score highly without producing a better answer, and what would catch it.
```
#### What This Part Should Cover
- Concrete RL tasks, such as correctness on questions with verifiable answers, helpfulness judged by a reward model, and deciding when to search or decline, each with its reward signal.
- Specific reward weaknesses: reward hacking, length and style bias, over-optimization, and sparse or noisy signals.
- Mitigations, and how you would verify that the reward still tracks answer quality.
### What a Strong Answer Covers
- A design that ties the meaning of "real time" to concrete choices in both serving and training.
- A data strategy that covers every training stage and respects user privacy.
- Rewards specific to the chosen tasks, with named failure modes and mitigations rather than generic warnings.
- Offline and online evaluation that catches regressions in correctness, freshness, latency and safety.
### Follow-up Questions
- How would you teach the model when to search for fresh information and when to answer from its own knowledge, and how would you reward that decision?
- If users rate answers with a thumbs up or down, how would you use that signal without letting it dominate the reward?
- How would you detect during training that the policy has started to exploit the reward model?
- How would you handle a question whose correct answer changed yesterday while the old answer is common in the training data?
Overview: An AI design question about a chatbot that answers user questions in real time. It tests the low-latency serving and retrieval path, training data collection for fine-tuning and preference learning, reinforcement-learning task and reward design, and the ability to name reward weaknesses and mitigate them.