Design an LLM Chatbot: Training Data, Reward Design and Deployment Considerations

Quick Overview

Design a chatbot built on a large language model, explaining how to build supervised, preference and safety training data and how to design and combine rewards for reinforcement learning without reward hacking. Also covers deployment concerns such as serving latency and cost, safety filtering, monitoring and staged rollout.

Design an LLM Chatbot: Training Data, Reward Design and Deployment Considerations

Company: Amazon

Role: Machine Learning Engineer

Category: ML System Design

Difficulty: medium

Interview Round: Onsite

Design a chatbot built on a large language model. Walk through the whole lifecycle, and explain three areas in depth: how you would build the training data, how you would design the rewards used to align the model, and what you would need to consider to deploy it. The chatbot's users and domain are not specified. Establish them first, because they shape every later decision. ### Clarifying Questions - Is this a general-purpose assistant or a chatbot for a specific domain or task, and who are its users? - Are we starting from a pretrained base model, from an instruction-tuned model, or from scratch? - Which languages and conversation lengths must it support, and does it need tools or access to documents? - What are the latency, cost and safety requirements at launch? ### Part 1 — Training data Explain how you would create the data that turns a pretrained model into a chatbot: which datasets you need, where each one comes from, and how you control its quality. ```hint Match data to training stages Each stage of alignment consumes a different kind of example. List the stages first, and the datasets follow from them. ``` #### What This Part Should Cover - The datasets needed for each training stage and how each is sourced: human-written, model-generated, or logged from users. - Quality control: annotator guidelines, agreement checks, filtering, deduplication, and decontamination against evaluation sets. - Coverage of multi-turn behavior, refusals and safety cases, and the privacy handling of any user data. ### Part 2 — Reward design Explain how the model is rewarded during reinforcement learning: where the reward signals come from, how they are combined, and how you stop the policy from exploiting them. ```hint Separate checkable from judged Some qualities of a reply can be verified by a program, while others need a learned judge. The two fail in different ways. ``` #### What This Part Should Cover - The sources of reward (a learned preference model, programmatic checks, safety and format signals) and how they are combined. - How the reward model is trained and validated. - Reward hacking: how it shows up and the controls against it. ### Part 3 — Deployment considerations Explain what must be decided and built to serve the chatbot to real users. ```hint Follow one request Trace a single user message from arrival to streamed reply, and list what can go wrong or cost money at each step. ``` #### What This Part Should Cover - Serving efficiency: latency targets, batching, caching, and model-size or quantization choices. - Safety and reliability in production: input and output filtering, fallbacks, abuse prevention and rate limits. - Pre-launch evaluation, staged rollout, monitoring, and the feedback loop back into training data. ### What a Strong Answer Covers - Requirements established before design, with later choices tied back to them. - A coherent pipeline in which data, reward and deployment decisions support one another. - Explicit trade-offs between helpfulness and safety, and between quality and serving cost. - Evaluation at every stage: offline benchmarks, human preference tests and online metrics. ### Follow-up Questions - The reward model's scores keep rising during training, but human raters prefer the new checkpoint less often. What is happening, and what do you do? - How would you use production conversations to improve the next version without violating user privacy? - How would you add retrieval over a document store to reduce made-up answers, and how would that change the training data? - How would you decide whether a smaller distilled model is good enough to replace the large one in production?

Overview: Design a chatbot built on a large language model, explaining how to build supervised, preference and safety training data and how to design and combine rewards for reinforcement learning without reward hacking. Also covers deployment concerns such as serving latency and cost, safety filtering, monitoring and staged rollout.

|Home/ML System Design/Amazon
Amazon logo
Amazon
Jan 3, 2026
mediumMachine Learning EngineerOnsiteML System Design
0
0

Design a chatbot built on a large language model. Walk through the whole lifecycle, and explain three areas in depth: how you would build the training data, how you would design the rewards used to align the model, and what you would need to consider to deploy it.

The chatbot's users and domain are not specified. Establish them first, because they shape every later decision.

Clarifying Questions Guidance

  • Is this a general-purpose assistant or a chatbot for a specific domain or task, and who are its users?
  • Are we starting from a pretrained base model, from an instruction-tuned model, or from scratch?
  • Which languages and conversation lengths must it support, and does it need tools or access to documents?
  • What are the latency, cost and safety requirements at launch?

Part 1 — Training data

Explain how you would create the data that turns a pretrained model into a chatbot: which datasets you need, where each one comes from, and how you control its quality.

What This Part Should Cover Guidance

  • The datasets needed for each training stage and how each is sourced: human-written, model-generated, or logged from users.
  • Quality control: annotator guidelines, agreement checks, filtering, deduplication, and decontamination against evaluation sets.
  • Coverage of multi-turn behavior, refusals and safety cases, and the privacy handling of any user data.

Part 2 — Reward design

Explain how the model is rewarded during reinforcement learning: where the reward signals come from, how they are combined, and how you stop the policy from exploiting them.

What This Part Should Cover Guidance

  • The sources of reward (a learned preference model, programmatic checks, safety and format signals) and how they are combined.
  • How the reward model is trained and validated.
  • Reward hacking: how it shows up and the controls against it.

Part 3 — Deployment considerations

Explain what must be decided and built to serve the chatbot to real users.

What This Part Should Cover Guidance

  • Serving efficiency: latency targets, batching, caching, and model-size or quantization choices.
  • Safety and reliability in production: input and output filtering, fallbacks, abuse prevention and rate limits.
  • Pre-launch evaluation, staged rollout, monitoring, and the feedback loop back into training data.

What a Strong Answer Covers Guidance

  • Requirements established before design, with later choices tied back to them.
  • A coherent pipeline in which data, reward and deployment decisions support one another.
  • Explicit trade-offs between helpfulness and safety, and between quality and serving cost.
  • Evaluation at every stage: offline benchmarks, human preference tests and online metrics.

Follow-up Questions Guidance

  • The reward model's scores keep rising during training, but human raters prefer the new checkpoint less often. What is happening, and what do you do?
  • How would you use production conversations to improve the next version without violating user privacy?
  • How would you add retrieval over a document store to reduce made-up answers, and how would that change the training data?
  • How would you decide whether a smaller distilled model is good enough to replace the large one in production?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...