Build a System to Manage Prompts and Model Behavior in Production

Read the full interview experience this question came from →

Quick Overview

A take-home asks you to build a system for managing prompts and model behavior in production, then defend its trade-offs in a product review and extend it for a gradual migration between model providers. Tests versioning, evaluation gates for open-ended output, safe rollout and rollback, observability, and scoping under a time box.

Build a System to Manage Prompts and Model Behavior in Production

Company: Luma

Role: Software Engineer

Category: ML System Design

Difficulty: medium

Interview Round: Onsite

This prompt is one of five offered in a take-home for a software engineering role on an AI product team. The candidate picks one prompt; across the set they cover frontend, backend and infrastructure work. The take-home is budgeted at 8–12 hours. The chat logs from the AI coding tools you use (such as Cursor, Codex or Claude Code) are submitted along with the code, and the work is judged on agentic development practice, 0-to-1 product scoping, and user or developer experience design. A later onsite round reviews the submission in detail. **Prompt:** build a system for managing prompts and model behavior in production. Take "model behavior" to mean everything that shapes what a large language model (LLM) feature does in production: the prompt template, the chosen model, its parameters, and any tools or output format it is given. Confirm this scope before you start. ### Constraints and Clarifications - Time box: 8–12 hours for design, code, tests and a written explanation. - AI coding tools are allowed, and their session logs are part of what is evaluated. - The review round treats the submission as a prototype, so anything you defer should be written down with a reason. ### Clarifying Questions - Who makes changes: only engineers, or also product and support staff who edit prompts through a UI? - Does "model behavior" cover only prompt text, or also model choice, parameters, tool definitions and safety settings? - Roughly how many prompts, features and models are in production, and how often do they change? - Is there an existing evaluation dataset or labeled feedback, or must the system build one from production traffic? - Does the application fetch configuration on every request, and what latency can that add? ### Part 1 — Scope and build the prototype Decide what the prototype must do within the time box and what it defers. Describe the data model, the API, how application code obtains the active prompt at request time, how a change is evaluated before it reaches users, and the interface people use to make and review changes. Explain how you will use AI coding agents so that the submitted logs show a disciplined workflow. ```hint Start from the incident Picture a prompt edit that quietly makes one feature's answers worse. Scope the prototype around noticing that before users do and undoing it quickly. ``` #### What This Part Should Cover - A versioning model that ties every production output to one exact configuration, and a way to switch versions without a code deploy. - How configurations reach the application at request time without making the management service a single point of failure. - An evaluation step that gates promotion, with checks suited to open-ended LLM output. - Agentic practice visible in the logs: a written spec as context, small tasks, tests, and reviewed diffs. ### Part 2 — Product review: trade-offs and scaling limits In the onsite review, the interviewer treats your submission as a prototype and probes it in detail; expect every shortcut to be found and asked about. Explain the trade-offs you made, where the system stops scaling first, which needs of teams running LLM features in production it does not yet meet, and where the interface for editing, comparing and promoting changes falls short. ```hint Follow one request Trace a single user request through your system and note every call it makes, then ask what happens to that request when each of those calls is slow or failing. ``` #### What This Part Should Cover - Trade-offs such as runtime configuration versus prompts kept in the code repository, and automated judging versus human review. - Load and cost estimates for configuration reads, trace logging and evaluation runs, and which limit is reached first. - Needs of larger or regulated teams: approvals, audit history, cost visibility, and detection of provider-side model changes. - Interface gaps for reviewers and non-engineers: diffs, a playground, readable evaluation results. ### Part 3 — A new customer requirement The interviewer then introduces requirements from other customers and asks how you would integrate them, how long the work would take, and how it changes the roadmap. The requirements vary; practice with this one: a customer wants to move a feature from one model provider to another, rolling the change out to a growing share of traffic, with automatic rollback if quality, error or cost metrics get worse. ```hint Decide what one user should see Before choosing percentages, decide whether a single end user may see both versions during the rollout, and what that implies for how traffic is split. ``` #### What This Part Should Cover - Integration: traffic splitting, a provider abstraction, per-version metrics and rollback triggers. - A task-level estimate with the riskiest unknowns called out. - A roadmap order and the measures that show the rollout worked. ### What a Strong Answer Covers - A prototype scope that fits the time box and still demonstrates the full loop: edit, evaluate, promote, observe, roll back. - Treating LLM output as non-deterministic in both evaluation and monitoring. - Continuity from Part 1 decisions to the Part 2 limits and the Part 3 plan. - Evidence of disciplined use of AI coding agents and clear written reasons for deferred features. ### Follow-up Questions - How would you evaluate a change when outputs are open-ended and no single answer is correct? - A provider updates a model behind the same name and behavior shifts. How would you detect it, and what would you do? - Logged prompts and completions can contain personal data. How would you store, redact and expire them? - How do you keep a prompt template and the application code that fills its variables compatible across versions?

Overview: A take-home asks you to build a system for managing prompts and model behavior in production, then defend its trade-offs in a product review and extend it for a gradual migration between model providers. Tests versioning, evaluation gates for open-ended output, safe rollout and rollback, observability, and scoping under a time box.

Read the full Luma Software Engineer interview experience this question came from

|Home/ML System Design/Luma
Luma logo
Luma
Sep 17, 2026
mediumSoftware EngineerOnsiteML System Design
0
0

This prompt is one of five offered in a take-home for a software engineering role on an AI product team. The candidate picks one prompt; across the set they cover frontend, backend and infrastructure work. The take-home is budgeted at 8–12 hours. The chat logs from the AI coding tools you use (such as Cursor, Codex or Claude Code) are submitted along with the code, and the work is judged on agentic development practice, 0-to-1 product scoping, and user or developer experience design. A later onsite round reviews the submission in detail.

Prompt: build a system for managing prompts and model behavior in production.

Take "model behavior" to mean everything that shapes what a large language model (LLM) feature does in production: the prompt template, the chosen model, its parameters, and any tools or output format it is given. Confirm this scope before you start.

Constraints and Clarifications

  • Time box: 8–12 hours for design, code, tests and a written explanation.
  • AI coding tools are allowed, and their session logs are part of what is evaluated.
  • The review round treats the submission as a prototype, so anything you defer should be written down with a reason.

Clarifying Questions Guidance

  • Who makes changes: only engineers, or also product and support staff who edit prompts through a UI?
  • Does "model behavior" cover only prompt text, or also model choice, parameters, tool definitions and safety settings?
  • Roughly how many prompts, features and models are in production, and how often do they change?
  • Is there an existing evaluation dataset or labeled feedback, or must the system build one from production traffic?
  • Does the application fetch configuration on every request, and what latency can that add?

Part 1 — Scope and build the prototype

Decide what the prototype must do within the time box and what it defers. Describe the data model, the API, how application code obtains the active prompt at request time, how a change is evaluated before it reaches users, and the interface people use to make and review changes. Explain how you will use AI coding agents so that the submitted logs show a disciplined workflow.

What This Part Should Cover Guidance

  • A versioning model that ties every production output to one exact configuration, and a way to switch versions without a code deploy.
  • How configurations reach the application at request time without making the management service a single point of failure.
  • An evaluation step that gates promotion, with checks suited to open-ended LLM output.
  • Agentic practice visible in the logs: a written spec as context, small tasks, tests, and reviewed diffs.

Part 2 — Product review: trade-offs and scaling limits

In the onsite review, the interviewer treats your submission as a prototype and probes it in detail; expect every shortcut to be found and asked about. Explain the trade-offs you made, where the system stops scaling first, which needs of teams running LLM features in production it does not yet meet, and where the interface for editing, comparing and promoting changes falls short.

What This Part Should Cover Guidance

  • Trade-offs such as runtime configuration versus prompts kept in the code repository, and automated judging versus human review.
  • Load and cost estimates for configuration reads, trace logging and evaluation runs, and which limit is reached first.
  • Needs of larger or regulated teams: approvals, audit history, cost visibility, and detection of provider-side model changes.
  • Interface gaps for reviewers and non-engineers: diffs, a playground, readable evaluation results.

Part 3 — A new customer requirement

The interviewer then introduces requirements from other customers and asks how you would integrate them, how long the work would take, and how it changes the roadmap. The requirements vary; practice with this one: a customer wants to move a feature from one model provider to another, rolling the change out to a growing share of traffic, with automatic rollback if quality, error or cost metrics get worse.

What This Part Should Cover Guidance

  • Integration: traffic splitting, a provider abstraction, per-version metrics and rollback triggers.
  • A task-level estimate with the riskiest unknowns called out.
  • A roadmap order and the measures that show the rollout worked.

What a Strong Answer Covers Guidance

  • A prototype scope that fits the time box and still demonstrates the full loop: edit, evaluate, promote, observe, roll back.
  • Treating LLM output as non-deterministic in both evaluation and monitoring.
  • Continuity from Part 1 decisions to the Part 2 limits and the Part 3 plan.
  • Evidence of disciplined use of AI coding agents and clear written reasons for deferred features.

Follow-up Questions Guidance

  • How would you evaluate a change when outputs are open-ended and no single answer is correct?
  • A provider updates a model behind the same name and behavior shifts. How would you detect it, and what would you do?
  • Logged prompts and completions can contain personal data. How would you store, redact and expire them?
  • How do you keep a prompt template and the application code that fills its variables compatible across versions?

Submit Your Answer to Earn 20XP

Sign in to leave a comment

Loading comments...