Detect Harmful Content in LLM Prompts and Responses
Company: Databricks
Role: Machine Learning Engineer
Category: ML System Design
Difficulty: medium
Interview Round: Onsite
Design a machine learning system that detects harmful content submitted to, or produced by, a large language model (LLM) product.
The report records only this prompt from an ML design round. The scope, the harm categories and the product constraints are yours to clarify and state.
```hint Where does the check sit?
List the points in a chat request's lifetime where content could be checked, and what each point can and cannot see.
```
```hint Not all mistakes cost the same
Think separately about the cost of a missed harmful response and the cost of refusing a benign one, and how each should shape the per-category thresholds.
```
### Clarifying Questions
- Should the system screen user prompts, model responses, or both?
- Which harm categories are in scope, and is there a written policy defining each one?
- What happens on a detection: block, warn, rewrite the response, or route to human review?
- How much latency can each request absorb, and are responses streamed token by token?
- Which languages and modalities (text only, code, images) must be covered?
### What a Strong Answer Covers
- Framing as classification against an explicit policy taxonomy, with an action tied to each label
- A data strategy: labeling guidelines, human annotation, adversarial and red-team data, and handling of rare categories
- Model choices, and a layered design that balances latency, cost and accuracy
- Offline evaluation per category: precision and recall at the operating thresholds, over-blocking of benign content, and adversarial robustness
- A serving design for prompts and streaming responses that fits the latency budget
- Monitoring, feedback loops, and fast response to new attack patterns
### Follow-up Questions
- How would you detect a harmful request split across many innocuous-looking turns of a conversation?
- How would you keep the classifier robust to jailbreak prompts that change quickly?
- How would you measure recall in production, when you cannot see most of the harmful content you missed?
Overview: Design a machine learning system that detects harmful content in prompts sent to, and responses produced by, a large language model. Tests policy-driven framing, labeling and adversarial data, layered classifiers under latency limits, per-category evaluation, streaming checks and monitoring.