Using, Evaluating and Adapting a Pretrained Large Language Model
Company: Wells Fargo
Role: Data Scientist
Category: Machine Learning
Difficulty: hard
Interview Round: Technical Screen
You are given a pretrained large language model. How would you use it, how would you evaluate it, and how would you adapt it to your needs? The interviewer asked the question in this open form, without naming a task; make the answer concrete by choosing an example task or by asking for one.
### Clarifying Questions
- What is the target task and domain, and what does a good output look like?
- Is the model a base model or already instruction-tuned, and are its weights available or only an API?
- How much labeled and unlabeled task data is available, and what compute budget?
- What latency, cost and data-privacy constraints apply in deployment?
### Part 1 — Using the model
Before changing any weights, how would you put the model to work on a task?
```hint Cheapest lever first
Rank the ways of getting value from the model by their cost and by how much each one changes the model itself.
```
#### What This Part Should Cover
- Inspecting what the model is: type, size, context window, tokenizer, license
- Prompting strategies and structured output
- Supplying missing knowledge without retraining
- A serving setup that accounts for latency and cost
### Part 2 — Evaluating the model
How would you test whether the model is good enough for the task, and how would you detect its problems?
```hint Your data, not the leaderboard
Think about why a public benchmark score may say little about your task, and what evaluation data you would build instead.
```
```hint Beyond accuracy
List the ways the model can fail other than by giving a wrong answer.
```
#### What This Part Should Cover
- A task-specific evaluation set, with metrics matched to the output type
- Baselines and contamination checks
- Hallucination, robustness, bias, safety and calibration checks
- Operational metrics and a regression process
### Part 3 — Adapting the model
If prompting is not enough, how would you adjust the model, and how would you choose among the options?
```hint Diagnose the gap
Ask whether the model lacks knowledge, the right behavior or format, or familiarity with the domain's language. Each gap calls for a different tool.
```
#### What This Part Should Cover
- Matching the adaptation method to the diagnosed gap
- Data and training choices for fine-tuning, including parameter-efficient methods
- Risks such as forgetting and overfitting, and how to detect them
- Re-evaluation and deployment after adaptation
### What a Strong Answer Covers
- An escalation path from prompting to retrieval to fine-tuning, justified by cost and need
- Evaluation designed around the task, including failure modes beyond accuracy
- Correct mechanics of at least one parameter-efficient fine-tuning method
- Trade-offs between retrieval and fine-tuning, and between model size and serving cost
- A closed loop: evaluate, adapt, re-evaluate and monitor
### Follow-up Questions
- What do you gain and lose with LoRA compared with full fine-tuning?
- After fine-tuning, the model is better at your task but worse at following general instructions. Why, and what do you do?
- How would you evaluate free-form generations that have no single correct answer?
- How would you check whether your evaluation set leaked into the model's pretraining data?
Overview: Given a pretrained large language model, explain how you would put it to use, how you would evaluate it on your task, and how you would adapt it when prompting is not enough. It tests practical judgment across prompting, retrieval, evaluation design, failure detection and fine-tuning choices.
Read the full Wells Fargo Data Scientist interview experience this question came from