LLM fundamentals: Transformer internals, beam search decoding, and LoRA fine-tuning
Company: UiPath
Role: Machine Learning Engineer
Category: Machine Learning
Difficulty: medium
Interview Round: Technical Screen
Explain three building blocks of modern large language models in enough depth to handle follow-up questions on the details. Answer the parts in order.
### Clarifying Questions
- Should the Transformer explanation focus on the decoder-only architecture used by most current LLMs, or on the original encoder-decoder design?
- For beam search, is the context open-ended text generation or tasks with a single expected output?
### Part 1 — Transformer principles and details
Explain how a Transformer processes a sequence: how self-attention is computed and why it is scaled, what multi-head attention adds, how position is represented, what the feed-forward sublayer, residual connections and normalization do, and how masking lets a decoder-only LLM generate text.
```hint Scaling the scores
Think about what happens to the softmax when the query and key vectors get wider.
```
#### What This Part Should Cover
- The attention computation with its projections, scaling and masking
- Multi-head attention and positional information
- Block structure (feed-forward sublayer, residuals, normalization), the training objective, and autoregressive generation with its cost in sequence length
### Part 2 — Beam search
Explain how beam search decodes text, how it compares with greedy decoding and with sampling, and when it is or is not a good choice for an LLM.
```hint What greedy throws away
Consider what greedy decoding discards at each step and what it would cost to keep some of it.
```
#### What This Part Should Cover
- The algorithm: beam width, cumulative log-probability scoring, finished hypotheses, stopping
- Length normalization and computational cost
- Comparison with greedy decoding and sampling, and suitable versus unsuitable tasks
### Part 3 — LoRA
Explain what LoRA is, why it makes fine-tuning cheaper, what its main hyperparameters are, and how a LoRA-tuned model is served.
```hint Shape of the update
Think about the rank of the change that fine-tuning makes to a large weight matrix.
```
#### What This Part Should Cover
- The low-rank update to frozen weights and its initialization
- Parameter and memory savings, with a worked count
- Rank, scaling and target modules, merging versus serving separate adapters, and limitations
### What a Strong Answer Covers
- Precise mechanics, including the formulas and the reasons behind design choices
- Awareness of computational cost: sequence length, beam width, trainable parameters
- Practical judgment about when each technique fits
- Correct handling of the details interviewers probe: the scaling factor, the causal mask, length normalization, adapter initialization
### Follow-up Questions
- Why does a KV cache speed up autoregressive generation, and what limits how far you can use it?
- Why can beam search produce repetitive or generic text in open-ended generation, and what mitigates it?
- When fine-tuning an agent model on tool-use traces with LoRA, how would you choose the rank and which weight matrices to adapt?
- How does QLoRA change the memory picture compared with LoRA?
Overview: An LLM fundamentals question that asks you to explain Transformer internals such as scaled self-attention, multi-head attention, positional information and causal masking. It also covers how beam search decoding compares with greedy decoding and sampling, and how LoRA reduces the cost of fine-tuning.
Read the full UiPath Machine Learning Engineer interview experience this question came from