LLM fundamentals: Transformer internals, beam search decoding, and LoRA fine-tuning

Read the full interview experience this question came from →

Quick Overview

An LLM fundamentals question that asks you to explain Transformer internals such as scaled self-attention, multi-head attention, positional information and causal masking. It also covers how beam search decoding compares with greedy decoding and sampling, and how LoRA reduces the cost of fine-tuning.

LLM fundamentals: Transformer internals, beam search decoding, and LoRA fine-tuning

Company: UiPath

Role: Machine Learning Engineer

Category: Machine Learning

Difficulty: medium

Interview Round: Technical Screen

Explain three building blocks of modern large language models in enough depth to handle follow-up questions on the details. Answer the parts in order. ### Clarifying Questions - Should the Transformer explanation focus on the decoder-only architecture used by most current LLMs, or on the original encoder-decoder design? - For beam search, is the context open-ended text generation or tasks with a single expected output? ### Part 1 — Transformer principles and details Explain how a Transformer processes a sequence: how self-attention is computed and why it is scaled, what multi-head attention adds, how position is represented, what the feed-forward sublayer, residual connections and normalization do, and how masking lets a decoder-only LLM generate text. ```hint Scaling the scores Think about what happens to the softmax when the query and key vectors get wider. ``` #### What This Part Should Cover - The attention computation with its projections, scaling and masking - Multi-head attention and positional information - Block structure (feed-forward sublayer, residuals, normalization), the training objective, and autoregressive generation with its cost in sequence length ### Part 2 — Beam search Explain how beam search decodes text, how it compares with greedy decoding and with sampling, and when it is or is not a good choice for an LLM. ```hint What greedy throws away Consider what greedy decoding discards at each step and what it would cost to keep some of it. ``` #### What This Part Should Cover - The algorithm: beam width, cumulative log-probability scoring, finished hypotheses, stopping - Length normalization and computational cost - Comparison with greedy decoding and sampling, and suitable versus unsuitable tasks ### Part 3 — LoRA Explain what LoRA is, why it makes fine-tuning cheaper, what its main hyperparameters are, and how a LoRA-tuned model is served. ```hint Shape of the update Think about the rank of the change that fine-tuning makes to a large weight matrix. ``` #### What This Part Should Cover - The low-rank update to frozen weights and its initialization - Parameter and memory savings, with a worked count - Rank, scaling and target modules, merging versus serving separate adapters, and limitations ### What a Strong Answer Covers - Precise mechanics, including the formulas and the reasons behind design choices - Awareness of computational cost: sequence length, beam width, trainable parameters - Practical judgment about when each technique fits - Correct handling of the details interviewers probe: the scaling factor, the causal mask, length normalization, adapter initialization ### Follow-up Questions - Why does a KV cache speed up autoregressive generation, and what limits how far you can use it? - Why can beam search produce repetitive or generic text in open-ended generation, and what mitigates it? - When fine-tuning an agent model on tool-use traces with LoRA, how would you choose the rank and which weight matrices to adapt? - How does QLoRA change the memory picture compared with LoRA?

Overview: An LLM fundamentals question that asks you to explain Transformer internals such as scaled self-attention, multi-head attention, positional information and causal masking. It also covers how beam search decoding compares with greedy decoding and sampling, and how LoRA reduces the cost of fine-tuning.

Read the full UiPath Machine Learning Engineer interview experience this question came from

|Home/Machine Learning/UiPath
UiPath logo
UiPath
Sep 9, 2026
mediumMachine Learning EngineerTechnical ScreenMachine Learning
0
0

Explain three building blocks of modern large language models in enough depth to handle follow-up questions on the details. Answer the parts in order.

Clarifying Questions Guidance

  • Should the Transformer explanation focus on the decoder-only architecture used by most current LLMs, or on the original encoder-decoder design?
  • For beam search, is the context open-ended text generation or tasks with a single expected output?

Part 1 — Transformer principles and details

Explain how a Transformer processes a sequence: how self-attention is computed and why it is scaled, what multi-head attention adds, how position is represented, what the feed-forward sublayer, residual connections and normalization do, and how masking lets a decoder-only LLM generate text.

What This Part Should Cover Guidance

  • The attention computation with its projections, scaling and masking
  • Multi-head attention and positional information
  • Block structure (feed-forward sublayer, residuals, normalization), the training objective, and autoregressive generation with its cost in sequence length

Explain how beam search decodes text, how it compares with greedy decoding and with sampling, and when it is or is not a good choice for an LLM.

What This Part Should Cover Guidance

  • The algorithm: beam width, cumulative log-probability scoring, finished hypotheses, stopping
  • Length normalization and computational cost
  • Comparison with greedy decoding and sampling, and suitable versus unsuitable tasks

Part 3 — LoRA

Explain what LoRA is, why it makes fine-tuning cheaper, what its main hyperparameters are, and how a LoRA-tuned model is served.

What This Part Should Cover Guidance

  • The low-rank update to frozen weights and its initialization
  • Parameter and memory savings, with a worked count
  • Rank, scaling and target modules, merging versus serving separate adapters, and limitations

What a Strong Answer Covers Guidance

  • Precise mechanics, including the formulas and the reasons behind design choices
  • Awareness of computational cost: sequence length, beam width, trainable parameters
  • Practical judgment about when each technique fits
  • Correct handling of the details interviewers probe: the scaling factor, the causal mask, length normalization, adapter initialization

Follow-up Questions Guidance

  • Why does a KV cache speed up autoregressive generation, and what limits how far you can use it?
  • Why can beam search produce repetitive or generic text in open-ended generation, and what mitigates it?
  • When fine-tuning an agent model on tool-use traces with LoRA, how would you choose the rank and which weight matrices to adapt?
  • How does QLoRA change the memory picture compared with LoRA?
Loading comments...