BERT Interview Questions: Pretraining Objectives, Attention Masks, and Task Heads
Quick Overview
Trace BERT pretraining and fine-tuning through an original token batch. Distinguish corruption, attention masks, ignored loss labels, and task-head outputs, with primary-source boundaries and debugging exercises.
BERT interview questions become easier when you separate three decisions: which input tokens were corrupted, which positions attention may use, and which outputs contribute to the loss. A [MASK] token is usually visible to BERT's attention. A padding position is excluded as context by the padding mask. Neither decision alone specifies the training labels.
This article traces those decisions through a small batch, then changes the task head without confusing hidden states with predictions. Use PracHub's BERT span-extraction exercise to practice turning a language-model discussion into a precise input, output, and label contract.
Evidence boundary: Architecture and pretraining facts come from the original BERT paper and Google's implementation; API conventions come from Hugging Face documentation checked on October 4, 2026. Worked tokens and calculations are original exercises, not actual tokenizer outputs or candidate reports. No employer-specific question frequency is claimed, and the numerical checks do not represent a pretrained-model training run.

Why can BERT use both left and right context?
Official research fact: The BERT paper introduces deep bidirectional encoder representations trained with masked language modeling and next sentence prediction. In the ordinary encoder setup, a real token may attend to real tokens on either side. This differs from the usual causal decoder restriction that prevents a position from reading future tokens.
The important follow-up is why token prediction does not become trivial copying. For selected training positions, the input is often corrupted while the target remains the original token. The model must use the supplied context to reconstruct the target. However, some selected tokens are deliberately left unchanged; the original recipe does not eliminate all copying opportunities at every selected position.
A concise interview answer should name architecture, context visibility, corruption, and supervision separately. Saying “BERT is bidirectional because it predicts masked words” leaves the attention rule implicit. Saying “BERT uses a mask” leaves unclear whether you mean input replacement or attention restrictions.
An interviewer can change the question to sentence classification without changing the encoder's contextual attention. The task head and training labels change; a triangular causal mask does not suddenly become necessary because the task has a label.
What do the 15% and 80/10/10 numbers mean?
Official implementation fact: Google's pretraining data generator selects prediction positions and applies replacement choices. In the original recipe, roughly 15% of eligible tokens are selected; among selected tokens, 80% receive [MASK], 10% a random token, and 10% remain unchanged.
Those percentages have different denominators. For an idealized collection of 1,000 eligible positions, the expected selected count is 150. Expected visible [MASK] replacements are 120, random replacements 15, and unchanged selected positions 15. These are expectations under the probabilities, not a guarantee about each short sentence or each implementation's rounding and caps.
All 150 selected positions are prediction targets in this example. Only 120 visibly contain [MASK]. Therefore, finding mask-token IDs after corruption is insufficient to reconstruct the complete selected-position set.
The Hugging Face language-modeling collator documents the corresponding default selection and replacement probabilities, along with special-token handling. A collator performs batch preparation; it does not make the encoder causal or choose a downstream classification head.
Preparation inference: Preserve a selection mask before changing input IDs. Use it to construct the loss labels. This makes random replacements and unchanged selected tokens behave correctly without guessing which tokens were selected after the fact.
Trace input IDs, attention masks, and labels independently
Use five schematic positions: [CLS], cats, sleep, [SEP], [PAD]. Select sleep as the prediction target and replace it with [MASK]. This is a teaching sequence; it does not assert a particular WordPiece segmentation or vocabulary ID.
| Position after corruption | Attention mask and MLM label |
|---|---|
[CLS] | 1; ignored loss label |
cats | 1; ignored loss label |
[MASK] replacing sleep | 1; original sleep token ID |
[SEP] | 1; ignored loss label |
[PAD] | 0; ignored loss label |
Official API convention: Hugging Face's BERT reference uses an attention mask with 1 for positions to attend to and 0 for masked padding context. For masked-language-model labels, the documented ignored value is -100.
The [MASK] position has attention value 1. It is a real position with a learned representation, and it must read useful context to predict sleep. Setting its attention value to 0 because its spelling contains “mask” confuses two independent mechanisms.
The loss label at cats is ignored even though attention may use that token. The padding label is also ignored, but for a different reason: it is not a useful prediction target. A zero attention entry is not a substitute for a correctly constructed loss label.
Now leave the selected sleep unchanged instead. Its attention value remains 1 and its target remains sleep. It still contributes to MLM loss. The visible token string changed, but the selection decision did not.

Does a padding mask force padding outputs to zero?
No. Excluding a padding position as an attention key does not automatically zero every hidden state or logit at that position. A padded query can still produce an output through the surrounding computation. Do not infer that a zero mask entry means a zero vector in the returned tensor.
This is an original debugging distinction derived from the mask's purpose. If your token-level metric includes padding predictions, it may report misleading accuracy even though the encoder received the correct attention mask. Exclude ignored labels explicitly when computing that metric.
Similarly, masking attention does not prevent a downstream aggregation from including padding outputs unless that aggregation applies its own valid-position rule. That issue belongs to embedding pooling and mask handling; here, the critical interview skill is tracing supervision and task outputs.
When moving from the high-level BERT interface to a low-level attention API, check Boolean semantics and broadcasting. Some APIs interpret a true value as “blocked,” while others use “allowed.” Explain the convention of the actual function before converting the mask. A mechanically inverted tensor can silently change which context the model reads.
What changes when you choose a task head?
An encoder hidden state is a representation, not a class probability. Let batch size be B, padded sequence length L, hidden width H, vocabulary size V, and task class count C. Trace the tensor through the chosen prediction boundary.
The current Transformers glossary describes task-dependent labels and output shapes. For this original shape exercise, use B=2 and L=5. Assume a hypothetical H=8, V=20, and C=3 so every dimension is easy to inspect.
| Output in the exercise | Shape and supervised target |
|---|---|
| Encoder hidden states | (2, 5, 8); representations before a task head |
| MLM vocabulary logits | (2, 5, 20); original token IDs at selected positions |
| Sequence classification logits | (2, 3); one class per sequence |
| Token classification logits | (2, 5, 3); one label per supervised token |
| Extractive QA start and end logits | Two (2, 5) tensors; answer boundary positions |
These are hypothetical dimensions, not a specification of a standard BERT checkpoint. A classification head with three outputs cannot directly implement vocabulary prediction with twenty possible targets. A tensor can have valid numerical values and still encode the wrong task.
For sequence classification, ask whether the problem is single-label classification, multilabel classification, or regression. That decision determines label representation, loss, and interpretation of outputs. Three numbers do not by themselves tell you whether to apply softmax or independent sigmoids.
For extractive QA, confirm that answers correspond to spans in the tokenized input. A start/end head does not generate arbitrary answer text. Truncating away the answer or misaligning character offsets with subword indices creates a label problem that a larger encoder cannot repair.
How do NSP and downstream classification differ?
Original BERT's next sentence prediction objective classifies whether a supplied second segment follows the first in the constructed training example. The original paper describes a binary pretraining objective. It is not a universal measure of semantic similarity and is not the same task as classifying the sentiment of a customer review.
Segment identifiers, separator tokens, and attention masks also have different roles. Segment identifiers distinguish parts of an input pair; they do not independently block cross-segment attention. Two segments can participate in bidirectional contextual computation while remaining identifiable as segment A and segment B.
During downstream fine-tuning, the supervised label should match the actual task. A new classification head usually requires task training; loading encoder weights does not establish that its new outputs already carry your business labels. Check initialization messages and the checkpoint's declared architecture.
Avoid generalizing the original two-objective recipe to every model with “BERT” in its family name. Discuss the specific checkpoint and training recipe when the interviewer asks about a later variant. Historical BERT facts should remain historical facts.
Calculate the selected-position loss before debugging a model
Here is an original numerical exercise independent of any library. Suppose two selected positions assign probabilities 0.8 and 0.25 to their original tokens. Mean negative log-likelihood over those targets is:
loss = (-ln(0.8) - ln(0.25)) / 2
≈ 0.80472
If you divide the same sum by ten total batch positions, you obtain about 0.16094. That smaller value results from a different denominator; it does not indicate better predictions. Define whether the reduction averages selected tokens, sequences, or another unit before comparing losses.
Now add eight ignored positions with arbitrary logits. Under the selected-position calculation, the loss must remain unchanged. This is a useful invariant for a small preprocessing or metric fixture. It tests the target-selection logic without requiring a pretrained model download.
Also handle a batch with no selected targets. Depending on your selection scheme, this can happen in short examples. Choose an explicit policy, such as resampling or skipping the loss contribution, rather than accidentally interpreting an undefined reduction as a valid zero-loss success. Check the behavior of the library version you actually use.
What makes a debugging answer convincing?
Preparation recommendation: Start with one printed batch and a shape ledger. Inspect original IDs, corrupted IDs, selected positions, attention mask, and labels before changing optimizer settings. Verify that special and padding positions follow the intended selection policy.
Then inspect the task head and label contract. For token classification, decide how labels align with multiple subwords; for MLM, preserve original targets before corruption. For classification, separate evaluation mode from gradient disabling: disabling gradients alone does not turn off training-time dropout.
Use a bounded diagnostic sequence. Confirm a tiny supervised example can be learned, inspect nonfinite loss and gradient values, and ensure evaluation examples were not used to construct training targets. A tiny overfit check can reveal wiring mistakes, but it does not establish generalization or production quality.
These PracHub records provide related practice. They are reported question records, not predictions of a particular employer's interview.
| PracHub question | What to make explicit |
|---|---|
| Clarify a BERT Keyphrase Span-Extraction Task | Span labels, token boundaries, and output contract |
| Optimizing BERT inference latency and throughput for a production NLP service | Batch length and serving behavior |
| Explain Transformer, GPT vs BERT, and PR metrics | Context visibility and task evaluation |
| Explain Transformer Encoder and Decoder Behavior | Architecture and attention restrictions |
| Debug and fix a PyTorch Transformer training loop | Inputs, targets, losses, and training state |
Practice the BERT span-extraction question by stating tensor shapes and label alignment before proposing a model change. Then explain how your answer would change for MLM or sequence classification.
Comments (0)