Write the Cross-Entropy Loss and Compute Precision, Recall and F1

Quick Overview

An ML fundamentals question in two parts: write the cross-entropy loss for binary and multi-class classification and explain why it is used, then compute precision, recall, F1 and accuracy for a described imbalanced classifier. It tests exact formulas, careful arithmetic and metric choice.

Write the Cross-Entropy Loss and Compute Precision, Recall and F1

Company: ByteDance

Role: Applied Scientist

Category: Machine Learning

Difficulty: medium

Interview Round: Technical Screen

Two ML fundamentals questions about classification: the cross-entropy loss used to train a classifier, and the F1 score used to evaluate one. ### Clarifying Questions - Should the loss be written for one example or averaged over a batch, and for binary classification, multi-class classification, or both? - Does the model output probabilities or raw scores (logits)? - In Part 2, which class is the positive class, and has a decision threshold already been applied to the predictions? ### Part 1 — Cross-entropy loss Write the cross-entropy loss for binary and for multi-class classification, defining every symbol. Explain why it is the standard training loss for classifiers. ```hint Start from the likelihood Write the probability the model assigns to the observed label of one example, then ask what maximizing it over the whole dataset looks like after taking logarithms. ``` #### What This Part Should Cover - Binary and multi-class formulas with every symbol defined - Its origin as the negative log-likelihood, and its link to the KL divergence - The gradient with respect to the logits, and why it trains better than squared error - A numerically stable way to compute it from logits ### Part 2 — Computing F1 Explain how precision, recall and the F1 score are computed. The numbers used in the interview are not known, so compute them for this example: a binary classifier is evaluated on 1,000 examples, 100 of which are positive. It predicts 80 examples as positive, and 60 of those are truly positive. Also compute the accuracy, and explain which metric you would report for this data. ```hint Fill the confusion matrix first Only one of the four confusion-matrix counts is stated directly. Derive the other three from the totals before computing any ratio. ``` #### What This Part Should Cover - The four confusion-matrix counts derived from the description - Precision, recall and F1 as their harmonic mean, with the computed values - Why accuracy misleads at this class balance - What F1 ignores, and how it depends on the decision threshold ### What a Strong Answer Covers - Exact formulas with consistent notation - Correct arithmetic, checked by a second route - The link between the two parts: a differentiable training loss versus a threshold-dependent evaluation metric - Awareness of class imbalance in both the loss and the metric ### Follow-up Questions - For a three-class problem, how do macro-averaged and micro-averaged F1 differ, and when can each one mislead? - F1 is not differentiable. Given a model trained with cross-entropy, how would you choose the decision threshold that maximizes F1? - How do class-weighted cross-entropy or focal loss change training on a heavily imbalanced dataset? - What happens to the cross-entropy if the model assigns probability 0 to the true class, and how do implementations guard against it?

Overview: An ML fundamentals question in two parts: write the cross-entropy loss for binary and multi-class classification and explain why it is used, then compute precision, recall, F1 and accuracy for a described imbalanced classifier. It tests exact formulas, careful arithmetic and metric choice.

|Home/Machine Learning/ByteDance
ByteDance logo
ByteDance
Oct 8, 2026
mediumApplied ScientistTechnical ScreenMachine Learning
0
0

Two ML fundamentals questions about classification: the cross-entropy loss used to train a classifier, and the F1 score used to evaluate one.

Clarifying Questions Guidance

  • Should the loss be written for one example or averaged over a batch, and for binary classification, multi-class classification, or both?
  • Does the model output probabilities or raw scores (logits)?
  • In Part 2, which class is the positive class, and has a decision threshold already been applied to the predictions?

Part 1 — Cross-entropy loss

Write the cross-entropy loss for binary and for multi-class classification, defining every symbol. Explain why it is the standard training loss for classifiers.

What This Part Should Cover Guidance

  • Binary and multi-class formulas with every symbol defined
  • Its origin as the negative log-likelihood, and its link to the KL divergence
  • The gradient with respect to the logits, and why it trains better than squared error
  • A numerically stable way to compute it from logits

Part 2 — Computing F1

Explain how precision, recall and the F1 score are computed. The numbers used in the interview are not known, so compute them for this example: a binary classifier is evaluated on 1,000 examples, 100 of which are positive. It predicts 80 examples as positive, and 60 of those are truly positive. Also compute the accuracy, and explain which metric you would report for this data.

What This Part Should Cover Guidance

  • The four confusion-matrix counts derived from the description
  • Precision, recall and F1 as their harmonic mean, with the computed values
  • Why accuracy misleads at this class balance
  • What F1 ignores, and how it depends on the decision threshold

What a Strong Answer Covers Guidance

  • Exact formulas with consistent notation
  • Correct arithmetic, checked by a second route
  • The link between the two parts: a differentiable training loss versus a threshold-dependent evaluation metric
  • Awareness of class imbalance in both the loss and the metric

Follow-up Questions Guidance

  • For a three-class problem, how do macro-averaged and micro-averaged F1 differ, and when can each one mislead?
  • F1 is not differentiable. Given a model trained with cross-entropy, how would you choose the decision threshold that maximizes F1?
  • How do class-weighted cross-entropy or focal loss change training on a heavily imbalanced dataset?
  • What happens to the cross-entropy if the model assigns probability 0 to the true class, and how do implementations guard against it?
Loading comments...