Write the Cross-Entropy Loss and Compute Precision, Recall and F1
Company: ByteDance
Role: Applied Scientist
Category: Machine Learning
Difficulty: medium
Interview Round: Technical Screen
Two ML fundamentals questions about classification: the cross-entropy loss used to train a classifier, and the F1 score used to evaluate one.
### Clarifying Questions
- Should the loss be written for one example or averaged over a batch, and for binary classification, multi-class classification, or both?
- Does the model output probabilities or raw scores (logits)?
- In Part 2, which class is the positive class, and has a decision threshold already been applied to the predictions?
### Part 1 — Cross-entropy loss
Write the cross-entropy loss for binary and for multi-class classification, defining every symbol. Explain why it is the standard training loss for classifiers.
```hint Start from the likelihood
Write the probability the model assigns to the observed label of one example, then ask what maximizing it over the whole dataset looks like after taking logarithms.
```
#### What This Part Should Cover
- Binary and multi-class formulas with every symbol defined
- Its origin as the negative log-likelihood, and its link to the KL divergence
- The gradient with respect to the logits, and why it trains better than squared error
- A numerically stable way to compute it from logits
### Part 2 — Computing F1
Explain how precision, recall and the F1 score are computed. The numbers used in the interview are not known, so compute them for this example: a binary classifier is evaluated on 1,000 examples, 100 of which are positive. It predicts 80 examples as positive, and 60 of those are truly positive. Also compute the accuracy, and explain which metric you would report for this data.
```hint Fill the confusion matrix first
Only one of the four confusion-matrix counts is stated directly. Derive the other three from the totals before computing any ratio.
```
#### What This Part Should Cover
- The four confusion-matrix counts derived from the description
- Precision, recall and F1 as their harmonic mean, with the computed values
- Why accuracy misleads at this class balance
- What F1 ignores, and how it depends on the decision threshold
### What a Strong Answer Covers
- Exact formulas with consistent notation
- Correct arithmetic, checked by a second route
- The link between the two parts: a differentiable training loss versus a threshold-dependent evaluation metric
- Awareness of class imbalance in both the loss and the metric
### Follow-up Questions
- For a three-class problem, how do macro-averaged and micro-averaged F1 differ, and when can each one mislead?
- F1 is not differentiable. Given a model trained with cross-entropy, how would you choose the decision threshold that maximizes F1?
- How do class-weighted cross-entropy or focal loss change training on a heavily imbalanced dataset?
- What happens to the cross-entropy if the model assigns probability 0 to the true class, and how do implementations guard against it?
Overview: An ML fundamentals question in two parts: write the cross-entropy loss for binary and multi-class classification and explain why it is used, then compute precision, recall, F1 and accuracy for a described imbalanced classifier. It tests exact formulas, careful arithmetic and metric choice.